Parent directory

AGENTS.md

5564 bytes

kaiwari — Japanese learning RPG (TUI + HTTP server)

Overview

A game for practicing spoken Japanese. The play loop is an adventure log: the player types English actions to explore an LLM-narrated world, speaks Japanese to people at locations through push-to-talk (each spoken line gets separate fluency feedback from a judge), and can ask scratch questions with ?question for how to say something in Japanese. New conversation partners get a hidden character sheet generated from their introducing reply so they stay consistent across visits. Only romaji is ever shown to the player; kana and kanji stay internal.

Two entry points share one game core (internal/game), the LLM layer (internal/llm), config, and bootstrap:

  • cmd/kaiwari — a Bubble Tea TUI. Runs the game in-process: it builds its own state + orchestrator with local adapters (mic capture via arecord, local audio playback) and talks straight to the model services.
  • cmd/kaiwari-server — an HTTP server exposing in-memory game sessions for a thin external client (e.g. iOS). The client records and plays audio; the server does ASR transcription, the game loop, the judge, and TTS synthesis (returns the NPC line as WAV bytes; replay is client-side). Turns serialize per session; sessions die with the process. All requests need a bearer token.

Both entry points self-host two model services as child subprocesses by default (spawned on loopback at startup, torn down on exit):

  • LLM — llama.cpp router (OpenAI-compatible chat completions) for NPC dialogue, the judge, compaction, character sheets, and scratch questions. Prompts live in internal/llm/prompt.go; strict FIELD|value output contracts are parsed in internal/llm/contract.go. Spawned at 127.0.0.1:9931.
  • Audio — one audio.cpp server: Qwen3-ASR for STT (the raw Japanese transcript stays internal; it is what the LLM sees as the player's lines) and an OpenAI-compatible speech endpoint that plays the NPC's kana. Spawned at 127.0.0.1:9932.

With --disable-model-loading, the entry points connect to external services at whatever URLs are configured instead of spawning them.

Turn flow (spoken): record → transcribe → judge + NPC reply in parallel → synthesize the NPC's kana to audio → record the turn in per-location history. Typed actions skip recording, transcription, and the judge; scratch questions only touch the scratch model and never enter world state. Per-location history is compacted into summaries when it grows past budget (internal/game/state.go, internal/llm/history.go). Raw model I/O for debugging is appended to kaiwari_raw.log in the working directory.

How things are done here

Model services

The entry points spawn llama-server and audiocpp_server as child processes by default (see internal/services/). GPU VRAM is normally almost fully allocated to the chat model, so a 16 GiB free-VRAM gate runs before spawning.

  • Never start, stop, or restart the llama-server or audiocpp_server yourself during development (the binaries manage their own children at runtime).
  • Never run llama-cli, or use a command that can load a model into VRAM.

Product invariants

  • Render romaji only; never expose kana or kanji to the player — not in the TUI and not across the kaiwari-server wire.
  • NPCs behave as normal people, not language teachers. Grading stays separate from NPC dialogue; judge output never enters the NPC prompt.
  • Persona data contains no game, player, NPC, quest, or scenario context.
  • Keep the Bubble Tea event loop non-blocking.

Code style

  • Small single-purpose packages under internal/; adapters implement the small interfaces defined at their point of use in internal/game/orchestrator.go.
  • Speech has two adapter pairs over the same interface: local (mic via arecord + player) in internal/adapters for the TUI, and remote (uploaded WAV in, returned WAV out) in internal/server/speech.go for kaiwari-server.
  • LLM replies are parsed from strict FIELD|value contracts, never freeform JSON.
  • Errors wrap with %w. All startup failures are fatal (VRAM check, service spawn, health check, LLM warmup).

Subagents

  • Run at most one worker subagent at a time. Never spawn two in parallel: the harness model does not support concurrent turns, so parallel agents overwrite each other's context cache and produce corrupt results.
  • Sequence work in dependency-ordered batches; wait for each agent to finish before starting the next.
  • Give workers a plain-language description of the change (goal, behavior, constraints, how to verify), not a full file to transcribe. The main context does the design and thinking; the worker turns that into code. Pasting whole files wastes context and defeats the purpose of offloading work.
  • Workers commit their changes when done: one commit per batch with a short, polished message describing the change.

Development

  • Clean out dead code by default: when a change leaves functions, types, fields, or imports unused, remove them — including anything newly orphaned by that removal.
  • Verify changes with make build, make fmt, and make lint (wraps go build, golangci-lint fmt, and golangci-lint run). Do not write any tests; there is no test suite in this project.
  • Only the user may prepare live services for manual integration.
  • Docs: README.md covers running and configuration only.
  • DO NOT prepend bash commands with cd for the directory you're already