AGENTS.md
kaiwari — Japanese learning RPG (TUI + HTTP server)
Overview
A game for practicing spoken Japanese. The play loop is an adventure log:
the player types English actions to explore an LLM-narrated world, speaks
Japanese to people at locations through push-to-talk (each spoken line gets
separate fluency feedback from a judge), and can ask scratch questions with
?question for how to say something in Japanese. New conversation partners get
a hidden character sheet generated from their introducing reply so they stay
consistent across visits. Only romaji is ever shown to the player; kana and kanji stay internal.
Two entry points share one game core (internal/game), the LLM layer
(internal/llm), config, and bootstrap:
cmd/kaiwari— a Bubble Tea TUI. Runs the game in-process: it builds its own state + orchestrator with local adapters (mic capture viaarecord, local audio playback) and talks straight to the model services.cmd/kaiwari-server— an HTTP server exposing in-memory game sessions for a thin external client (e.g. iOS). The client records and plays audio; the server does ASR transcription, the game loop, the judge, and TTS synthesis (returns the NPC line as WAV bytes; replay is client-side). Turns serialize per session; sessions die with the process. All requests need a bearer token.
Both entry points self-host two model services as child subprocesses by default (spawned on loopback at startup, torn down on exit):
- LLM — llama.cpp router (OpenAI-compatible chat completions) for NPC
dialogue, the judge, compaction, character sheets, and scratch questions.
Prompts live in
internal/llm/prompt.go; strictFIELD|valueoutput contracts are parsed ininternal/llm/contract.go. Spawned at127.0.0.1:9931. - Audio — one audio.cpp server: Qwen3-ASR for STT (the raw Japanese
transcript stays internal; it is what the LLM sees as the player's lines) and
an OpenAI-compatible speech endpoint that plays the NPC's kana. Spawned at
127.0.0.1:9932.
With --disable-model-loading, the entry points connect to external services
at whatever URLs are configured instead of spawning them.
Turn flow (spoken): record → transcribe → judge + NPC reply in parallel → synthesize the NPC's kana to audio → record
the turn in per-location history. Typed actions skip recording,
transcription, and the judge; scratch questions only touch the scratch model
and never enter world state. Per-location history is compacted into summaries
when it grows past budget (internal/game/state.go,
internal/llm/history.go). Raw model I/O for debugging is appended to
kaiwari_raw.log in the working directory.
How things are done here
Model services
The entry points spawn llama-server and audiocpp_server as child processes
by default (see internal/services/). GPU VRAM is normally almost fully
allocated to the chat model, so a 16 GiB free-VRAM gate runs before spawning.
- Never start, stop, or restart the
llama-serveroraudiocpp_serveryourself during development (the binaries manage their own children at runtime). - Never run
llama-cli, or use a command that can load a model into VRAM.
Product invariants
- Render romaji only; never expose kana or kanji to the player — not in the TUI and not across the kaiwari-server wire.
- NPCs behave as normal people, not language teachers. Grading stays separate from NPC dialogue; judge output never enters the NPC prompt.
- Persona data contains no game, player, NPC, quest, or scenario context.
- Keep the Bubble Tea event loop non-blocking.
Code style
- Small single-purpose packages under
internal/; adapters implement the small interfaces defined at their point of use ininternal/game/orchestrator.go. - Speech has two adapter pairs over the same interface: local (mic via
arecord+ player) ininternal/adaptersfor the TUI, and remote (uploaded WAV in, returned WAV out) ininternal/server/speech.gofor kaiwari-server. - LLM replies are parsed from strict
FIELD|valuecontracts, never freeform JSON. - Errors wrap with
%w. All startup failures are fatal (VRAM check, service spawn, health check, LLM warmup).
Subagents
- Run at most one worker subagent at a time. Never spawn two in parallel: the harness model does not support concurrent turns, so parallel agents overwrite each other's context cache and produce corrupt results.
- Sequence work in dependency-ordered batches; wait for each agent to finish before starting the next.
- Give workers a plain-language description of the change (goal, behavior, constraints, how to verify), not a full file to transcribe. The main context does the design and thinking; the worker turns that into code. Pasting whole files wastes context and defeats the purpose of offloading work.
- Workers commit their changes when done: one commit per batch with a short, polished message describing the change.
Development
- Clean out dead code by default: when a change leaves functions, types, fields, or imports unused, remove them — including anything newly orphaned by that removal.
- Verify changes with
make build,make fmt, andmake lint(wrapsgo build,golangci-lint fmt, andgolangci-lint run). Do not write any tests; there is no test suite in this project. - Only the user may prepare live services for manual integration.
- Docs:
README.mdcovers running and configuration only. - DO NOT prepend bash commands with cd for the directory you're already