README.md
kaiwari — Japanese learning RPG (TUI)
A terminal game for practicing spoken Japanese. You type English actions to explore an LLM-narrated world and push-to-talk lines of Japanese; your words are recorded, transcribed, answered by the world, and scored by a separate judge. Every person you talk to gets a hidden character sheet before their first line so they stay consistent across visits. The interface shows romaji only; it never displays kana or kanji.
Both entry points self-host the two model services (llama.cpp router +
audio.cpp) as child subprocesses by default. On startup they check free VRAM,
spawn the services on loopback, wait until healthy, and tear everything down
on exit or Ctrl-C / SIGTERM. Use --disable-model-loading to connect to
already-running external services instead.
Requirements
- Go 1.27+
- AMD GPU with at least 16 GiB free VRAM (checked via sysfs at startup; fatal if not met)
arecord(ALSA) for microphone capture (TUI only). The game invokes barearecord(16 kHz mono S16_LE WAV written to the path given as its last argument)- Model binaries on
$PATH:- LLM —
llama-server(llama.cpp router, OpenAI-compatiblePOST /v1/chat/completions). Spawned at127.0.0.1:9931 - Audio —
audiocpp_server: TTS speech endpoint (POST /v1/audio/speech, returns WAV) and ASR transcriptions (multipartPOST /v1/audio/transcriptions, JSONtextresponse). Spawned at127.0.0.1:9932
- LLM —
When using --disable-model-loading, you run those services yourself and the
entry points connect to whatever URLs the flags/env vars point at.
Build and run
go build ./... # or: make build
kaiwari # or: make run (go run ./cmd/kaiwari)
Startup sequence (model loading ON, the default):
- Check free VRAM ≥ 16 GiB via AMD sysfs (fatal on failure).
- Spawn
llama-serverandaudiocpp_serveras child processes on loopback (9931 / 9932). If either fails to start, both are killed. - Poll until all three endpoints are healthy (LLM chat completion, TTS speech, ASR transcription) — up to 5 minutes total (fatal on timeout).
- Warm up the core prompt slots once so first replies are fast (fatal on failure).
On exit, Ctrl-C, or SIGTERM both child processes receive SIGTERM, wait up to 10 s, then SIGKILL if still alive.
Configuration
Every setting is a flag or an environment variable; flags win over env vars, which win over defaults. Defaults match a local setup.
| Setting | Flag | Env var | Default | Use |
| ------------------- | --------------------------- | --------------------------- | -------------------------------------- | ------------------------------------------------------- |
| LLM router URL | --llm.url | JP_LLM_BASE_URL | https://llama.home.theedgeofrage.com | llama.cpp router; POST /v1/chat/completions |
| Temperature | --llm.temperature | JP_LLM_TEMPERATURE | 1.0 | model sampling temperature |
| Max tokens | --llm.max-tokens | JP_LLM_MAX_TOKENS | 256 | max output tokens per reply |
| Thinking mode | --llm.enable-thinking | JP_LLM_ENABLE_THINKING | false | model thinking (chat_template_kwargs.enable_thinking) |
| Audio base URL | --audio.url | JP_AUDIO_BASE_URL | http://127.0.0.1:9932 | audio.cpp server (TTS + ASR) |
| Scenario brief | --scenario | JP_SCENARIO_PATH | assets/scenarios/small_city.md | plain-text scenario brief file |
| Disable model loading | --disable-model-loading | JP_DISABLE_MODEL_LOADING | false | connect to external services instead of spawning them |
When model loading is ON (default), the LLM and audio base URLs are pinned to
the local loopback endpoints (127.0.0.1:9931 / 127.0.0.1:9932) regardless
of what the URL flags say.
Run kaiwari -help for the full flag list.
Controls
- Type an action in English and press
enterto send it - Ask how to say something in Japanese: type
?your questionand pressenter. The answer is a learning aid; it does not affect the world - Generate flashcards: type
!instructionsand pressenter. The LLM generates a Japanese sentence matching the instructions and appends the vocabulary toflashcards.csv - Push-to-talk:
F2(press to start recording, press again to stop and send the turn) - Reveal NPC romaji:
F3 - Reveal NPC English:
F4 - Replay the last spoken line:
F5 - Quit:
esc
Input is ignored while a turn is recording or processing. Each spoken turn is graded by a judge that stays separate from NPC dialogue; NPCs behave as ordinary people, not language teachers.