TheEdgeOfRage/kaiwari

Default ref refs/heads/main

kaiwari — Japanese learning RPG (TUI)

A terminal game for practicing spoken Japanese. You type English actions to explore an LLM-narrated world and push-to-talk lines of Japanese; your words are recorded, transcribed, answered by the world, and scored by a separate judge. Every person you talk to gets a hidden character sheet before their first line so they stay consistent across visits. The interface shows romaji only; it never displays kana or kanji.

Both entry points self-host the two model services (llama.cpp router + audio.cpp) as child subprocesses by default. On startup they check free VRAM, spawn the services on loopback, wait until healthy, and tear everything down on exit or Ctrl-C / SIGTERM. Use --disable-model-loading to connect to already-running external services instead.

Requirements

  • Go 1.27+
  • AMD GPU with at least 16 GiB free VRAM (checked via sysfs at startup; fatal if not met)
  • arecord (ALSA) for microphone capture (TUI only). The game invokes bare arecord (16 kHz mono S16_LE WAV written to the path given as its last argument)
  • Model binaries on $PATH:
    • LLM — llama-server (llama.cpp router, OpenAI-compatible POST /v1/chat/completions). Spawned at 127.0.0.1:9931
    • Audio — audiocpp_server: TTS speech endpoint (POST /v1/audio/speech, returns WAV) and ASR transcriptions (multipart POST /v1/audio/transcriptions, JSON text response). Spawned at 127.0.0.1:9932

When using --disable-model-loading, you run those services yourself and the entry points connect to whatever URLs the flags/env vars point at.

Build and run

go build ./...             # or: make build
kaiwari                    # or: make run (go run ./cmd/kaiwari)

Startup sequence (model loading ON, the default):

  1. Check free VRAM ≥ 16 GiB via AMD sysfs (fatal on failure).
  2. Spawn llama-server and audiocpp_server as child processes on loopback (9931 / 9932). If either fails to start, both are killed.
  3. Poll until all three endpoints are healthy (LLM chat completion, TTS speech, ASR transcription) — up to 5 minutes total (fatal on timeout).
  4. Warm up the core prompt slots once so first replies are fast (fatal on failure).

On exit, Ctrl-C, or SIGTERM both child processes receive SIGTERM, wait up to 10 s, then SIGKILL if still alive.

Configuration

Every setting is a flag or an environment variable; flags win over env vars, which win over defaults. Defaults match a local setup.

| Setting | Flag | Env var | Default | Use | | ------------------- | --------------------------- | --------------------------- | -------------------------------------- | ------------------------------------------------------- | | LLM router URL | --llm.url | JP_LLM_BASE_URL | https://llama.home.theedgeofrage.com | llama.cpp router; POST /v1/chat/completions | | Temperature | --llm.temperature | JP_LLM_TEMPERATURE | 1.0 | model sampling temperature | | Max tokens | --llm.max-tokens | JP_LLM_MAX_TOKENS | 256 | max output tokens per reply | | Thinking mode | --llm.enable-thinking | JP_LLM_ENABLE_THINKING | false | model thinking (chat_template_kwargs.enable_thinking) | | Audio base URL | --audio.url | JP_AUDIO_BASE_URL | http://127.0.0.1:9932 | audio.cpp server (TTS + ASR) | | Scenario brief | --scenario | JP_SCENARIO_PATH | assets/scenarios/small_city.md | plain-text scenario brief file | | Disable model loading | --disable-model-loading | JP_DISABLE_MODEL_LOADING | false | connect to external services instead of spawning them |

When model loading is ON (default), the LLM and audio base URLs are pinned to the local loopback endpoints (127.0.0.1:9931 / 127.0.0.1:9932) regardless of what the URL flags say.

Run kaiwari -help for the full flag list.

Controls

  • Type an action in English and press enter to send it
  • Ask how to say something in Japanese: type ?your question and press enter. The answer is a learning aid; it does not affect the world
  • Generate flashcards: type !instructions and press enter. The LLM generates a Japanese sentence matching the instructions and appends the vocabulary to flashcards.csv
  • Push-to-talk: F2 (press to start recording, press again to stop and send the turn)
  • Reveal NPC romaji: F3
  • Reveal NPC English: F4
  • Replay the last spoken line: F5
  • Quit: esc

Input is ignored while a turn is recording or processing. Each spoken turn is graded by a judge that stays separate from NPC dialogue; NPCs behave as ordinary people, not language teachers.