c313645d2d533c5b95ac920d064ad74c8f3a1f23

Author
TheEdgeOfRage <git@theedgeofrage.com>
Committer
TheEdgeOfRage <git@theedgeofrage.com>
Date

Message

docs: update README and AGENTS for self-hosted model services

Diff

  1diff --git a/AGENTS.md b/AGENTS.md
  2index 2cda111e89a163f78359f9c1ea1450df4a6295ca..06b2d7de6d0dceeffbd2065ec4f66ca5b0428df4 100644
  3--- a/AGENTS.md
  4+++ b/AGENTS.md
  5@@ -22,15 +22,21 @@ Two entry points share one game core (`internal/game`), the LLM layer
  6   (returns the NPC line as WAV bytes; replay is client-side). Turns serialize
  7   per session; sessions die with the process. All requests need a bearer token.
  8 
  9-Both entry points are HTTP clients for two externally managed model services:
 10+Both entry points self-host two model services as child subprocesses by
 11+default (spawned on loopback at startup, torn down on exit):
 12 
 13 - **LLM** — llama.cpp router (OpenAI-compatible chat completions) for NPC
 14   dialogue, the judge, compaction, character sheets, and scratch questions.
 15   Prompts live in `internal/llm/prompt.go`; strict `FIELD|value` output
 16-  contracts are parsed in `internal/llm/contract.go`.
 17+  contracts are parsed in `internal/llm/contract.go`. Spawned at
 18+  `127.0.0.1:9931`.
 19 - **Audio** — one audio.cpp server: Qwen3-ASR for STT (the raw Japanese
 20   transcript stays internal; it is what the LLM sees as the player's lines) and
 21-  an OpenAI-compatible speech endpoint that plays the NPC's kana.
 22+  an OpenAI-compatible speech endpoint that plays the NPC's kana. Spawned at
 23+  `127.0.0.1:9932`.
 24+
 25+With `--disable-model-loading`, the entry points connect to external services
 26+at whatever URLs are configured instead of spawning them.
 27 
 28 Turn flow (spoken): record → transcribe → judge + NPC reply in parallel → synthesize the NPC's kana to audio → record
 29 the turn in per-location history. Typed actions skip recording,
 30@@ -44,10 +50,13 @@ when it grows past budget (`internal/game/state.go`,
 31 
 32 ### Model services
 33 
 34-Model services are external, user-managed processes. GPU VRAM is normally
 35-almost fully allocated to the chat model.
 36+The entry points spawn `llama-server` and `audiocpp_server` as child processes
 37+by default (see `internal/services/`). GPU VRAM is normally almost fully
 38+allocated to the chat model, so a 16 GiB free-VRAM gate runs before spawning.
 39 
 40-- Never start, stop, or restart the `llama-server` or `audiocpp_server`.
 41+- Never start, stop, or restart the `llama-server` or `audiocpp_server`
 42+  yourself during development (the binaries manage their own children at
 43+  runtime).
 44 - Never run `llama-cli`, or use a command that can load a model into VRAM.
 45 
 46 ### Product invariants
 47@@ -67,8 +76,8 @@ almost fully allocated to the chat model.
 48   WAV in, returned WAV out) in `internal/server/speech.go` for kaiwari-server.
 49 - LLM replies are parsed from strict `FIELD|value` contracts, never freeform
 50   JSON.
 51-- Errors wrap with `%w`. Startup failures print one stderr line naming the
 52-  service and continue; only a bad scenario file is fatal.
 53+- Errors wrap with `%w`. All startup failures are fatal (VRAM check, service
 54+  spawn, health check, LLM warmup).
 55 
 56 ## Subagents
 57 
 58diff --git a/README.md b/README.md
 59index cd5b691b319da6461cf9fcd3b87f36bb84389bc3..ae47923a6c1992b9565fbab34b8883ea829e5c0a 100644
 60--- a/README.md
 61+++ b/README.md
 62@@ -7,25 +7,29 @@ judge. Every person you talk to gets a hidden character sheet before their
 63 first line so they stay consistent across visits. The interface shows romaji
 64 only; it never displays kana or kanji.
 65 
 66-The game is an HTTP client only. It connects to two externally managed model
 67-services and **never** starts, stops, restarts, kills, reconfigures, or
 68-downloads a model for any of them. You run those services yourself.
 69+Both entry points self-host the two model services (llama.cpp router +
 70+audio.cpp) as child subprocesses by default. On startup they check free VRAM,
 71+spawn the services on loopback, wait until healthy, and tear everything down
 72+on exit or Ctrl-C / SIGTERM. Use `--disable-model-loading` to connect to
 73+already-running external services instead.
 74 
 75 ## Requirements
 76 
 77 - Go 1.27+
 78-- `arecord` (ALSA) for microphone capture. The game invokes bare `arecord`
 79-  (16 kHz mono S16_LE WAV written to the path given as its last argument)
 80-- Two external model services, both run by you:
 81-  - **LLM** — llama.cpp router (OpenAI-compatible `POST /v1/chat/completions`).
 82-    Start it with `make llama`
 83-  - **Audio** — audio.cpp server (`JP_AUDIO_BASE_URL`): TTS speech endpoint
 84-    (`POST /v1/audio/speech`, returns WAV) and ASR transcriptions (multipart
 85-    `POST /v1/audio/transcriptions`, JSON `text` response). Start it from the
 86-    repository root with `audiocpp_server --config audio.cpp.json`; keep
 87-    `lazy_load: true` in that config so both speech models stay resident —
 88-    the game alternates TTS and transcription every turn, so unloading would
 89-    thrash
 90+- AMD GPU with at least **16 GiB free VRAM** (checked via sysfs at startup;
 91+  fatal if not met)
 92+- `arecord` (ALSA) for microphone capture (TUI only). The game invokes bare
 93+  `arecord` (16 kHz mono S16_LE WAV written to the path given as its last
 94+  argument)
 95+- Model binaries on `$PATH`:
 96+  - **LLM** — `llama-server` (llama.cpp router, OpenAI-compatible
 97+    `POST /v1/chat/completions`). Spawned at `127.0.0.1:9931`
 98+  - **Audio** — `audiocpp_server`: TTS speech endpoint (`POST /v1/audio/speech`,
 99+    returns WAV) and ASR transcriptions (multipart `POST /v1/audio/transcriptions`,
100+    JSON `text` response). Spawned at `127.0.0.1:9932`
101+
102+When using `--disable-model-loading`, you run those services yourself and the
103+entry points connect to whatever URLs the flags/env vars point at.
104 
105 ## Build and run
106 
107@@ -34,26 +38,37 @@ go build ./...             # or: make build
108 kaiwari                    # or: make run (go run ./cmd/kaiwari)
109 ```
110 
111-On startup the game waits up to 30s for the LLM model to finish loading
112-(polling every second), runs a bounded readiness check against the audio
113-service, then warms the core system prompts once so first replies
114-are fast. A failed step prints one error line naming the affected service and
115-its configured URL; startup continues either way. The game never launches a
116-service to make a check pass.
117+Startup sequence (model loading ON, the default):
118+
119+1. Check free VRAM ≥ 16 GiB via AMD sysfs (fatal on failure).
120+2. Spawn `llama-server` and `audiocpp_server` as child processes on loopback
121+   (9931 / 9932). If either fails to start, both are killed.
122+3. Poll until all three endpoints are healthy (LLM chat completion, TTS speech,
123+   ASR transcription) — up to 5 minutes total (fatal on timeout).
124+4. Warm up the core prompt slots once so first replies are fast (fatal on
125+   failure).
126+
127+On exit, Ctrl-C, or SIGTERM both child processes receive SIGTERM, wait up to
128+10 s, then SIGKILL if still alive.
129 
130 ## Configuration
131 
132 Every setting is a flag or an environment variable; flags win over env vars,
133 which win over defaults. Defaults match a local setup.
134 
135-| Setting        | Flag                    | Env var                  | Default                                | Use                                                     |
136-| -------------- | ----------------------- | ------------------------ | -------------------------------------- | ------------------------------------------------------- |
137-| LLM router URL | `--llm.url`             | `JP_LLM_BASE_URL`        | `https://llama.home.theedgeofrage.com` | llama.cpp router; `POST /v1/chat/completions`           |
138-| Temperature    | `--llm.temperature`     | `JP_LLM_TEMPERATURE`     | `1.0`                                  | model sampling temperature                              |
139-| Max tokens     | `--llm.max-tokens`      | `JP_LLM_MAX_TOKENS`      | `256`                                  | max output tokens per reply                             |
140-| Thinking mode  | `--llm.enable-thinking` | `JP_LLM_ENABLE_THINKING` | `false`                                | model thinking (`chat_template_kwargs.enable_thinking`) |
141-| Audio base URL | `--audio.url`           | `JP_AUDIO_BASE_URL`      | `http://127.0.0.1:8080`                | audio.cpp server (TTS + ASR)                            |
142-| Scenario brief | `--scenario`            | `JP_SCENARIO_PATH`       | `assets/scenarios/small_city.md`       | plain-text scenario brief file                          |
143+| Setting             | Flag                        | Env var                     | Default                                | Use                                                     |
144+| ------------------- | --------------------------- | --------------------------- | -------------------------------------- | ------------------------------------------------------- |
145+| LLM router URL      | `--llm.url`                 | `JP_LLM_BASE_URL`           | `https://llama.home.theedgeofrage.com` | llama.cpp router; `POST /v1/chat/completions`           |
146+| Temperature         | `--llm.temperature`         | `JP_LLM_TEMPERATURE`        | `1.0`                                  | model sampling temperature                              |
147+| Max tokens          | `--llm.max-tokens`          | `JP_LLM_MAX_TOKENS`         | `256`                                  | max output tokens per reply                             |
148+| Thinking mode       | `--llm.enable-thinking`     | `JP_LLM_ENABLE_THINKING`    | `false`                                | model thinking (`chat_template_kwargs.enable_thinking`) |
149+| Audio base URL      | `--audio.url`               | `JP_AUDIO_BASE_URL`         | `http://127.0.0.1:9932`                | audio.cpp server (TTS + ASR)                            |
150+| Scenario brief      | `--scenario`                | `JP_SCENARIO_PATH`          | `assets/scenarios/small_city.md`       | plain-text scenario brief file                          |
151+| Disable model loading | `--disable-model-loading` | `JP_DISABLE_MODEL_LOADING`  | `false`                                | connect to external services instead of spawning them   |
152+
153+When model loading is ON (default), the LLM and audio base URLs are pinned to
154+the local loopback endpoints (`127.0.0.1:9931` / `127.0.0.1:9932`) regardless
155+of what the URL flags say.
156 
157 Run `kaiwari -help` for the full flag list.
158