Diff
1diff --git a/AGENTS.md b/AGENTS.md
2index 2cda111e89a163f78359f9c1ea1450df4a6295ca..06b2d7de6d0dceeffbd2065ec4f66ca5b0428df4 100644
3--- a/AGENTS.md
4+++ b/AGENTS.md
5@@ -22,15 +22,21 @@ Two entry points share one game core (`internal/game`), the LLM layer
6 (returns the NPC line as WAV bytes; replay is client-side). Turns serialize
7 per session; sessions die with the process. All requests need a bearer token.
8
9-Both entry points are HTTP clients for two externally managed model services:
10+Both entry points self-host two model services as child subprocesses by
11+default (spawned on loopback at startup, torn down on exit):
12
13 - **LLM** — llama.cpp router (OpenAI-compatible chat completions) for NPC
14 dialogue, the judge, compaction, character sheets, and scratch questions.
15 Prompts live in `internal/llm/prompt.go`; strict `FIELD|value` output
16- contracts are parsed in `internal/llm/contract.go`.
17+ contracts are parsed in `internal/llm/contract.go`. Spawned at
18+ `127.0.0.1:9931`.
19 - **Audio** — one audio.cpp server: Qwen3-ASR for STT (the raw Japanese
20 transcript stays internal; it is what the LLM sees as the player's lines) and
21- an OpenAI-compatible speech endpoint that plays the NPC's kana.
22+ an OpenAI-compatible speech endpoint that plays the NPC's kana. Spawned at
23+ `127.0.0.1:9932`.
24+
25+With `--disable-model-loading`, the entry points connect to external services
26+at whatever URLs are configured instead of spawning them.
27
28 Turn flow (spoken): record → transcribe → judge + NPC reply in parallel → synthesize the NPC's kana to audio → record
29 the turn in per-location history. Typed actions skip recording,
30@@ -44,10 +50,13 @@ when it grows past budget (`internal/game/state.go`,
31
32 ### Model services
33
34-Model services are external, user-managed processes. GPU VRAM is normally
35-almost fully allocated to the chat model.
36+The entry points spawn `llama-server` and `audiocpp_server` as child processes
37+by default (see `internal/services/`). GPU VRAM is normally almost fully
38+allocated to the chat model, so a 16 GiB free-VRAM gate runs before spawning.
39
40-- Never start, stop, or restart the `llama-server` or `audiocpp_server`.
41+- Never start, stop, or restart the `llama-server` or `audiocpp_server`
42+ yourself during development (the binaries manage their own children at
43+ runtime).
44 - Never run `llama-cli`, or use a command that can load a model into VRAM.
45
46 ### Product invariants
47@@ -67,8 +76,8 @@ almost fully allocated to the chat model.
48 WAV in, returned WAV out) in `internal/server/speech.go` for kaiwari-server.
49 - LLM replies are parsed from strict `FIELD|value` contracts, never freeform
50 JSON.
51-- Errors wrap with `%w`. Startup failures print one stderr line naming the
52- service and continue; only a bad scenario file is fatal.
53+- Errors wrap with `%w`. All startup failures are fatal (VRAM check, service
54+ spawn, health check, LLM warmup).
55
56 ## Subagents
57
58diff --git a/README.md b/README.md
59index cd5b691b319da6461cf9fcd3b87f36bb84389bc3..ae47923a6c1992b9565fbab34b8883ea829e5c0a 100644
60--- a/README.md
61+++ b/README.md
62@@ -7,25 +7,29 @@ judge. Every person you talk to gets a hidden character sheet before their
63 first line so they stay consistent across visits. The interface shows romaji
64 only; it never displays kana or kanji.
65
66-The game is an HTTP client only. It connects to two externally managed model
67-services and **never** starts, stops, restarts, kills, reconfigures, or
68-downloads a model for any of them. You run those services yourself.
69+Both entry points self-host the two model services (llama.cpp router +
70+audio.cpp) as child subprocesses by default. On startup they check free VRAM,
71+spawn the services on loopback, wait until healthy, and tear everything down
72+on exit or Ctrl-C / SIGTERM. Use `--disable-model-loading` to connect to
73+already-running external services instead.
74
75 ## Requirements
76
77 - Go 1.27+
78-- `arecord` (ALSA) for microphone capture. The game invokes bare `arecord`
79- (16 kHz mono S16_LE WAV written to the path given as its last argument)
80-- Two external model services, both run by you:
81- - **LLM** — llama.cpp router (OpenAI-compatible `POST /v1/chat/completions`).
82- Start it with `make llama`
83- - **Audio** — audio.cpp server (`JP_AUDIO_BASE_URL`): TTS speech endpoint
84- (`POST /v1/audio/speech`, returns WAV) and ASR transcriptions (multipart
85- `POST /v1/audio/transcriptions`, JSON `text` response). Start it from the
86- repository root with `audiocpp_server --config audio.cpp.json`; keep
87- `lazy_load: true` in that config so both speech models stay resident —
88- the game alternates TTS and transcription every turn, so unloading would
89- thrash
90+- AMD GPU with at least **16 GiB free VRAM** (checked via sysfs at startup;
91+ fatal if not met)
92+- `arecord` (ALSA) for microphone capture (TUI only). The game invokes bare
93+ `arecord` (16 kHz mono S16_LE WAV written to the path given as its last
94+ argument)
95+- Model binaries on `$PATH`:
96+ - **LLM** — `llama-server` (llama.cpp router, OpenAI-compatible
97+ `POST /v1/chat/completions`). Spawned at `127.0.0.1:9931`
98+ - **Audio** — `audiocpp_server`: TTS speech endpoint (`POST /v1/audio/speech`,
99+ returns WAV) and ASR transcriptions (multipart `POST /v1/audio/transcriptions`,
100+ JSON `text` response). Spawned at `127.0.0.1:9932`
101+
102+When using `--disable-model-loading`, you run those services yourself and the
103+entry points connect to whatever URLs the flags/env vars point at.
104
105 ## Build and run
106
107@@ -34,26 +38,37 @@ go build ./... # or: make build
108 kaiwari # or: make run (go run ./cmd/kaiwari)
109 ```
110
111-On startup the game waits up to 30s for the LLM model to finish loading
112-(polling every second), runs a bounded readiness check against the audio
113-service, then warms the core system prompts once so first replies
114-are fast. A failed step prints one error line naming the affected service and
115-its configured URL; startup continues either way. The game never launches a
116-service to make a check pass.
117+Startup sequence (model loading ON, the default):
118+
119+1. Check free VRAM ≥ 16 GiB via AMD sysfs (fatal on failure).
120+2. Spawn `llama-server` and `audiocpp_server` as child processes on loopback
121+ (9931 / 9932). If either fails to start, both are killed.
122+3. Poll until all three endpoints are healthy (LLM chat completion, TTS speech,
123+ ASR transcription) — up to 5 minutes total (fatal on timeout).
124+4. Warm up the core prompt slots once so first replies are fast (fatal on
125+ failure).
126+
127+On exit, Ctrl-C, or SIGTERM both child processes receive SIGTERM, wait up to
128+10 s, then SIGKILL if still alive.
129
130 ## Configuration
131
132 Every setting is a flag or an environment variable; flags win over env vars,
133 which win over defaults. Defaults match a local setup.
134
135-| Setting | Flag | Env var | Default | Use |
136-| -------------- | ----------------------- | ------------------------ | -------------------------------------- | ------------------------------------------------------- |
137-| LLM router URL | `--llm.url` | `JP_LLM_BASE_URL` | `https://llama.home.theedgeofrage.com` | llama.cpp router; `POST /v1/chat/completions` |
138-| Temperature | `--llm.temperature` | `JP_LLM_TEMPERATURE` | `1.0` | model sampling temperature |
139-| Max tokens | `--llm.max-tokens` | `JP_LLM_MAX_TOKENS` | `256` | max output tokens per reply |
140-| Thinking mode | `--llm.enable-thinking` | `JP_LLM_ENABLE_THINKING` | `false` | model thinking (`chat_template_kwargs.enable_thinking`) |
141-| Audio base URL | `--audio.url` | `JP_AUDIO_BASE_URL` | `http://127.0.0.1:8080` | audio.cpp server (TTS + ASR) |
142-| Scenario brief | `--scenario` | `JP_SCENARIO_PATH` | `assets/scenarios/small_city.md` | plain-text scenario brief file |
143+| Setting | Flag | Env var | Default | Use |
144+| ------------------- | --------------------------- | --------------------------- | -------------------------------------- | ------------------------------------------------------- |
145+| LLM router URL | `--llm.url` | `JP_LLM_BASE_URL` | `https://llama.home.theedgeofrage.com` | llama.cpp router; `POST /v1/chat/completions` |
146+| Temperature | `--llm.temperature` | `JP_LLM_TEMPERATURE` | `1.0` | model sampling temperature |
147+| Max tokens | `--llm.max-tokens` | `JP_LLM_MAX_TOKENS` | `256` | max output tokens per reply |
148+| Thinking mode | `--llm.enable-thinking` | `JP_LLM_ENABLE_THINKING` | `false` | model thinking (`chat_template_kwargs.enable_thinking`) |
149+| Audio base URL | `--audio.url` | `JP_AUDIO_BASE_URL` | `http://127.0.0.1:9932` | audio.cpp server (TTS + ASR) |
150+| Scenario brief | `--scenario` | `JP_SCENARIO_PATH` | `assets/scenarios/small_city.md` | plain-text scenario brief file |
151+| Disable model loading | `--disable-model-loading` | `JP_DISABLE_MODEL_LOADING` | `false` | connect to external services instead of spawning them |
152+
153+When model loading is ON (default), the LLM and audio base URLs are pinned to
154+the local loopback endpoints (`127.0.0.1:9931` / `127.0.0.1:9932`) regardless
155+of what the URL flags say.
156
157 Run `kaiwari -help` for the full flag list.
158