Diff
1diff --git a/README.md b/README.md
2index 02364a5c1af0c25464ce046909cc817b6b0b0e67..ea58e313f560f1b00586a48aa0fefdce49e27b8c 100644
3--- a/README.md
4+++ b/README.md
5@@ -1,9 +1,11 @@
6 # jp — Japanese learning RPG (TUI)
7
8-A terminal game for practicing spoken Japanese. You walk a small map and talk
9-with people; your words are recorded, transcribed, and sent to an NPC, and a
10-separate judge scores each line. The interface shows romaji only; it never
11-displays kana or kanji.
12+A terminal game for practicing spoken Japanese. You type English actions to
13+explore an LLM-narrated world and push-to-talk lines of Japanese; your words
14+are recorded, transcribed, answered by the world, and scored by a separate
15+judge. Every person you talk to gets a hidden character sheet before their
16+first line so they stay consistent across visits. The interface shows romaji
17+only; it never displays kana or kanji.
18
19 The game is an HTTP client only. It connects to three externally managed model
20 services and **never** starts, stops, restarts, kills, reconfigures, or
21@@ -29,47 +31,43 @@ go build ./...
22 jp # or: go run ./cmd/jp
23 ```
24
25-On startup the game preloads the LLM model, runs a bounded readiness check per
26-service, then warms each conversation's system prompt once so first replies are
27-fast. The TUI appears immediately; each service line shows `checking…`, then
28-flips to `up` or `down` in place as its result arrives. A failed check names
29-the affected service and its configured URL. The game never launches a service
30-to make a check pass.
31+On startup the game preloads the LLM model, runs bounded readiness checks
32+against TTS and STT, then warms each conversation's system prompt once so first
33+replies are fast. A failed step prints one error line naming the affected
34+service and its configured URL; startup continues either way. The game never
35+launches a service to make a check pass.
36
37 ## Configuration
38
39-Every endpoint is set with a flag or an environment variable. Defaults match a
40-local setup.
41-
42-| Service | Flag | Env var | Default | Use |
43-| ----------------- | ------------------ | ----------------- | ----------------------------------------- | ----------------------------------------------- |
44-| LLM router URL | `--llm-url` | `JP_LLM_BASE_URL` | `https://llama.home.theedgeofrage.com` | `POST /v1/chat/completions` |
45-| LLM game slot | `--llm-game-slot` | `JP_LLM_GAME_SLOT` | `0` | llama-server slot for game-loop requests (`-1` = server decides) |
46-| LLM scratch slot | `--llm-scratch-slot` | `JP_LLM_SCRATCH_SLOT` | `1` | llama-server slot for judge and compaction requests (`-1` = server decides) |
47-| TTS base URL | `--tts-url` | `JP_TTS_BASE_URL` | `http://127.0.0.1:8080/v1` | `POST /audio/speech` |
48-| STT server URL | `--stt-url` | `JP_STT_URL` | `http://127.0.0.1:8178` | multipart `POST /inference` |
49-| STT language | `--stt-language` | `JP_STT_LANGUAGE` | `ja` | Whisper request language (ISO 639-1) |
50-| Recorder command | `--record-command` | — | `arecord` | 16 kHz mono S16_LE WAV capture (see note above) |
51-
52-Content paths:
53-
54-| Flag | Env var | Default |
55-| --------------- | ---------------- | ----------------------- |
56-| `--map` | `JP_MAP_PATH` | `assets/maps/city.json` |
57-| `--persona-dir` | `JP_PERSONA_DIR` | `assets/personas` |
58+Every setting is a flag or an environment variable; flags win over env vars,
59+which win over defaults. Defaults match a local setup.
60+
61+| Setting | Flag | Env var | Default | Use |
62+| ---------------- | -------------------- | --------------------- | ----------------------------------------- | ----------------------------------------------- |
63+| LLM router URL | `--url` | `JP_LLM_BASE_URL` | `https://llama.home.theedgeofrage.com` | llama.cpp router; `POST /v1/chat/completions` |
64+| LLM model | `--model` | `MODEL` | `jp` | model name sent in every LLM request |
65+| Temperature | `--temperature` | `TEMPERATURE` | `1.0` | sampling temperature |
66+| Max tokens | `--max-tokens` | `MAX_TOKENS` | `256` | max output tokens per reply |
67+| Thinking mode | `--enable-thinking` | `ENABLE_THINKING` | `false` | Qwen3 thinking (`chat_template_kwargs.enable_thinking`) |
68+| LLM game slot | `--llm-game-slot` | `JP_LLM_GAME_SLOT` | `0` | llama-server slot for game-loop requests (`-1` = server decides) |
69+| LLM scratch slot | `--llm-scratch-slot` | `JP_LLM_SCRATCH_SLOT` | `1` | llama-server slot for judge, compaction, and sheet requests (`-1` = server decides) |
70+| TTS base URL | `--tts-url` | `JP_TTS_BASE_URL` | `http://127.0.0.1:8080/v1` | `POST /audio/speech` |
71+| STT server URL | `--stt-url` | `JP_STT_URL` | `http://127.0.0.1:8178` | multipart `POST /inference` |
72+| STT language | `--stt-language` | `JP_STT_LANGUAGE` | `ja` | Whisper request language (ISO 639-1) |
73+| Recorder command | `--record-command` | — | `arecord` | 16 kHz mono S16_LE WAV capture (see note above) |
74+| Scenario brief | `--scenario` | `JP_SCENARIO_PATH` | `assets/scenarios/small_city.md` | plain-text scenario brief file |
75
76 Run `jp -help` for the full flag list.
77
78 ## Controls
79
80-- Move: arrow keys or `w`/`a`/`s`/`d`
81-- Push-to-talk: `space` or `enter` (one toggle key: press to start recording,
82- press again to stop and send the turn)
83-- Reveal NPC romaji: `r`
84-- Reveal NPC English (only after `r`): `t`
85-- Quit: `q` or `esc`
86+- Type an action in English and press `enter` to send it
87+- Push-to-talk: `F2` (press to start recording, press again to stop and send
88+ the turn)
89+- Reveal NPC romaji: `F3`
90+- Reveal NPC English: `F4`
91+- Quit: `esc`
92
93-Movement, talk, and reveal keys are ignored while a turn is recording or
94-processing. Each spoken turn is graded by a judge that stays separate from NPC
95-dialogue: the judge scores your attempt in the learning panel; NPCs behave as
96-ordinary people, not language teachers.
97+Input is ignored while a turn is recording or processing. Each spoken turn is
98+graded by a judge that stays separate from NPC dialogue; NPCs behave as ordinary
99+people, not language teachers.
100diff --git a/docs/services.md b/docs/services.md
101index 3c7d4ae798af5aa0b7aede4fabe3cb073e4643a0..f943b76aed10ec84b36bc89ba616b0daebc4fa63 100644
102--- a/docs/services.md
103+++ b/docs/services.md
104@@ -11,7 +11,7 @@ match the local setup:
105
106 | Service | Flag | Env var | Default | Use |
107 | ----------------- | ------------------ | ----------------- | ----------------------------------------- | --------------------------------- |
108-| LLM router URL | `--llm-url` | `JP_LLM_BASE_URL` | `https://llama.home.theedgeofrage.com` | `POST /v1/chat/completions` |
109+| LLM router URL | `--url` | `JP_LLM_BASE_URL` | `https://llama.home.theedgeofrage.com` | `POST /v1/chat/completions` |
110 | TTS base URL | `--tts-url` | `JP_TTS_BASE_URL` | `http://127.0.0.1:8080/v1` | `POST /audio/speech` |
111 | STT server URL | `--stt-url` | `JP_STT_URL` | `http://127.0.0.1:8178` | multipart `POST /inference` |
112 | STT language | `--stt-language` | `JP_STT_LANGUAGE` | `ja` | Whisper request language (ISO 639-1) |
113@@ -20,7 +20,7 @@ match the local setup:
114 At startup, `jp` first preloads the LLM model with
115 `POST {JP_LLM_BASE_URL}/models/load` (30s timeout) before anything else loads,
116 then makes only safe, bounded HTTP readiness checks (a few seconds each, no
117-inference requests). After persona data loads, it sends one best-effort warmup
118+inference requests). It then sends one best-effort warmup
119 completion per system prompt (`max_tokens: 1`) so the server's prompt cache
120 keeps those prefixes hot. It stays silent when every service answers; a down
121 service prints one error line naming the service and its URL. The game never
122@@ -33,12 +33,12 @@ llama.cpp server in router mode (started without a model argument) at
123 `POST {JP_LLM_BASE_URL}/v1/chat/completions`.
124
125 Model: `unsloth/Qwen3-8B-GGUF:UD-Q4_K_XL`, resolvable by the router as `jp`. The client
126-sends OpenAI-compatible requests with `model` set to `jp`, disables Qwen3
127-thinking on every request (`chat_template_kwargs.enable_thinking: false`),
128-sets `cache_prompt: true` so the server keeps prompt prefixes cached between
129-turns, pins each request to a llama-server slot via `id_slot` (see below), and
130-reads streaming SSE. The model chat template (`--jinja`) is used; ChatML is
131-not built manually by the client.
132+sends OpenAI-compatible requests with `model` set to `jp`, sends
133+`chat_template_kwargs.enable_thinking` per the `--enable-thinking` flag
134+(default false, which disables Qwen3 thinking), sets `cache_prompt: true` so
135+the server keeps prompt prefixes cached between turns, pins each request to a
136+llama-server slot via `id_slot` (see below), and reads streaming SSE. The model
137+chat template (`--jinja`) is used; ChatML is not built manually by the client.
138
139 ### Slot pinning
140
141@@ -50,11 +50,11 @@ the game loop's cached context:
142 | Client use | Flag | Env var | Default |
143 | -------------------------------------------- | ---------------------- | --------------------- | ------- |
144 | Game loop (world replies) | `--llm-game-slot` | `JP_LLM_GAME_SLOT` | `0` |
145-| Scratch (judge, compaction; warmup included) | `--llm-scratch-slot` | `JP_LLM_SCRATCH_SLOT` | `1` |
146+| Scratch (judge, compaction, character sheets; warmup included) | `--llm-scratch-slot` | `JP_LLM_SCRATCH_SLOT` | `1` |
147
148 Warmup follows the same split: the game system prompt warms on the game
149-client, the judge and compaction prompts warm on the scratch client, so initial
150-cache placement is deterministic per slot.
151+client, the judge, compaction, and sheet prompts warm on the scratch client,
152+so initial cache placement is deterministic per slot.
153
154 **Fallback:** if the router does not honor `id_slot`, set both env vars to `-1`
155 (`JP_LLM_GAME_SLOT=-1 JP_LLM_SCRATCH_SLOT=-1`). The server then auto-selects a