08b33b914c4ac72c0f6151016340eec1f627ecee

Author
TheEdgeOfRage <git@theedgeofrage.com>
Committer
TheEdgeOfRage <git@theedgeofrage.com>
Date

Message

Align README and services docs with final flags and behavior

Diff

  1diff --git a/README.md b/README.md
  2index 02364a5c1af0c25464ce046909cc817b6b0b0e67..ea58e313f560f1b00586a48aa0fefdce49e27b8c 100644
  3--- a/README.md
  4+++ b/README.md
  5@@ -1,9 +1,11 @@
  6 # jp — Japanese learning RPG (TUI)
  7 
  8-A terminal game for practicing spoken Japanese. You walk a small map and talk
  9-with people; your words are recorded, transcribed, and sent to an NPC, and a
 10-separate judge scores each line. The interface shows romaji only; it never
 11-displays kana or kanji.
 12+A terminal game for practicing spoken Japanese. You type English actions to
 13+explore an LLM-narrated world and push-to-talk lines of Japanese; your words
 14+are recorded, transcribed, answered by the world, and scored by a separate
 15+judge. Every person you talk to gets a hidden character sheet before their
 16+first line so they stay consistent across visits. The interface shows romaji
 17+only; it never displays kana or kanji.
 18 
 19 The game is an HTTP client only. It connects to three externally managed model
 20 services and **never** starts, stops, restarts, kills, reconfigures, or
 21@@ -29,47 +31,43 @@ go build ./...
 22 jp                         # or: go run ./cmd/jp
 23 ```
 24 
 25-On startup the game preloads the LLM model, runs a bounded readiness check per
 26-service, then warms each conversation's system prompt once so first replies are
 27-fast. The TUI appears immediately; each service line shows `checking…`, then
 28-flips to `up` or `down` in place as its result arrives. A failed check names
 29-the affected service and its configured URL. The game never launches a service
 30-to make a check pass.
 31+On startup the game preloads the LLM model, runs bounded readiness checks
 32+against TTS and STT, then warms each conversation's system prompt once so first
 33+replies are fast. A failed step prints one error line naming the affected
 34+service and its configured URL; startup continues either way. The game never
 35+launches a service to make a check pass.
 36 
 37 ## Configuration
 38 
 39-Every endpoint is set with a flag or an environment variable. Defaults match a
 40-local setup.
 41-
 42-| Service           | Flag               | Env var           | Default                                   | Use                                             |
 43-| ----------------- | ------------------ | ----------------- | ----------------------------------------- | ----------------------------------------------- |
 44-| LLM router URL    | `--llm-url`        | `JP_LLM_BASE_URL` | `https://llama.home.theedgeofrage.com`    | `POST /v1/chat/completions`                     |
 45-| LLM game slot     | `--llm-game-slot`  | `JP_LLM_GAME_SLOT`    | `0`   | llama-server slot for game-loop requests (`-1` = server decides) |
 46-| LLM scratch slot  | `--llm-scratch-slot` | `JP_LLM_SCRATCH_SLOT` | `1`   | llama-server slot for judge and compaction requests (`-1` = server decides) |
 47-| TTS base URL      | `--tts-url`        | `JP_TTS_BASE_URL` | `http://127.0.0.1:8080/v1`                | `POST /audio/speech`                            |
 48-| STT server URL    | `--stt-url`        | `JP_STT_URL`      | `http://127.0.0.1:8178`                  | multipart `POST /inference`                     |
 49-| STT language      | `--stt-language`   | `JP_STT_LANGUAGE` | `ja`                                      | Whisper request language (ISO 639-1)           |
 50-| Recorder command  | `--record-command` | —                 | `arecord`                                 | 16 kHz mono S16_LE WAV capture (see note above) |
 51-
 52-Content paths:
 53-
 54-| Flag            | Env var          | Default                 |
 55-| --------------- | ---------------- | ----------------------- |
 56-| `--map`         | `JP_MAP_PATH`    | `assets/maps/city.json` |
 57-| `--persona-dir` | `JP_PERSONA_DIR` | `assets/personas`       |
 58+Every setting is a flag or an environment variable; flags win over env vars,
 59+which win over defaults. Defaults match a local setup.
 60+
 61+| Setting          | Flag                 | Env var               | Default                                   | Use                                             |
 62+| ---------------- | -------------------- | --------------------- | ----------------------------------------- | ----------------------------------------------- |
 63+| LLM router URL   | `--url`              | `JP_LLM_BASE_URL`     | `https://llama.home.theedgeofrage.com`    | llama.cpp router; `POST /v1/chat/completions`   |
 64+| LLM model        | `--model`            | `MODEL`               | `jp`                                      | model name sent in every LLM request            |
 65+| Temperature      | `--temperature`      | `TEMPERATURE`         | `1.0`                                     | sampling temperature                            |
 66+| Max tokens       | `--max-tokens`       | `MAX_TOKENS`          | `256`                                     | max output tokens per reply                     |
 67+| Thinking mode    | `--enable-thinking`  | `ENABLE_THINKING`     | `false`                                   | Qwen3 thinking (`chat_template_kwargs.enable_thinking`) |
 68+| LLM game slot    | `--llm-game-slot`    | `JP_LLM_GAME_SLOT`    | `0`                                       | llama-server slot for game-loop requests (`-1` = server decides) |
 69+| LLM scratch slot | `--llm-scratch-slot` | `JP_LLM_SCRATCH_SLOT` | `1`                                       | llama-server slot for judge, compaction, and sheet requests (`-1` = server decides) |
 70+| TTS base URL     | `--tts-url`          | `JP_TTS_BASE_URL`     | `http://127.0.0.1:8080/v1`               | `POST /audio/speech`                            |
 71+| STT server URL   | `--stt-url`          | `JP_STT_URL`          | `http://127.0.0.1:8178`                  | multipart `POST /inference`                     |
 72+| STT language     | `--stt-language`     | `JP_STT_LANGUAGE`     | `ja`                                      | Whisper request language (ISO 639-1)           |
 73+| Recorder command | `--record-command`   | —                     | `arecord`                                 | 16 kHz mono S16_LE WAV capture (see note above) |
 74+| Scenario brief   | `--scenario`         | `JP_SCENARIO_PATH`    | `assets/scenarios/small_city.md`          | plain-text scenario brief file                  |
 75 
 76 Run `jp -help` for the full flag list.
 77 
 78 ## Controls
 79 
 80-- Move: arrow keys or `w`/`a`/`s`/`d`
 81-- Push-to-talk: `space` or `enter` (one toggle key: press to start recording,
 82-  press again to stop and send the turn)
 83-- Reveal NPC romaji: `r`
 84-- Reveal NPC English (only after `r`): `t`
 85-- Quit: `q` or `esc`
 86+- Type an action in English and press `enter` to send it
 87+- Push-to-talk: `F2` (press to start recording, press again to stop and send
 88+  the turn)
 89+- Reveal NPC romaji: `F3`
 90+- Reveal NPC English: `F4`
 91+- Quit: `esc`
 92 
 93-Movement, talk, and reveal keys are ignored while a turn is recording or
 94-processing. Each spoken turn is graded by a judge that stays separate from NPC
 95-dialogue: the judge scores your attempt in the learning panel; NPCs behave as
 96-ordinary people, not language teachers.
 97+Input is ignored while a turn is recording or processing. Each spoken turn is
 98+graded by a judge that stays separate from NPC dialogue; NPCs behave as ordinary
 99+people, not language teachers.
100diff --git a/docs/services.md b/docs/services.md
101index 3c7d4ae798af5aa0b7aede4fabe3cb073e4643a0..f943b76aed10ec84b36bc89ba616b0daebc4fa63 100644
102--- a/docs/services.md
103+++ b/docs/services.md
104@@ -11,7 +11,7 @@ match the local setup:
105 
106 | Service           | Flag               | Env var           | Default                                   | Use                               |
107 | ----------------- | ------------------ | ----------------- | ----------------------------------------- | --------------------------------- |
108-| LLM router URL    | `--llm-url`        | `JP_LLM_BASE_URL` | `https://llama.home.theedgeofrage.com`    | `POST /v1/chat/completions`       |
109+| LLM router URL    | `--url`            | `JP_LLM_BASE_URL` | `https://llama.home.theedgeofrage.com`    | `POST /v1/chat/completions`       |
110 | TTS base URL      | `--tts-url`        | `JP_TTS_BASE_URL` | `http://127.0.0.1:8080/v1`                | `POST /audio/speech`              |
111 | STT server URL    | `--stt-url`        | `JP_STT_URL`      | `http://127.0.0.1:8178`                  | multipart `POST /inference`       |
112 | STT language      | `--stt-language`   | `JP_STT_LANGUAGE` | `ja`                                      | Whisper request language (ISO 639-1) |
113@@ -20,7 +20,7 @@ match the local setup:
114 At startup, `jp` first preloads the LLM model with
115 `POST {JP_LLM_BASE_URL}/models/load` (30s timeout) before anything else loads,
116 then makes only safe, bounded HTTP readiness checks (a few seconds each, no
117-inference requests). After persona data loads, it sends one best-effort warmup
118+inference requests). It then sends one best-effort warmup
119 completion per system prompt (`max_tokens: 1`) so the server's prompt cache
120 keeps those prefixes hot. It stays silent when every service answers; a down
121 service prints one error line naming the service and its URL. The game never
122@@ -33,12 +33,12 @@ llama.cpp server in router mode (started without a model argument) at
123 `POST {JP_LLM_BASE_URL}/v1/chat/completions`.
124 
125 Model: `unsloth/Qwen3-8B-GGUF:UD-Q4_K_XL`, resolvable by the router as `jp`. The client
126-sends OpenAI-compatible requests with `model` set to `jp`, disables Qwen3
127-thinking on every request (`chat_template_kwargs.enable_thinking: false`),
128-sets `cache_prompt: true` so the server keeps prompt prefixes cached between
129-turns, pins each request to a llama-server slot via `id_slot` (see below), and
130-reads streaming SSE. The model chat template (`--jinja`) is used; ChatML is
131-not built manually by the client.
132+sends OpenAI-compatible requests with `model` set to `jp`, sends
133+`chat_template_kwargs.enable_thinking` per the `--enable-thinking` flag
134+(default false, which disables Qwen3 thinking), sets `cache_prompt: true` so
135+the server keeps prompt prefixes cached between turns, pins each request to a
136+llama-server slot via `id_slot` (see below), and reads streaming SSE. The model
137+chat template (`--jinja`) is used; ChatML is not built manually by the client.
138 
139 ### Slot pinning
140 
141@@ -50,11 +50,11 @@ the game loop's cached context:
142 | Client use                                   | Flag                   | Env var               | Default |
143 | -------------------------------------------- | ---------------------- | --------------------- | ------- |
144 | Game loop (world replies)                    | `--llm-game-slot`      | `JP_LLM_GAME_SLOT`    | `0`     |
145-| Scratch (judge, compaction; warmup included) | `--llm-scratch-slot`   | `JP_LLM_SCRATCH_SLOT` | `1`     |
146+| Scratch (judge, compaction, character sheets; warmup included) | `--llm-scratch-slot`   | `JP_LLM_SCRATCH_SLOT` | `1`     |
147 
148 Warmup follows the same split: the game system prompt warms on the game
149-client, the judge and compaction prompts warm on the scratch client, so initial
150-cache placement is deterministic per slot.
151+client, the judge, compaction, and sheet prompts warm on the scratch client,
152+so initial cache placement is deterministic per slot.
153 
154 **Fallback:** if the router does not honor `id_slot`, set both env vars to `-1`
155 (`JP_LLM_GAME_SLOT=-1 JP_LLM_SCRATCH_SLOT=-1`). The server then auto-selects a