aa2b2233f08118e110c50971887086e0b0a22966
- Author
- TheEdgeOfRage <git@theedgeofrage.com>
- Committer
- TheEdgeOfRage <git@theedgeofrage.com>
- Date
Message
Diff
This diff is truncated to protect this page.
1diff --git a/.gitignore b/.gitignore
2index b778eb32d4dd9d21424336a0b796ad589494668f..c93b908c3a51b6e85deb995598f4bb10adbcd589 100644
3--- a/.gitignore
4+++ b/.gitignore
5@@ -7,3 +7,6 @@
6 /tmp/
7
8 .pi/mcp.json
9+
10+# Local debug output
11+jp_raw.log
12diff --git a/AGENTS.md b/AGENTS.md
13index 4dfc895fccb52ac2820b8bd48370f67d7d48f0d2..e82f399a3220637eeadf451edf971dcc1c86fc83 100644
14--- a/AGENTS.md
15+++ b/AGENTS.md
16@@ -1,20 +1,53 @@
17-# Project instructions
18+# jp — Japanese learning RPG (TUI)
19
20-## Model services
21+## Overview
22
23-Model services are external, user-managed processes. GPU VRAM is normally almost fully allocated to the chat model.
24+A terminal game for practicing spoken Japanese. The player walks a small ASCII
25+map, talks to people at locations through push-to-talk, and gets separate
26+fluency feedback on each line. The UI renders romaji only; kana and kanji are
27+never shown.
28
29-- Never start, stop, restart, kill, supervise, or reconfigure `llama-server`, `audiocpp_server`, or `whisper-server`.
30-- Never run `llama-cli`, download a model, or use a command that can load a model into VRAM.
31-- Do not add server-launcher or model-management code to this repository.
32-- The application is an HTTP client only. Configure endpoints with flags or environment variables and report unavailable services clearly.
33-- Do not write any tests. There is no test suite in this project; verify changes with `go build ./...` and `go vet ./...`. Only the user may prepare live services for manual integration.
34+The app is an HTTP client for three externally managed model services:
35+
36+- **LLM** — llama.cpp router (OpenAI-compatible chat completions) for NPC
37+ dialogue and the judge. Prompts live in `internal/llm/prompt.go`; strict
38+ `FIELD|value` output contracts are parsed in `internal/llm/contract.go`.
39+- **STT** — whisper.cpp server. The raw Japanese transcript stays internal; it
40+ is what the LLM sees as the player's lines.
41+- **TTS** — OpenAI-compatible speech endpoint that plays the NPC's kana.
42+
43+Turn flow: record → transcribe → judge + NPC reply in parallel → play kana →
44+record the turn in per-location history. Raw model I/O for debugging is
45+appended to `jp_raw.log` in the working directory.
46
47-See `PLAN.md` for the endpoint contract, expected externally managed models, and implementation sequence.
48+## How things are done here
49
50-## Product invariants
51+### Model services
52+
53+Model services are external, user-managed processes. GPU VRAM is normally
54+almost fully allocated to the chat model.
55+
56+- Never start, stop, restart, kill, supervise, or reconfigure `llama-server`,
57+ `audiocpp_server`, or `whisper-server`.
58+- Never run `llama-cli`, download a model, or use a command that can load a
59+ model into VRAM.
60+- Do not add server-launcher or model-management code to this repository.
61+- The application is an HTTP client only. Configure endpoints with flags or
62+ environment variables and report unavailable services clearly.
63+
64+### Product invariants
65
66 - Render romaji only; never show kana or kanji in the TUI.
67-- NPCs behave as normal people, not language teachers. Grading stays separate from NPC dialogue.
68-- Persona-authoring prompts contain no game, player, NPC, quest, or scenario context.
69+- NPCs behave as normal people, not language teachers. Grading stays separate
70+ from NPC dialogue; judge output never enters the NPC prompt.
71+- Persona data contains no game, player, NPC, quest, or scenario context
72+ (enforced by `persona.Validate`).
73 - Keep the Bubble Tea event loop non-blocking.
74+
75+## Development
76+
77+- Verify changes with `go build ./...` and `go vet ./...`. Do not write any
78+ tests; there is no test suite in this project.
79+- Only the user may prepare live services for manual integration.
80+- Docs: `README.md` covers running and configuration only. Operator reference
81+ for the external services lives in `docs/services.md`.
82diff --git a/PLAN.md b/PLAN.md
83deleted file mode 100644
84index 241f7cc556d1379bc859c9515c92a2c41207903f..0000000000000000000000000000000000000000
85--- a/PLAN.md
86+++ /dev/null
87@@ -1,206 +0,0 @@
88-# Japanese learning RPG — implementation plan
89-
90-## Outcome
91-
92-Create a lightweight local Japanese conversation game in Go. The player moves through a small ASCII city in a Bubble Tea TUI, speaks to workers through a microphone, hears Japanese responses, and gets separate fluency feedback.
93-
94-The first demo has three free-roam locations: a ramen shop, a station, and tourist information. There is no save, XP, unlock system, or model-server lifecycle management.
95-
96-## Fixed product rules
97-
98-- UI text is romaji only. Never display kana or kanji in the TUI.
99-- NPC speech plays immediately, but its text stays hidden. `R` reveals romaji; `T` then reveals the already-generated English translation.
100-- Each player turn gets a separate score and short correction in romaji.
101diff --git a/README.md b/README.md
102index fae6c0af88d345fafd7b139aa29cb1055e26f7f7..e0eee2b72b7b76d3951d0ce511966c27b8a3592c 100644
103--- a/README.md
104+++ b/README.md
105@@ -1,14 +1,14 @@
106 # jp — Japanese learning RPG (TUI)
107
108 A terminal game for practicing spoken Japanese. You walk a small map and talk
109-with people to you in romaji. Your words are recorded, transcribed, romanized,
110-and sent to an NPC. A separate judge scores your attempt. The interface shows
111-romaji only; it never displays kana or kanji.
112+with people; your words are recorded, transcribed, and sent to an NPC, and a
113+separate judge scores each line. The interface shows romaji only; it never
114+displays kana or kanji.
115
116 The game is an HTTP client only. It connects to three externally managed model
117 services and **never** starts, stops, restarts, kills, reconfigures, or
118 downloads a model for any of them. You run those services yourself (see
119-`docs/services.md` for operator reference commands).
120+[`docs/services.md`](docs/services.md) for operator reference commands).
121
122 ## Requirements
123
124@@ -18,7 +18,7 @@ downloads a model for any of them. You run those services yourself (see
125 to get the required 16 kHz mono S16_LE WAV. Any command works as long as it
126 writes a 16 kHz mono S16_LE WAV file to the path given as its last argument
127 - Three external model services, all run by you:
128- - **LLM** — OpenAI-compatible chat completions (`POST /chat/completions`)
129+ - **LLM** — llama.cpp router (OpenAI-compatible `POST /v1/chat/completions`)
130 - **TTS** — OpenAI-compatible speech (`POST /audio/speech`, returns WAV)
131 - **STT** — Whisper inference (multipart `POST`, JSON `text` response)
132
133@@ -29,11 +29,12 @@ go build ./...
134 jp # or: go run ./cmd/jp
135 ```
136
137-On startup the TUI appears immediately. A bounded HTTP readiness check runs per
138-service in the background; each service line shows `checking…`, then flips to
139-`up` or `down` in place as its result arrives. A failed check names the affected
140-service and its configured URL. The game never launches a service to make a
141-check pass.
142+On startup the game preloads the LLM model, runs a bounded readiness check per
143+service, then warms each conversation's system prompt once so first replies are
144+fast. The TUI appears immediately; each service line shows `checking…`, then
145+flips to `up` or `down` in place as its result arrives. A failed check names
146+the affected service and its configured URL. The game never launches a service
147+to make a check pass.
148
149 ## Configuration
150
151@@ -42,10 +43,10 @@ local setup.
152
153 | Service | Flag | Env var | Default | Use |
154 | ----------------- | ------------------ | ----------------- | ----------------------------------------- | ----------------------------------------------- |
155-| LLM base URL | `--llm-url` | `JP_LLM_BASE_URL` | `https://llama.home.theedgeofrage.com/v1` | `POST /chat/completions` |
156+| LLM router URL | `--llm-url` | `JP_LLM_BASE_URL` | `https://llama.home.theedgeofrage.com` | `POST /v1/chat/completions` |
157 | TTS base URL | `--tts-url` | `JP_TTS_BASE_URL` | `http://127.0.0.1:8080/v1` | `POST /audio/speech` |
158 | STT server URL | `--stt-url` | `JP_STT_URL` | `http://127.0.0.1:8178` | multipart `POST /inference` |
159-| STT language | `--stt-language` | `JP_STT_LANGUAGE` | `auto` | optional Whisper request language |
160+| STT language | `--stt-language` | `JP_STT_LANGUAGE` | `ja` | Whisper request language (ISO 639-1) |
161 | Recorder command | `--record-command` | — | `arecord` | 16 kHz mono S16_LE WAV capture (see note above) |
162
163 Content paths:
164@@ -57,14 +58,6 @@ Content paths:
165
166 Run `jp -help` for the full flag list.
167
168-## External services
169-
170-The game talks to LLM, TTS, and STT over HTTP and does nothing else with them.
171-Each endpoint is configuration, not a hard-coded process assumption. See
172-[`docs/services.md`](docs/services.md) for what each service must expose and the
173-operator-run reference commands. The game never invokes `llama-server`,
174-`whisper-server`, or any model binary.
175-
176 ## Controls
177
178 - Move: arrow keys or `w`/`a`/`s`/`d`
179@@ -78,33 +71,3 @@ Movement, talk, and reveal keys are ignored while a turn is recording or
180 processing. Each spoken turn is graded by a judge that stays separate from NPC
181 dialogue: the judge scores your attempt in the learning panel; NPCs behave as
182 ordinary people, not language teachers.
183-
184-## Development
185-
186-```bash
187-go build ./...
188-go vet ./...
189-```
190-
191-## Manual end-to-end check
192-
193-Prerequisites, all run by you (the game never starts or stops them):
194-
195-- llama.cpp serving the Unsloth Qwen3 model on port 8081
196-- whisper-server with `ggml-large-v3-turbo` on `127.0.0.1:8178`
197-- audio.cpp serving Qwen3-TTS on `127.0.0.1:8080`
198-
199-Exact reference commands are in [`docs/services.md`](docs/services.md).
200-
201-1. Run `jp`. The TUI appears immediately; each service line flips from
202- `checking` to `up` as its probe result arrives.
203-2. Walk to the ramen shop marker and confirm it is marked active in the legend.
204-3. Press `space`, speak a Japanese line, press `space` again. Your transcript
205diff --git a/cmd/jp/main.go b/cmd/jp/main.go
206index 3e41e7aa265617c5dd680ca7938b6ba7bc745a4f..549b26135d85b7c5fb858eaace87982866728b3d 100644
207--- a/cmd/jp/main.go
208+++ b/cmd/jp/main.go
209@@ -1,7 +1,7 @@
210-// Command jp is the entry point for the Japanese learning RPG.
211 package main
212
213 import (
214+ "context"
215 "fmt"
216 "net/http"
217 "os"
218@@ -23,13 +23,19 @@ import (
219 // recordCap bounds a single push-to-talk capture. It is passed to the recorder
220 // (which enforces it) and to the UI (which shows it as the cap indicator).
221 const recordCap = 10 * time.Second
222+const requestTimeout = 10 * time.Second
223+const warmupTimeout = 60 * time.Second
224
225 func main() {
226 cfg := config.ParseConfig()
227
228- hc := &http.Client{Timeout: 5 * time.Second}
229- if err := availability.CheckLLM(hc, cfg.LLMConfig); err != nil {
230+ hc := &http.Client{Timeout: requestTimeout}
231+
232+ llmClient := llm.NewClient(cfg, hc)
233+ llmUp := true
234+ if err := llmClient.LoadModel(); err != nil {
235 fmt.Fprintf(os.Stderr, "jp: LLM unavailable: %v\n", err)
236+ llmUp = false
237 }
238 if err := availability.CheckTTS(hc, cfg.TTSBaseURL); err != nil {
239 fmt.Fprintf(os.Stderr, "jp: TTS unavailable: %v\n", err)
240@@ -51,9 +57,19 @@ func main() {
241 fatalf("load personas: %v", err)
242 }
243
244+ views := buildPersonaViews(personas)
245+ if llmUp {
246+ ctx, cancel := context.WithTimeout(context.Background(), warmupTimeout)
247+ err := llmClient.Warmup(ctx, warmupSystems(views))
248+ cancel()
249+ if err != nil {
250+ fmt.Fprintf(os.Stderr, "jp: LLM warmup failed: %v\n", err)
251+ }
252+ }
253+
254 state := game.NewState(mapData)
255 playDone := make(chan struct{}, 1)
256- orch := buildOrchestrator(cfg, state, personas, playDone, hc)
257+ orch := buildOrchestrator(cfg, state, views, playDone, hc, llmClient)
258
259 m := ui.NewModel(state, orch, serviceURLs(cfg), playDone, int(recordCap.Seconds()))
260 if _, err := tea.NewProgram(m).Run(); err != nil {
261@@ -71,8 +87,7 @@ func serviceURLs(cfg *config.Config) map[string]string {
262 }
263 }
264
265-func buildOrchestrator(cfg *config.Config, state *game.State, personas []persona.Persona, playDone chan struct{}, hc *http.Client) *game.Orchestrator {
266- llmClient := llm.NewClient(cfg, hc)
267+func buildOrchestrator(cfg *config.Config, state *game.State, views map[string]game.PersonaView, playDone chan struct{}, hc *http.Client, llmClient *llm.Client) *game.Orchestrator {
268 npcModel := &adapters.NPCModel{Client: llmClient, Policy: llm.DefaultHistoryPolicy()}
269 judgeModel := &adapters.JudgeModel{Client: llmClient}
270
271@@ -91,7 +106,18 @@ func buildOrchestrator(cfg *config.Config, state *game.State, personas []persona
272 },
273 }
274
275- return game.NewOrchestrator(state, buildPersonaViews(personas), speechIn, npcModel, judgeModel, speechOut)
276+ return game.NewOrchestrator(state, views, speechIn, npcModel, judgeModel, speechOut)
277+}
278+
279+// warmupSystems lists every system prompt the game will send: one per location
280+// plus the judge. Warming them pins each prefix in the server's KV cache.
281+func warmupSystems(views map[string]game.PersonaView) []string {
282+ out := make([]string, 0, len(views)+1)
283+ for _, v := range views {
284+ out = append(out, llm.NPCSystemPrompt(v.Description, v.Situation))
285+ }
286+ out = append(out, llm.JudgeSystemPrompt())
287+ return out
288 }
289
290 func buildPersonaViews(personas []persona.Persona) map[string]game.PersonaView {
291diff --git a/docs/services.md b/docs/services.md
292index ecd29991542b553688c27bd9779db3ecfdd83303..e5b902bbb0fb74884346e4695c63c0c207dfb29b 100644
293--- a/docs/services.md
294+++ b/docs/services.md
295@@ -6,46 +6,56 @@ downloads a model for any of them. Operators run these services themselves; the
296 game only talks to the configured endpoints over HTTP.
297
298 All endpoints are configuration, not hard-coded process assumptions. Each can be
299-set with a flag (see `jp run -help`) or an environment variable, and defaults
300+set with a flag (see `jp -help`) or an environment variable, and defaults
301 match the local setup:
302
303 | Service | Flag | Env var | Default | Use |
304 | ----------------- | ------------------ | ----------------- | ----------------------------------------- | --------------------------------- |
305-| LLM base URL | `--llm-url` | `JP_LLM_BASE_URL` | `https://llama.home.theedgeofrage.com/v1` | `POST /chat/completions` |
306+| LLM router URL | `--llm-url` | `JP_LLM_BASE_URL` | `https://llama.home.theedgeofrage.com` | `POST /v1/chat/completions` |
307 | TTS base URL | `--tts-url` | `JP_TTS_BASE_URL` | `http://127.0.0.1:8080/v1` | `POST /audio/speech` |
308 | STT server URL | `--stt-url` | `JP_STT_URL` | `http://127.0.0.1:8178` | multipart `POST /inference` |
309-| STT language | `--stt-language` | `JP_STT_LANGUAGE` | `auto` | optional Whisper request language |
310+| STT language | `--stt-language` | `JP_STT_LANGUAGE` | `ja` | Whisper request language (ISO 639-1) |
311 | Recorder command | `--record-command` | - | `arecord` | 16kHz mono S16_LE WAV capture |
312
313-At startup, `jp run` makes only safe, bounded HTTP readiness checks (a few
314-seconds each, no inference requests). It stays silent when every service
315-answers; a down service prints one error line naming the service and its URL.
316-The game never launches or supervises a service to make a check pass.
317+At startup, `jp` first preloads the LLM model with
318+`POST {JP_LLM_BASE_URL}/models/load` (30s timeout) before anything else loads,
319+then makes only safe, bounded HTTP readiness checks (a few seconds each, no
320+inference requests). After persona data loads, it sends one best-effort warmup
321+completion per system prompt (`max_tokens: 1`) so the server's prompt cache
322+keeps those prefixes hot. It stays silent when every service answers; a down
323+service prints one error line naming the service and its URL. The game never
324+launches or supervises a service to make a check pass.
325
326 ## LLM - llama.cpp (Unsloth Qwen3-8B)
327
328-OpenAI-compatible chat completions at `POST {JP_LLM_BASE_URL}/chat/completions`.
329+llama.cpp server in router mode (started without a model argument) at
330+`JP_LLM_BASE_URL`, with OpenAI-compatible chat completions at
331+`POST {JP_LLM_BASE_URL}/v1/chat/completions`.
332
333-Model: `unsloth/Qwen3-8B-GGUF:UD-Q4_K_XL`, served with alias `jp`. The client
334+Model: `unsloth/Qwen3-8B-GGUF:UD-Q4_K_XL`, resolvable by the router as `jp`. The client
335 sends OpenAI-compatible requests with `model` set to `jp`, disables Qwen3
336-thinking on every request (`chat_template_kwargs.enable_thinking: false`), and
337-reads streaming SSE. The model chat template (`--jinja`) is used; ChatML is not
338-built manually by the client.
339+thinking on every request (`chat_template_kwargs.enable_thinking: false`),
340+sets `cache_prompt: true` so the server keeps prompt prefixes cached between
341+turns, and reads streaming SSE. The model chat template (`--jinja`) is used;
342+ChatML is not built manually by the client.
343+
344+At startup, before anything else loads, the game preloads the model with
345+`POST {JP_LLM_BASE_URL}/models/load` and body `{"model": "jp"}` (30s timeout).
346+A failed preload prints one error line and startup continues.
347
348 > **Operator-run only - the game never starts this.**
349 >
350 > ```bash
351-> llama-server -hf unsloth/Qwen3-8B-GGUF:UD-Q4_K_XL \
352-> --alias jp --jinja --reasoning-format deepseek --ctx-size 8192 --port 8081
353+> llama-server # router mode; `jp` must resolve to the Qwen3 model above
354 > ```
355
356 ## STT - Whisper server (ggml-large-v3-turbo)
357
358 Multilingual `ggml-large-v3-turbo.bin` served on `127.0.0.1:8178`. Each turn,
359 the game uploads a 16kHz mono S16_LE WAV as multipart form data to
360-`JP_STT_URL/inference` with the recording as a form file (`file=@recording.wav`). The optional
361-`language` field is included only when it is non-empty. The response is JSON
362-with a `text` field.
363+`JP_STT_URL/inference` with the recording as a form file (`file=@recording.wav`). The
364+`language` field is included only when it is non-empty (default `ja`, which
365+forces Japanese transcription). The response is JSON with a `text` field.
366
367 > **Operator-run only - the game never starts this.**
368 >
369diff --git a/internal/adapters/llm.go b/internal/adapters/llm.go
370index d98f1a30d0f379738c861f88d5ee5d311a903fd2..20f1220ba395987ef8d96d8ba874ecc792333f88 100644
371--- a/internal/adapters/llm.go
372+++ b/internal/adapters/llm.go
373@@ -7,11 +7,27 @@ package adapters
374
375 import (
376 "context"
377+ "fmt"
378+ "os"
379+ "time"
380
381 "japanese/internal/game"
382 "japanese/internal/llm"
383 )
384
385+// rawLogPath is where raw model outputs are dumped for debugging, since the
386+// TUI owns the terminal and direct stderr writes get mangled by redraws.
387+const rawLogPath = "jp_raw.log"
388+
389+func dumpRaw(label, transcript, raw string) {
390+ f, err := os.OpenFile(rawLogPath, os.O_APPEND|os.O_CREATE|os.O_WRONLY, 0o644)
391+ if err != nil {
392+ return
393+ }
394+ defer f.Close()
395+ fmt.Fprintf(f, "--- %s %s ---\ntranscript: %s\n%s\n\n", label, time.Now().Format(time.RFC3339), transcript, raw)
396+}
397+
398 // NPCModel adapts an llm.Client to game.NPCModel. It builds the persona and
399 // situation prompt with bounded history and parses the strict three-field reply.
400 type NPCModel struct {
401@@ -29,6 +45,7 @@ func (m *NPCModel) Reply(ctx context.Context, req game.NPCRequest) (llm.NPCReply
402 if err != nil {
403 return llm.NPCReply{}, err
404 }
405+ dumpRaw("npc", req.Transcript, raw)
406 return llm.ParseNPCReply(raw)
407 }
408
409@@ -44,5 +61,6 @@ func (m *JudgeModel) Judge(ctx context.Context, req game.JudgeRequest) (llm.Judg
410 if err != nil {
411 return llm.JudgeResult{}, err
412 }
413+ dumpRaw("judge", req.Transcript, raw)
414 return llm.ParseJudge(raw)
415 }
416diff --git a/internal/availability/availability.go b/internal/availability/availability.go
417index c21f6d0ac1fee50d4f92e94b43508addd96621da..4ecd7dad2737ca35429c41c3bd2f1ecaebe3818e 100644
418--- a/internal/availability/availability.go
419+++ b/internal/availability/availability.go
420@@ -4,44 +4,12 @@
421 package availability
422
423 import (
424- "encoding/json"
425 "fmt"
426 "io"
427 "net/http"
428- "slices"
429 "strings"
430-
431- "japanese/internal/config"
432 )
433
434-// CheckLLM verifies the LLM endpoint answers and lists the configured model.
435-func CheckLLM(client *http.Client, cfg config.LLMConfig) error {
436- modelsURL := strings.TrimRight(cfg.BaseURL, "/") + "/models"
437- status, body, err := doGet(client, modelsURL)
438- if err != nil {
439- return fmt.Errorf("GET %s: %w", modelsURL, err)
440- }
441- if status != http.StatusOK {
442- return fmt.Errorf("GET %s: HTTP %d", modelsURL, status)
443- }
444-
445- var m struct {
446- Data []struct {
447- ID string `json:"id"`
448- Aliases []string `json:"aliases"`
449- } `json:"data"`
450- }
451- if err := json.Unmarshal(body, &m); err != nil {
452- return fmt.Errorf("decode models: %w", err)
453- }
454- for _, d := range m.Data {
455- if d.ID == cfg.Model || slices.Contains(d.Aliases, cfg.Model) {
456- return nil
457- }
458- }
459- return fmt.Errorf("model %q not found", cfg.Model)
460-}
461-
462 // CheckTTS verifies the TTS server answers. A 404 on /models is fine: some
463 // servers do not expose it.
464 func CheckTTS(client *http.Client, base string) error {
465diff --git a/internal/config/config.go b/internal/config/config.go
466index e7f1e2120371adf594c03c8f37bd52832c0ddfc2..b38a1eb021debe2b3bc67e648349332af64f3e68 100644
467--- a/internal/config/config.go
468+++ b/internal/config/config.go
469@@ -12,7 +12,7 @@ import (
470 )
471
472 type LLMConfig struct {
473- BaseURL string `long:"url" env:"JP_LLM_BASE_URL" default:"https://llama.home.theedgeofrage.com/v1" description:"OpenAI-compatible LLM base URL"`
474+ BaseURL string `long:"url" env:"JP_LLM_BASE_URL" default:"https://llama.home.theedgeofrage.com" description:"llama.cpp router base URL"`
475 Model string `long:"model" env:"MODEL" default:"jp"`
476 Temperature float64 `long:"temperature" env:"TEMPERATURE" default:"1.0"`
477 MaxTokens int `long:"max-tokens" env:"MAX_TOKENS" default:"256"`
478@@ -25,7 +25,7 @@ type Config struct {
479
480 TTSBaseURL string `long:"tts-url" env:"JP_TTS_BASE_URL" default:"http://127.0.0.1:8080/v1" description:"OpenAI-compatible TTS base URL"`
481 STTURL string `long:"stt-url" env:"JP_STT_URL" default:"http://127.0.0.1:8178" description:"Whisper server URL"`
482- STTLanguage string `long:"stt-language" env:"JP_STT_LANGUAGE" default:"auto" description:"Optional Whisper request language"`
483+ STTLanguage string `long:"stt-language" env:"JP_STT_LANGUAGE" default:"ja" description:"Whisper request language (ISO 639-1 code)"`
484 RecordCommand string `long:"record-command" default:"arecord" description:"Recorder command for 16kHz mono S16_LE WAV capture"`
485
486 MapPath string `long:"map" env:"JP_MAP_PATH" default:"assets/maps/city.json" description:"path to the city map JSON"`
487diff --git a/internal/game/orchestrator.go b/internal/game/orchestrator.go
488index 5eb41d43afde5c94c885343ba780f2acd33cb018..4a4b5231a2a140f4057eaf36c1b1a160410f8150 100644
489--- a/internal/game/orchestrator.go
490+++ b/internal/game/orchestrator.go
491@@ -153,14 +153,14 @@ func (o *Orchestrator) Finish(ctx context.Context) TurnResult {
492 res.SpeakErr = err
493 }
494
495- o.state.RecordTurn(loc, raw, Reply{Romaji: npcRes.Romaji, English: npcRes.English})
496+ o.state.RecordTurn(loc, raw, npcRes.Kana, Reply{Romaji: npcRes.Romaji, English: npcRes.English})
497 return res
498 }
499
500 func toLLMTurns(turns []Turn) []llm.Turn {
501 out := make([]llm.Turn, len(turns))
502 for i, t := range turns {
503- out[i] = llm.Turn{User: t.PlayerRaw, Assistant: t.NPC.Romaji}
504+ out[i] = llm.Turn{User: t.PlayerRaw, Assistant: t.NPCSpoken}
505 }
506 return out
507 }
508diff --git a/internal/game/state.go b/internal/game/state.go
509index 80aa9a5b2b1afc50776857230b28987d4021f66e..6c46fc7028f90439e02c49f7943a826a94dec731 100644
510--- a/internal/game/state.go
511+++ b/internal/game/state.go
512@@ -52,10 +52,11 @@ type Reply struct {
513 }
514
515 // Turn is one exchange in a location's history: the player's spoken line and
516-// the NPC reply to it. PlayerRaw is the original transcript, kept for LLM
517-// prompts only.
518+// the NPC reply to it. PlayerRaw and NPCSpoken are the original Japanese
519+// utterances, kept for LLM prompts only; NPC holds the display fields.
520 type Turn struct {
521 PlayerRaw string
522+ NPCSpoken string
523 NPC Reply
524 }
525
526@@ -131,10 +132,11 @@ func (s *State) History(id string) []Turn {
527
528 // RecordTurn appends one exchange to location id's history. A recorded reply
529 // always starts with both text fields hidden.
530-func (s *State) RecordTurn(id string, playerRaw string, npc Reply) {
531+func (s *State) RecordTurn(id string, playerRaw, npcSpoken string, npc Reply) {
532 npc.Reveal = Hidden
533 s.history[id] = append(s.history[id], Turn{
534 PlayerRaw: playerRaw,
535+ NPCSpoken: npcSpoken,
536 NPC: npc,
537 })
538 }
539diff --git a/internal/llm/client.go b/internal/llm/client.go
540index b07869206dd36d6f58813746bdeab6002a746b9f..607c822546d856629082cfaf9111651be566ef8f 100644
541--- a/internal/llm/client.go
542+++ b/internal/llm/client.go
543@@ -14,6 +14,7 @@ import (
544 "japanese/internal/config"
545 "net/http"
546 "strings"
547+ "time"
548 )
549
550 type Role string
551@@ -59,6 +60,7 @@ type chatRequest struct {
552 Temperature float64 `json:"temperature"`
553 MaxTokens int `json:"max_tokens"`
554 Stream bool `json:"stream"`
555+ CachePrompt bool `json:"cache_prompt"`
556 }
557
558 type sseDelta struct {
559@@ -73,6 +75,23 @@ type sseDelta struct {
560 // content, consuming SSE data events until [DONE]. The context bounds the whole
561 // call, including the read.
562 func (c *Client) Generate(ctx context.Context, msgs []Message) (string, error) {
563+ return c.chat(ctx, msgs, c.MaxTokens)
564+}
565+
566+// Warmup sends each system prompt once with a minimal user line and one output
567+// token so the server keeps those prompt prefixes in its KV cache before the
568+// first real turn. It is best-effort; callers report failures as warnings.
569+func (c *Client) Warmup(ctx context.Context, systems []string) error {
570+ for _, s := range systems {
571+ msgs := []Message{{Role: RoleSystem, Content: s}, {Role: RoleUser, Content: "Ready."}}
572+ if _, err := c.chat(ctx, msgs, 1); err != nil {
573+ return fmt.Errorf("llm: warmup: %w", err)
574+ }
575+ }
576+ return nil
577+}
578+
579+func (c *Client) chat(ctx context.Context, msgs []Message, maxTokens int) (string, error) {
580 if c.baseURL == "" {
581 return "", fmt.Errorf("llm: no base URL configured")
582 }
583@@ -81,15 +100,16 @@ func (c *Client) Generate(ctx context.Context, msgs []Message) (string, error) {
584 Messages: msgs,
585 ChatTemplateKwargs: map[string]any{"enable_thinking": c.EnableThinking},
586 Temperature: c.Temperature,
587- MaxTokens: c.MaxTokens,
588+ MaxTokens: maxTokens,
589 Stream: true,
590+ CachePrompt: true,
591 }
592 payload, err := json.Marshal(body)
593 if err != nil {
594 return "", fmt.Errorf("llm: encode request: %w", err)
595 }
596
597- url := strings.TrimRight(c.baseURL, "/") + "/chat/completions"
598+ url := strings.TrimRight(c.baseURL, "/") + "/v1/chat/completions"
599 req, err := http.NewRequestWithContext(ctx, http.MethodPost, url, bytes.NewReader(payload))
600 if err != nil {
601 return "", fmt.Errorf("llm: build request: %w", err)
602@@ -114,6 +134,44 @@ func (c *Client) Generate(ctx context.Context, msgs []Message) (string, error) {
603 return text, nil
604 }
605
606+// loadTimeout bounds the model preload request. Loading a large model into
607+// VRAM takes much longer than a normal API call.
608+const loadTimeout = 30 * time.Second
609+
610+// LoadModel asks the llama.cpp router to load the configured model so the
611+// first turn does not pay the loading cost. It is safe to call when the model
612+// is already loaded; the router keeps it loaded.
613+func (c *Client) LoadModel() error {
614+ if c.baseURL == "" {
615+ return fmt.Errorf("llm: no base URL configured")
616+ }
617+ ctx, cancel := context.WithTimeout(context.Background(), loadTimeout)
618+ defer cancel()
619+
620+ payload, err := json.Marshal(map[string]string{"model": c.Model})
621+ if err != nil {
622+ return fmt.Errorf("llm: encode request: %w", err)
623+ }
624+ url := strings.TrimRight(c.baseURL, "/") + "/models/load"
625+ req, err := http.NewRequestWithContext(ctx, http.MethodPost, url, bytes.NewReader(payload))
626+ if err != nil {
627+ return fmt.Errorf("llm: build request: %w", err)
628+ }
629+ req.Header.Set("Content-Type", "application/json")
630+
631+ resp, err := c.httpClient.Do(req)
632+ if err != nil {
633+ return fmt.Errorf("llm: request to %s: %w", url, err)
634+ }
635+ defer func() { _ = resp.Body.Close() }()
636+
637+ if resp.StatusCode < 200 || resp.StatusCode >= 300 {
638+ detail, _ := io.ReadAll(io.LimitReader(resp.Body, 4096))
639+ return fmt.Errorf("llm: %s returned HTTP %d: %s", url, resp.StatusCode, strings.TrimSpace(string(detail)))
640+ }
641+ return nil
642+}
643diff --git a/internal/llm/prompt.go b/internal/llm/prompt.go
644index 5cad800b04e88d965e5f19d6d14a0f475d7e20bd..9fb65176de2db58a95f0f1e9864ffc9ddbd4fbbf 100644
645--- a/internal/llm/prompt.go
646+++ b/internal/llm/prompt.go
647@@ -15,10 +15,16 @@ Situation: %s
648
649 Speak naturally, as yourself, in Japanese. Keep your reply short and conversational. Do not teach language, correct mistakes, or mention any evaluation.
650
651-Respond with exactly three lines and nothing else, in this order:
652+Your reply must be exactly three lines of plain text, in this order. No extra lines, no markdown, no code fences. Every line starts with its field key, then a pipe character |, then the value. Never omit, rename, or reorder the keys.
653+
654 ROMAJI|<romaji of the sentence you say>
655 KANA|<the same sentence written in kana>
656-ENGLISH|<a natural English translation of that sentence>`
657+ENGLISH|<a natural English translation of that sentence>
658+
659+Example reply:
660+ROMAJI|nee san, chuumon wa?
661+KANA|ねえさん、注文は?
662+ENGLISH|Miss, what is your order?`
663
664 // judgeSystemTemplate is persona-free and fixes the three-field judge contract.
665 // It scores fluency, naturalness, and fit to the situation, accepting any
666@@ -28,16 +34,31 @@ const judgeSystemTemplate = `You evaluate one spoken Japanese sentence for a lan
667
668 Also transcribe the given Japanese line into Hepburn romaji, keeping its meaning intact. Write only romaji in every field: never kana or kanji.
669
670-Respond with exactly three lines and nothing else, in this order:
671+Your reply must be exactly three lines of plain text, in this order. No extra lines, no markdown, no code fences. Every line starts with its field key, then a pipe character |, then the value. Never omit, rename, or reorder the keys.
672+
673 SCORE|<integer from 0 to 100>
674 ROMAJI|<hepburn romaji transcription of the given Japanese line>
675-FEEDBACK|<one or two concise sentences in romaji>`
676+FEEDBACK|<one or two concise sentences in english>
677+
678+Example reply:
679+SCORE|85
680+ROMAJI|konnichiwa
681+FEEDBACK|It's a polite greeting, but you can use a more specific greeting for the time period e.g. "ohayoo gozaimasu"`
682+
683+// NPCSystemPrompt renders the NPC system message for one location. Every NPC
684+// request for that location starts with this exact text.
685+func NPCSystemPrompt(persona, situation string) string {
686+ return fmt.Sprintf(npcSystemTemplate, fallback(persona, "a friendly local"), fallback(situation, "everyday conversation"))
687+}
688+
689+// JudgeSystemPrompt returns the static judge system message.
690+func JudgeSystemPrompt() string { return judgeSystemTemplate }
691
692 // BuildNPCMessages assembles the NPC request: a system message holding the
693 // stable persona and situation, bounded prior history, and the current raw
694 // transcript. Judge output never enters this prompt.
695 func BuildNPCMessages(persona, situation, transcript string, prior []Turn, policy HistoryPolicy) []Message {
696- system := fmt.Sprintf(npcSystemTemplate, fallback(persona, "a friendly local"), fallback(situation, "everyday conversation"))
697+ system := NPCSystemPrompt(persona, situation)
698 historyBudget := policy.PromptCharBudget - runeCount(system) - runeCount(transcript)
699 if historyBudget < 0 {
700 historyBudget = 0