10acfb9d8827a6ece181c33381f7dd8f666f62f4

Author
TheEdgeOfRage <git@theedgeofrage.com>
Committer
TheEdgeOfRage <git@theedgeofrage.com>
Date

Message

Drop docs/services.md, fold service run instructions into README

Diff

This diff is truncated to protect this page.

  1diff --git a/AGENTS.md b/AGENTS.md
  2index 5212071c86234eb66653b33dbacd6a13ebe64171..11023772477a0d8d92168657988e7e2e1bfb1ed1 100644
  3--- a/AGENTS.md
  4+++ b/AGENTS.md
  5@@ -79,6 +79,5 @@ almost fully allocated to the chat model.
  6   `go build`, `golangci-lint fmt`, and `golangci-lint run`). Do not write any
  7   tests; there is no test suite in this project.
  8 - Only the user may prepare live services for manual integration.
  9-- Docs: `README.md` covers running and configuration only. Operator reference
 10-  for the external services lives in `docs/services.md`.
 11+- Docs: `README.md` covers running and configuration only.
 12 - DO NOT cd into the directory you're already in when running the bash tool
 13diff --git a/README.md b/README.md
 14index 3df5ad7210f324d8f68e68d4e2ddb6559c693a7b..b16f01b536b52b19d1230bbe89ed2ba39d7be67a 100644
 15--- a/README.md
 16+++ b/README.md
 17@@ -9,8 +9,7 @@ only; it never displays kana or kanji.
 18 
 19 The game is an HTTP client only. It connects to two externally managed model
 20 services and **never** starts, stops, restarts, kills, reconfigures, or
 21-downloads a model for any of them. You run those services yourself (see
 22-[`docs/services.md`](docs/services.md) for operator reference commands).
 23+downloads a model for any of them. You run those services yourself.
 24 
 25 ## Requirements
 26 
 27@@ -18,10 +17,15 @@ downloads a model for any of them. You run those services yourself (see
 28 - `arecord` (ALSA) for microphone capture. The game invokes bare `arecord`
 29   (16 kHz mono S16_LE WAV written to the path given as its last argument)
 30 - Two external model services, both run by you:
 31-  - **LLM** — llama.cpp router (OpenAI-compatible `POST /v1/chat/completions`)
 32+  - **LLM** — llama.cpp router (OpenAI-compatible `POST /v1/chat/completions`).
 33+    Start it with `make llama`
 34   - **Audio** — audio.cpp server (`JP_AUDIO_BASE_URL`): TTS speech endpoint
 35     (`POST /v1/audio/speech`, returns WAV) and ASR transcriptions (multipart
 36-    `POST /v1/audio/transcriptions`, JSON `text` response)
 37+    `POST /v1/audio/transcriptions`, JSON `text` response). Start it from the
 38+    repository root with `audiocpp_server --config audio.cpp.json`; keep
 39+    `lazy_load: true` in that config so both speech models stay resident —
 40+    the game alternates TTS and transcription every turn, so unloading would
 41+    thrash
 42 
 43 ## Build and run
 44 
 45@@ -58,10 +62,14 @@ Run `jp -help` for the full flag list.
 46 - Type an action in English and press `enter` to send it
 47 - Ask how to say something in Japanese: type `?your question` and press
 48   `enter`. The answer is a learning aid; it does not affect the world
 49+- Generate flashcards: type `!instructions` and press `enter`. The LLM
 50+  generates a Japanese sentence matching the instructions and appends the
 51+  vocabulary to `flashcards.csv`
 52 - Push-to-talk: `F2` (press to start recording, press again to stop and send
 53   the turn)
 54 - Reveal NPC romaji: `F3`
 55 - Reveal NPC English: `F4`
 56+- Replay the last spoken line: `F5`
 57 - Quit: `esc`
 58 
 59 Input is ignored while a turn is recording or processing. Each spoken turn is
 60diff --git a/docs/services.md b/docs/services.md
 61deleted file mode 100644
 62index 882484ab8683cc8655af1132027efcd75c46eb81..0000000000000000000000000000000000000000
 63--- a/docs/services.md
 64+++ /dev/null
 65@@ -1,76 +0,0 @@
 66-# External services
 67-
 68-The game is an HTTP client only. It connects to two externally managed model
 69-services and **never** starts, stops, restarts, kills, reconfigures, or
 70-downloads a model for any of them. Operators run these services themselves; the
 71-game only talks to the configured endpoints over HTTP.
 72-
 73-All endpoints are configuration, not hard-coded process assumptions. Each can be
 74-set with a flag (see `jp -help`) or an environment variable, and defaults
 75-match the local setup:
 76-
 77-| Service          | Flag               | Env var             | Default                                | Use                                                               |
 78-| ---------------- | ------------------ | ------------------- | -------------------------------------- | ----------------------------------------------------------------- |
 79-| LLM router URL   | `--url`            | `JP_LLM_BASE_URL`   | `https://llama.home.theedgeofrage.com` | `POST /v1/chat/completions`                                       |
 80-| Audio base URL   | `--audio-url`      | `JP_AUDIO_BASE_URL` | `http://127.0.0.1:8080`                | TTS `POST /v1/audio/speech` + ASR `POST /v1/audio/transcriptions` |
 81-
 82-At startup, `jp` waits for the LLM router to accept completions: a minimal
 83-chat completion (`max_tokens: 1`) polled once per second, up to 30s, so a
 84-freshly started model has time to load. It then checks the audio server at
 85-`/v1/models` (bounded, no inference requests) and sends one best-effort
 86-warmup completion per system prompt (`max_tokens: 1`) so the server's prompt
 87-cache keeps those prefixes hot. It stays silent when every service answers; a
 88-down or slow service prints one error line naming the service and its URL.
 89-The game never launches or supervises a service to make a check pass.
 90-
 91-## LLM - llama.cpp (Unsloth Gemma 4 12B)
 92-
 93-llama.cpp server in single-instance mode at `JP_LLM_BASE_URL`, with
 94-OpenAI-compatible chat completions at `POST {JP_LLM_BASE_URL}/v1/chat/completions`.
 95-
 96-Model: `unsloth/gemma-4-12B-it-qat-GGUF:UD-Q4_K_XL`. The client sends
 97-OpenAI-compatible requests without a model field (the single-instance server
 98-runs exactly one model), sends `chat_template_kwargs.enable_thinking` per the
 99-`--enable-thinking` flag (default false), sets `cache_prompt: true` so the
100-server keeps prompt prefixes cached between turns, pins each request to a
101-llama-server slot via `id_slot` (see below), and reads streaming SSE. The model
102-chat template (`--jinja`) is used; ChatML is not built manually by the client.
103-
104-### Slot pinning
105-
106-Each request body carries an `id_slot` field that assigns the task to a