Diff
1diff --git a/docs/services.md b/docs/services.md
2new file mode 100644
3index 0000000000000000000000000000000000000000..7139cb5a0c002c8af1890170e3779a68a746c786
4--- /dev/null
5+++ b/docs/services.md
6@@ -0,0 +1,68 @@
7+# External services
8+
9+The game is an HTTP client only. It connects to three externally managed model
10+services and **never** starts, stops, restarts, kills, reconfigures, or
11+downloads a model for any of them. Operators run these services themselves; the
12+game only talks to the configured endpoints over HTTP.
13+
14+All endpoints are configuration, not hard-coded process assumptions. Each can be
15+set with a flag (see `jp run -help`) or an environment variable, and defaults
16+match the local setup:
17+
18+| Service | Flag | Env var | Default | Use |
19+| --- | --- | --- | --- | --- |
20+| LLM base URL | `--llm-url` | `JP_LLM_BASE_URL` | `http://127.0.0.1:8081/v1` | `POST /chat/completions` |
21+| TTS base URL | `--tts-url` | `JP_TTS_BASE_URL` | `http://127.0.0.1:8080/v1` | `POST /audio/speech` |
22+| STT inference URL | `--stt-url` | `JP_STT_URL` | `http://127.0.0.1:8178/inference` | multipart `POST` |
23+| STT language | `--stt-language` | `JP_STT_LANGUAGE` | `auto` | optional Whisper request language |
24+| Recorder command | `--record-command` | - | `arecord` | 16kHz mono S16_LE WAV capture |
25+
26+At startup, `jp run` makes only safe, bounded HTTP readiness checks (a few
27+seconds each, no inference requests) and prints a per-service status line with
28+the configured URL and, when down, the connection error. A failed check names
29+the affected service and its URL. The game never launches or supervises a
30+service to make a check pass.
31+
32+## LLM - llama.cpp (Unsloth Qwen3-8B)
33+
34+OpenAI-compatible chat completions at `POST {JP_LLM_BASE_URL}/chat/completions`.
35+
36+Model: `unsloth/Qwen3-8B-GGUF:UD-Q4_K_XL`, served with alias `jp`. The client
37+sends OpenAI-compatible requests with `model` set to `jp`, disables Qwen3
38+thinking on every request (`chat_template_kwargs.enable_thinking: false`), and
39+reads streaming SSE. The model chat template (`--jinja`) is used; ChatML is not
40+built manually by the client.
41+
42+> **Operator-run only - the game never starts this.**
43+>
44+> ```bash
45+> llama-server -hf unsloth/Qwen3-8B-GGUF:UD-Q4_K_XL \
46+> --alias jp --jinja --reasoning-format deepseek --ctx-size 8192 --port 8081
47+> ```
48+
49+## STT - Whisper server (ggml-large-v3-turbo)
50+
51+Multilingual `ggml-large-v3-turbo.bin` served on `127.0.0.1:8178`. Each turn,
52+the game uploads a 16kHz mono S16_LE WAV as multipart form data to `JP_STT_URL`
53+with the recording as a form file (`file=@recording.wav`). The optional
54+`language` field is included only when it is non-empty. The response is JSON
55+with a `text` field.
56+
57+> **Operator-run only - the game never starts this.**
58+>
59+> ```bash
60+> whisper-server --host 127.0.0.1 --port 8178 \
61+> --model ~/.local/share/llama-server/models/ggml-large-v3-turbo.bin \
62+> --language auto --threads 8
63+> ```
64+
65+The game does not invoke `whisper-server` or `whisper-cli`.
66+
67+## TTS - audio.cpp (Qwen3-TTS)
68+
69+OpenAI-compatible speech endpoint at `POST {JP_TTS_BASE_URL}/audio/speech`,
70+returning WAV bytes. The service uses the existing configuration in
71+`~/dev/tts/audio.cpp.qwen3.json` and the existing Qwen3-TTS 1.7B CustomVoice
72+GGUF. Reference only; the operator runs this service.
73+
74+> **Operator-run only - the game never starts this.**