2fa2c4ae80eab50831abc847c837629547c1be6d

Author
TheEdgeOfRage <git@theedgeofrage.com>
Committer
TheEdgeOfRage <git@theedgeofrage.com>
Date

Message

Update default LLM endpoint

Diff

This diff is truncated to protect this page.

  1diff --git a/PLAN.md b/PLAN.md
  2index c9285a74073ab313ec36fda1c54f58e0897652c5..035aa5ca3dee0947e4722eecfc9d54bb43d75063 100644
  3--- a/PLAN.md
  4+++ b/PLAN.md
  5@@ -1,11 +1,13 @@
  6 # Japanese learning RPG — implementation plan
  7 
  8 ## Outcome
  9+
 10 Create a lightweight local Japanese conversation game in Go. The player moves through a small ASCII city in a Bubble Tea TUI, speaks to workers through a microphone, hears Japanese responses, and gets separate fluency feedback.
 11 
 12 The first demo has three free-roam locations: a ramen shop, a station, and tourist information. There is no save, XP, unlock system, or model-server lifecycle management.
 13 
 14 ## Fixed product rules
 15+
 16 - UI text is romaji only. Never display kana or kanji in the TUI.
 17 - NPC speech plays immediately, but its text stays hidden. `R` reveals romaji; `T` then reveals the already-generated English translation.
 18 - Each player turn gets a separate score and short correction in romaji.
 19@@ -14,24 +16,27 @@ The first demo has three free-roam locations: a ramen shop, a station, and touri
 20 - In-game prompts describe a real situation only: never mention a game, NPC, player, or quest.
 21 
 22 ## Non-negotiable runtime rule
 23+
 24 The game is an HTTP client of externally managed model services. It must **never** start, stop, restart, kill, configure, or download a model for `llama-server`, `audiocpp_server`, or `whisper-server`.
 25 
 26diff --git a/README.md b/README.md
 27index 3373762676b99831f2d4f1b38067926a384ecc87..e3a51b5c3dcf3345e551cc0721bd339a76bd6515 100644
 28--- a/README.md
 29+++ b/README.md
 30@@ -40,13 +40,13 @@ check pass.
 31 Every endpoint is set with a flag or an environment variable. Defaults match a
 32 local setup.
 33 
 34-| Service           | Flag               | Env var           | Default                           | Use                                             |
 35-| ----------------- | ------------------ | ----------------- | --------------------------------- | ----------------------------------------------- |
 36-| LLM base URL      | `--llm-url`        | `JP_LLM_BASE_URL` | `http://127.0.0.1:8081/v1`        | `POST /chat/completions`                        |
 37-| TTS base URL      | `--tts-url`        | `JP_TTS_BASE_URL` | `http://127.0.0.1:8080/v1`        | `POST /audio/speech`                            |
 38-| STT inference URL | `--stt-url`        | `JP_STT_URL`      | `http://127.0.0.1:8178/inference` | multipart `POST`                                |
 39-| STT language      | `--stt-language`   | `JP_STT_LANGUAGE` | `auto`                            | optional Whisper request language               |
 40-| Recorder command  | `--record-command` | —                 | `arecord`                         | 16 kHz mono S16_LE WAV capture (see note above) |
 41+| Service           | Flag               | Env var           | Default                                   | Use                                             |
 42+| ----------------- | ------------------ | ----------------- | ----------------------------------------- | ----------------------------------------------- |
 43+| LLM base URL      | `--llm-url`        | `JP_LLM_BASE_URL` | `https://llama.home.theedgeofrage.com/v1` | `POST /chat/completions`                        |
 44+| TTS base URL      | `--tts-url`        | `JP_TTS_BASE_URL` | `http://127.0.0.1:8080/v1`                | `POST /audio/speech`                            |
 45+| STT inference URL | `--stt-url`        | `JP_STT_URL`      | `http://127.0.0.1:8178/inference`         | multipart `POST`                                |
 46+| STT language      | `--stt-language`   | `JP_STT_LANGUAGE` | `auto`                                    | optional Whisper request language               |
 47+| Recorder command  | `--record-command` | —                 | `arecord`                                 | 16 kHz mono S16_LE WAV capture (see note above) |
 48 
 49 Content paths:
 50 
 51diff --git a/docs/services.md b/docs/services.md
 52index 7139cb5a0c002c8af1890170e3779a68a746c786..39a4edd3c228d5b47ca4b93693fe7b1083e3e096 100644
 53--- a/docs/services.md
 54+++ b/docs/services.md
 55@@ -9,13 +9,13 @@ All endpoints are configuration, not hard-coded process assumptions. Each can be
 56 set with a flag (see `jp run -help`) or an environment variable, and defaults
 57 match the local setup:
 58 
 59-| Service | Flag | Env var | Default | Use |
 60-| --- | --- | --- | --- | --- |
 61-| LLM base URL | `--llm-url` | `JP_LLM_BASE_URL` | `http://127.0.0.1:8081/v1` | `POST /chat/completions` |
 62-| TTS base URL | `--tts-url` | `JP_TTS_BASE_URL` | `http://127.0.0.1:8080/v1` | `POST /audio/speech` |
 63-| STT inference URL | `--stt-url` | `JP_STT_URL` | `http://127.0.0.1:8178/inference` | multipart `POST` |
 64-| STT language | `--stt-language` | `JP_STT_LANGUAGE` | `auto` | optional Whisper request language |
 65-| Recorder command | `--record-command` | - | `arecord` | 16kHz mono S16_LE WAV capture |
 66+| Service           | Flag               | Env var           | Default                                   | Use                               |
 67+| ----------------- | ------------------ | ----------------- | ----------------------------------------- | --------------------------------- |
 68+| LLM base URL      | `--llm-url`        | `JP_LLM_BASE_URL` | `https://llama.home.theedgeofrage.com/v1` | `POST /chat/completions`          |
 69+| TTS base URL      | `--tts-url`        | `JP_TTS_BASE_URL` | `http://127.0.0.1:8080/v1`                | `POST /audio/speech`              |
 70+| STT inference URL | `--stt-url`        | `JP_STT_URL`      | `http://127.0.0.1:8178/inference`         | multipart `POST`                  |
 71+| STT language      | `--stt-language`   | `JP_STT_LANGUAGE` | `auto`                                    | optional Whisper request language |
 72+| Recorder command  | `--record-command` | -                 | `arecord`                                 | 16kHz mono S16_LE WAV capture     |
 73 
 74 At startup, `jp run` makes only safe, bounded HTTP readiness checks (a few
 75 seconds each, no inference requests) and prints a per-service status line with
 76diff --git a/internal/config/config.go b/internal/config/config.go
 77index c0027bb55b90914609c6644edb11b5ec7eed6a53..6dcbc0c8b6acd3e32c2cf8bf212b1063aec4957d 100644
 78--- a/internal/config/config.go
 79+++ b/internal/config/config.go
 80@@ -7,7 +7,7 @@ package config
 81 // Config is the shared option group. Subcommands embed it in their own options
 82 // struct and parse with github.com/jessevdk/go-flags.
 83 type Config struct {
 84-	LLMBaseURL    string `long:"llm-url" env:"JP_LLM_BASE_URL" default:"http://127.0.0.1:8081/v1" description:"OpenAI-compatible LLM base URL"`
 85+	LLMBaseURL    string `long:"llm-url" env:"JP_LLM_BASE_URL" default:"https://llama.home.theedgeofrage.com/v1" description:"OpenAI-compatible LLM base URL"`
 86 	TTSBaseURL    string `long:"tts-url" env:"JP_TTS_BASE_URL" default:"http://127.0.0.1:8080/v1" description:"OpenAI-compatible TTS base URL"`
 87 	STTURL        string `long:"stt-url" env:"JP_STT_URL" default:"http://127.0.0.1:8178/inference" description:"Whisper inference URL"`
 88 	STTLanguage   string `long:"stt-language" env:"JP_STT_LANGUAGE" default:"auto" description:"Optional Whisper request language"`
 89diff --git a/internal/llm/client.go b/internal/llm/client.go
 90index fc6551c9454fb7afcd17a950b329ea6782e9c5cc..0a239cebc69cad200c44eaacdc72e23d56a61803 100644
 91--- a/internal/llm/client.go
 92+++ b/internal/llm/client.go
 93@@ -15,6 +15,8 @@ import (
 94 	"strings"
 95 )
 96 
 97+const llmModel = "qwen3.8-27b-q3"
 98+
 99 type Role string
100 
101 const (
102@@ -38,7 +40,7 @@ type Options struct {
103 }
104 
105 func Qwen3Options() Options {
106-	return Options{Model: "jp", Temperature: 0.4, MaxTokens: 256, EnableThinking: false}
107+	return Options{Model: llmModel, Temperature: 1.0, MaxTokens: 256, EnableThinking: false}
108 }
109 
110 type Client struct {