README.md
pi-llama
llama.cpp Pi extension. Auto-discovers models from a running llama.cpp server and
registers them as the llama-cpp provider in pi.
Install
From the shell:
pi install git:github.com/huggingface/pi-llama
This clones to ~/.pi/agent/packages/pi-llama/ and adds an entry to your pi
settings. Every future pi invocation auto-loads it.
From inside an interactive pi session:
!pi install git:github.com/huggingface/pi-llama
Then run /reload (or restart pi) to load the extension.
Dev mode:
git clone https://github.com/huggingface/pi-llama ~/code/pi-llama
pi -e ~/code/pi-llama/index.ts
-e loads the extension only for the current session, useful while
developing.
Configuration
The extension does not autoload a model when you select it by default. It does autoload the selected model when you send a message.
Use ~/.pi/agent/extensions/pi-llama.json for global settings or
.pi/pi-llama.json for project settings:
{
"autoloadOnSelect": false,
"baseUrl": "http://localhost:8080/v1"
}
Project settings override global settings. Set autoloadOnSelect to true to
restore loading on model selection. LLAMA_BASE_URL overrides the configured
baseUrl. Restart pi or run /reload after changing configuration.
Environment Variables
This extension supports the following environment variables:
- LLAMA_BASE_URL (Default:
http://localhost:8080/v1) - LLAMA_API_KEY (Default:
no-key)
Usage
# 1. Install llama.cpp
curl -LsSf https://llama.app/install.sh | bash
# 2. Start the server
llama serve
# Optional: Use a remote instance of llama.cpp server instead of local
export LLAMA_BASE_URL="https://llama.example.com/v1"
# 3. Launch pi in another terminal
pi
# 4. Inside pi - search "llama-cpp" to browse your local models
/model