TheEdgeOfRage/pi-llama

Default ref refs/heads/main

pi-llama

llama.cpp Pi extension. Auto-discovers models from a running llama.cpp server and registers them as the llama-cpp provider in pi.

Install

From the shell:

pi install git:github.com/huggingface/pi-llama

This clones to ~/.pi/agent/packages/pi-llama/ and adds an entry to your pi settings. Every future pi invocation auto-loads it.

From inside an interactive pi session:

!pi install git:github.com/huggingface/pi-llama

Then run /reload (or restart pi) to load the extension.

Dev mode:

git clone https://github.com/huggingface/pi-llama ~/code/pi-llama
pi -e ~/code/pi-llama/index.ts

-e loads the extension only for the current session, useful while developing.

Configuration

The extension does not autoload a model when you select it by default. It does autoload the selected model when you send a message.

Use ~/.pi/agent/extensions/pi-llama.json for global settings or .pi/pi-llama.json for project settings:

{
  "autoloadOnSelect": false,
  "baseUrl": "http://localhost:8080/v1"
}

Project settings override global settings. Set autoloadOnSelect to true to restore loading on model selection. LLAMA_BASE_URL overrides the configured baseUrl. Restart pi or run /reload after changing configuration.

Environment Variables

This extension supports the following environment variables:

  • LLAMA_BASE_URL (Default: http://localhost:8080/v1)
  • LLAMA_API_KEY (Default: no-key)

Usage

# 1. Install llama.cpp
curl -LsSf https://llama.app/install.sh | bash

# 2. Start the server
llama serve

# Optional: Use a remote instance of llama.cpp server instead of local
export LLAMA_BASE_URL="https://llama.example.com/v1"

# 3. Launch pi in another terminal
pi

# 4. Inside pi - search "llama-cpp" to browse your local models
/model