4.5 KiB
4.5 KiB
Local LLM Guide
Two LLM models run locally on Tabitha via llama.cpp with Metal (Apple Silicon GPU) acceleration. Both expose OpenAI-compatible REST APIs — drop-in replacements for api.openai.com endpoints.
llama.cpp GitHub: https://github.com/ggml-org/llama.cpp llama.cpp API docs: https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md
Running the Servers
cd /Users/Tabitha/.openclaw
npm run api:llm
This starts both models in the background. Logs go to stdout/stderr.
Models
Gemma-4-E4B — Fast / Light
| Field | Value |
|---|---|
| Model | Google Gemma 4 E4B (4-billion parameter, efficient variant) |
| Endpoint | http://localhost:8080 |
| API | OpenAI-compatible (/v1/chat/completions, /v1/completions, /v1/models) |
| Use case | Quick tasks, code generation, summarization — fast response |
| Acceleration | Apple Metal (GPU) |
Gemma-4-26B — Capable / Reasoning
| Field | Value |
|---|---|
| Model | Google Gemma 4 27B |
| Endpoint | http://localhost:8081 |
| API | OpenAI-compatible |
| Use case | Complex reasoning, architecture questions, longer context tasks |
| Acceleration | Apple Metal (GPU) |
API Usage
Both models speak the OpenAI API format:
# Chat completion — Gemma-4-E4B (fast)
curl -s http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "gemma-4-e4b",
"messages": [{"role": "user", "content": "Explain this code: ..."}]
}' | python3 -m json.tool
# Chat completion — Gemma-4-26B (capable)
curl -s http://localhost:8081/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "gemma-4-26b",
"messages": [{"role": "user", "content": "Design a REST API for..."}]
}' | python3 -m json.tool
# List available models
curl -s http://localhost:8080/v1/models
curl -s http://localhost:8081/v1/models
Python (openai SDK)
from openai import OpenAI
# Fast model
fast = OpenAI(base_url="http://localhost:8080/v1", api_key="local")
response = fast.chat.completions.create(
model="gemma-4-e4b",
messages=[{"role": "user", "content": "Write a Python function to..."}]
)
print(response.choices[0].message.content)
# Capable model
smart = OpenAI(base_url="http://localhost:8081/v1", api_key="local")
response = smart.chat.completions.create(
model="gemma-4-26b",
messages=[{"role": "user", "content": "Architect a microservice system for..."}],
temperature=0.7,
max_tokens=2048
)
print(response.choices[0].message.content)
Using with AI Coding Tools
Claude Code CLI
export ANTHROPIC_BASE_URL=http://localhost:8080/v1
export ANTHROPIC_API_KEY=local
claude --model gemma-4-e4b
Codex CLI
OPENAI_BASE_URL=http://localhost:8080/v1 \
OPENAI_API_KEY=local \
codex --model gemma-4-e4b "explain this codebase"
Cursor (IDE)
Settings → Models → Add Model:
- Base URL:
http://localhost:8080/v1 - API Key:
local - Model name:
gemma-4-e4b
Repeat for port 8081 / gemma-4-26b.
LangChain / any OpenAI-compatible client
from langchain_openai import ChatOpenAI
llm = ChatOpenAI(
base_url="http://localhost:8080/v1",
api_key="local",
model="gemma-4-e4b"
)
Health Check
# Check if servers are responding
curl -s http://localhost:8080/health && echo " — E4B OK"
curl -s http://localhost:8081/health && echo " — 26B OK"
# Or check model list
curl -s http://localhost:8080/v1/models | python3 -c "import sys,json; [print(m['id']) for m in json.load(sys.stdin)['data']]"
Model Storage
Models are stored on Theodora (hot SSD cache):
~/.cache/huggingface → /Volumes/Theodora/models/huggingface (symlink)
~/.cache/llama.cpp → /Volumes/Theodora/models/llama-cpp (symlink)
Archive copies on Hagia:
/Volumes/Hagia/models/huggingface/
/Volumes/Hagia/models/llama-cpp/
Notes for AI Agents
- The servers run on the host (not in Docker/k8s) — they are accessible from any process on Tabitha
- Theodora (SSD) must be mounted for fast model loading —
ls /Volumes/Theodorato verify - If a server is not responding, run
cd /Users/Tabitha/.openclaw && npm run api:llmto start it - Gemma-4-E4B is the default fast option — use it for code tasks, summarization, quick Q&A
- Gemma-4-26B costs more compute — use it for complex reasoning, architecture, long context
- Both are private and local — no data leaves the machine