# Local LLM Guide Two LLM models run locally on Tabitha via llama.cpp with Metal (Apple Silicon GPU) acceleration. Both expose OpenAI-compatible REST APIs — drop-in replacements for `api.openai.com` endpoints. **llama.cpp GitHub:** https://github.com/ggml-org/llama.cpp **llama.cpp API docs:** https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md --- ## Running the Servers ```bash cd /Users/Tabitha/.openclaw npm run api:llm ``` This starts both models in the background. Logs go to stdout/stderr. --- ## Models ### Gemma-4-E4B — Fast / Light | Field | Value | |-------|-------| | **Model** | Google Gemma 4 E4B (4-billion parameter, efficient variant) | | **Endpoint** | `http://localhost:8080` | | **API** | OpenAI-compatible (`/v1/chat/completions`, `/v1/completions`, `/v1/models`) | | **Use case** | Quick tasks, code generation, summarization — fast response | | **Acceleration** | Apple Metal (GPU) | ### Gemma-4-26B — Capable / Reasoning | Field | Value | |-------|-------| | **Model** | Google Gemma 4 27B | | **Endpoint** | `http://localhost:8081` | | **API** | OpenAI-compatible | | **Use case** | Complex reasoning, architecture questions, longer context tasks | | **Acceleration** | Apple Metal (GPU) | --- ## API Usage Both models speak the OpenAI API format: ```bash # Chat completion — Gemma-4-E4B (fast) curl -s http://localhost:8080/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ "model": "gemma-4-e4b", "messages": [{"role": "user", "content": "Explain this code: ..."}] }' | python3 -m json.tool # Chat completion — Gemma-4-26B (capable) curl -s http://localhost:8081/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ "model": "gemma-4-26b", "messages": [{"role": "user", "content": "Design a REST API for..."}] }' | python3 -m json.tool # List available models curl -s http://localhost:8080/v1/models curl -s http://localhost:8081/v1/models ``` --- ## Python (openai SDK) ```python from openai import OpenAI # Fast model fast = OpenAI(base_url="http://localhost:8080/v1", api_key="local") response = fast.chat.completions.create( model="gemma-4-e4b", messages=[{"role": "user", "content": "Write a Python function to..."}] ) print(response.choices[0].message.content) # Capable model smart = OpenAI(base_url="http://localhost:8081/v1", api_key="local") response = smart.chat.completions.create( model="gemma-4-26b", messages=[{"role": "user", "content": "Architect a microservice system for..."}], temperature=0.7, max_tokens=2048 ) print(response.choices[0].message.content) ``` --- ## Using with AI Coding Tools ### Claude Code CLI ```bash export ANTHROPIC_BASE_URL=http://localhost:8080/v1 export ANTHROPIC_API_KEY=local claude --model gemma-4-e4b ``` ### Codex CLI ```bash OPENAI_BASE_URL=http://localhost:8080/v1 \ OPENAI_API_KEY=local \ codex --model gemma-4-e4b "explain this codebase" ``` ### Cursor (IDE) Settings → Models → Add Model: - Base URL: `http://localhost:8080/v1` - API Key: `local` - Model name: `gemma-4-e4b` Repeat for port 8081 / `gemma-4-26b`. ### LangChain / any OpenAI-compatible client ```python from langchain_openai import ChatOpenAI llm = ChatOpenAI( base_url="http://localhost:8080/v1", api_key="local", model="gemma-4-e4b" ) ``` --- ## Health Check ```bash # Check if servers are responding curl -s http://localhost:8080/health && echo " — E4B OK" curl -s http://localhost:8081/health && echo " — 26B OK" # Or check model list curl -s http://localhost:8080/v1/models | python3 -c "import sys,json; [print(m['id']) for m in json.load(sys.stdin)['data']]" ``` --- ## Model Storage Models are stored on Theodora (hot SSD cache): ``` ~/.cache/huggingface → /Volumes/Theodora/models/huggingface (symlink) ~/.cache/llama.cpp → /Volumes/Theodora/models/llama-cpp (symlink) ``` Archive copies on Hagia: ``` /Volumes/Hagia/models/huggingface/ /Volumes/Hagia/models/llama-cpp/ ``` --- ## Notes for AI Agents - The servers run **on the host** (not in Docker/k8s) — they are accessible from any process on Tabitha - Theodora (SSD) must be mounted for fast model loading — `ls /Volumes/Theodora` to verify - If a server is not responding, run `cd /Users/Tabitha/.openclaw && npm run api:llm` to start it - **Gemma-4-E4B** is the default fast option — use it for code tasks, summarization, quick Q&A - **Gemma-4-26B** costs more compute — use it for complex reasoning, architecture, long context - Both are private and local — no data leaves the machine