# AGENTS.md ## Project `mp-ai-module-controller` is a Rust CLI scaffold for local llama.cpp-driven action dispatch. It starts a local `llama-server`, generates verbose and compact action dictionaries for small-model training/runtime lookup, queries the server through its OpenAI-compatible chat endpoint, and dispatches exact-match Rust scripts. ## Commands - Build: `cargo build` - Test: `cargo test` - Generate dictionary: `cargo run -- dictionary generate` - Start llama.cpp: `cargo run -- serve` - Query local model: `cargo run -- query "log hello from codex"` - Direct proof-of-concept dispatch: `cargo run -- run logger --payload '{"message":"hello from codex"}'` - Direct math dispatch: `cargo run -- run calculate --payload '{"expression":"2 + 2 * 3"}'` - Direct text dispatch: `cargo run -- run to_slug --payload '{"text":"Hello, World!"}'` - Direct file dispatch: `cargo run -- run write_note --payload '{"filename":"hello.txt","content":"hello"}'` - Codex setup: `scripts/codex-setup.sh` - Codex logger smoke test: `scripts/codex-run-logger.sh` - Precommit/CI check: `scripts/precommit-check.sh` - Codex environment: `.codex/environments/environment.toml` - Training data dry run: `python3 scripts/generate_training_data.py --dry-run` - Local training data bridge: `python3 scripts/generate_training_data.py` - Deploy training data to Zoe: `scripts/deploy_training_data.sh` - Deploy training data with Mimir SDG: `ENABLE_MIMIR_SDG=1 scripts/deploy_training_data.sh` ## Environment Configuration is read from environment variables: - `LLAMA_MODEL_PATH`: required GGUF model path for `cargo run -- serve`. - `LLAMA_HOST`: llama.cpp host. Default: `127.0.0.1`. - `LLAMA_PORT`: llama.cpp port. Default: `8080`. - `LLAMA_CONTEXT_SIZE`: optional value passed as `--ctx-size`. - `LLAMA_EXTRA_ARGS`: optional shell-split arguments appended to `llama-server`. - `CONTROLLER_LOG_LEVEL`: tracing log level. Default: `info`. - `ENABLE_MIMIR_SDG`: set to `1` to include Mimir/Data Designer expansion in training-data deploys. - `MIMIR_SDG_HOST`, `MIMIR_SDG_USER`, `MIMIR_SDG_KEY`, `MIMIR_SDG_MODEL_ALIAS`, `MIMIR_SDG_NUM_RECORDS`: optional SDG deploy overrides. - `NVIDIA_API_KEY`: optional local Data Designer provider credential; keep only in ignored `.env` files or Mimir's Data Designer env file. Use `.env` for local values and keep it out of git. Update `.env.example`, `README.md`, `manifest.llm.json`, and `llm.txt` when environment variables or commands change. ## Dictionary Rules - `src/registry.rs` is the source of truth for executable action definitions. - `dictionary/static-base.json` is the tracked static compact-code base. Change it only when the compact protocol changes. - `dictionary/actions.jsonl` is generated canonical verbose training/action data. - `dictionary/actions.index.json` is the generated verbose runtime lookup file. - `dictionary/model.codebook.json` is the generated model-facing compact codebook. - `dictionary/model.examples.jsonl` is generated template/example data for future synthetic-data tooling. - Registered action suites are Math & Conversion, Text Utilities, File Operations, and Logger. - Regenerate dictionary files with `cargo run -- dictionary generate` after changing registered actions. - Do not manually edit generated dictionary files. - Generated metadata includes `dictionary_version`, `registry_checksum`, `static_base_checksum`, and `codebook_checksum`. ## SDG Bridge Rules - `SDG_BRIDGE_PLAN.md` documents the local-to-Data-Designer bridge design. - `scripts/generate_training_data.py` reads generated dictionary artifacts and produces instruction-format seed datasets. - Default output is `/Volumes/Zoe/custom-local-llm/training-data/module_controller_intents_/`. - Default manifest patch target is `/Volumes/Zoe/custom-local-llm/manifest.llm.json`. - `/Volumes/Zoe/custom-local-llm` has its own `README.md`, `AGENTS.md`, `manifest.llm.json`, and `llm.txt`. - Use `--dry-run` for validation; it must not write output data or patch manifests. - Mimir/Data Designer expansion only runs when `--sdg-host`, `--sdg-user`, and `--sdg-key` are all provided. - `scripts/deploy_training_data.sh` enables Mimir expansion when `ENABLE_MIMIR_SDG=1`; Woodpecker uses this mode on main-branch pushes. - Mimir expansion defaults to `MIMIR_SDG_MODEL_ALIAS=nvidia-text`; set `MIMIR_SDG_MODEL_ALIAS=openrouter-text` only if the direct NVIDIA provider needs a fallback. - Mimir expansion sources `/home/aaron-pressey/.config/home-grown-llm-data/data-designer.env`, writes `expanded.raw.jsonl`, and normalizes usable instruction rows into `expanded.jsonl`. - Do not start or restart Mimir llama.cpp/Keiro from this bridge. - If Mimir expansion fails, the bridge keeps the local seed dataset, writes `SDG_ERROR.txt`, patches the manifest with `generated_local_remote_failed`, and exits non-zero. - Seed-only `train.jsonl` and `validation.jsonl` are routing smoke-test quality; use `expanded.jsonl` from Mimir for serious training-data enrichment. - Woodpecker deploys training data on main-branch pushes with `ENABLE_MIMIR_SDG=1 scripts/deploy_training_data.sh`. ## Dispatch Rules - Dispatch only exact `action_id` values present in `dictionary/actions.index.json`. - Compact model outputs must include matching `v` and `c` values from `dictionary/model.codebook.json` before expansion. - Aliases and intent examples are training hints only; they are not executable ids. - Unknown, malformed, ambiguous, missing, or empty action ids must be rejected without running scripts. - Script processes are detached. The controller spawns them and does not monitor completion. - Keep script payload validation strict and local to the script module. - File operation scripts must remain confined to project-local `notes/` paths and reject absolute paths or path traversal. ## Generated Paths - `target/` - `logs/` - `runtime/` - `notes/` - `dictionary/actions.jsonl` - `dictionary/actions.index.json` - `dictionary/model.codebook.json` - `dictionary/model.examples.jsonl` ## CI And Hooks - Codex environment actions live at `.codex/environments/environment.toml`. - Gitea Actions workflow lives at `.gitea/workflows/ci.yml`. - Woodpecker workflow lives at `.woodpecker.yml`. - The committed pre-commit hook lives at `.githooks/pre-commit`. - Configure local hooks with `git config core.hooksPath .githooks`. - All CI and pre-commit checks should call `scripts/precommit-check.sh` so full dictionary generation stays validated consistently. ## Coding Guidelines - Keep the CLI Rust-first and small. - Prefer adding scripts under `src/scripts/` and registering them in `src/registry.rs`. - Add focused tests for parsing, dictionary generation, action lookup, and payload validation. - Add stdlib Python tests for bridge behavior in `tests/` when changing `scripts/generate_training_data.py`. - Run `cargo test` before handing off code changes.