Files
mp-ai-module-controller/SDG_BRIDGE_PLAN.md
T
2026-07-21 13:14:52 -05:00

233 lines
7.9 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# SDG Bridge Plan
## Bridge Script Filename
```
scripts/generate_training_data.py
```
This script reads the generated dictionary artifacts from this project, optionally runs a
NeMo Data Designer expansion pass on Mimir via SSH, and writes a timestamped raw training
dataset into the `workspace_Data` pipeline, then registers it in `manifest.llm.json` so the
Training Monitor picks it up automatically.
---
## Mimir — Data Designer Status
Mimir (`100.80.52.47`) has **Data Designer 0.8.0** installed and ready:
| Component | Location | Status |
|-----------|----------|--------|
| CLI | `/home/aaron-pressey/.venvs/data-designer/bin/data-designer` | ✅ installed |
| Backend adapter | `/home/aaron-pressey/.local/share/home-grown-llm-data/nemo_data_designer_backend.py` | ✅ present |
| Storage | `/mnt/storage/data-designer/` | ✅ mounted |
| Model configs | `/mnt/storage/data-designer/model_configs.yaml` | ✅ configured |
| Config env | `/home/aaron-pressey/.config/home-grown-llm-data/data-designer.env` | ✅ present |
**Current model provider:** NVIDIA API (`nvidia-text` alias → `nvidia/nemotron-3-nano-30b-a3b`).
Local llama.cpp on port 8081 is configured but not currently running — cloud provider is used
for generation unless a local inference server is started separately.
---
## What the Bridge Does
```
Step 1 (local — this project):
dictionary/actions.index.json ─┐
dictionary/model.codebook.json ├─► generate_training_data.py ─► seed JSONL
dictionary/model.examples.jsonl ─┘ (20 records/action, 260 total)
Step 2 (remote — Mimir, optional): │
SCP seed JSONL + DD config to Mimir ◄────────┘
SSH: data-designer create module_controller.yaml
--num-records <N>
--dataset-name module_controller_intents_<timestamp>
SCP expanded dataset back
Step 3 (local — workspace_Data): │
Write to data/module_controller_intents_<timestamp>/ ◄──────────────┘
Patch manifest.llm.json with timestamped entry
Training Monitor picks it up as a new Data Source
```
---
## Timestamp Tagging
Every dataset generated by this script is tagged with a **Unix timestamp at generation time**.
The timestamp is embedded in:
1. **Output directory name:** `data/module_controller_intents_<unix_ts>/`
2. **Manifest entry key:** `"module_controller_intents_<unix_ts>"`
3. **Manifest entry field:** `"generated_at": <unix_ts>`
4. **Every JSONL record's metadata:** `"generated_at": <unix_ts>`
This means you can run the bridge multiple times (after adding actions, changing examples,
or expanding via NeMo) and the Training Monitor will show each run as a distinct, traceable
Data Source. No clobbering, no ambiguity.
---
## Output Record Schema
Each output record uses the `instruction` schema profile:
```json
{
"instruction": "Map the following natural language request to the correct action and payload.",
"input": "hash hello with sha256",
"output": "{\"action_id\": \"hash_string\", \"payload\": {\"text\": \"hello\", \"algorithm\": \"sha256\"}}",
"metadata": {
"action_id": "hash_string",
"compact": "{\"v\":\"1.0.0\",\"c\":\"<codebook_checksum>\",\"a\":\"A007\",\"p\":{\"F010\":\"hello\",\"F001\":\"sha256\"}}",
"source": "seed_intent_example",
"split": "train",
"generated_at": 1784660000
}
}
```
---
## Generation Strategy
For each action in `actions.index.json`:
1. **Seed examples** — emit each `intent_example` directly (source: `seed_intent_example`)
2. **Alias variants** — one phrase per alias using the alias as the verb (source: `alias_variant`)
3. **Payload field variants** — for enum fields, one record per enum value; for numeric fields,
low/mid/high values (source: `field_variant`)
4. **Template expansion** — 5 records per action from `synthetic_template` with filled placeholders
(source: `template_expansion`)
**Minimum seed target:** ~20 records × 13 actions = 260 records.
**After NeMo expansion:** configurable via `--num-records` (default: 1000).
Split: 90% train / 10% validation, partitioned by action so all 13 actions appear in both splits.
---
## Mimir Data Designer Config (auto-generated by bridge)
The bridge writes this config to Mimir before running generation:
```yaml
# auto-generated by generate_training_data.py — do not edit by hand
# generated_at: <unix_ts>
model_config_path: /mnt/storage/data-designer/model_configs.yaml
model_providers_path: /mnt/storage/data-designer/model_providers.yaml
dataset:
name: module_controller_intents_<unix_ts>
schema_profile: instruction
seed_file: /mnt/storage/data-designer/managed-assets/module_controller_seed_<unix_ts>.jsonl
columns:
- name: input
type: seed_passthrough
- name: output
type: llm_text
model_alias: nvidia-text
prompt: |
You are generating training data for a local action dispatcher.
Given the natural language request below, produce valid JSON with
exactly two fields: "action_id" (string) and "payload" (object).
The action_id must be one of: {action_ids}.
Vary the phrasing of the input naturally but keep the output schema strict.
Request: {{input}}
output_schema:
type: object
required: [action_id, payload]
additionalProperties: false
properties:
action_id:
type: string
enum: {action_ids}
payload:
type: object
```
---
## Manifest Entry (auto-patched by bridge)
```json
"module_controller_intents_<unix_ts>": {
"name": "Module Controller Intents",
"category": "synthetic_sft",
"schema_profile": "instruction",
"source_url": null,
"acquisition_method": "local_synthetic_generator",
"license_spdx": "UNLICENSED",
"redistribution_allowed": false,
"estimated_tokens": null,
"status": "generated_local",
"local_path": "data/module_controller_intents_<unix_ts>",
"local_format": "jsonl",
"generator": "mp-ai-module-controller/scripts/generate_training_data.py",
"sdg_engine": "data-designer",
"sdg_host": "mimir (100.80.52.47)",
"sdg_model": "nvidia/nemotron-3-nano-30b-a3b",
"codebook_version": "1.0.0",
"codebook_checksum": "<from model.codebook.json>",
"registry_checksum": "<from actions.index.json>",
"generated_at": <unix_ts>,
"doc": "data/module_controller_intents_<unix_ts>/README.md"
}
```
---
## CLI Usage
```bash
# From mp-ai-module-controller project root:
# Seed only (no NeMo expansion — fast, local, 260 records):
python3 scripts/generate_training_data.py \
--dictionary-dir dictionary/ \
--output-dir ../workspace_Data/data/ \
--manifest ../workspace_Data/manifest.llm.json
# Full SDG run via Mimir (requires NVIDIA_API_KEY on Mimir):
python3 scripts/generate_training_data.py \
--dictionary-dir dictionary/ \
--output-dir ../workspace_Data/data/ \
--manifest ../workspace_Data/manifest.llm.json \
--sdg-host 100.80.52.47 \
--sdg-user aaron-pressey \
--sdg-key ~/.ssh/silma_orson_ed25519 \
--num-records 1000
# Dry run (prints what would be generated, writes nothing):
python3 scripts/generate_training_data.py --dry-run
```
---
## Prerequisites
- Python 3.11+ (stdlib only: `json`, `pathlib`, `hashlib`, `time`, `argparse`, `subprocess`)
- SSH access to Mimir via `~/.ssh/silma_orson_ed25519` (for NeMo expansion pass)
- `NVIDIA_API_KEY` set in Mimir's environment (for cloud Nemotron generation)
- `dictionary/` files current — run `cargo run -- dictionary generate` first if registry changed
---
## Staleness Check
The manifest entry embeds both `codebook_checksum` and `generated_at`. If you add actions:
```bash
cargo run -- dictionary generate # updates codebook_checksum
python3 scripts/generate_training_data.py # new timestamp → new manifest entry
```
The Training Monitor will show both the old and new dataset. You can deprecate the old one
by setting its `status` to `"superseded"` in the manifest.