7.9 KiB
SDG Bridge Plan
Bridge Script Filename
scripts/generate_training_data.py
This script reads the generated dictionary artifacts from this project, optionally runs a
NeMo Data Designer expansion pass on Mimir via SSH, and writes a timestamped raw training
dataset into the workspace_Data pipeline, then registers it in manifest.llm.json so the
Training Monitor picks it up automatically.
Mimir — Data Designer Status
Mimir (100.80.52.47) has Data Designer 0.8.0 installed and ready:
| Component | Location | Status |
|---|---|---|
| CLI | /home/aaron-pressey/.venvs/data-designer/bin/data-designer |
✅ installed |
| Backend adapter | /home/aaron-pressey/.local/share/home-grown-llm-data/nemo_data_designer_backend.py |
✅ present |
| Storage | /mnt/storage/data-designer/ |
✅ mounted |
| Model configs | /mnt/storage/data-designer/model_configs.yaml |
✅ configured |
| Config env | /home/aaron-pressey/.config/home-grown-llm-data/data-designer.env |
✅ present |
Current model provider: NVIDIA API (nvidia-text alias → nvidia/nemotron-3-nano-30b-a3b).
Local llama.cpp on port 8081 is configured but not currently running — cloud provider is used
for generation unless a local inference server is started separately.
What the Bridge Does
Step 1 (local — this project):
dictionary/actions.index.json ─┐
dictionary/model.codebook.json ├─► generate_training_data.py ─► seed JSONL
dictionary/model.examples.jsonl ─┘ (20 records/action, 260 total)
│
Step 2 (remote — Mimir, optional): │
SCP seed JSONL + DD config to Mimir ◄────────┘
SSH: data-designer create module_controller.yaml
--num-records <N>
--dataset-name module_controller_intents_<timestamp>
SCP expanded dataset back
│
Step 3 (local — workspace_Data): │
Write to data/module_controller_intents_<timestamp>/ ◄──────────────┘
Patch manifest.llm.json with timestamped entry
Training Monitor picks it up as a new Data Source
Timestamp Tagging
Every dataset generated by this script is tagged with a Unix timestamp at generation time. The timestamp is embedded in:
- Output directory name:
data/module_controller_intents_<unix_ts>/ - Manifest entry key:
"module_controller_intents_<unix_ts>" - Manifest entry field:
"generated_at": <unix_ts> - Every JSONL record's metadata:
"generated_at": <unix_ts>
This means you can run the bridge multiple times (after adding actions, changing examples, or expanding via NeMo) and the Training Monitor will show each run as a distinct, traceable Data Source. No clobbering, no ambiguity.
Output Record Schema
Each output record uses the instruction schema profile:
{
"instruction": "Map the following natural language request to the correct action and payload.",
"input": "hash hello with sha256",
"output": "{\"action_id\": \"hash_string\", \"payload\": {\"text\": \"hello\", \"algorithm\": \"sha256\"}}",
"metadata": {
"action_id": "hash_string",
"compact": "{\"v\":\"1.0.0\",\"c\":\"<codebook_checksum>\",\"a\":\"A007\",\"p\":{\"F010\":\"hello\",\"F001\":\"sha256\"}}",
"source": "seed_intent_example",
"split": "train",
"generated_at": 1784660000
}
}
Generation Strategy
For each action in actions.index.json:
- Seed examples — emit each
intent_exampledirectly (source:seed_intent_example) - Alias variants — one phrase per alias using the alias as the verb (source:
alias_variant) - Payload field variants — for enum fields, one record per enum value; for numeric fields,
low/mid/high values (source:
field_variant) - Template expansion — 5 records per action from
synthetic_templatewith filled placeholders (source:template_expansion)
Minimum seed target: ~20 records × 13 actions = 260 records.
After NeMo expansion: configurable via --num-records (default: 1000).
Split: 90% train / 10% validation, partitioned by action so all 13 actions appear in both splits.
Mimir Data Designer Config (auto-generated by bridge)
The bridge writes this config to Mimir before running generation:
# auto-generated by generate_training_data.py — do not edit by hand
# generated_at: <unix_ts>
model_config_path: /mnt/storage/data-designer/model_configs.yaml
model_providers_path: /mnt/storage/data-designer/model_providers.yaml
dataset:
name: module_controller_intents_<unix_ts>
schema_profile: instruction
seed_file: /mnt/storage/data-designer/managed-assets/module_controller_seed_<unix_ts>.jsonl
columns:
- name: input
type: seed_passthrough
- name: output
type: llm_text
model_alias: nvidia-text
prompt: |
You are generating training data for a local action dispatcher.
Given the natural language request below, produce valid JSON with
exactly two fields: "action_id" (string) and "payload" (object).
The action_id must be one of: {action_ids}.
Vary the phrasing of the input naturally but keep the output schema strict.
Request: {{input}}
output_schema:
type: object
required: [action_id, payload]
additionalProperties: false
properties:
action_id:
type: string
enum: {action_ids}
payload:
type: object
Manifest Entry (auto-patched by bridge)
"module_controller_intents_<unix_ts>": {
"name": "Module Controller Intents",
"category": "synthetic_sft",
"schema_profile": "instruction",
"source_url": null,
"acquisition_method": "local_synthetic_generator",
"license_spdx": "UNLICENSED",
"redistribution_allowed": false,
"estimated_tokens": null,
"status": "generated_local",
"local_path": "data/module_controller_intents_<unix_ts>",
"local_format": "jsonl",
"generator": "mp-ai-module-controller/scripts/generate_training_data.py",
"sdg_engine": "data-designer",
"sdg_host": "mimir (100.80.52.47)",
"sdg_model": "nvidia/nemotron-3-nano-30b-a3b",
"codebook_version": "1.0.0",
"codebook_checksum": "<from model.codebook.json>",
"registry_checksum": "<from actions.index.json>",
"generated_at": <unix_ts>,
"doc": "data/module_controller_intents_<unix_ts>/README.md"
}
CLI Usage
# From mp-ai-module-controller project root:
# Seed only (no NeMo expansion — fast, local, 260 records):
python3 scripts/generate_training_data.py \
--dictionary-dir dictionary/ \
--output-dir ../workspace_Data/data/ \
--manifest ../workspace_Data/manifest.llm.json
# Full SDG run via Mimir (requires NVIDIA_API_KEY on Mimir):
python3 scripts/generate_training_data.py \
--dictionary-dir dictionary/ \
--output-dir ../workspace_Data/data/ \
--manifest ../workspace_Data/manifest.llm.json \
--sdg-host 100.80.52.47 \
--sdg-user aaron-pressey \
--sdg-key ~/.ssh/silma_orson_ed25519 \
--num-records 1000
# Dry run (prints what would be generated, writes nothing):
python3 scripts/generate_training_data.py --dry-run
Prerequisites
- Python 3.11+ (stdlib only:
json,pathlib,hashlib,time,argparse,subprocess) - SSH access to Mimir via
~/.ssh/silma_orson_ed25519(for NeMo expansion pass) NVIDIA_API_KEYset in Mimir's environment (for cloud Nemotron generation)dictionary/files current — runcargo run -- dictionary generatefirst if registry changed
Staleness Check
The manifest entry embeds both codebook_checksum and generated_at. If you add actions:
cargo run -- dictionary generate # updates codebook_checksum
python3 scripts/generate_training_data.py # new timestamp → new manifest entry
The Training Monitor will show both the old and new dataset. You can deprecate the old one
by setting its status to "superseded" in the manifest.