Feature/DAIP2-2314 DAIP2 1687 hybrid * remove -files from s3 prefix requirements * Resolve input paths * fix: VendorProcessor.process_file returns (df, None) tuple runner.safe_process_file unpacks the result as (cc_df, dashboard_df), so returning a single DataFrame caused every vendor/generic file to fail with "too many values to unpack (expected 2)" — Python iterates DataFrame columns during unpacking. Vendor pipelines have no dashboard variant; second slot is None and the existing `dashboard_result is not None` guard in runner.py already handles it. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * DAIP2-2314 + DAIP2-1687: pad DYNAMIC_PRIMARY + DYNAMIC_PRIMARY_ENTITY_CLASSIFICATION over 1024-token cache floor - Pad DYNAMIC_PRIMARY_INSTRUCTION with three new sections: [SCOPE BOUNDARIES], [SOURCE TEXT INTERPRETATION], [REASONING DISCIPLINE], plus a [WORKED EXAMPLES] block. Estimated tokens: 447 -> 1117 (Sonnet 4.5 1024-min, +93 margin). All additions reinforce existing rules (sibling-field separation, alias mapping, pricing-vs-LOB distinction, contrastive-clause exclusion, exhibit-header binding) — no new directives that could bias extraction. - Pad DYNAMIC_PRIMARY_ENTITY_CLASSIFICATION_INSTRUCTION with a [FINAL CHECKLIST BEFORE OUTPUT] block. Estimated tokens: 956 -> 1101 (Sonnet 4.5 1024-min, +77 margin). Reinforces the existing 4-step anti-duplication protocol and JSON shape requirements. - Register both new entries in cache_registry: DYNAMIC_PRIMARY_ENTITY_CLASSIFICATION as INSTRUCTION_PLUS_CONTEXT (caches at warm-up), DYNAMIC_PRIMARY_ENTITIES as CONTEXT (instruction is intentionally short; CONTEXT c… * black format fix * Merged dev into feature/DAIP2-2314-DAIP2-1687-hybrid * fixed raw lob values in base lob field mapping and composite entities fix * black format fix * fixed LOB Program output issues * issue fixes * remove debugging code * Updated prompts * updated additional instructions * Update Program-->LOB * LLM-based AD-Program/Product mapping to LOB even when there is a crosswalk * black format fix * Merged dev into feature/DAIP2-2314-DAIP2-1687-hybrid * added logging in prompt call tracking * added updated logging in prompt call tracking * aaded min cache token per usage label * added cache registry for dynamic primary mapping prompt calls * reolved mapping prompts ambiguities * black format fix * Phase 2 modifications added * reverted phase 2 modifications Approved-by: Katon Minhas
7.9 KiB
Prompt Caching System Proposal
Branch: feature/DAIP2-2314-expand-caching-for-short-prompts
Goal
Ensure every stable instruction prompt that is intended to be cached is actually cacheable, warmed on the same model used at runtime, and observable in runtime logs.
Recommended Caching System
-
Introduce a central prompt cache registry.
Each cacheable prompt should be declared once with:
usage_label- instruction function
- model family used at runtime
- minimum cache token policy
- runtime call sites
- whether the prompt also uses
context_for_caching
This avoids the current split where cache warming is in
prompt_templates.get_cacheable_instructions()but runtime model choice is scattered through prompt call sites. -
Enforce the 1024-token minimum before runtime.
Prompt caching should not be treated as enabled just because
cache=Trueis passed. At startup, the registry should validate every cached instruction with the same tokenizer/estimate used by reporting. If a prompt is below the minimum, it should either:- fail fast in a cache audit mode, or
- automatically append approved static cache padding/instruction context until it crosses the threshold.
The padding must be stable and domain-relevant, not dynamic request data.
-
Warm caches per model, not globally.
Cache entries are model-specific. The runner currently warms all instructions with
sonnet_latest, butSPLIT_SERVICE_TERMruns withhaiku_latest. The cache warming layer should warm(usage_label, resolved_model_id)pairs derived from the registry. -
Only mark supported models as cacheable.
_supports_prompt_cache()currently excludes Haiku. If Haiku prompt caching is not supported in this Bedrock setup, Haiku prompts should not be registered as cached. Either move those prompt calls tosonnet_latestwhen caching matters, or keep them uncached and report them as intentionally uncached. -
Keep static instruction caching separate from dynamic context caching.
Instruction prompts can be cached predictably.
context_for_cachingis different because it changes by contract, exhibit, or page. It should be tracked as a separate cache class:instruction_cache: expected to read after warm-upcontext_cache: expected to create/read per repeated context
This prevents context variability from making instruction caching look broken.
-
Add a cache contract test.
A local test should inspect the built request body for every registered prompt and assert:
- the instruction block has
cache_control - the resolved model supports caching
- the static instruction token estimate is at least 1024
- warm-up model equals runtime model
- the instruction block has
-
Upgrade runtime reporting.
Add these fields to prompt-call logs:
resolved_model_idcache_policy:instruction,context,instruction+context,noneinstruction_cache_eligiblecontext_cache_eligiblecache_miss_reason
This makes root cause visible without manual investigation.
Prompt Application Report
Apply Instruction Caching, Already Over 1024 Tokens
These should remain registered and warmed. They are good candidates for strict cache contract tests.
| Prompt | Estimated Instruction Tokens | Action |
|---|---|---|
OUTLIER_BREAKOUT |
2483 | Keep cached |
METHODOLOGY_BREAKOUT |
1927 | Keep cached |
SPLIT_SERVICE_TERM |
1600 | Cache only if runtime model supports caching; currently runs on Haiku |
DYNAMIC_ASSIGNMENT |
1578 | Keep cached |
CARVEOUT_CHECK |
1498 | Keep cached |
SERVICE_ENRICHMENT |
1459 | Keep cached |
REIMBURSEMENT_PRIMARY |
1390 | Keep cached |
EXHIBIT_LEVEL |
1108 | Keep cached |
Apply Instruction Caching After Padding/Expansion
These are wired or partially wired but below the 1024-token minimum. They should be padded or expanded with stable domain guidance before being considered cached.
| Prompt | Estimated Instruction Tokens | Tokens Needed | Action |
|---|---|---|---|
fill_bill_type |
968 | 56 | Add small static guidance |
DYNAMIC_CODE_ASSIGNMENT |
966 | 58 | Add small static guidance |
CHECK_PROVIDER_NAME_MATCH |
957 | 67 | Add small static guidance |
GROUPER_BREAKOUT |
939 | 85 | Add small static guidance |
EXHIBIT_HEADER |
922 | 102 | Add static header examples/rules |
LESSER_OF_DISTRIBUTION |
914 | 110 | Add static decision examples |
LESSER_OF_CHECK |
900 | 124 | Add static classification examples |
CODE_EXPLICIT |
871 | 153 | Add static code-system rules |
REIMB_DATES_ASSIGNMENT |
803 | 221 | Add date extraction examples |
DYNAMIC_PRIMARY |
441 | 583 | Expand instruction substantially |
SPLIT_REIMB_DATES |
410 | 614 | Expand instruction substantially |
prompt_lob_relationship |
408 | 616 | Move field prompt into instruction or add static LOB rules |
code_implicit_arbitration |
391 | 633 | Expand arbitration rules |
AARETE_DERIVED_PAYER_NAME |
372 | 652 | Expand payer naming rules |
FEE_SCHEDULE_BREAKOUT |
352 | 672 | Expand fee schedule rules |
code_implicit_special |
329 | 695 | Expand special code rules |
DATE_FIX |
322 | 702 | Expand date normalization rules |
SPECIAL_CASE_ASSIGNMENT |
251 | 773 | Expand special case rules |
DERIVED_TERM_DATE |
247 | 777 | Expand term date rules |
code_last_check |
179 | 845 | Expand or leave uncached if low value |
EXTRACT_AMENDMENT_NUM_FROM_FILENAME |
152 | 872 | Expand or leave uncached if low value |
EXHIBIT_HEADER_DEDUP |
35 | 989 | Do not cache unless redesigned |
Wire for Caching or Mark Intentionally Uncached
These showed runtime volume but no instruction function match in the audit. They need explicit registry decisions.
| Prompt | Runtime Finding | Action |
|---|---|---|
prompt_hsc_single_field |
109 calls, cache_enabled=False, large dynamic prompt |
Split stable HSC instructions from dynamic chunk and cache the instruction |
SPECIAL_CASE_BREAKOUT |
6 calls, no instruction cache | Wire only if repeated volume justifies expansion |
code_implicit_rag |
Cached in API, audit alias missing | Add registry alias to the actual instruction function |
prompt_provider_info |
Cached in API, audit alias missing | Add registry alias to TIN_NPI_TEMPLATE_INSTRUCTION or rename usage label |
validate_reimbursements_for_llm |
Cached in API, audit alias missing | Add registry alias to VALIDATE_REIMBURSEMENTS_INSTRUCTION |
Immediate Fixes
-
Change
SPLIT_SERVICE_TERMto use a cache-supported model or mark it intentionally uncached. Current behavior is misleading:cache=Trueis passed, but no Bedrock cache tokens are recorded. -
Add a registry alias map for usage labels that do not match instruction names:
code_implicit_rag->CODE_IMPLICITprompt_provider_info->TIN_NPI_TEMPLATEvalidate_reimbursements_for_llm->VALIDATE_REIMBURSEMENTSprompt_lob_relationship-> its real instruction function
-
Expand the near-threshold instructions first because they need the least work and have current evidence of API cache reads:
fill_bill_typeDYNAMIC_CODE_ASSIGNMENTCHECK_PROVIDER_NAME_MATCHGROUPER_BREAKOUTEXHIBIT_HEADERLESSER_OF_DISTRIBUTIONLESSER_OF_CHECKCODE_EXPLICIT
-
Redesign
prompt_hsc_single_fieldinto(instruction, prompt)form. It is high-volume and currently sends all prompt text dynamically.
Expected Behavior After Fix
For instruction-only prompts, after warm-up the API should report cache reads on normal calls. The cached percentage will still not be 100% of total billed input because the dynamic user prompt remains uncached. The correct target is:
cache_creation_tokensonly during warm-up or first uncached usecache_read_tokens > 0on every eligible runtime callcache_miss_reasonempty for registered instruction cache prompts
For context-cached prompts, 100% cache reads should not be expected unless the same context block repeats. The report should evaluate those separately.