Merged in dev (pull request #1001) * Merged in feature/fixPlaceholder2 (pull request #979) Updated feature> dev gate to print variables of echo statements * Updated feature> dev gate to print variables of echo statement * Reverted placeholder changes Approved-by: Sujit Deokar * Merged in bugfix/prov_info_json_fixes (pull request #981) fix: fall back to PROVIDER_NAME in PROV_INFO_JSON when no TIN/NPI extracted * fix: fall back to PROVIDER_NAME in PROV_INFO_JSON when no TIN/NPI extracted When get_prov_info_json short-circuits due to no TIN/NPI regex matches, PROV_INFO_JSON was left as [] even when PROVIDER_NAME was successfully extracted via the one-to-one pipeline. This caused inconsistent output across contracts with the same provider — some files produced a NAME-only entry (via a false-positive regex hit triggering the LLM), others produced []. Reconcile at add_group_and_other, the first point where both extraction streams' results are available. When PROV_INFO_JSON is empty but PROVIDER_NAME is known, synthesize a NAME-only entry with IS_GROUP:"Y" and populate PROV_GROUP_NAME_FULL directly — skipping the provider_name_match_check LLM call since the match is tautological by construction. Adds 5 unit tests covering the str, list, already-populated, empty-name, and all-empty-lis… * Merged in feature/DAIP2-pacificsource-reimbursements-issues (pull request #982) Feature/DAIP2 pacificsource reimbursements issues * Tighten PREMIUM_TERM and DISCOUNT_TERM classifier prompts CARVEOUT_CHECK was misrouting table rate rows into special-case fields, dropping them from the reimbursement output: - "110% of CMS allowed" (base fee-schedule rates) was being classified as PREMIUM_TERM because the prompt treated "above 100% of reference" as an implicit premium. Seen on PacificSource Medicare_Attachment_A1 and A2 Facility contracts where Inpatient/Outpatient rows were missing (2556) or silently fell back to 100% fee-schedule (2557). - Per-service discount rates like "Progressive Lenses: 15% discount", "Contact Lenses: 2% discount", "Frame: 20% discount" were being classified as DISCOUNT_TERM because the prompt only required the word "discount" to appear. Seen on PacificSource Commercial_Attachment_A2 and A5 Professional contracts (2558, 2559). Prompts now require the literal keyword … * Merged in bugfix/filter-docusign-lines (pull request #983) remove docusign lines * remove docusign lines * add unit tests for clean_header_footer docusign/deleted_lines changes Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * black forrmatting Approved-by: Katon Minhas * Merged in feature/PC_logic_cleanup_output (pull request #984) Feature/PC logic cleanup output * Few tweaks PC_logics * Fixed orphan_ranking * black formatting * Changes on output field and ranking method * updated few hotfixes * black format fix Approved-by: Katon Minhas * Merged in hotfix/provider_name_group_fix (pull request #988) Hotfix/provider name group fix * Fixes done in GROUP column * black format fix Approved-by: Katon Minhas * Merged in bugfix/DAIP2-2524-carveout-code-optimization (pull request #980) CARVEOUT_CD issue fixed * CARVEOUT_CD issue fixed * pipeline error fixed * Merged dev into bugfix/DAIP2-2524-carveout-code-optimization * Merged dev into bugfix/DAIP2-2524-carveout-code-optimization * Merged dev into bugfix/DAIP2-2524-carveout-code-optimization * trigger cap issue fixed * trigger cap prompt updated * Merged dev into bugfix/DAIP2-2524-carveout-code-optimization Approved-by: Katon Minhas * Merged in bugfix/exhibit-smart-chunking-cost-improvements (pull request #985) Bugfix/exhibit smart chunking cost improvements * Add opt-in instrumentation for per-call token and row-count tracing Introduce src/utils/instrumentation.py (thread-safe CSV logger) and src/utils/instrumentation_context.py (ContextVar scope plus submit_with_context / map_with_context helpers for propagating context into ThreadPoolExecutor workers). Emit events at every Bedrock call in llm_utils.invoke_claude, including in-memory claude_cache hits, with full input/output/cache-read/cache-write token breakdown. Emit row-count events at each row-mutating stage in the one-to-N pipeline (clean_reimbursement_primary, filter_services_without_reimbursements, methodology_breakout, split_service_terms, carveout, dynamic_code_assignment, lesser_of_distribution, dynamic_assignment) and chunking / retrieval events in exhibit smart chunking (chunking_done, retrieval_done) plus exhibit lifecycle events (exhibit_start, exhibit_gate_skip, stage_… * Merged in hotfix/fileextension_issue (pull request #989) Hotfix/fileextension issue * fixed strip_ext issue * black format Approved-by: Katon Minhas * Merged in hotfix/exhibit-header-in-tables (pull request #987) Hotfix/exhibit header in tables * Merged in feature/FixplaceholderIssue (pull request #977) Remove curly braces from echo statements in dev->stg * Remove curly brances from echo statements in dev->stg * Removed curly braces in echo statements in feature-> dev gate Approved-by: Sujit Deokar * Merged in feature/standardized-services (pull request #958) Feature/standardized services * service term standardization * prompt update * prompt updates for standardization * only service standardization * new file * add supporting files and test scripts for standardization work Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * remove old files * Merge remote-tracking branch 'origin/dev' into feature/standardized-services * final fixes * Merged dev into feature/standardized-services * additional features * removed unwanted files * remove unwanted files * Merge branch 'dev' into feature/standardized-services * Merge remote-tr… * Merged in hotfix/postprocess_csv (pull request #990) Hotfix/postprocess csv * date issue_fix * black_format * Merged dev into hotfix/postprocess_csv Approved-by: Katon Minhas * Merged in bugfix/molina_ut_dynamic_primary (pull request #991) Bugfix/molina ut dynamic primary * Prompt changes for dynamic primary * Route cover-sheet-only files to ERRORS.csv instead of leaking phantom rows When every page of a contract was filtered out as a cover sheet / quick-review form, process_file silently returned a FILE_NAME-only DataFrame. Because the runner routes by checking for an "error" column, that file landed in RESULTS.csv as a near-empty row and no ERRORS.csv was generated for the run. - saas/file_processing.py: raise ValueError when text_dict is empty after cover-sheet filtering, so safe_process_file produces a proper error row. - runner.py: add _is_phantom_result defense-in-depth — promote any result with no extracted fields beyond FILE_NAME to error_results with error_type=PhantomSuccess. * Merged dev into bugfix/molina_ut_dynamic_primary * Tighten PRODUCT prompt: restrict to valid_values, prune LOB/PROGRAM examples * Merge branch 'bugfix/molina_ut_dynamic_primary' of bit… * Merged in bugfix/black-format (pull request #994) Black format for pipeline pass * Black format for pipeline pass * Merged in bugfix/DAIP2-2679-fix-nebraska-issues (pull request #995) incorrect inclusion of CPT4_PROC_CD fixed * incorrect inclusion of CPT4_PROC_CD fixed Approved-by: Katon Minhas * Merged in feature/DAIP2-2314-DAIP2-1687-hybrid (pull request #993) Feature/DAIP2-2314 DAIP2 1687 hybrid * remove -files from s3 prefix requirements * Resolve input paths * fix: VendorProcessor.process_file returns (df, None) tuple runner.safe_process_file unpacks the result as (cc_df, dashboard_df), so returning a single DataFrame caused every vendor/generic file to fail with "too many values to unpack (expected 2)" — Python iterates DataFrame columns during unpacking. Vendor pipelines have no dashboard variant; second slot is None and the existing `dashboard_result is not None` guard in runner.py already handles it. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * DAIP2-2314 + DAIP2-1687: pad DYNAMIC_PRIMARY + DYNAMIC_PRIMARY_ENTITY_CLASSIFICATION over 1024-token cache floor - Pad DYNAMIC_PRIMARY_INSTRUCTION with three new sections: [SCOPE BOUNDARIES], [SOURCE TEXT INTERPRETATION], [REASONING DISCIPLINE], plus a [WORKED EXAMPLES] block. Estimated tokens: 447 -> 1117 (Sonnet 4.5 … * Merged in bugfix/postprocessing_date_fix (pull request #996) Date formatting changes * Date formatting changes * Merged dev into bugfix/postprocessing_date_fix Approved-by: Katon Minhas * Merged in bugfix/parent-child-rank-orphan-uniqueness (pull request #999) PC_logic bugfix * PC_logic bugfix Approved-by: Katon Minhas * Merged in feature/document-index (pull request #1005) Feature/document index * Add Document Index preprocessing — Layers 1, 2, and 3 wiring Parse the Textract-emitted Document Index block at the top of each contract with a single cached LLM call (prompt_document_index) instead of one per-page call per page. Layer 2 verifies parsed entries via literal string match and structural regex sweep, escalating suspect pages back to the existing per-page path. Layer 1+2 failure triggers a full fallback to today's per-page flow. New symbols: - preprocessing_funcs.extract_document_index_block — regex slice of index prefix - preprocessing_funcs.verify_index_against_pages — structural verifier (plain dict return) - prompt_templates.DOCUMENT_INDEX_INSTRUCTION / DOCUMENT_INDEX — cached prompt pair - prompt_calls.prompt_document_index — LLM wrapper (usage_label DOCUMENT_INDEX_PARSE) - config: DOCUMENT_INDEX_PARSE_ENABLED and three threshold flags - instrumentation: DOCUMENT_INDEX_PARSE mapped to preprocessing segment one… * Merged in feature/active-rates (pull request #1004) Feature/active rates * initial commit * Merged in feature/FixplaceholderIssue (pull request #977) Remove curly braces from echo statements in dev->stg * Remove curly brances from echo statements in dev->stg * Removed curly braces in echo statements in feature-> dev gate Approved-by: Sujit Deokar * Merged in feature/standardized-services (pull request #958) Feature/standardized services * service term standardization * prompt update * prompt updates for standardization * only service standardization * new file * add supporting files and test scripts for standardization work Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * remove old files * Merge remote-tracking branch 'origin/dev' into feature/standardized-services * final fixes * Merged dev into feature/standardized-services * additional features * removed unwanted files * remove unwanted files * Merge branch 'dev' into feature/standardized-services * Merge remote-track… * Merged in feature/one-to-one-confidence-scoring (pull request #1002) Feature/one to one confidence scoring * T1 plumbing: capture per-field confidence + retrieved-chunk metadata for 1:1 HSC fields Prep work for the 1:1 confidence-scoring stage. No scoring logic yet — this just collects the inputs the next ticket (rule-based scorer) will consume. - ONE_TO_ONE_SINGLE_FIELD_TEMPLATE: ask the LLM for confidence (0.0-1.0), verdict (correct/uncertain/not_found), and supporting_snippet alongside the field value. Existing field parser passes the extra keys through unchanged. - prompt_hsc_single_field: now returns a 4-tuple (name, value, field, metadata) where metadata holds the confidence/verdict/snippet plus a lightweight summary of which chunks the LLM saw (count + ids). _extract_hsc_metadata is defensive: clamps out-of-range confidences, defaults a missing/garbage verdict, caps the snippet at 500 chars, returns _empty_hsc_metadata() on every bail-out path. - run_hybrid_smart_chunked_fields: opt… * Merged in bugfix/confidence-flagged-missing-field-col (pull request #1006) Fix KeyError in compute_flagged when a *_CONF column has no value sibling * Fix KeyError in compute_flagged when a *_CONF column has no value sibling Production hit a hard crash at the end of every run: KeyError: "['DYNAMIC_PRIMARY_ENTITIES'] not in index" src/qc_qa/confidence/summary.py:143 compute_flagged was iterating over *_CONF columns and unconditionally indexing the dataframe with both the FILE_NAME column and the stripped value column. That assumed every <FIELD>_CONF column has a sibling <FIELD> value column in final_df. That isn't always true: dynamic-primary features carry only the _CONF side (their value side is dropped by reorder_columns since it isn't in FIELD_FORMAT_MAPPING but its _CONF suffix matches the explicit _CONF carve-out). When the model produced a below-threshold score for one of these and the value column was absent, pandas .loc raised KeyError and the runner crashed. Fix: - Build the .loc colu… * Merged in feature/active-rates (pull request #1007) Feature/active rates * amendment intent tag * prompt update * Merge remote-tracking branch 'origin/dev' into feature/active-rates * Merge remote-tracking branch 'origin/dev' into feature/active-rates * Merge branch 'dev' into feature/active-rates * exhibit standardization updates * use AARETE_DERIVED_EXHIBIT_TITLE for intent * active rates stuff * Merge branch 'dev' into feature/active-rates * prompt update * amendment intent types * active rates logic update * unit tests * Merged dev into feature/active-rates * logic updates * Merge remote-tracking branch 'origin/dev' into feature/active-rates * Merge remote-tracking branch 'origin/dev' into feature/active-rates * updated instrumentation cost logs * cost loging * caching updates * Merge remote-tracking branch 'origin/dev' into feature/active-rates * Merge remote-tracking branch 'origin/dev' into feature/active-rates * logging fix * null check issue fixes * caching fix Approved-by… * Merged in bugfix/stg-to-main-prep (pull request #1012) Bugfix/stg to main prep * Merged in dev (pull request #1001) Dev * Revert premature merge of bugfix/retire_stale_client_file_processing PR #959 was merged into dev without approval. This reverts commits5e143c10,63f32c41,849aa626, and927abcaeto restore dev to its pre-merge state. The changes will be re-submitted via a new PR after proper review. * Merged in bugfix/retire_stale_client_file_processing (pull request #961) Return None for dashboard output when dashboard postprocessing is off * Return None for dashboard output when dashboard postprocessing is off FINAL_RESULT_DF_DASHBOARD was initialized as an empty DataFrame even when RUN_DASHBOARD_POSTPROCESSING was False, causing downstream code to needlessly process it (reorder_columns, etc). Now returns None when dashboard is not requested, matching the postprocess() contract. * Merged dev into bugfix/retire_stale_client_file_processing * Merge dev (with revert) into feature branch * Re-ap… * Merged in bugfix/sync-stg-into-dev-20260518 (pull request #1015) Merged in dev (pull request #1001) * Merged in dev (pull request #1001) Dev * Revert premature merge of bugfix/retire_stale_client_file_processing PR #959 was merged into dev without approval. This reverts commits5e143c10,63f32c41,849aa626, and927abcaeto restore dev to its pre-merge state. The changes will be re-submitted via a new PR after proper review. * Merged in bugfix/retire_stale_client_file_processing (pull request #961) Return None for dashboard output when dashboard postprocessing is off * Return None for dashboard output when dashboard postprocessing is off FINAL_RESULT_DF_DASHBOARD was initialized as an empty DataFrame even when RUN_DASHBOARD_POSTPROCESSING was False, causing downstream code to needlessly process it (reorder_columns, etc). Now returns None when dashboard is not requested, matching the postprocess() contract. * Merged dev into bugfix/retire_stale_client_file_processing * Merge dev (with revert) into fe… * Merged in dev (pull request #1016) Dev * Merged in bugfix/DAIP2-2823-generic-issue-fixes-one-to-one (pull request #1011) Bugfix/DAIP2-2823 generic issue fixes one to one * updated auto renewal ind prompt * updated CONTRACT_AMENDMENT_NUM prompt * Merged dev into bugfix/DAIP2-2823-generic-issue-fixes-one-to-one Approved-by: Praneel Panchigar Approved-by: Siddhant Medar * Merged in feature/DAIP2-2698-phase-3-program-product-lob-mapping (pull request #1000) Feature/DAIP2-2698 phase 3 program product lob mapping * added missing phase 2 modifications * added phase 3 modifications * Fixed acronym issues * black format fix * Merged dev into feature/DAIP2-2698-phase-3-program-product-lob-mapping * standardization fixes * Merged dev into feature/DAIP2-2698-phase-3-program-product-lob-mapping * added hyphenated suffix names fix * black format fix * standardization fix * dynamic primary fix * Merge Dev into feature/DAIP2-2698-phase-3-program-product-lob-mapping * updated Program Product Standardiza… Approved-by: Praneel Panchigar
7.9 KiB
Prompt Caching System Proposal
Branch: feature/DAIP2-2314-expand-caching-for-short-prompts
Goal
Ensure every stable instruction prompt that is intended to be cached is actually cacheable, warmed on the same model used at runtime, and observable in runtime logs.
Recommended Caching System
-
Introduce a central prompt cache registry.
Each cacheable prompt should be declared once with:
usage_label- instruction function
- model family used at runtime
- minimum cache token policy
- runtime call sites
- whether the prompt also uses
context_for_caching
This avoids the current split where cache warming is in
prompt_templates.get_cacheable_instructions()but runtime model choice is scattered through prompt call sites. -
Enforce the 1024-token minimum before runtime.
Prompt caching should not be treated as enabled just because
cache=Trueis passed. At startup, the registry should validate every cached instruction with the same tokenizer/estimate used by reporting. If a prompt is below the minimum, it should either:- fail fast in a cache audit mode, or
- automatically append approved static cache padding/instruction context until it crosses the threshold.
The padding must be stable and domain-relevant, not dynamic request data.
-
Warm caches per model, not globally.
Cache entries are model-specific. The runner currently warms all instructions with
sonnet_latest, butSPLIT_SERVICE_TERMruns withhaiku_latest. The cache warming layer should warm(usage_label, resolved_model_id)pairs derived from the registry. -
Only mark supported models as cacheable.
_supports_prompt_cache()currently excludes Haiku. If Haiku prompt caching is not supported in this Bedrock setup, Haiku prompts should not be registered as cached. Either move those prompt calls tosonnet_latestwhen caching matters, or keep them uncached and report them as intentionally uncached. -
Keep static instruction caching separate from dynamic context caching.
Instruction prompts can be cached predictably.
context_for_cachingis different because it changes by contract, exhibit, or page. It should be tracked as a separate cache class:instruction_cache: expected to read after warm-upcontext_cache: expected to create/read per repeated context
This prevents context variability from making instruction caching look broken.
-
Add a cache contract test.
A local test should inspect the built request body for every registered prompt and assert:
- the instruction block has
cache_control - the resolved model supports caching
- the static instruction token estimate is at least 1024
- warm-up model equals runtime model
- the instruction block has
-
Upgrade runtime reporting.
Add these fields to prompt-call logs:
resolved_model_idcache_policy:instruction,context,instruction+context,noneinstruction_cache_eligiblecontext_cache_eligiblecache_miss_reason
This makes root cause visible without manual investigation.
Prompt Application Report
Apply Instruction Caching, Already Over 1024 Tokens
These should remain registered and warmed. They are good candidates for strict cache contract tests.
| Prompt | Estimated Instruction Tokens | Action |
|---|---|---|
OUTLIER_BREAKOUT |
2483 | Keep cached |
METHODOLOGY_BREAKOUT |
1927 | Keep cached |
SPLIT_SERVICE_TERM |
1600 | Cache only if runtime model supports caching; currently runs on Haiku |
DYNAMIC_ASSIGNMENT |
1578 | Keep cached |
CARVEOUT_CHECK |
1498 | Keep cached |
SERVICE_ENRICHMENT |
1459 | Keep cached |
REIMBURSEMENT_PRIMARY |
1390 | Keep cached |
EXHIBIT_LEVEL |
1108 | Keep cached |
Apply Instruction Caching After Padding/Expansion
These are wired or partially wired but below the 1024-token minimum. They should be padded or expanded with stable domain guidance before being considered cached.
| Prompt | Estimated Instruction Tokens | Tokens Needed | Action |
|---|---|---|---|
fill_bill_type |
968 | 56 | Add small static guidance |
DYNAMIC_CODE_ASSIGNMENT |
966 | 58 | Add small static guidance |
CHECK_PROVIDER_NAME_MATCH |
957 | 67 | Add small static guidance |
GROUPER_BREAKOUT |
939 | 85 | Add small static guidance |
EXHIBIT_HEADER |
922 | 102 | Add static header examples/rules |
LESSER_OF_DISTRIBUTION |
914 | 110 | Add static decision examples |
LESSER_OF_CHECK |
900 | 124 | Add static classification examples |
CODE_EXPLICIT |
871 | 153 | Add static code-system rules |
REIMB_DATES_ASSIGNMENT |
803 | 221 | Add date extraction examples |
DYNAMIC_PRIMARY |
441 | 583 | Expand instruction substantially |
SPLIT_REIMB_DATES |
410 | 614 | Expand instruction substantially |
prompt_lob_relationship |
408 | 616 | Move field prompt into instruction or add static LOB rules |
code_implicit_arbitration |
391 | 633 | Expand arbitration rules |
AARETE_DERIVED_PAYER_NAME |
372 | 652 | Expand payer naming rules |
FEE_SCHEDULE_BREAKOUT |
352 | 672 | Expand fee schedule rules |
code_implicit_special |
329 | 695 | Expand special code rules |
DATE_FIX |
322 | 702 | Expand date normalization rules |
SPECIAL_CASE_ASSIGNMENT |
251 | 773 | Expand special case rules |
DERIVED_TERM_DATE |
247 | 777 | Expand term date rules |
code_last_check |
179 | 845 | Expand or leave uncached if low value |
EXTRACT_AMENDMENT_NUM_FROM_FILENAME |
152 | 872 | Expand or leave uncached if low value |
EXHIBIT_HEADER_DEDUP |
35 | 989 | Do not cache unless redesigned |
Wire for Caching or Mark Intentionally Uncached
These showed runtime volume but no instruction function match in the audit. They need explicit registry decisions.
| Prompt | Runtime Finding | Action |
|---|---|---|
prompt_hsc_single_field |
109 calls, cache_enabled=False, large dynamic prompt |
Split stable HSC instructions from dynamic chunk and cache the instruction |
SPECIAL_CASE_BREAKOUT |
6 calls, no instruction cache | Wire only if repeated volume justifies expansion |
code_implicit_rag |
Cached in API, audit alias missing | Add registry alias to the actual instruction function |
prompt_provider_info |
Cached in API, audit alias missing | Add registry alias to TIN_NPI_TEMPLATE_INSTRUCTION or rename usage label |
validate_reimbursements_for_llm |
Cached in API, audit alias missing | Add registry alias to VALIDATE_REIMBURSEMENTS_INSTRUCTION |
Immediate Fixes
-
Change
SPLIT_SERVICE_TERMto use a cache-supported model or mark it intentionally uncached. Current behavior is misleading:cache=Trueis passed, but no Bedrock cache tokens are recorded. -
Add a registry alias map for usage labels that do not match instruction names:
code_implicit_rag->CODE_IMPLICITprompt_provider_info->TIN_NPI_TEMPLATEvalidate_reimbursements_for_llm->VALIDATE_REIMBURSEMENTSprompt_lob_relationship-> its real instruction function
-
Expand the near-threshold instructions first because they need the least work and have current evidence of API cache reads:
fill_bill_typeDYNAMIC_CODE_ASSIGNMENTCHECK_PROVIDER_NAME_MATCHGROUPER_BREAKOUTEXHIBIT_HEADERLESSER_OF_DISTRIBUTIONLESSER_OF_CHECKCODE_EXPLICIT
-
Redesign
prompt_hsc_single_fieldinto(instruction, prompt)form. It is high-volume and currently sends all prompt text dynamically.
Expected Behavior After Fix
For instruction-only prompts, after warm-up the API should report cache reads on normal calls. The cached percentage will still not be 100% of total billed input because the dynamic user prompt remains uncached. The correct target is:
cache_creation_tokensonly during warm-up or first uncached usecache_read_tokens > 0on every eligible runtime callcache_miss_reasonempty for registered instruction cache prompts
For context-cached prompts, 100% cache reads should not be expected unless the same context block repeats. The report should evaluate those separately.