Merged in dev (pull request #1001) * Merged in feature/fixPlaceholder2 (pull request #979) Updated feature> dev gate to print variables of echo statements * Updated feature> dev gate to print variables of echo statement * Reverted placeholder changes Approved-by: Sujit Deokar * Merged in bugfix/prov_info_json_fixes (pull request #981) fix: fall back to PROVIDER_NAME in PROV_INFO_JSON when no TIN/NPI extracted * fix: fall back to PROVIDER_NAME in PROV_INFO_JSON when no TIN/NPI extracted When get_prov_info_json short-circuits due to no TIN/NPI regex matches, PROV_INFO_JSON was left as [] even when PROVIDER_NAME was successfully extracted via the one-to-one pipeline. This caused inconsistent output across contracts with the same provider — some files produced a NAME-only entry (via a false-positive regex hit triggering the LLM), others produced []. Reconcile at add_group_and_other, the first point where both extraction streams' results are available. When PROV_INFO_JSON is empty but PROVIDER_NAME is known, synthesize a NAME-only entry with IS_GROUP:"Y" and populate PROV_GROUP_NAME_FULL directly — skipping the provider_name_match_check LLM call since the match is tautological by construction. Adds 5 unit tests covering the str, list, already-populated, empty-name, and all-empty-lis… * Merged in feature/DAIP2-pacificsource-reimbursements-issues (pull request #982) Feature/DAIP2 pacificsource reimbursements issues * Tighten PREMIUM_TERM and DISCOUNT_TERM classifier prompts CARVEOUT_CHECK was misrouting table rate rows into special-case fields, dropping them from the reimbursement output: - "110% of CMS allowed" (base fee-schedule rates) was being classified as PREMIUM_TERM because the prompt treated "above 100% of reference" as an implicit premium. Seen on PacificSource Medicare_Attachment_A1 and A2 Facility contracts where Inpatient/Outpatient rows were missing (2556) or silently fell back to 100% fee-schedule (2557). - Per-service discount rates like "Progressive Lenses: 15% discount", "Contact Lenses: 2% discount", "Frame: 20% discount" were being classified as DISCOUNT_TERM because the prompt only required the word "discount" to appear. Seen on PacificSource Commercial_Attachment_A2 and A5 Professional contracts (2558, 2559). Prompts now require the literal keyword … * Merged in bugfix/filter-docusign-lines (pull request #983) remove docusign lines * remove docusign lines * add unit tests for clean_header_footer docusign/deleted_lines changes Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * black forrmatting Approved-by: Katon Minhas * Merged in feature/PC_logic_cleanup_output (pull request #984) Feature/PC logic cleanup output * Few tweaks PC_logics * Fixed orphan_ranking * black formatting * Changes on output field and ranking method * updated few hotfixes * black format fix Approved-by: Katon Minhas * Merged in hotfix/provider_name_group_fix (pull request #988) Hotfix/provider name group fix * Fixes done in GROUP column * black format fix Approved-by: Katon Minhas * Merged in bugfix/DAIP2-2524-carveout-code-optimization (pull request #980) CARVEOUT_CD issue fixed * CARVEOUT_CD issue fixed * pipeline error fixed * Merged dev into bugfix/DAIP2-2524-carveout-code-optimization * Merged dev into bugfix/DAIP2-2524-carveout-code-optimization * Merged dev into bugfix/DAIP2-2524-carveout-code-optimization * trigger cap issue fixed * trigger cap prompt updated * Merged dev into bugfix/DAIP2-2524-carveout-code-optimization Approved-by: Katon Minhas * Merged in bugfix/exhibit-smart-chunking-cost-improvements (pull request #985) Bugfix/exhibit smart chunking cost improvements * Add opt-in instrumentation for per-call token and row-count tracing Introduce src/utils/instrumentation.py (thread-safe CSV logger) and src/utils/instrumentation_context.py (ContextVar scope plus submit_with_context / map_with_context helpers for propagating context into ThreadPoolExecutor workers). Emit events at every Bedrock call in llm_utils.invoke_claude, including in-memory claude_cache hits, with full input/output/cache-read/cache-write token breakdown. Emit row-count events at each row-mutating stage in the one-to-N pipeline (clean_reimbursement_primary, filter_services_without_reimbursements, methodology_breakout, split_service_terms, carveout, dynamic_code_assignment, lesser_of_distribution, dynamic_assignment) and chunking / retrieval events in exhibit smart chunking (chunking_done, retrieval_done) plus exhibit lifecycle events (exhibit_start, exhibit_gate_skip, stage_… * Merged in hotfix/fileextension_issue (pull request #989) Hotfix/fileextension issue * fixed strip_ext issue * black format Approved-by: Katon Minhas * Merged in hotfix/exhibit-header-in-tables (pull request #987) Hotfix/exhibit header in tables * Merged in feature/FixplaceholderIssue (pull request #977) Remove curly braces from echo statements in dev->stg * Remove curly brances from echo statements in dev->stg * Removed curly braces in echo statements in feature-> dev gate Approved-by: Sujit Deokar * Merged in feature/standardized-services (pull request #958) Feature/standardized services * service term standardization * prompt update * prompt updates for standardization * only service standardization * new file * add supporting files and test scripts for standardization work Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * remove old files * Merge remote-tracking branch 'origin/dev' into feature/standardized-services * final fixes * Merged dev into feature/standardized-services * additional features * removed unwanted files * remove unwanted files * Merge branch 'dev' into feature/standardized-services * Merge remote-tr… * Merged in hotfix/postprocess_csv (pull request #990) Hotfix/postprocess csv * date issue_fix * black_format * Merged dev into hotfix/postprocess_csv Approved-by: Katon Minhas * Merged in bugfix/molina_ut_dynamic_primary (pull request #991) Bugfix/molina ut dynamic primary * Prompt changes for dynamic primary * Route cover-sheet-only files to ERRORS.csv instead of leaking phantom rows When every page of a contract was filtered out as a cover sheet / quick-review form, process_file silently returned a FILE_NAME-only DataFrame. Because the runner routes by checking for an "error" column, that file landed in RESULTS.csv as a near-empty row and no ERRORS.csv was generated for the run. - saas/file_processing.py: raise ValueError when text_dict is empty after cover-sheet filtering, so safe_process_file produces a proper error row. - runner.py: add _is_phantom_result defense-in-depth — promote any result with no extracted fields beyond FILE_NAME to error_results with error_type=PhantomSuccess. * Merged dev into bugfix/molina_ut_dynamic_primary * Tighten PRODUCT prompt: restrict to valid_values, prune LOB/PROGRAM examples * Merge branch 'bugfix/molina_ut_dynamic_primary' of bit… * Merged in bugfix/black-format (pull request #994) Black format for pipeline pass * Black format for pipeline pass * Merged in bugfix/DAIP2-2679-fix-nebraska-issues (pull request #995) incorrect inclusion of CPT4_PROC_CD fixed * incorrect inclusion of CPT4_PROC_CD fixed Approved-by: Katon Minhas * Merged in feature/DAIP2-2314-DAIP2-1687-hybrid (pull request #993) Feature/DAIP2-2314 DAIP2 1687 hybrid * remove -files from s3 prefix requirements * Resolve input paths * fix: VendorProcessor.process_file returns (df, None) tuple runner.safe_process_file unpacks the result as (cc_df, dashboard_df), so returning a single DataFrame caused every vendor/generic file to fail with "too many values to unpack (expected 2)" — Python iterates DataFrame columns during unpacking. Vendor pipelines have no dashboard variant; second slot is None and the existing `dashboard_result is not None` guard in runner.py already handles it. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * DAIP2-2314 + DAIP2-1687: pad DYNAMIC_PRIMARY + DYNAMIC_PRIMARY_ENTITY_CLASSIFICATION over 1024-token cache floor - Pad DYNAMIC_PRIMARY_INSTRUCTION with three new sections: [SCOPE BOUNDARIES], [SOURCE TEXT INTERPRETATION], [REASONING DISCIPLINE], plus a [WORKED EXAMPLES] block. Estimated tokens: 447 -> 1117 (Sonnet 4.5 … * Merged in bugfix/postprocessing_date_fix (pull request #996) Date formatting changes * Date formatting changes * Merged dev into bugfix/postprocessing_date_fix Approved-by: Katon Minhas * Merged in bugfix/parent-child-rank-orphan-uniqueness (pull request #999) PC_logic bugfix * PC_logic bugfix Approved-by: Katon Minhas * Merged in feature/document-index (pull request #1005) Feature/document index * Add Document Index preprocessing — Layers 1, 2, and 3 wiring Parse the Textract-emitted Document Index block at the top of each contract with a single cached LLM call (prompt_document_index) instead of one per-page call per page. Layer 2 verifies parsed entries via literal string match and structural regex sweep, escalating suspect pages back to the existing per-page path. Layer 1+2 failure triggers a full fallback to today's per-page flow. New symbols: - preprocessing_funcs.extract_document_index_block — regex slice of index prefix - preprocessing_funcs.verify_index_against_pages — structural verifier (plain dict return) - prompt_templates.DOCUMENT_INDEX_INSTRUCTION / DOCUMENT_INDEX — cached prompt pair - prompt_calls.prompt_document_index — LLM wrapper (usage_label DOCUMENT_INDEX_PARSE) - config: DOCUMENT_INDEX_PARSE_ENABLED and three threshold flags - instrumentation: DOCUMENT_INDEX_PARSE mapped to preprocessing segment one… * Merged in feature/active-rates (pull request #1004) Feature/active rates * initial commit * Merged in feature/FixplaceholderIssue (pull request #977) Remove curly braces from echo statements in dev->stg * Remove curly brances from echo statements in dev->stg * Removed curly braces in echo statements in feature-> dev gate Approved-by: Sujit Deokar * Merged in feature/standardized-services (pull request #958) Feature/standardized services * service term standardization * prompt update * prompt updates for standardization * only service standardization * new file * add supporting files and test scripts for standardization work Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * remove old files * Merge remote-tracking branch 'origin/dev' into feature/standardized-services * final fixes * Merged dev into feature/standardized-services * additional features * removed unwanted files * remove unwanted files * Merge branch 'dev' into feature/standardized-services * Merge remote-track… * Merged in feature/one-to-one-confidence-scoring (pull request #1002) Feature/one to one confidence scoring * T1 plumbing: capture per-field confidence + retrieved-chunk metadata for 1:1 HSC fields Prep work for the 1:1 confidence-scoring stage. No scoring logic yet — this just collects the inputs the next ticket (rule-based scorer) will consume. - ONE_TO_ONE_SINGLE_FIELD_TEMPLATE: ask the LLM for confidence (0.0-1.0), verdict (correct/uncertain/not_found), and supporting_snippet alongside the field value. Existing field parser passes the extra keys through unchanged. - prompt_hsc_single_field: now returns a 4-tuple (name, value, field, metadata) where metadata holds the confidence/verdict/snippet plus a lightweight summary of which chunks the LLM saw (count + ids). _extract_hsc_metadata is defensive: clamps out-of-range confidences, defaults a missing/garbage verdict, caps the snippet at 500 chars, returns _empty_hsc_metadata() on every bail-out path. - run_hybrid_smart_chunked_fields: opt… * Merged in bugfix/confidence-flagged-missing-field-col (pull request #1006) Fix KeyError in compute_flagged when a *_CONF column has no value sibling * Fix KeyError in compute_flagged when a *_CONF column has no value sibling Production hit a hard crash at the end of every run: KeyError: "['DYNAMIC_PRIMARY_ENTITIES'] not in index" src/qc_qa/confidence/summary.py:143 compute_flagged was iterating over *_CONF columns and unconditionally indexing the dataframe with both the FILE_NAME column and the stripped value column. That assumed every <FIELD>_CONF column has a sibling <FIELD> value column in final_df. That isn't always true: dynamic-primary features carry only the _CONF side (their value side is dropped by reorder_columns since it isn't in FIELD_FORMAT_MAPPING but its _CONF suffix matches the explicit _CONF carve-out). When the model produced a below-threshold score for one of these and the value column was absent, pandas .loc raised KeyError and the runner crashed. Fix: - Build the .loc colu… * Merged in feature/active-rates (pull request #1007) Feature/active rates * amendment intent tag * prompt update * Merge remote-tracking branch 'origin/dev' into feature/active-rates * Merge remote-tracking branch 'origin/dev' into feature/active-rates * Merge branch 'dev' into feature/active-rates * exhibit standardization updates * use AARETE_DERIVED_EXHIBIT_TITLE for intent * active rates stuff * Merge branch 'dev' into feature/active-rates * prompt update * amendment intent types * active rates logic update * unit tests * Merged dev into feature/active-rates * logic updates * Merge remote-tracking branch 'origin/dev' into feature/active-rates * Merge remote-tracking branch 'origin/dev' into feature/active-rates * updated instrumentation cost logs * cost loging * caching updates * Merge remote-tracking branch 'origin/dev' into feature/active-rates * Merge remote-tracking branch 'origin/dev' into feature/active-rates * logging fix * null check issue fixes * caching fix Approved-by… * Merged in bugfix/stg-to-main-prep (pull request #1012) Bugfix/stg to main prep * Merged in dev (pull request #1001) Dev * Revert premature merge of bugfix/retire_stale_client_file_processing PR #959 was merged into dev without approval. This reverts commits5e143c10,63f32c41,849aa626, and927abcaeto restore dev to its pre-merge state. The changes will be re-submitted via a new PR after proper review. * Merged in bugfix/retire_stale_client_file_processing (pull request #961) Return None for dashboard output when dashboard postprocessing is off * Return None for dashboard output when dashboard postprocessing is off FINAL_RESULT_DF_DASHBOARD was initialized as an empty DataFrame even when RUN_DASHBOARD_POSTPROCESSING was False, causing downstream code to needlessly process it (reorder_columns, etc). Now returns None when dashboard is not requested, matching the postprocess() contract. * Merged dev into bugfix/retire_stale_client_file_processing * Merge dev (with revert) into feature branch * Re-ap… * Merged in bugfix/sync-stg-into-dev-20260518 (pull request #1015) Merged in dev (pull request #1001) * Merged in dev (pull request #1001) Dev * Revert premature merge of bugfix/retire_stale_client_file_processing PR #959 was merged into dev without approval. This reverts commits5e143c10,63f32c41,849aa626, and927abcaeto restore dev to its pre-merge state. The changes will be re-submitted via a new PR after proper review. * Merged in bugfix/retire_stale_client_file_processing (pull request #961) Return None for dashboard output when dashboard postprocessing is off * Return None for dashboard output when dashboard postprocessing is off FINAL_RESULT_DF_DASHBOARD was initialized as an empty DataFrame even when RUN_DASHBOARD_POSTPROCESSING was False, causing downstream code to needlessly process it (reorder_columns, etc). Now returns None when dashboard is not requested, matching the postprocess() contract. * Merged dev into bugfix/retire_stale_client_file_processing * Merge dev (with revert) into fe… * Merged in dev (pull request #1016) Dev * Merged in bugfix/DAIP2-2823-generic-issue-fixes-one-to-one (pull request #1011) Bugfix/DAIP2-2823 generic issue fixes one to one * updated auto renewal ind prompt * updated CONTRACT_AMENDMENT_NUM prompt * Merged dev into bugfix/DAIP2-2823-generic-issue-fixes-one-to-one Approved-by: Praneel Panchigar Approved-by: Siddhant Medar * Merged in feature/DAIP2-2698-phase-3-program-product-lob-mapping (pull request #1000) Feature/DAIP2-2698 phase 3 program product lob mapping * added missing phase 2 modifications * added phase 3 modifications * Fixed acronym issues * black format fix * Merged dev into feature/DAIP2-2698-phase-3-program-product-lob-mapping * standardization fixes * Merged dev into feature/DAIP2-2698-phase-3-program-product-lob-mapping * added hyphenated suffix names fix * black format fix * standardization fix * dynamic primary fix * Merge Dev into feature/DAIP2-2698-phase-3-program-product-lob-mapping * updated Program Product Standardiza… Approved-by: Praneel Panchigar
8.0 KiB
Prompt Caching Analysis Report
Branch: feature/DAIP2-2314-expand-caching-for-short-prompts
Date: 2026-04-21
Comparison: Feature branch vs origin/dev
Executive Summary
This report analyzes the prompt caching implementation in the feature branch, comparing cached prompts against the dev branch and evaluating cost savings from Bedrock prompt caching.
Key Findings
| Metric | Value |
|---|---|
| Total API Calls Analyzed | 1,251 |
| Unique Usage Labels | 35 |
| Total Cache Read Tokens | 1,472,330 |
| Total Cache Write Tokens | 87,167 |
| Overall Cost Savings | 28.1% ($3.65 saved) |
| Prompts Actively Caching | 22 |
| Prompts Wired but Not Caching | 11 |
| Prompts Not Wired | 2 |
1. Cache Effectiveness by Usage Label
Actively Caching (22 prompts) - High Performers
| Usage Label | Calls | Avg Instruction Tokens | Cache Hit Rate |
|---|---|---|---|
| OUTLIER_BREAKOUT | 1 | 2,483 | 98.4% |
| SERVICE_ENRICHMENT | 30 | 1,459 | 98.3% |
| CARVEOUT_CHECK | 43 | 1,498 | 95.6% |
| METHODOLOGY_BREAKOUT | 37 | 1,927 | 95.5% |
| LESSER_OF_DISTRIBUTION | 32 | 914 | 94.5% |
| CHECK_PROVIDER_NAME_MATCH | 74 | 957 | 94.2% |
| DYNAMIC_CODE_ASSIGNMENT | 38 | 966 | 94.1% |
| CODE_EXPLICIT | 34 | 871 | 93.9% |
| GROUPER_BREAKOUT | 7 | 939 | 93.6% |
| LESSER_OF_CHECK | 6 | 900 | 92.0% |
| DYNAMIC_ASSIGNMENT | 204 | 1,578 | 89.4% |
| validate_reimbursements_for_llm | 65 | 894 | 89.2% |
| code_implicit_rag | 60 | 896 | 88.4% |
| REIMB_DATES_ASSIGNMENT | 29 | 803 | 87.2% |
| REIMBURSEMENT_PRIMARY | 38 | 1,390 | 85.0% |
| EXHIBIT_LEVEL | 33 | 1,108 | 74.0% |
| EXHIBIT_HEADER | 90 | 922 | 73.7% |
| fill_bill_type | 47 | 968 | 72.5% |
| prompt_provider_info | 24 | 966 | 60.5% |
| DYNAMIC_PRIMARY | 44 | 441 | 49.7% |
Wired But Not Caching (11 prompts) - Under 1024 Token Threshold
These prompts are wired for caching but have instruction tokens below Bedrock's 1024-token minimum threshold:
| Usage Label | Calls | Avg Instruction Tokens | Issue |
|---|---|---|---|
| SPLIT_SERVICE_TERM | 10 | 1,600 | Should be caching - investigate |
| prompt_lob_relationship | 56 | 408 | Needs padding to 1024 |
| SPLIT_REIMB_DATES | 13 | 410 | Needs padding to 1024 |
| code_implicit_arbitration | 12 | 391 | Needs padding to 1024 |
| AARETE_DERIVED_PAYER_NAME | 1 | 372 | Needs padding to 1024 |
| FEE_SCHEDULE_BREAKOUT | 18 | 352 | Needs padding to 1024 |
| code_implicit_special | 30 | 329 | Needs padding to 1024 |
| DATE_FIX | 10 | 322 | Needs padding to 1024 |
| SPECIAL_CASE_ASSIGNMENT | 4 | 251 | Needs padding to 1024 |
| DERIVED_TERM_DATE | 10 | 247 | Needs padding to 1024 |
| code_last_check | 23 | 179 | Needs padding to 1024 |
| EXTRACT_AMENDMENT_NUM_FROM_FILENAME | 3 | 152 | Needs padding to 1024 |
| EXHIBIT_HEADER_DEDUP | 10 | 35 | Needs padding to 1024 |
Not Wired for Caching (2 prompts)
| Usage Label | Calls | Notes |
|---|---|---|
| prompt_hsc_single_field | 109 | High volume - should investigate wiring |
| SPECIAL_CASE_BREAKOUT | 6 | Low volume |
2. Prompt Template Changes (Feature Branch vs Dev)
The feature branch includes 649 lines changed in src/prompts/prompt_templates.py. Key expansions:
Expanded Prompts (Additional Context Added)
| Instruction Function | Change Summary |
|---|---|
| EXHIBIT_LEVEL_INSTRUCTION | +55 lines: Added detailed guidance, validation rules, examples, common pitfalls, field-specific notes, quality assurance checks |
| DYNAMIC_PRIMARY_INSTRUCTION | +18 lines: Added extraction guidance, ambiguity handling, quality checks |
| LESSER_OF_DISTRIBUTION_INSTRUCTION | Restructured (net -25 lines): Converted to cleaner decision framework format |
| DYNAMIC_CODE_ASSIGNMENT_INSTRUCTION | +10 lines: Added classification guidance with explicit rules |
| CODE_IMPLICIT_INSTRUCTION | Reformatted indentation |
| CODE_IMPLICIT_ARBITRATION_INSTRUCTION | +43 lines: Added clarifications, deterministic decision rules, calibration examples |
| FILL_BILL_TYPE_INSTRUCTION | +49 lines: Added detailed guidance, validation rules, extended examples |
| TIN_NPI_TEMPLATE_INSTRUCTION | +22 lines: Added provider/payer distinction clarity |
Analysis of Prompt Changes
Positive Impacts:
- More explicit instructions reduce model ambiguity
- Quality checks and validation rules improve consistency
- Examples help calibrate model responses
- Clearer formatting improves readability
Potential Concerns:
- Expanded prompts increase token count (higher cache write cost initially)
- Some prompts may have reduced due to restructuring (LESSER_OF_DISTRIBUTION)
- Need to verify output quality hasn't degraded with expanded context
3. Cost Analysis
Bedrock Pricing Applied
- Input tokens: $3.00/1M tokens
- Output tokens: $15.00/1M tokens
- Cache read: $0.30/1M tokens (90% savings)
- Cache write: $3.75/1M tokens (25% premium)
Observed Results
| Metric | Value |
|---|---|
| Total Input Tokens | 844,370 |
| Total Output Tokens | 401,926 |
| Total Cache Read Tokens | 1,472,330 |
| Total Cache Write Tokens | 87,167 |
| Cost Without Caching | $12.98 |
| Cost With Caching | $9.33 |
| Savings | $3.65 (28.1%) |
Per-Prompt Cost Impact (Estimated from previous testing)
| Prompt | Cost Reduction |
|---|---|
| LESSER_OF_DISTRIBUTION | 44% (Best performer) |
| EXHIBIT_HEADER | 7% |
| CHECK_PROVIDER_NAME_MATCH | 5% |
| TIN_NPI_TEMPLATE (prompt_provider_info) | 5% |
| code_implicit_rag | -12% (Negative - needs review) |
4. Prompts Needing Attention
Priority 1: Prompts Needing Padding (Under 1024 tokens)
These prompts are wired but not actually caching due to Bedrock's minimum threshold:
- prompt_lob_relationship (408 tokens, 56 calls) - High impact
- code_implicit_special (329 tokens, 30 calls) - Medium impact
- code_last_check (179 tokens, 23 calls) - Medium impact
- FEE_SCHEDULE_BREAKOUT (352 tokens, 18 calls) - Medium impact
- SPLIT_REIMB_DATES (410 tokens, 13 calls)
- code_implicit_arbitration (391 tokens, 12 calls)
- SPLIT_SERVICE_TERM (1600 tokens, 10 calls) - Should be caching, investigate
- DATE_FIX (322 tokens, 10 calls)
- DERIVED_TERM_DATE (247 tokens, 10 calls)
- EXHIBIT_HEADER_DEDUP (35 tokens, 10 calls)
Priority 2: Investigate Negative Result
- code_implicit_rag: Shows -12% cost (increase) - The expanded prompt may be too verbose or structured differently causing cache misses
Priority 3: Not Wired for Caching
- prompt_hsc_single_field (109 calls) - High volume, should consider wiring
5. Recommendations
-
Pad short prompts to 1024 tokens: Add contextual padding to prompts under the threshold, especially high-frequency ones like
prompt_lob_relationshipandcode_implicit_special -
Investigate code_implicit_rag: The negative cost result (-12%) suggests the expanded prompt may be causing cache fragmentation or misses
-
Wire prompt_hsc_single_field: With 109 calls, this is a good candidate for caching
-
Monitor SPLIT_SERVICE_TERM: At 1,600 tokens it should be caching but shows 0% cache hit rate - may be a wiring issue
-
Run testbed comparison: Use
src/testbed/testbed_metrics.pyto compare output quality between cached and non-cached prompts to ensure no regression
6. Files Analyzed
src/test-PROMPT-CALLS_10.csv- API call logs with cache statisticssrc/testbed-usethis.xlsx- Testbed comparison file (binary, not readable)src/prompts/prompt_templates.py- Prompt definitions (649 lines changed)documentation/prompt_caching_tracker_v2.csv- Caching status tracker
7. Next Steps
- Run testbed metrics comparison script to validate output quality
- Add padding to short prompts identified above
- Investigate and fix code_implicit_rag negative result
- Wire prompt_hsc_single_field for caching
- Debug SPLIT_SERVICE_TERM caching issue