Files
doczyai-pipelines/documentation/PROMPT_CACHING_ANALYSIS.md
T
Venkatakrishna Reddy Avula 2da639e59b Merged in feature/DAIP2-2314-DAIP2-1687-hybrid (pull request #993)
Feature/DAIP2-2314 DAIP2 1687 hybrid

* remove -files from s3 prefix requirements

* Resolve input paths

* fix: VendorProcessor.process_file returns (df, None) tuple

runner.safe_process_file unpacks the result as (cc_df, dashboard_df), so
returning a single DataFrame caused every vendor/generic file to fail with
"too many values to unpack (expected 2)" — Python iterates DataFrame columns
during unpacking. Vendor pipelines have no dashboard variant; second slot is
None and the existing `dashboard_result is not None` guard in runner.py
already handles it.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* DAIP2-2314 + DAIP2-1687: pad DYNAMIC_PRIMARY + DYNAMIC_PRIMARY_ENTITY_CLASSIFICATION over 1024-token cache floor

- Pad DYNAMIC_PRIMARY_INSTRUCTION with three new sections: [SCOPE BOUNDARIES], [SOURCE TEXT INTERPRETATION], [REASONING DISCIPLINE], plus a [WORKED EXAMPLES] block. Estimated tokens: 447 -> 1117 (Sonnet 4.5 1024-min, +93 margin). All additions reinforce existing rules (sibling-field separation, alias mapping, pricing-vs-LOB distinction, contrastive-clause exclusion, exhibit-header binding) — no new directives that could bias extraction.
- Pad DYNAMIC_PRIMARY_ENTITY_CLASSIFICATION_INSTRUCTION with a [FINAL CHECKLIST BEFORE OUTPUT] block. Estimated tokens: 956 -> 1101 (Sonnet 4.5 1024-min, +77 margin). Reinforces the existing 4-step anti-duplication protocol and JSON shape requirements.
- Register both new entries in cache_registry: DYNAMIC_PRIMARY_ENTITY_CLASSIFICATION as INSTRUCTION_PLUS_CONTEXT (caches at warm-up), DYNAMIC_PRIMARY_ENTITIES as CONTEXT (instruction is intentionally short; CONTEXT c…
* black format fix

* Merged dev into feature/DAIP2-2314-DAIP2-1687-hybrid

* fixed raw lob values in base lob field mapping and composite entities fix

* black format fix

* fixed LOB Program output issues

* issue fixes

* remove debugging code

* Updated prompts

* updated additional instructions

* Update Program-->LOB

* LLM-based AD-Program/Product mapping to LOB even when there is a crosswalk

* black format fix

* Merged dev into feature/DAIP2-2314-DAIP2-1687-hybrid

* added logging in prompt call tracking

* added updated logging in prompt call tracking

* aaded min cache token per usage label

* added cache registry for dynamic primary mapping prompt calls

* reolved mapping prompts ambiguities

* black format fix

* Phase 2 modifications added

* reverted phase 2 modifications


Approved-by: Katon Minhas
2026-05-08 15:52:32 +00:00

205 lines
8.0 KiB
Markdown

# Prompt Caching Analysis Report
**Branch:** `feature/DAIP2-2314-expand-caching-for-short-prompts`
**Date:** 2026-04-21
**Comparison:** Feature branch vs `origin/dev`
---
## Executive Summary
This report analyzes the prompt caching implementation in the feature branch, comparing cached prompts against the dev branch and evaluating cost savings from Bedrock prompt caching.
### Key Findings
| Metric | Value |
|--------|-------|
| Total API Calls Analyzed | 1,251 |
| Unique Usage Labels | 35 |
| Total Cache Read Tokens | 1,472,330 |
| Total Cache Write Tokens | 87,167 |
| Overall Cost Savings | **28.1%** ($3.65 saved) |
| Prompts Actively Caching | 22 |
| Prompts Wired but Not Caching | 11 |
| Prompts Not Wired | 2 |
---
## 1. Cache Effectiveness by Usage Label
### Actively Caching (22 prompts) - High Performers
| Usage Label | Calls | Avg Instruction Tokens | Cache Hit Rate |
|-------------|-------|------------------------|----------------|
| OUTLIER_BREAKOUT | 1 | 2,483 | 98.4% |
| SERVICE_ENRICHMENT | 30 | 1,459 | 98.3% |
| CARVEOUT_CHECK | 43 | 1,498 | 95.6% |
| METHODOLOGY_BREAKOUT | 37 | 1,927 | 95.5% |
| LESSER_OF_DISTRIBUTION | 32 | 914 | 94.5% |
| CHECK_PROVIDER_NAME_MATCH | 74 | 957 | 94.2% |
| DYNAMIC_CODE_ASSIGNMENT | 38 | 966 | 94.1% |
| CODE_EXPLICIT | 34 | 871 | 93.9% |
| GROUPER_BREAKOUT | 7 | 939 | 93.6% |
| LESSER_OF_CHECK | 6 | 900 | 92.0% |
| DYNAMIC_ASSIGNMENT | 204 | 1,578 | 89.4% |
| validate_reimbursements_for_llm | 65 | 894 | 89.2% |
| code_implicit_rag | 60 | 896 | 88.4% |
| REIMB_DATES_ASSIGNMENT | 29 | 803 | 87.2% |
| REIMBURSEMENT_PRIMARY | 38 | 1,390 | 85.0% |
| EXHIBIT_LEVEL | 33 | 1,108 | 74.0% |
| EXHIBIT_HEADER | 90 | 922 | 73.7% |
| fill_bill_type | 47 | 968 | 72.5% |
| prompt_provider_info | 24 | 966 | 60.5% |
| DYNAMIC_PRIMARY | 44 | 441 | 49.7% |
### Wired But Not Caching (11 prompts) - Under 1024 Token Threshold
These prompts are wired for caching but have instruction tokens below Bedrock's 1024-token minimum threshold:
| Usage Label | Calls | Avg Instruction Tokens | Issue |
|-------------|-------|------------------------|-------|
| SPLIT_SERVICE_TERM | 10 | 1,600 | Should be caching - investigate |
| prompt_lob_relationship | 56 | 408 | **Needs padding to 1024** |
| SPLIT_REIMB_DATES | 13 | 410 | **Needs padding to 1024** |
| code_implicit_arbitration | 12 | 391 | **Needs padding to 1024** |
| AARETE_DERIVED_PAYER_NAME | 1 | 372 | **Needs padding to 1024** |
| FEE_SCHEDULE_BREAKOUT | 18 | 352 | **Needs padding to 1024** |
| code_implicit_special | 30 | 329 | **Needs padding to 1024** |
| DATE_FIX | 10 | 322 | **Needs padding to 1024** |
| SPECIAL_CASE_ASSIGNMENT | 4 | 251 | **Needs padding to 1024** |
| DERIVED_TERM_DATE | 10 | 247 | **Needs padding to 1024** |
| code_last_check | 23 | 179 | **Needs padding to 1024** |
| EXTRACT_AMENDMENT_NUM_FROM_FILENAME | 3 | 152 | **Needs padding to 1024** |
| EXHIBIT_HEADER_DEDUP | 10 | 35 | **Needs padding to 1024** |
### Not Wired for Caching (2 prompts)
| Usage Label | Calls | Notes |
|-------------|-------|-------|
| prompt_hsc_single_field | 109 | High volume - should investigate wiring |
| SPECIAL_CASE_BREAKOUT | 6 | Low volume |
---
## 2. Prompt Template Changes (Feature Branch vs Dev)
The feature branch includes **649 lines changed** in `src/prompts/prompt_templates.py`. Key expansions:
### Expanded Prompts (Additional Context Added)
| Instruction Function | Change Summary |
|---------------------|----------------|
| **EXHIBIT_LEVEL_INSTRUCTION** | +55 lines: Added detailed guidance, validation rules, examples, common pitfalls, field-specific notes, quality assurance checks |
| **DYNAMIC_PRIMARY_INSTRUCTION** | +18 lines: Added extraction guidance, ambiguity handling, quality checks |
| **LESSER_OF_DISTRIBUTION_INSTRUCTION** | Restructured (net -25 lines): Converted to cleaner decision framework format |
| **DYNAMIC_CODE_ASSIGNMENT_INSTRUCTION** | +10 lines: Added classification guidance with explicit rules |
| **CODE_IMPLICIT_INSTRUCTION** | Reformatted indentation |
| **CODE_IMPLICIT_ARBITRATION_INSTRUCTION** | +43 lines: Added clarifications, deterministic decision rules, calibration examples |
| **FILL_BILL_TYPE_INSTRUCTION** | +49 lines: Added detailed guidance, validation rules, extended examples |
| **TIN_NPI_TEMPLATE_INSTRUCTION** | +22 lines: Added provider/payer distinction clarity |
### Analysis of Prompt Changes
**Positive Impacts:**
- More explicit instructions reduce model ambiguity
- Quality checks and validation rules improve consistency
- Examples help calibrate model responses
- Clearer formatting improves readability
**Potential Concerns:**
- Expanded prompts increase token count (higher cache write cost initially)
- Some prompts may have reduced due to restructuring (LESSER_OF_DISTRIBUTION)
- Need to verify output quality hasn't degraded with expanded context
---
## 3. Cost Analysis
### Bedrock Pricing Applied
- **Input tokens:** $3.00/1M tokens
- **Output tokens:** $15.00/1M tokens
- **Cache read:** $0.30/1M tokens (90% savings)
- **Cache write:** $3.75/1M tokens (25% premium)
### Observed Results
| Metric | Value |
|--------|-------|
| Total Input Tokens | 844,370 |
| Total Output Tokens | 401,926 |
| Total Cache Read Tokens | 1,472,330 |
| Total Cache Write Tokens | 87,167 |
| **Cost Without Caching** | $12.98 |
| **Cost With Caching** | $9.33 |
| **Savings** | $3.65 (28.1%) |
### Per-Prompt Cost Impact (Estimated from previous testing)
| Prompt | Cost Reduction |
|--------|---------------|
| LESSER_OF_DISTRIBUTION | 44% (Best performer) |
| EXHIBIT_HEADER | 7% |
| CHECK_PROVIDER_NAME_MATCH | 5% |
| TIN_NPI_TEMPLATE (prompt_provider_info) | 5% |
| code_implicit_rag | -12% (Negative - needs review) |
---
## 4. Prompts Needing Attention
### Priority 1: Prompts Needing Padding (Under 1024 tokens)
These prompts are wired but not actually caching due to Bedrock's minimum threshold:
1. **prompt_lob_relationship** (408 tokens, 56 calls) - High impact
2. **code_implicit_special** (329 tokens, 30 calls) - Medium impact
3. **code_last_check** (179 tokens, 23 calls) - Medium impact
4. **FEE_SCHEDULE_BREAKOUT** (352 tokens, 18 calls) - Medium impact
5. **SPLIT_REIMB_DATES** (410 tokens, 13 calls)
6. **code_implicit_arbitration** (391 tokens, 12 calls)
7. **SPLIT_SERVICE_TERM** (1600 tokens, 10 calls) - Should be caching, investigate
8. **DATE_FIX** (322 tokens, 10 calls)
9. **DERIVED_TERM_DATE** (247 tokens, 10 calls)
10. **EXHIBIT_HEADER_DEDUP** (35 tokens, 10 calls)
### Priority 2: Investigate Negative Result
- **code_implicit_rag**: Shows -12% cost (increase) - The expanded prompt may be too verbose or structured differently causing cache misses
### Priority 3: Not Wired for Caching
- **prompt_hsc_single_field** (109 calls) - High volume, should consider wiring
---
## 5. Recommendations
1. **Pad short prompts to 1024 tokens**: Add contextual padding to prompts under the threshold, especially high-frequency ones like `prompt_lob_relationship` and `code_implicit_special`
2. **Investigate code_implicit_rag**: The negative cost result (-12%) suggests the expanded prompt may be causing cache fragmentation or misses
3. **Wire prompt_hsc_single_field**: With 109 calls, this is a good candidate for caching
4. **Monitor SPLIT_SERVICE_TERM**: At 1,600 tokens it should be caching but shows 0% cache hit rate - may be a wiring issue
5. **Run testbed comparison**: Use `src/testbed/testbed_metrics.py` to compare output quality between cached and non-cached prompts to ensure no regression
---
## 6. Files Analyzed
- `src/test-PROMPT-CALLS_10.csv` - API call logs with cache statistics
- `src/testbed-usethis.xlsx` - Testbed comparison file (binary, not readable)
- `src/prompts/prompt_templates.py` - Prompt definitions (649 lines changed)
- `documentation/prompt_caching_tracker_v2.csv` - Caching status tracker
---
## 7. Next Steps
1. Run testbed metrics comparison script to validate output quality
2. Add padding to short prompts identified above
3. Investigate and fix code_implicit_rag negative result
4. Wire prompt_hsc_single_field for caching
5. Debug SPLIT_SERVICE_TERM caching issue