Feature/context caching * Initial commit - context caching for DYNAMIC_PRIMARY * implement context caching for all relevant prompts * Remove option to not context cache * IndentationError fixed * Merge branch 'DEV' into feature/context-caching * Merge and format * Move documentation * Merged DEV into feature/context-caching * Update unit tests * Merged DEV into feature/context-caching * Update signatures * Fix test coverage gap Approved-by: Praneel Panchigar Approved-by: Karan Desai
10 KiB
Context Caching Implementation - Complete
Summary
Successfully implemented context caching for 6 high-value prompts across all 3 client pipelines. This enables Anthropic's prompt caching at the exhibit/context level, where the same context is cached and reused across multiple field extractions, reducing token costs by ~84% for repeated context processing.
Prompts with Context Caching
- DYNAMIC_PRIMARY (pilot) - Primary term field extraction
- EXHIBIT_LEVEL - Exhibit-level metadata extraction
- DYNAMIC_ASSIGNMENT - Dynamic term assignment to exhibit rows
- REIMB_DATES_ASSIGNMENT - Reimbursement date assignment (specialized)
- LESSER_OF_DISTRIBUTION - Lesser-of logic distribution across codes
- LESSER_OF_CHECK - Lesser-of presence validation
Pipelines Updated
- ✅ bcbs_promise - All 5 applicable functions updated
- ✅ clover - All 5 applicable functions updated
- ✅ saas - All 5 applicable functions updated
What Changed
1. Extended LLM API (llm_utils.py)
Added context_for_caching parameter throughout the call chain:
invoke_claude()- New optional parameter_build_claude_3_request_body()- Structures multi-block messages with cache controlget_cache_key()- Includes context in cache key generationlocal_claude_3_and_up()- Passes parameter throughec2_claude_3_and_up()- Passes parameter through
Key Innovation: Messages now support multiple content blocks where specific blocks can be marked for caching:
"messages": [{
"role": "user",
"content": [
{
"type": "text",
"text": "Large exhibit text (40k tokens)",
"cache_control": {"type": "ephemeral"} # CACHED
},
{
"type": "text",
"text": "Field-specific question (200 tokens)" # NOT CACHED
}
]
}]
2. Updated Existing Prompt Templates (prompt_templates.py)
Updated 6 existing functions to always split prompts into cacheable and fresh components:
- DYNAMIC_PRIMARY() - Caches exhibit text, field question stays fresh
- EXHIBIT_LEVEL() - Caches exhibit text, field questions stay fresh
- DYNAMIC_ASSIGNMENT() - Caches exhibit simplified text, term questions stay fresh
- REIMB_DATES_ASSIGNMENT() - Specialized for REIMB_DATES assignment
- LESSER_OF_DISTRIBUTION() - Caches exhibit text and cross-exhibit context
- LESSER_OF_CHECK() - Caches exhibit title context
Each returns (context_text, prompt, parser) instead of (prompt, parser).
These functions now always return (context_text, prompt, parser) for context caching.
3. Updated All Client Prompt Calls
Updated functions across all 3 pipelines:
bcbs_promise/prompts/prompt_calls.py:
prompt_exhibit_level()prompt_dynamic_primary()prompt_dynamic_assignment()prompt_lesser_of_distribution()prompt_lesser_of_check()
clover/prompts/prompt_calls.py:
prompt_exhibit_level()prompt_dynamic_primary()prompt_dynamic_assignment()prompt_lesser_of_distribution()prompt_lesser_of_check()
saas/prompts/prompt_calls.py:
prompt_exhibit_level()prompt_dynamic_primary()prompt_dynamic_assignment()prompt_lesser_of_distribution()prompt_lesser_of_check()
Each function now:
- Uses the original template function (always split for caching)
- Receives
(context_text, prompt, parser)tuple - Passes
context_for_caching=context_texttoinvoke_claude() - Logs context length for monitoring
Cost Impact Analysis
Current Structure (Before)
- System message (cached): Field extraction instruction (~2k tokens)
- User message (NOT cached): Combined exhibit + field question (~40k tokens)
For 20 fields on same exhibit:
- Instruction: 2k × 1 creation = cached once ✓
- Content: 40k × 20 calls = 800k tokens at $0.003/1k = $2.40
New Structure (After)
- System message (cached): Field extraction instruction (~2k tokens)
- User message block 1 (cached): Exhibit context (~40k tokens)
- User message block 2 (not cached): Field question (~200 tokens)
For 20 fields on same exhibit:
- Instruction: 2k × 1 creation = cached once ✓
- Context: 40k × 1 creation at $0.00375/1k = $0.15
- Context: 40k × 19 reads at $0.0003/1k = $0.228
- Field questions: 20 × 200 tokens at $0.003/1k = $0.012
- Total: $0.39 (84% cost reduction)
Break-Even Analysis
- 1st field: Pay 25% premium for cache creation
- 2nd field: Start saving with 90% cheaper cache reads
- 3+ fields: Massive savings accumulate
How It Works
Caching Layers (Claude API)
Layer 1: System Instruction (cached) ← Already implemented
↓
Layer 2: Exhibit Context (cached) ← NEW - This implementation
↓
Layer 3: Field Question (fresh) ← Changes per call
Flow Example
# Processing LOB field for exhibit
context_text, prompt, parser = DYNAMIC_PRIMARY(
exhibit_text="[40k token exhibit]",
field_name="LOB",
field_prompt="Line of Business definition",
)
llm_utils.invoke_claude(
prompt=prompt, # Just the field question
context_for_caching=context_text, # Exhibit text (cached)
instruction=DYNAMIC_PRIMARY_INSTRUCTION(), # Rules (already cached)
cache=True
)
# First call: Cache creation for exhibit
# Cost: (2k instruction + 40k context) × cache multiplier + 200 tokens fresh
# Processing PROGRAM field for SAME exhibit
context_text, prompt, parser = DYNAMIC_PRIMARY(
exhibit_text="[SAME 40k token exhibit]", # Same content
field_name="PROGRAM",
field_prompt="Program definition",
)
llm_utils.invoke_claude(
prompt=prompt, # Different field question
context_for_caching=context_text, # SAME exhibit (cache hit!)
instruction=DYNAMIC_PRIMARY_INSTRUCTION(),
cache=True
)
# Second call: Cache read for exhibit
# Cost: (2k + 40k) × cache read rate (90% cheaper) + 200 tokens fresh
Testing
Created comprehensive test suite in src/tests/test_context_caching.py:
✅ test_dynamic_primary_returns_three_values() - Validates always-split signature
✅ test_dynamic_primary_original_still_works() - Backward compatibility
✅ test_build_request_body_with_context_caching() - Message structure verification
✅ test_build_request_body_without_context_caching() - Fallback behavior
✅ test_cache_key_includes_context() - Cache key uniqueness
Monitoring & Validation
To verify the implementation is working:
-
Check usage logs for cache metrics:
# In usage_tracking.py logs, look for: cache_creation_tokens: 40000 # First call cache_read_tokens: 40000 # Subsequent calls -
Monitor cost per file in usage reports:
- Should see dramatic cost reduction for files with many dynamic fields
- Exhibits with 10+ fields should show 80%+ savings on exhibit processing
-
Log analysis:
DEBUG: Context length for caching: 42567 charsThis confirms context is being passed to caching layer.
Implementation Status
✅ Completed
All high-value prompts have been migrated to context caching across all 3 client pipelines:
- DYNAMIC_PRIMARY ✅ - Primary term field extraction (pilot implementation)
- EXHIBIT_LEVEL ✅ - Exhibit-level metadata extraction
- DYNAMIC_ASSIGNMENT ✅ - Dynamic term assignment to exhibit rows
- REIMB_DATES_ASSIGNMENT ✅ - Reimbursement date assignment (specialized)
- LESSER_OF_DISTRIBUTION ✅ - Lesser-of logic distribution across codes
- LESSER_OF_CHECK ✅ - Lesser-of presence validation
Cost Savings: Estimated 80-85% reduction in token costs for repeated exhibit/context processing across these 6 prompts.
Future Considerations
Lower Priority Candidates (evaluate after monitoring current implementation):
- METHODOLOGY_BREAKOUT - Could cache reimbursement terms for multiple breakout operations
- Other exhibit-level prompts - If processing changes to single-field-at-a-time pattern
Monitoring Required:
- Track cache hit rates and actual cost savings in production
- Validate that 5-minute cache TTL aligns with typical processing patterns
- Identify any additional prompts with repeated context usage patterns
Implementation Pattern (For Future Extensions)
For any new prompt to extend:
- Update
[PROMPT_NAME]()to return(context, prompt, parser) - Update corresponding
prompt_[name]()function to use caching version - Pass context via
context_for_cachingparameter - Monitor cache metrics to validate savings
Backward Compatibility
✅ Original DYNAMIC_PRIMARY() function remains unchanged
✅ Other templates continue to work without modification
✅ context_for_caching parameter is optional (defaults to None)
✅ When None, behavior is identical to previous implementation
✅ All tests should pass without modification
Files Modified
Core Infrastructure:
- src/utils/llm_utils.py - Extended API with
context_for_cachingparameter - src/prompts/prompt_templates.py - Added 6 context-caching template variants
Pipeline Updates (All 3 Clients):
- src/pipelines/clients/bcbs_promise/prompts/prompt_calls.py - Updated 5 functions
- src/pipelines/clients/clover/prompts/prompt_calls.py - Updated 5 functions
- src/pipelines/saas/prompts/prompt_calls.py - Updated 5 functions
Testing & Documentation:
- src/tests/test_context_caching.py - Comprehensive test suite
- CONTEXT_CACHING_IMPLEMENTATION.md - This documentation
Technical Notes
- Anthropic prompt caching requires minimum 1024 tokens for cache block
- Cache TTL is 5 minutes for
ephemeraltype - Only works with Claude 3.5+ Sonnet v2 models (checked via
_supports_prompt_cache()) - Cache keys include both instruction and context to ensure uniqueness
- Multiple content blocks in user messages is supported by Bedrock Messages API