afb6d5185d
Feature/lesser table caching refactor hybrid * chore: Remove unused duplicate main.py from shared pipeline * fix: Correct crosswalk paths in aarete_derived.py * chore: Remove unused documentation files from fieldExtraction * docs: Add documentation files to documentation folder * docs: Update README with uv setup, expanded project structure, and branching conventions * docs: Add uv installation steps with Ubuntu/WSL emphasis * Enable prompt caching for all remaining LLM calls - Add _INSTRUCTION() functions for: EXHIBIT_HEADER, EXHIBIT_LINKAGE, EXHIBIT_TITLE_MATCH, DATE_FIX, DERIVED_TERM_DATE, CHECK_PROVIDER_NAME_MATCH, SPECIAL_CASE_ASSIGNMENT - Update all invoke_claude() calls in saas and clover pipelines to use cache=True with corresponding _INSTRUCTION() functions - Add new instructions to get_cacheable_instructions() for cache warming - Update tests for new instruction functions Functions now using caching: - prompt_exhibit_level - prompt_exhibit_lesser (EXHIBIT_LEVEL_LESSER_OF) - prompt_fee_schedule_breakout - prompt_grouper_breakout - prompt_special_case_assignment - prompt_exhibit_linkage - prompt_exhibit_header - prompt_smart_chunked (ONE_TO_ONE templates) - prompt_date_fix - prompt_derived_term_date - prompt_exhibit_title_match - provider_name_match_check 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * Reorder * feat: Add bcbs_promise client pipeline with OFFSET_TERM extraction - Add new bcbs_promise client with HSC-based OFFSET_TERM field extraction - Extract full paragraph text of offset/recoupment provisions from contracts - Derive OFFSET_INDICATOR (Y/N) from OFFSET_TERM presence - Fix reorder_columns to preserve extra columns not in COLUMN_ORDER - Update QC/QA output path to outputs/qc_qa/ * fix: Update dev deps and test assertions for QC/QA output path - Add pytest/pytest-mock to dev dependencies for mypy type checking - Update test assertions to expect outputs/qc_qa instead of qa_qc_output * style: Apply black formatting to prompt_templates.py * Merge main, move scripts * Archive some scripts * update py version * remove .py version file * Remove ASCII characters * Restore testbed code * restore tracking * Update testbed metrics * Enable prompt caching for CODE_LAST_CHECK, FILL_BILL_TYPE, DUAL_LOB_CHECK, and GROUPER_BREAKOUT - Add CODE_LAST_CHECK_INSTRUCTION() for service specificity classification - Add FILL_BILL_TYPE_INSTRUCTION() for bill type code determination - Add DUAL_LOB_CHECK_INSTRUCTION() for Medicare/Medicaid classification - Update code_funcs.py to use caching for CODE_LAST_CHECK, FILL_BILL_TYPE, GROUPER_BREAKOUT - Update postprocessing_funcs.py to use caching for DUAL_LOB_CHECK - Add new instructions to get_cacheable_instructions() for cache warming - Add unit tests for new instruction functions 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * Fix postprocessing_funcs to remove invalid columns * Merge branch 'main' into feature/lesser-table-caching-refactor-hybrid * Revert prompt caching changes from aed1b73c * update formatting * Update imports Approved-by: Sha Brown Approved-by: Praneel Panchigar
8.8 KiB
8.8 KiB
Performance Analysis Guide
Overview
This guide explains the timing instrumentation added to the pipeline and how to interpret the results to identify bottlenecks and parallelization opportunities.
Instrumented Components
1. Main Pipeline (main.py)
read_constants: Loading configuration and constantsread_input: Reading input files from S3 or local storagewarm_prompt_caches: Pre-warming Claude prompt cachesparallel_file_processing: Main parallel processing of all files (CRITICAL METRIC)concat_results: Concatenating all file results into final dataframesqc_qa_validation: Quality control and QA validationwrite_outputs: Writing final outputs to S3/local storageparent_child_mapping: Parent-child relationship mapping (if enabled)
2. Per-File Processing (file_processing.py)
Each file goes through these stages (tracked with context=filename):
{filename}.preprocess: Text cleaning, splitting, header/footer detection{filename}.one_to_n_exhibit_chunking: Chunking document into exhibits{filename}.one_to_n_extraction: Extracting one-to-many relationships (parallel exhibits){filename}.one_to_n_generate_reimb_ids: Generating reimbursement IDs{filename}.one_to_one_extraction: Extracting one-to-one fields (see below){filename}.merge_one_to_one_into_one_to_n: Merging results{filename}.code_processing: Code breakout and grouper processing{filename}.postprocess: Final postprocessing and validation{filename}.write_individual: Writing individual file results
3. One-to-One Processing (file_processing.py -> run_one_to_one_prompts)
{filename}.one_to_one.provider_info: Extracting TIN/NPI provider information{filename}.one_to_one.hybrid_smart_chunking: RAG-based field extraction (CRITICAL){filename}.one_to_one.full_context: Full-context field extraction (fallback)
4. Hybrid Smart Chunking (RAG) (hybrid_smart_chunking_funcs.py)
This is often the bottleneck. Detailed timing:
{filename}.hsc.setup: Setting up embeddings, reranker, and text splitter{filename}.hsc.preprocessing: Document preprocessing and chunking{filename}.hsc.vectorstore_creation: Creating FAISS vectorstore + hybrid retriever{filename}.hsc.parallel_field_processing: Processing all fields in parallel (max 5 workers)- Per-field timing (inside parallel processing):
{filename}.hsc.field.{field_name}.retrieval: BM25 + semantic retrieval + reranking{filename}.hsc.field.{field_name}.llm_call: Claude API call for field extraction
How to Analyze Results
Step 1: Run Your Pipeline
poetry run python -m src.investment.main run_mode=local read_mode=local input_dir=test-multi write_to_s3=False state=MS-IL-NY-CA-TX-WI-AR-KY-TN-ME-OH-OR-FL-AZ-IN-NJ-MI max_workers=50 client=Centene-Parkland-Molina-Moda-Healthnet-Trillium
Step 2: Review Timing Summary
At the end of execution, you'll see:
- Total pipeline duration
- Hierarchical timing summary organized by component
Look for the output sections:
================================================================================
PIPELINE COMPLETE - Total time: XXX.XXs
================================================================================
HIERARCHICAL TIMING SUMMARY
================================================================================
Step 3: Identify Bottlenecks
Key Metrics to Check:
-
parallel_file_processingtotal time- This should be the largest component
- Compare to total pipeline time to see overhead
- If this is much shorter than pipeline time: Look for overhead in setup, I/O, or postprocessing
-
Per-file timings (look at individual file contexts)
- Find the slowest files: Look for files with longest total time
- Identify slow components within files: Check which stage takes longest
- Is it
one_to_one.hybrid_smart_chunking? - Is it
one_to_n_extraction? - Is it
postprocess?
- Is it
-
Hybrid Smart Chunking (HSC) breakdown
hsc.vectorstore_creation: Should be relatively quick (<5s per file)hsc.parallel_field_processing: Often the bottleneck- Check if fields are processing in parallel (should see ~5 fields at a time)
- Look at per-field retrieval vs llm_call times
hsc.field.{field_name}.llm_call: LLM latency (network + Claude processing)- High variance? Could be prompt cache issues or rate limiting
- Consistently slow? Could be context size or model load
-
One-to-N parallel processing
- Check if exhibits and pages are processing in parallel
- Look for imbalanced workload (some exhibits much slower than others)
Step 4: Optimization Opportunities
Scenario 1: HSC Vectorstore Creation is Slow
- Problem: Document preprocessing/chunking is slow
- Solution:
- Consider caching vectorstores per document
- Reduce chunk_size or chunk_overlap in HSC_CONFIG
- Optimize preprocessing pipeline
Scenario 2: HSC LLM Calls are Slow
- Problem: Too many sequential LLM calls or high latency
- Current: max_workers=5 for field processing
- Solution:
- Increase
max_workersinhybrid_smart_chunking_funcs.py:304(currently 5) - Check Bedrock connection pool (now set to 60, should be sufficient)
- Verify prompt caching is working (check cache hits in logs)
- Increase
Scenario 3: HSC Retrieval/Reranking is Slow
- Problem: RAG retrieval taking too long per field
- Solution:
- Reduce
ENSEMBLE_TOP_KorRERANKER_TOP_Kin RAG config - Consider using faster reranker model
- Optimize BM25 index size
- Reduce
Scenario 4: Parallel File Processing Not Saturated
- Problem: Not all
max_workersare being used - Check:
- Connection pool exhaustion (should be fixed with
max_pool_connections=60) - Memory constraints (monitor memory usage)
- I/O bottlenecks (local disk or S3 bandwidth)
- Connection pool exhaustion (should be fixed with
- Solution:
- Increase
max_workerscarefully (currently 50) - Profile memory usage to find leaks
- Increase
Scenario 5: One-to-N Exhibits Not Parallel
- Current: max_workers=5 for exhibit processing, max_workers=5 for page processing
- Solution:
- Increase these limits in
file_processing.py:244(exhibits) and:192(pages) - Balance against total file-level parallelism to avoid oversubscription
- Increase these limits in
Scenario 6: Postprocessing is Slow
- Problem: Code processing or postprocessing taking too long
- Solution:
- Profile specific postprocessing functions
- Look for opportunities to batch operations
- Consider caching expensive lookups
Current Parallelization Strategy
Three Levels of Parallelism:
-
File Level (main.py)
max_workers=50(from command line arg)- 50 files processed simultaneously
-
Exhibit/Page Level (within each file)
- Exhibits:
max_workers=5per file - Pages:
max_workers=5per exhibit - This means potentially 50 files × 5 exhibits × 5 pages = 1,250 threads (too many!)
- Currently limited by nested executors
- Exhibits:
-
Field Level (within HSC per file)
max_workers=5for field extraction- 5 fields extracted in parallel per file
Total Theoretical Concurrency:
- Bedrock API calls: Up to 60 concurrent (connection pool limit)
- Thread pools: Nested executors create complex concurrency
- Bottleneck: Likely the 60-connection Bedrock pool with 50 file workers
Recommendations After First Run
- Run with timing enabled and review the hierarchical summary
- Identify the top 3 slowest components by total time
- Check if parallelism is effective:
- Are all max_workers being utilized?
- Is there thread contention or blocking?
- Look for outliers: Files that take much longer than others
- Profile memory usage: Ensure no memory leaks with many threads
- Monitor Bedrock throttling: Check for rate limit errors in logs
Next Steps
Based on timing results, we can:
- Increase parallelism where bottlenecks are found
- Add caching for repeated operations (vectorstores, embeddings)
- Optimize prompts to reduce token count and latency
- Batch operations where sequential processing is required
- Add more granular timing to specific slow functions
Timing Log Format
Timing logs appear as:
INFO - ⏱️ read_constants: 0.50s
INFO - ⏱️ parallel_file_processing: 450.23s
DEBUG - ⏱️ file1.pdf.preprocess: 2.34s
DEBUG - ⏱️ file1.pdf.one_to_one.hybrid_smart_chunking: 45.67s
DEBUG - ⏱️ file1.pdf.hsc.setup: 0.12s
DEBUG - ⏱️ file1.pdf.hsc.vectorstore_creation: 3.45s
DEBUG - ⏱️ file1.pdf.hsc.parallel_field_processing: 40.12s
DEBUG - ⏱️ file1.pdf.hsc.field.EFFECTIVE_DT.retrieval: 0.89s
DEBUG - ⏱️ file1.pdf.hsc.field.EFFECTIVE_DT.llm_call: 2.34s
Use these to drill down into specific bottlenecks!