Files
doczyai-pipelines/documentation/PERFORMANCE_ANALYSIS_GUIDE.md
T
Katon Minhas afb6d5185d Merged in feature/lesser-table-caching-refactor-hybrid (pull request #847)
Feature/lesser table caching refactor hybrid

* chore: Remove unused duplicate main.py from shared pipeline

* fix: Correct crosswalk paths in aarete_derived.py

* chore: Remove unused documentation files from fieldExtraction

* docs: Add documentation files to documentation folder

* docs: Update README with uv setup, expanded project structure, and branching conventions

* docs: Add uv installation steps with Ubuntu/WSL emphasis

* Enable prompt caching for all remaining LLM calls

- Add _INSTRUCTION() functions for: EXHIBIT_HEADER, EXHIBIT_LINKAGE,
  EXHIBIT_TITLE_MATCH, DATE_FIX, DERIVED_TERM_DATE, CHECK_PROVIDER_NAME_MATCH,
  SPECIAL_CASE_ASSIGNMENT
- Update all invoke_claude() calls in saas and clover pipelines to use
  cache=True with corresponding _INSTRUCTION() functions
- Add new instructions to get_cacheable_instructions() for cache warming
- Update tests for new instruction functions

Functions now using caching:
- prompt_exhibit_level
- prompt_exhibit_lesser (EXHIBIT_LEVEL_LESSER_OF)
- prompt_fee_schedule_breakout
- prompt_grouper_breakout
- prompt_special_case_assignment
- prompt_exhibit_linkage
- prompt_exhibit_header
- prompt_smart_chunked (ONE_TO_ONE templates)
- prompt_date_fix
- prompt_derived_term_date
- prompt_exhibit_title_match
- provider_name_match_check

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* Reorder

* feat: Add bcbs_promise client pipeline with OFFSET_TERM extraction

- Add new bcbs_promise client with HSC-based OFFSET_TERM field extraction
- Extract full paragraph text of offset/recoupment provisions from contracts
- Derive OFFSET_INDICATOR (Y/N) from OFFSET_TERM presence
- Fix reorder_columns to preserve extra columns not in COLUMN_ORDER
- Update QC/QA output path to outputs/qc_qa/

* fix: Update dev deps and test assertions for QC/QA output path

- Add pytest/pytest-mock to dev dependencies for mypy type checking
- Update test assertions to expect outputs/qc_qa instead of qa_qc_output

* style: Apply black formatting to prompt_templates.py

* Merge main, move scripts

* Archive some scripts

* update py version

* remove .py version file

* Remove ASCII characters

* Restore testbed code

* restore tracking

* Update testbed metrics

* Enable prompt caching for CODE_LAST_CHECK, FILL_BILL_TYPE, DUAL_LOB_CHECK, and GROUPER_BREAKOUT

- Add CODE_LAST_CHECK_INSTRUCTION() for service specificity classification
- Add FILL_BILL_TYPE_INSTRUCTION() for bill type code determination
- Add DUAL_LOB_CHECK_INSTRUCTION() for Medicare/Medicaid classification
- Update code_funcs.py to use caching for CODE_LAST_CHECK, FILL_BILL_TYPE, GROUPER_BREAKOUT
- Update postprocessing_funcs.py to use caching for DUAL_LOB_CHECK
- Add new instructions to get_cacheable_instructions() for cache warming
- Add unit tests for new instruction functions

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* Fix postprocessing_funcs to remove invalid columns

* Merge branch 'main' into feature/lesser-table-caching-refactor-hybrid

* Revert prompt caching changes from aed1b73c

* update formatting

* Update imports


Approved-by: Sha Brown
Approved-by: Praneel Panchigar
2026-01-26 16:52:55 +00:00

8.8 KiB
Raw Blame History

Performance Analysis Guide

Overview

This guide explains the timing instrumentation added to the pipeline and how to interpret the results to identify bottlenecks and parallelization opportunities.

Instrumented Components

1. Main Pipeline (main.py)

  • read_constants: Loading configuration and constants
  • read_input: Reading input files from S3 or local storage
  • warm_prompt_caches: Pre-warming Claude prompt caches
  • parallel_file_processing: Main parallel processing of all files (CRITICAL METRIC)
  • concat_results: Concatenating all file results into final dataframes
  • qc_qa_validation: Quality control and QA validation
  • write_outputs: Writing final outputs to S3/local storage
  • parent_child_mapping: Parent-child relationship mapping (if enabled)

2. Per-File Processing (file_processing.py)

Each file goes through these stages (tracked with context=filename):

  • {filename}.preprocess: Text cleaning, splitting, header/footer detection
  • {filename}.one_to_n_exhibit_chunking: Chunking document into exhibits
  • {filename}.one_to_n_extraction: Extracting one-to-many relationships (parallel exhibits)
  • {filename}.one_to_n_generate_reimb_ids: Generating reimbursement IDs
  • {filename}.one_to_one_extraction: Extracting one-to-one fields (see below)
  • {filename}.merge_one_to_one_into_one_to_n: Merging results
  • {filename}.code_processing: Code breakout and grouper processing
  • {filename}.postprocess: Final postprocessing and validation
  • {filename}.write_individual: Writing individual file results

3. One-to-One Processing (file_processing.py -> run_one_to_one_prompts)

  • {filename}.one_to_one.provider_info: Extracting TIN/NPI provider information
  • {filename}.one_to_one.hybrid_smart_chunking: RAG-based field extraction (CRITICAL)
  • {filename}.one_to_one.full_context: Full-context field extraction (fallback)

4. Hybrid Smart Chunking (RAG) (hybrid_smart_chunking_funcs.py)

This is often the bottleneck. Detailed timing:

  • {filename}.hsc.setup: Setting up embeddings, reranker, and text splitter
  • {filename}.hsc.preprocessing: Document preprocessing and chunking
  • {filename}.hsc.vectorstore_creation: Creating FAISS vectorstore + hybrid retriever
  • {filename}.hsc.parallel_field_processing: Processing all fields in parallel (max 5 workers)
  • Per-field timing (inside parallel processing):
    • {filename}.hsc.field.{field_name}.retrieval: BM25 + semantic retrieval + reranking
    • {filename}.hsc.field.{field_name}.llm_call: Claude API call for field extraction

How to Analyze Results

Step 1: Run Your Pipeline

poetry run python -m src.investment.main run_mode=local read_mode=local input_dir=test-multi write_to_s3=False state=MS-IL-NY-CA-TX-WI-AR-KY-TN-ME-OH-OR-FL-AZ-IN-NJ-MI max_workers=50 client=Centene-Parkland-Molina-Moda-Healthnet-Trillium

Step 2: Review Timing Summary

At the end of execution, you'll see:

  1. Total pipeline duration
  2. Hierarchical timing summary organized by component

Look for the output sections:

================================================================================
PIPELINE COMPLETE - Total time: XXX.XXs
================================================================================
HIERARCHICAL TIMING SUMMARY
================================================================================

Step 3: Identify Bottlenecks

Key Metrics to Check:

  1. parallel_file_processing total time

    • This should be the largest component
    • Compare to total pipeline time to see overhead
    • If this is much shorter than pipeline time: Look for overhead in setup, I/O, or postprocessing
  2. Per-file timings (look at individual file contexts)

    • Find the slowest files: Look for files with longest total time
    • Identify slow components within files: Check which stage takes longest
      • Is it one_to_one.hybrid_smart_chunking?
      • Is it one_to_n_extraction?
      • Is it postprocess?
  3. Hybrid Smart Chunking (HSC) breakdown

    • hsc.vectorstore_creation: Should be relatively quick (<5s per file)
    • hsc.parallel_field_processing: Often the bottleneck
      • Check if fields are processing in parallel (should see ~5 fields at a time)
      • Look at per-field retrieval vs llm_call times
    • hsc.field.{field_name}.llm_call: LLM latency (network + Claude processing)
      • High variance? Could be prompt cache issues or rate limiting
      • Consistently slow? Could be context size or model load
  4. One-to-N parallel processing

    • Check if exhibits and pages are processing in parallel
    • Look for imbalanced workload (some exhibits much slower than others)

Step 4: Optimization Opportunities

Scenario 1: HSC Vectorstore Creation is Slow

  • Problem: Document preprocessing/chunking is slow
  • Solution:
    • Consider caching vectorstores per document
    • Reduce chunk_size or chunk_overlap in HSC_CONFIG
    • Optimize preprocessing pipeline

Scenario 2: HSC LLM Calls are Slow

  • Problem: Too many sequential LLM calls or high latency
  • Current: max_workers=5 for field processing
  • Solution:
    • Increase max_workers in hybrid_smart_chunking_funcs.py:304 (currently 5)
    • Check Bedrock connection pool (now set to 60, should be sufficient)
    • Verify prompt caching is working (check cache hits in logs)

Scenario 3: HSC Retrieval/Reranking is Slow

  • Problem: RAG retrieval taking too long per field
  • Solution:
    • Reduce ENSEMBLE_TOP_K or RERANKER_TOP_K in RAG config
    • Consider using faster reranker model
    • Optimize BM25 index size

Scenario 4: Parallel File Processing Not Saturated

  • Problem: Not all max_workers are being used
  • Check:
    • Connection pool exhaustion (should be fixed with max_pool_connections=60)
    • Memory constraints (monitor memory usage)
    • I/O bottlenecks (local disk or S3 bandwidth)
  • Solution:
    • Increase max_workers carefully (currently 50)
    • Profile memory usage to find leaks

Scenario 5: One-to-N Exhibits Not Parallel

  • Current: max_workers=5 for exhibit processing, max_workers=5 for page processing
  • Solution:
    • Increase these limits in file_processing.py:244 (exhibits) and :192 (pages)
    • Balance against total file-level parallelism to avoid oversubscription

Scenario 6: Postprocessing is Slow

  • Problem: Code processing or postprocessing taking too long
  • Solution:
    • Profile specific postprocessing functions
    • Look for opportunities to batch operations
    • Consider caching expensive lookups

Current Parallelization Strategy

Three Levels of Parallelism:

  1. File Level (main.py)

    • max_workers=50 (from command line arg)
    • 50 files processed simultaneously
  2. Exhibit/Page Level (within each file)

    • Exhibits: max_workers=5 per file
    • Pages: max_workers=5 per exhibit
    • This means potentially 50 files × 5 exhibits × 5 pages = 1,250 threads (too many!)
    • Currently limited by nested executors
  3. Field Level (within HSC per file)

    • max_workers=5 for field extraction
    • 5 fields extracted in parallel per file

Total Theoretical Concurrency:

  • Bedrock API calls: Up to 60 concurrent (connection pool limit)
  • Thread pools: Nested executors create complex concurrency
  • Bottleneck: Likely the 60-connection Bedrock pool with 50 file workers

Recommendations After First Run

  1. Run with timing enabled and review the hierarchical summary
  2. Identify the top 3 slowest components by total time
  3. Check if parallelism is effective:
    • Are all max_workers being utilized?
    • Is there thread contention or blocking?
  4. Look for outliers: Files that take much longer than others
  5. Profile memory usage: Ensure no memory leaks with many threads
  6. Monitor Bedrock throttling: Check for rate limit errors in logs

Next Steps

Based on timing results, we can:

  1. Increase parallelism where bottlenecks are found
  2. Add caching for repeated operations (vectorstores, embeddings)
  3. Optimize prompts to reduce token count and latency
  4. Batch operations where sequential processing is required
  5. Add more granular timing to specific slow functions

Timing Log Format

Timing logs appear as:

INFO - ⏱️  read_constants: 0.50s
INFO - ⏱️  parallel_file_processing: 450.23s
DEBUG - ⏱️  file1.pdf.preprocess: 2.34s
DEBUG - ⏱️  file1.pdf.one_to_one.hybrid_smart_chunking: 45.67s
DEBUG - ⏱️  file1.pdf.hsc.setup: 0.12s
DEBUG - ⏱️  file1.pdf.hsc.vectorstore_creation: 3.45s
DEBUG - ⏱️  file1.pdf.hsc.parallel_field_processing: 40.12s
DEBUG - ⏱️  file1.pdf.hsc.field.EFFECTIVE_DT.retrieval: 0.89s
DEBUG - ⏱️  file1.pdf.hsc.field.EFFECTIVE_DT.llm_call: 2.34s

Use these to drill down into specific bottlenecks!