623799f2a6a39c3c54371516b83ec7342d86188f
Merged in dev (pull request #1001) * Merged in feature/fixPlaceholder2 (pull request #979) Updated feature> dev gate to print variables of echo statements * Updated feature> dev gate to print variables of echo statement * Reverted placeholder changes Approved-by: Sujit Deokar * Merged in bugfix/prov_info_json_fixes (pull request #981) fix: fall back to PROVIDER_NAME in PROV_INFO_JSON when no TIN/NPI extracted * fix: fall back to PROVIDER_NAME in PROV_INFO_JSON when no TIN/NPI extracted When get_prov_info_json short-circuits due to no TIN/NPI regex matches, PROV_INFO_JSON was left as [] even when PROVIDER_NAME was successfully extracted via the one-to-one pipeline. This caused inconsistent output across contracts with the same provider — some files produced a NAME-only entry (via a false-positive regex hit triggering the LLM), others produced []. Reconcile at add_group_and_other, the first point where both extraction streams' results are available. When PROV_INFO_JSON is empty but PROVIDER_NAME is known, synthesize a NAME-only entry with IS_GROUP:"Y" and populate PROV_GROUP_NAME_FULL directly — skipping the provider_name_match_check LLM call since the match is tautological by construction. Adds 5 unit tests covering the str, list, already-populated, empty-name, and all-empty-lis… * Merged in feature/DAIP2-pacificsource-reimbursements-issues (pull request #982) Feature/DAIP2 pacificsource reimbursements issues * Tighten PREMIUM_TERM and DISCOUNT_TERM classifier prompts CARVEOUT_CHECK was misrouting table rate rows into special-case fields, dropping them from the reimbursement output: - "110% of CMS allowed" (base fee-schedule rates) was being classified as PREMIUM_TERM because the prompt treated "above 100% of reference" as an implicit premium. Seen on PacificSource Medicare_Attachment_A1 and A2 Facility contracts where Inpatient/Outpatient rows were missing (2556) or silently fell back to 100% fee-schedule (2557). - Per-service discount rates like "Progressive Lenses: 15% discount", "Contact Lenses: 2% discount", "Frame: 20% discount" were being classified as DISCOUNT_TERM because the prompt only required the word "discount" to appear. Seen on PacificSource Commercial_Attachment_A2 and A5 Professional contracts (2558, 2559). Prompts now require the literal keyword … * Merged in bugfix/filter-docusign-lines (pull request #983) remove docusign lines * remove docusign lines * add unit tests for clean_header_footer docusign/deleted_lines changes Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * black forrmatting Approved-by: Katon Minhas * Merged in feature/PC_logic_cleanup_output (pull request #984) Feature/PC logic cleanup output * Few tweaks PC_logics * Fixed orphan_ranking * black formatting * Changes on output field and ranking method * updated few hotfixes * black format fix Approved-by: Katon Minhas * Merged in hotfix/provider_name_group_fix (pull request #988) Hotfix/provider name group fix * Fixes done in GROUP column * black format fix Approved-by: Katon Minhas * Merged in bugfix/DAIP2-2524-carveout-code-optimization (pull request #980) CARVEOUT_CD issue fixed * CARVEOUT_CD issue fixed * pipeline error fixed * Merged dev into bugfix/DAIP2-2524-carveout-code-optimization * Merged dev into bugfix/DAIP2-2524-carveout-code-optimization * Merged dev into bugfix/DAIP2-2524-carveout-code-optimization * trigger cap issue fixed * trigger cap prompt updated * Merged dev into bugfix/DAIP2-2524-carveout-code-optimization Approved-by: Katon Minhas * Merged in bugfix/exhibit-smart-chunking-cost-improvements (pull request #985) Bugfix/exhibit smart chunking cost improvements * Add opt-in instrumentation for per-call token and row-count tracing Introduce src/utils/instrumentation.py (thread-safe CSV logger) and src/utils/instrumentation_context.py (ContextVar scope plus submit_with_context / map_with_context helpers for propagating context into ThreadPoolExecutor workers). Emit events at every Bedrock call in llm_utils.invoke_claude, including in-memory claude_cache hits, with full input/output/cache-read/cache-write token breakdown. Emit row-count events at each row-mutating stage in the one-to-N pipeline (clean_reimbursement_primary, filter_services_without_reimbursements, methodology_breakout, split_service_terms, carveout, dynamic_code_assignment, lesser_of_distribution, dynamic_assignment) and chunking / retrieval events in exhibit smart chunking (chunking_done, retrieval_done) plus exhibit lifecycle events (exhibit_start, exhibit_gate_skip, stage_… * Merged in hotfix/fileextension_issue (pull request #989) Hotfix/fileextension issue * fixed strip_ext issue * black format Approved-by: Katon Minhas * Merged in hotfix/exhibit-header-in-tables (pull request #987) Hotfix/exhibit header in tables * Merged in feature/FixplaceholderIssue (pull request #977) Remove curly braces from echo statements in dev->stg * Remove curly brances from echo statements in dev->stg * Removed curly braces in echo statements in feature-> dev gate Approved-by: Sujit Deokar * Merged in feature/standardized-services (pull request #958) Feature/standardized services * service term standardization * prompt update * prompt updates for standardization * only service standardization * new file * add supporting files and test scripts for standardization work Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * remove old files * Merge remote-tracking branch 'origin/dev' into feature/standardized-services * final fixes * Merged dev into feature/standardized-services * additional features * removed unwanted files * remove unwanted files * Merge branch 'dev' into feature/standardized-services * Merge remote-tr… * Merged in hotfix/postprocess_csv (pull request #990) Hotfix/postprocess csv * date issue_fix * black_format * Merged dev into hotfix/postprocess_csv Approved-by: Katon Minhas * Merged in bugfix/molina_ut_dynamic_primary (pull request #991) Bugfix/molina ut dynamic primary * Prompt changes for dynamic primary * Route cover-sheet-only files to ERRORS.csv instead of leaking phantom rows When every page of a contract was filtered out as a cover sheet / quick-review form, process_file silently returned a FILE_NAME-only DataFrame. Because the runner routes by checking for an "error" column, that file landed in RESULTS.csv as a near-empty row and no ERRORS.csv was generated for the run. - saas/file_processing.py: raise ValueError when text_dict is empty after cover-sheet filtering, so safe_process_file produces a proper error row. - runner.py: add _is_phantom_result defense-in-depth — promote any result with no extracted fields beyond FILE_NAME to error_results with error_type=PhantomSuccess. * Merged dev into bugfix/molina_ut_dynamic_primary * Tighten PRODUCT prompt: restrict to valid_values, prune LOB/PROGRAM examples * Merge branch 'bugfix/molina_ut_dynamic_primary' of bit… * Merged in bugfix/black-format (pull request #994) Black format for pipeline pass * Black format for pipeline pass * Merged in bugfix/DAIP2-2679-fix-nebraska-issues (pull request #995) incorrect inclusion of CPT4_PROC_CD fixed * incorrect inclusion of CPT4_PROC_CD fixed Approved-by: Katon Minhas * Merged in feature/DAIP2-2314-DAIP2-1687-hybrid (pull request #993) Feature/DAIP2-2314 DAIP2 1687 hybrid * remove -files from s3 prefix requirements * Resolve input paths * fix: VendorProcessor.process_file returns (df, None) tuple runner.safe_process_file unpacks the result as (cc_df, dashboard_df), so returning a single DataFrame caused every vendor/generic file to fail with "too many values to unpack (expected 2)" — Python iterates DataFrame columns during unpacking. Vendor pipelines have no dashboard variant; second slot is None and the existing `dashboard_result is not None` guard in runner.py already handles it. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * DAIP2-2314 + DAIP2-1687: pad DYNAMIC_PRIMARY + DYNAMIC_PRIMARY_ENTITY_CLASSIFICATION over 1024-token cache floor - Pad DYNAMIC_PRIMARY_INSTRUCTION with three new sections: [SCOPE BOUNDARIES], [SOURCE TEXT INTERPRETATION], [REASONING DISCIPLINE], plus a [WORKED EXAMPLES] block. Estimated tokens: 447 -> 1117 (Sonnet 4.5 … * Merged in bugfix/postprocessing_date_fix (pull request #996) Date formatting changes * Date formatting changes * Merged dev into bugfix/postprocessing_date_fix Approved-by: Katon Minhas * Merged in bugfix/parent-child-rank-orphan-uniqueness (pull request #999) PC_logic bugfix * PC_logic bugfix Approved-by: Katon Minhas * Merged in feature/document-index (pull request #1005) Feature/document index * Add Document Index preprocessing — Layers 1, 2, and 3 wiring Parse the Textract-emitted Document Index block at the top of each contract with a single cached LLM call (prompt_document_index) instead of one per-page call per page. Layer 2 verifies parsed entries via literal string match and structural regex sweep, escalating suspect pages back to the existing per-page path. Layer 1+2 failure triggers a full fallback to today's per-page flow. New symbols: - preprocessing_funcs.extract_document_index_block — regex slice of index prefix - preprocessing_funcs.verify_index_against_pages — structural verifier (plain dict return) - prompt_templates.DOCUMENT_INDEX_INSTRUCTION / DOCUMENT_INDEX — cached prompt pair - prompt_calls.prompt_document_index — LLM wrapper (usage_label DOCUMENT_INDEX_PARSE) - config: DOCUMENT_INDEX_PARSE_ENABLED and three threshold flags - instrumentation: DOCUMENT_INDEX_PARSE mapped to preprocessing segment one… * Merged in feature/active-rates (pull request #1004) Feature/active rates * initial commit * Merged in feature/FixplaceholderIssue (pull request #977) Remove curly braces from echo statements in dev->stg * Remove curly brances from echo statements in dev->stg * Removed curly braces in echo statements in feature-> dev gate Approved-by: Sujit Deokar * Merged in feature/standardized-services (pull request #958) Feature/standardized services * service term standardization * prompt update * prompt updates for standardization * only service standardization * new file * add supporting files and test scripts for standardization work Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * remove old files * Merge remote-tracking branch 'origin/dev' into feature/standardized-services * final fixes * Merged dev into feature/standardized-services * additional features * removed unwanted files * remove unwanted files * Merge branch 'dev' into feature/standardized-services * Merge remote-track… * Merged in feature/one-to-one-confidence-scoring (pull request #1002) Feature/one to one confidence scoring * T1 plumbing: capture per-field confidence + retrieved-chunk metadata for 1:1 HSC fields Prep work for the 1:1 confidence-scoring stage. No scoring logic yet — this just collects the inputs the next ticket (rule-based scorer) will consume. - ONE_TO_ONE_SINGLE_FIELD_TEMPLATE: ask the LLM for confidence (0.0-1.0), verdict (correct/uncertain/not_found), and supporting_snippet alongside the field value. Existing field parser passes the extra keys through unchanged. - prompt_hsc_single_field: now returns a 4-tuple (name, value, field, metadata) where metadata holds the confidence/verdict/snippet plus a lightweight summary of which chunks the LLM saw (count + ids). _extract_hsc_metadata is defensive: clamps out-of-range confidences, defaults a missing/garbage verdict, caps the snippet at 500 chars, returns _empty_hsc_metadata() on every bail-out path. - run_hybrid_smart_chunked_fields: opt… * Merged in bugfix/confidence-flagged-missing-field-col (pull request #1006) Fix KeyError in compute_flagged when a *_CONF column has no value sibling * Fix KeyError in compute_flagged when a *_CONF column has no value sibling Production hit a hard crash at the end of every run: KeyError: "['DYNAMIC_PRIMARY_ENTITIES'] not in index" src/qc_qa/confidence/summary.py:143 compute_flagged was iterating over *_CONF columns and unconditionally indexing the dataframe with both the FILE_NAME column and the stripped value column. That assumed every <FIELD>_CONF column has a sibling <FIELD> value column in final_df. That isn't always true: dynamic-primary features carry only the _CONF side (their value side is dropped by reorder_columns since it isn't in FIELD_FORMAT_MAPPING but its _CONF suffix matches the explicit _CONF carve-out). When the model produced a below-threshold score for one of these and the value column was absent, pandas .loc raised KeyError and the runner crashed. Fix: - Build the .loc colu… * Merged in feature/active-rates (pull request #1007) Feature/active rates * amendment intent tag * prompt update * Merge remote-tracking branch 'origin/dev' into feature/active-rates * Merge remote-tracking branch 'origin/dev' into feature/active-rates * Merge branch 'dev' into feature/active-rates * exhibit standardization updates * use AARETE_DERIVED_EXHIBIT_TITLE for intent * active rates stuff * Merge branch 'dev' into feature/active-rates * prompt update * amendment intent types * active rates logic update * unit tests * Merged dev into feature/active-rates * logic updates * Merge remote-tracking branch 'origin/dev' into feature/active-rates * Merge remote-tracking branch 'origin/dev' into feature/active-rates * updated instrumentation cost logs * cost loging * caching updates * Merge remote-tracking branch 'origin/dev' into feature/active-rates * Merge remote-tracking branch 'origin/dev' into feature/active-rates * logging fix * null check issue fixes * caching fix Approved-by… * Merged in bugfix/stg-to-main-prep (pull request #1012) Bugfix/stg to main prep * Merged in dev (pull request #1001) Dev * Revert premature merge of bugfix/retire_stale_client_file_processing PR #959 was merged into dev without approval. This reverts commits5e143c10,63f32c41,849aa626, and927abcaeto restore dev to its pre-merge state. The changes will be re-submitted via a new PR after proper review. * Merged in bugfix/retire_stale_client_file_processing (pull request #961) Return None for dashboard output when dashboard postprocessing is off * Return None for dashboard output when dashboard postprocessing is off FINAL_RESULT_DF_DASHBOARD was initialized as an empty DataFrame even when RUN_DASHBOARD_POSTPROCESSING was False, causing downstream code to needlessly process it (reorder_columns, etc). Now returns None when dashboard is not requested, matching the postprocess() contract. * Merged dev into bugfix/retire_stale_client_file_processing * Merge dev (with revert) into feature branch * Re-ap… * Merged in bugfix/sync-stg-into-dev-20260518 (pull request #1015) Merged in dev (pull request #1001) * Merged in dev (pull request #1001) Dev * Revert premature merge of bugfix/retire_stale_client_file_processing PR #959 was merged into dev without approval. This reverts commits5e143c10,63f32c41,849aa626, and927abcaeto restore dev to its pre-merge state. The changes will be re-submitted via a new PR after proper review. * Merged in bugfix/retire_stale_client_file_processing (pull request #961) Return None for dashboard output when dashboard postprocessing is off * Return None for dashboard output when dashboard postprocessing is off FINAL_RESULT_DF_DASHBOARD was initialized as an empty DataFrame even when RUN_DASHBOARD_POSTPROCESSING was False, causing downstream code to needlessly process it (reorder_columns, etc). Now returns None when dashboard is not requested, matching the postprocess() contract. * Merged dev into bugfix/retire_stale_client_file_processing * Merge dev (with revert) into fe… * Merged in dev (pull request #1016) Dev * Merged in bugfix/DAIP2-2823-generic-issue-fixes-one-to-one (pull request #1011) Bugfix/DAIP2-2823 generic issue fixes one to one * updated auto renewal ind prompt * updated CONTRACT_AMENDMENT_NUM prompt * Merged dev into bugfix/DAIP2-2823-generic-issue-fixes-one-to-one Approved-by: Praneel Panchigar Approved-by: Siddhant Medar * Merged in feature/DAIP2-2698-phase-3-program-product-lob-mapping (pull request #1000) Feature/DAIP2-2698 phase 3 program product lob mapping * added missing phase 2 modifications * added phase 3 modifications * Fixed acronym issues * black format fix * Merged dev into feature/DAIP2-2698-phase-3-program-product-lob-mapping * standardization fixes * Merged dev into feature/DAIP2-2698-phase-3-program-product-lob-mapping * added hyphenated suffix names fix * black format fix * standardization fix * dynamic primary fix * Merge Dev into feature/DAIP2-2698-phase-3-program-product-lob-mapping * updated Program Product Standardiza… Approved-by: Praneel Panchigar
Field Extraction Pipeline
Contract field extraction using LLMs.
Setup
Install uv (if not already installed)
# Ubuntu/WSL (recommended)
curl -LsSf https://astral.sh/uv/install.sh | sh
source ~/.bashrc # or restart terminal
# macOS
curl -LsSf https://astral.sh/uv/install.sh | sh
# Windows (PowerShell - native, not WSL)
powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"
Verify installation: uv --version
Install dependencies
uv sync
Running the Code
Use uv run with Python's module flag from the project root:
# Run the default SaaS pipeline
uv run python -m src.pipelines.saas.main
# Run with unified runner (supports client selection)
uv run python -m src.pipelines.runner --client saas --input-dir /path/to/input
# Run with specific client
uv run python -m src.pipelines.runner --client clover --input-dir /path/to/input
Note: Always use uv run to ensure correct virtual environment. Always use -m flag for proper imports.
Development
Before pushing, always run these checks to avoid breaking the CI pipeline:
# Format code
uv run black src/
# Type checking
uv run mypy src/
# Run tests
uv run pytest
All checks must pass before merging to main.
Project Structure
├── src/
│ ├── pipelines/ # Pipeline implementations
│ │ ├── runner.py # Unified CLI entry point with client routing
│ │ ├── saas/ # Default SaaS pipeline
│ │ │ └── main.py # Main entry point for SaaS
│ │ ├── shared/ # Shared pipeline components
│ │ │ ├── preprocessing/ # Document preprocessing
│ │ │ ├── extraction/ # Field extraction logic
│ │ │ └── postprocessing/ # Result postprocessing
│ │ └── clients/ # Client-specific overrides
│ │ └── clover/ # Clover client customizations
│ ├── core/ # Core utilities (registry, fieldset)
│ ├── constants/ # Constants, mappings, and field definitions
│ │ ├── mappings/ # Crosswalk JSON files
│ │ └── lists/ # Lookup lists
│ ├── prompts/ # LLM prompt templates
│ ├── utils/ # Shared utilities (IO, string, logging, etc.)
│ ├── codes/ # Medical code extraction utilities
│ ├── crosswalk/ # Crosswalk mapping logic
│ ├── embeddings/ # Pre-computed embeddings for code matching
│ ├── qc_qa/ # QC/QA validation pipeline
│ ├── parent_child/ # Parent-child relationship mapping
│ ├── document_classification/ # Document type classification (DTC)
│ └── tests/ # Unit tests
├── documentation/ # Project documentation
├── outputs/ # Pipeline output files (gitignored)
├── logs/ # Log files (gitignored)
└── pyproject.toml # Project dependencies (uv/pip)
Branching
Naming Conventions
| Type | Pattern | Use Case | Example |
|---|---|---|---|
| Feature | feature/<ticket>-<description> |
New functionality | feature/PROJ-123-add-export-csv |
| Bugfix | bugfix/<ticket>-<description> |
Bug fixes | bugfix/PROJ-456-fix-null-handling |
| Hotfix | hotfix/<ticket>-<description> |
Urgent production fixes | hotfix/PROJ-789-critical-parse-error |
| Test | test/<description> |
Testing/experimentation | test/lesser-table-caching-refactor |
| Release | release/<version> |
Release preparation | release/v1.2.0 |
Branch Guidelines
- Use lowercase with hyphens (kebab-case) for descriptions
- Include ticket number when applicable (e.g., JIRA, GitHub issue)
- Keep branch names concise but descriptive
- Delete branches after merging
Workflow
- Create branch from
main - Make changes and commit with clear messages
- Run checks before pushing:
uv run black src/ && uv run mypy src/ && uv run pytest - Create PR to
main - Ensure all CI checks pass
- Get code review approval
- Squash and merge
Description
Languages
Python
80.5%
Jupyter Notebook
13%
HCL
3.1%
HTML
1.6%
PLpgSQL
1.5%
Other
0.2%