- Updated all prompt_calls.py files (saas, bcbs_promise, clover) to use new parser pattern
- Updated prompt_templates.py with JSON format instructions
- Updated crosswalk_utils.py to handle JSON lists
- Updated json_utils.py with improved parsing
- Updated aarete_derived.py for JSON compatibility
- Updated tin_npi_funcs.py and qa_qc_utils.py for JSON parsing
- All changes align with Phase 4 completion of JSON standardization
- Removed all debug print statements from prompt_calls.py and one_to_n_funcs.py
- Updated PRD to reflect Phase 4 completion and recent bug fixes
- Documented REIMB_TERM normalization fixes and cleanup work
- Updated Phase 6 status to 25% complete (debug removal done)
- Fixed prompt_lesser_of_distribution to extract string from list when parser returns list
- Added debug print statements to trace REIMB_TERM format through processing pipeline
- Removed defensive normalization checks that are no longer needed upstream
- Added debug output in reimbursement_level and methodology_breakout_single_row to track data flow
- Updated progress table: Phase 4 is now 100% complete
- Marked model_evaluation_utils.py as complete (commit 13f3cbea)
- Verified testbed/QC files process internal data (out of scope)
- Updated remaining work section
- Updated overall status to Phase 4 of 6 Complete
- Fixed DYNAMIC_PRIMARY_TEXT: now unpacks (prompt, parser) tuple and uses parser
- Fixed REIMBURSEMENT_PRIMARY: corrected function signature (only takes context)
- Replaced extract_text_from_delimiters with JSON parser
- Replaced universal_json_load with parser from tuple
- Added proper instruction caching for both prompts
Other testbed/QC files verified: they process internal data formats (out of scope per PRD)
- Updated progress table: Phase 3 is now 100% complete
- Marked code_funcs.py as complete (commit a7a32b90)
- Updated remaining work section
- Added latest commit to commit history
- Line 498: Changed from universal_json_load to use _parser from FILL_BILL_TYPE tuple
- Line 856: Updated GROUPER_BREAKOUT call to unpack tuple and use grouper_parser
- Removed unused GROUPER_QUESTIONS variable (fields now in INSTRUCTION function for caching)
- All universal_json_load calls in code_funcs.py have been replaced
- Phase 3 is now 100% complete
- Updated EFFECTIVE_DT field: Replaced all pipe format examples (|date|) with JSON dictionary format
- Changed examples from 'return |3-1 , , 03|' to 'return "3-1 , , 03"'
- Updated final answer format from 'ENCLOSED in |pipes|' to JSON dict: {"EFFECTIVE_DT": "date"}
- Updated PROVIDER_NAME field: Replaced pipe-separated format with JSON list format
- Changed from 'separated by |' to JSON list: ["Provider A", "Provider B"]
- Updated to return JSON dict: {"PROVIDER_NAME": [...]} or {"PROVIDER_NAME": "single"}
- All pipe references removed from investment_prompts.json
- Updated PRD to document this change in Phase 2 checklist
- JSON file validated and syntax confirmed correct
- Create json_utils.py with parse_json_dict() and parse_json_list() functions
- Add comprehensive unit tests (49 tests, all passing)
- Add deprecation warnings to string_utils.py for pipe-delimited parsing
- Follows PRD Phase 1: Parser Infrastructure
This is the first step in replacing pipe-delimited LLM output format with
structured JSON format as per PRD_STANDARDIZE_LIST_FORMATS.md
Feature/standardize list formats
* updated unit test for code funcs
* pipeline fixes
* Merge branch 'main' into feature/standardize-list-formats
* black format
* removed duplicates
* Merge remote-tracking branch 'origin/main' into feature/standardize-list-formats
* updated prov other tin npi names to list
* prompt calls code made common for cleints and saas
* final format fixes
* removed pipes
* crosswalk fixes
* Merged main into feature/standardize-list-formats
* dubug dynamic fields
* dynamic format fixes
* dyanmic fields fixes
* removed pipes from the whole code
* fixed provider TIN NPI Name format issues
* bill code issue fix
* pipeline error fixes
* Merge branch 'feature/standardize-list-formats' of https://bitbucket.org/aarete/doczy.ai into feature/standardize-list-formats
* Merged main into feature/standardize-list-formats
* Black format
* pipeline errors fixed
* pipeline fixes page funcs and tinnpi funcs
* qa_qc_utils fix
Approved-by: Katon Minhas
Bugfix/generic dynamic jan26
* Refactor: Extract dynamic primary metrics logic to separate module
- Created src/testbed/dynamic_primary_metrics.py with all dynamic primary field analysis logic
- Moved tally-based analysis functions to new module for better modularity
- Updated testbed_utils.py to import and use functions from dynamic_primary_metrics
- Simplified testbed_metrics_dynamic_only.py to use new module
- Removed unused analyze_dynamic_primary_fields function (400+ lines)
- Added field filtering in match_rows to handle missing columns gracefully
- Minimal changes to testbed_utils.py (~100 lines vs 845 before)
* feat: Add debug print blocks and prompt changes for dynamic primary fields in 1:1 processing
This commit adds prompt changes for when we pass dynamic primary to 1:1 and comprehensive debug logging for dynamic primary fields when they are escalated to 1:1 processing via Hybrid Smart Chunking (HSC) or Full Context methods.
Changes:
- Added debug print blocks in hybrid_smart_chunking_funcs.py to log:
* Retrieval question, context chunks, and final prompt for HSC processing
* Raw LLM output (reasoning + answer) for dynamic primary fields
- Added debug print blocks in prompt_calls.py to log:
* Contract context, final prompt for Full Context processing
* Raw LLM output (reasoning + answer) for dynamic primary fields
- Added FULL_CONTEXT_DYNAMIC_PRIMARY_INSTRUCTION() to prompt_templates.py:
* Provides specific guidance for dynamic primary fields in full context
* Emphasizes extracting values only if they refer to entire contract
* Includes pipe-delimited formatting instructions
- Updated dynamic_fun…
* Ran black, for formatting
* Refine dynamic primary field prompts and add Duals auto-detection
- Enhanced DYNAMIC_ASSIGNMENT_INSTRUCTION with structural boundary rules
- Added exhibit header and subsection context hierarchy for LOB assignment
- Implemented Duals auto-detection in dynamic_primary discovery (Medicare+Medicaid -> Duals)
- Added Duals value formatting instruction to prevent oversimplification
- Commented out DUAL_LOB_CHECK for performance testing
- Removed verbose debug print blocks from HSC and Full Context processing
- Added focused debug print for dynamic assignment raw LLM output
* Remove debug print blocks from pipeline files
- Removed all debug print blocks from file_processing.py (8 blocks)
- Removed all debug print blocks from dynamic_funcs.py (8 blocks)
- Removed debug print block from prompt_calls.py (dynamic assignment)
- DUAL_LOB_CHECK remains commented as requested
- Total: 188 lines of debug code removed
* Merge remote-tracking branch 'origin/main' into bugfix/generic_dynamic_jan26
* Ran black for CI
* Remove update_lob_for_duals
* Merge branch 'DEV' into bugfix/generic_dynamic_jan26
Approved-by: Katon Minhas