Merged in bugfix/parser-downstream-improvements (pull request #875)
Bugfix/parser downstream improvements * Refactor: Implement field-aware JSON parsers with centralized normalization This refactor introduces a robust system for normalizing LLM output based on field format mappings, ensuring consistent data types throughout the pipeline. Key Changes: - Add FIELD_FORMAT_MAPPING constant defining expected formats for all fields - Create format_normalization.py utility for type-aware normalization - Update json_utils.py parsers to accept field_names/field_name parameters - Refactor prompt_templates.py to use parser factories (_create_json_dict_parser, _create_json_list_parser) that bind field metadata for automatic normalization - Update prompt_calls.py to pass field names to parsers, eliminating redundant normalization logic - Remove parse_json_dict_or_list (unused, ambiguous function) - Simplify METHODOLOGY_BREAKOUT and REIMBURSEMENT_PRIMARY to use helper functions - Add comprehensive integration tests verifying normalization works end-to-end Benefits: - Single source of truth for field formats (FIELD_FORMAT_… * refactor: normalize helper prompt outputs at prompt_calls level - Update CARVEOUT_CHECK to use field-aware parser for CARVEOUT_CD normalization - Update LOB_RELATIONSHIP to normalize to string format in prompt_calls.py - Update SPLIT_REIMB_DATES to normalize date values in prompt_calls.py - Remove defensive normalization from one_to_n_funcs.py for LOB relationships - Remove manual normalization from split_reimb_dates() - values now normalized upstream - All helper prompts that populate fields now normalize at prompt_calls.py level - Downstream functions receive correctly formatted values without additional processing * refactor: remove band-aid normalization functions and migrate HSC to field-aware parsers - Update ONE_TO_ONE_SINGLE_FIELD_TEMPLATE to use field-aware parser with field_name parameter - Remove list wrapping logic in hybrid_smart_chunking_funcs (field-aware parser handles normalization) - Remove normalize_one_to_one_field_value and normalize_one_to_one_answers_dict from string_utils.py - Remove all debug print statements from HSC processing - Remove commented-out normalization calls from client-specific files (clover, bcbs_promise) - All normalization now handled exclusively through FIELD_FORMAT_MAPPING via field-aware parsers * Ran Black for formatting * Print Statements removed, more cleaning * Merge branch 'DEV' into bugfix/parser-downstream-improvements * refactor: combine FIELD_FORMAT_MAPPING into investment_columns.py - Merged field_format_mapping.py into investment_columns.py to create single source of truth - FIELD_FORMAT_MAPPING now ordered by COLUMN_ORDER (161 fields) - Added 5 missing fields from COLUMN_ORDER with default format types - Updated all imports across codebase to use investment_columns - Python dict preserves insertion order (3.7+), maintaining COLUMN_ORDER sequence - All tests passing (38 field-aware tests verified) * Deprecate COLUMN_ORDER, rely on Mapping only * Merged DEV into bugfix/parser-downstream-improvements Approved-by: Katon Minhas
This commit is contained in:
committed by
Katon Minhas
parent
4587042c43
commit
9b0a344b13
@@ -311,7 +311,9 @@ def generate_reimb_ids(df: pd.DataFrame) -> pd.DataFrame:
|
||||
return df_temp
|
||||
|
||||
|
||||
def reorder_columns(df: pd.DataFrame, column_order: list[str]) -> pd.DataFrame:
|
||||
def reorder_columns(
|
||||
df: pd.DataFrame, field_format_mapping: dict[str, str]
|
||||
) -> pd.DataFrame:
|
||||
"""
|
||||
Reorders the columns of the DataFrame based on the given column order.
|
||||
Steps:
|
||||
@@ -326,6 +328,9 @@ def reorder_columns(df: pd.DataFrame, column_order: list[str]) -> pd.DataFrame:
|
||||
Returns:
|
||||
pd.DataFrame: The reordered DataFrame.
|
||||
"""
|
||||
# Get a list of column names, preserving order
|
||||
column_order = list(field_format_mapping.keys())
|
||||
|
||||
# Create a copy to avoid fragmentation
|
||||
df_copy = df.copy()
|
||||
|
||||
|
||||
Reference in New Issue
Block a user