Files
doczyai-pipelines/documentation/PROGRAM_PRODUCT_CANONICAL_DERIVATION.md
T
Katon Minhas 623799f2a6 Merged in stg (pull request #1003)
Merged in dev (pull request #1001)

* Merged in feature/fixPlaceholder2 (pull request #979)

Updated feature> dev gate to print variables of echo statements

* Updated feature> dev gate to print variables of echo statement

* Reverted placeholder changes


Approved-by: Sujit Deokar

* Merged in bugfix/prov_info_json_fixes (pull request #981)

fix: fall back to PROVIDER_NAME in PROV_INFO_JSON when no TIN/NPI extracted

* fix: fall back to PROVIDER_NAME in PROV_INFO_JSON when no TIN/NPI extracted

When get_prov_info_json short-circuits due to no TIN/NPI regex matches,
PROV_INFO_JSON was left as [] even when PROVIDER_NAME was successfully
extracted via the one-to-one pipeline. This caused inconsistent output
across contracts with the same provider — some files produced a NAME-only
entry (via a false-positive regex hit triggering the LLM), others produced [].

Reconcile at add_group_and_other, the first point where both extraction
streams' results are available. When PROV_INFO_JSON is empty but
PROVIDER_NAME is known, synthesize a NAME-only entry with IS_GROUP:"Y" and
populate PROV_GROUP_NAME_FULL directly — skipping the provider_name_match_check
LLM call since the match is tautological by construction.

Adds 5 unit tests covering the str, list, already-populated, empty-name,
and all-empty-lis…
* Merged in feature/DAIP2-pacificsource-reimbursements-issues (pull request #982)

Feature/DAIP2 pacificsource reimbursements issues

* Tighten PREMIUM_TERM and DISCOUNT_TERM classifier prompts

CARVEOUT_CHECK was misrouting table rate rows into special-case fields,
dropping them from the reimbursement output:

- "110% of CMS allowed" (base fee-schedule rates) was being classified as
  PREMIUM_TERM because the prompt treated "above 100% of reference" as an
  implicit premium. Seen on PacificSource Medicare_Attachment_A1 and A2
  Facility contracts where Inpatient/Outpatient rows were missing (2556)
  or silently fell back to 100% fee-schedule (2557).

- Per-service discount rates like "Progressive Lenses: 15% discount",
  "Contact Lenses: 2% discount", "Frame: 20% discount" were being
  classified as DISCOUNT_TERM because the prompt only required the word
  "discount" to appear. Seen on PacificSource Commercial_Attachment_A2
  and A5 Professional contracts (2558, 2559).

Prompts now require the literal keyword …
* Merged in bugfix/filter-docusign-lines (pull request #983)

remove docusign lines

* remove docusign lines

* add unit tests for clean_header_footer docusign/deleted_lines changes

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* black forrmatting


Approved-by: Katon Minhas

* Merged in feature/PC_logic_cleanup_output (pull request #984)

Feature/PC logic cleanup output

* Few tweaks PC_logics

* Fixed orphan_ranking

* black formatting

* Changes on output field and ranking method

* updated few hotfixes

* black format fix


Approved-by: Katon Minhas

* Merged in hotfix/provider_name_group_fix (pull request #988)

Hotfix/provider name group fix

* Fixes done in GROUP column

* black format fix


Approved-by: Katon Minhas

* Merged in bugfix/DAIP2-2524-carveout-code-optimization (pull request #980)

CARVEOUT_CD issue fixed

* CARVEOUT_CD issue fixed

* pipeline error fixed

* Merged dev into bugfix/DAIP2-2524-carveout-code-optimization

* Merged dev into bugfix/DAIP2-2524-carveout-code-optimization

* Merged dev into bugfix/DAIP2-2524-carveout-code-optimization

* trigger cap issue fixed

* trigger cap prompt updated

* Merged dev into bugfix/DAIP2-2524-carveout-code-optimization


Approved-by: Katon Minhas

* Merged in bugfix/exhibit-smart-chunking-cost-improvements (pull request #985)

Bugfix/exhibit smart chunking cost improvements

* Add opt-in instrumentation for per-call token and row-count tracing

Introduce src/utils/instrumentation.py (thread-safe CSV logger) and
src/utils/instrumentation_context.py (ContextVar scope plus
submit_with_context / map_with_context helpers for propagating context
into ThreadPoolExecutor workers).

Emit events at every Bedrock call in llm_utils.invoke_claude, including
in-memory claude_cache hits, with full input/output/cache-read/cache-write
token breakdown. Emit row-count events at each row-mutating stage in the
one-to-N pipeline (clean_reimbursement_primary,
filter_services_without_reimbursements, methodology_breakout,
split_service_terms, carveout, dynamic_code_assignment,
lesser_of_distribution, dynamic_assignment) and chunking / retrieval
events in exhibit smart chunking (chunking_done, retrieval_done) plus
exhibit lifecycle events (exhibit_start, exhibit_gate_skip,
stage_…
* Merged in hotfix/fileextension_issue (pull request #989)

Hotfix/fileextension issue

* fixed strip_ext issue

* black format


Approved-by: Katon Minhas

* Merged in hotfix/exhibit-header-in-tables (pull request #987)

Hotfix/exhibit header in tables

* Merged in feature/FixplaceholderIssue (pull request #977)

Remove curly braces from echo statements in dev->stg

* Remove curly brances from echo statements in dev->stg

* Removed curly braces in echo statements in feature-> dev gate


Approved-by: Sujit Deokar

* Merged in feature/standardized-services (pull request #958)

Feature/standardized services

* service term standardization

* prompt update

* prompt updates for standardization

* only service standardization

* new file

* add supporting files and test scripts for standardization work

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* remove old files

* Merge remote-tracking branch 'origin/dev' into feature/standardized-services

* final fixes

* Merged dev into feature/standardized-services

* additional features

* removed  unwanted files

* remove unwanted files

* Merge branch 'dev' into feature/standardized-services

* Merge remote-tr…
* Merged in hotfix/postprocess_csv (pull request #990)

Hotfix/postprocess csv

* date issue_fix

* black_format

* Merged dev into hotfix/postprocess_csv


Approved-by: Katon Minhas

* Merged in bugfix/molina_ut_dynamic_primary (pull request #991)

Bugfix/molina ut dynamic primary

* Prompt changes for dynamic primary

* Route cover-sheet-only files to ERRORS.csv instead of leaking phantom rows

When every page of a contract was filtered out as a cover sheet / quick-review
form, process_file silently returned a FILE_NAME-only DataFrame. Because the
runner routes by checking for an "error" column, that file landed in
RESULTS.csv as a near-empty row and no ERRORS.csv was generated for the run.

- saas/file_processing.py: raise ValueError when text_dict is empty after
  cover-sheet filtering, so safe_process_file produces a proper error row.
- runner.py: add _is_phantom_result defense-in-depth — promote any result
  with no extracted fields beyond FILE_NAME to error_results with
  error_type=PhantomSuccess.

* Merged dev into bugfix/molina_ut_dynamic_primary

* Tighten PRODUCT prompt: restrict to valid_values, prune LOB/PROGRAM examples

* Merge branch 'bugfix/molina_ut_dynamic_primary' of bit…
* Merged in bugfix/black-format (pull request #994)

Black format for pipeline pass

* Black format for pipeline pass

* Merged in bugfix/DAIP2-2679-fix-nebraska-issues (pull request #995)

incorrect inclusion of CPT4_PROC_CD fixed

* incorrect inclusion of CPT4_PROC_CD fixed


Approved-by: Katon Minhas

* Merged in feature/DAIP2-2314-DAIP2-1687-hybrid (pull request #993)

Feature/DAIP2-2314 DAIP2 1687 hybrid

* remove -files from s3 prefix requirements

* Resolve input paths

* fix: VendorProcessor.process_file returns (df, None) tuple

runner.safe_process_file unpacks the result as (cc_df, dashboard_df), so
returning a single DataFrame caused every vendor/generic file to fail with
"too many values to unpack (expected 2)" — Python iterates DataFrame columns
during unpacking. Vendor pipelines have no dashboard variant; second slot is
None and the existing `dashboard_result is not None` guard in runner.py
already handles it.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* DAIP2-2314 + DAIP2-1687: pad DYNAMIC_PRIMARY + DYNAMIC_PRIMARY_ENTITY_CLASSIFICATION over 1024-token cache floor

- Pad DYNAMIC_PRIMARY_INSTRUCTION with three new sections: [SCOPE BOUNDARIES], [SOURCE TEXT INTERPRETATION], [REASONING DISCIPLINE], plus a [WORKED EXAMPLES] block. Estimated tokens: 447 -> 1117 (Sonnet 4.5 …
* Merged in bugfix/postprocessing_date_fix (pull request #996)

Date formatting changes

* Date formatting changes

* Merged dev into bugfix/postprocessing_date_fix


Approved-by: Katon Minhas

* Merged in bugfix/parent-child-rank-orphan-uniqueness (pull request #999)

PC_logic bugfix

* PC_logic bugfix


Approved-by: Katon Minhas

* Merged in feature/document-index (pull request #1005)

Feature/document index

* Add Document Index preprocessing — Layers 1, 2, and 3 wiring

Parse the Textract-emitted Document Index block at the top of each contract
with a single cached LLM call (prompt_document_index) instead of one per-page
call per page. Layer 2 verifies parsed entries via literal string match and
structural regex sweep, escalating suspect pages back to the existing per-page
path. Layer 1+2 failure triggers a full fallback to today's per-page flow.

New symbols:
- preprocessing_funcs.extract_document_index_block — regex slice of index prefix
- preprocessing_funcs.verify_index_against_pages — structural verifier (plain dict return)
- prompt_templates.DOCUMENT_INDEX_INSTRUCTION / DOCUMENT_INDEX — cached prompt pair
- prompt_calls.prompt_document_index — LLM wrapper (usage_label DOCUMENT_INDEX_PARSE)
- config: DOCUMENT_INDEX_PARSE_ENABLED and three threshold flags
- instrumentation: DOCUMENT_INDEX_PARSE mapped to preprocessing segment

one…
* Merged in feature/active-rates (pull request #1004)

Feature/active rates

* initial commit

* Merged in feature/FixplaceholderIssue (pull request #977)

Remove curly braces from echo statements in dev->stg

* Remove curly brances from echo statements in dev->stg

* Removed curly braces in echo statements in feature-> dev gate


Approved-by: Sujit Deokar

* Merged in feature/standardized-services (pull request #958)

Feature/standardized services

* service term standardization

* prompt update

* prompt updates for standardization

* only service standardization

* new file

* add supporting files and test scripts for standardization work

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* remove old files

* Merge remote-tracking branch 'origin/dev' into feature/standardized-services

* final fixes

* Merged dev into feature/standardized-services

* additional features

* removed  unwanted files

* remove unwanted files

* Merge branch 'dev' into feature/standardized-services

* Merge remote-track…
* Merged in feature/one-to-one-confidence-scoring (pull request #1002)

Feature/one to one confidence scoring

* T1 plumbing: capture per-field confidence + retrieved-chunk metadata for 1:1 HSC fields

Prep work for the 1:1 confidence-scoring stage. No scoring logic yet — this
just collects the inputs the next ticket (rule-based scorer) will consume.

- ONE_TO_ONE_SINGLE_FIELD_TEMPLATE: ask the LLM for confidence (0.0-1.0),
  verdict (correct/uncertain/not_found), and supporting_snippet alongside
  the field value. Existing field parser passes the extra keys through
  unchanged.
- prompt_hsc_single_field: now returns a 4-tuple (name, value, field,
  metadata) where metadata holds the confidence/verdict/snippet plus a
  lightweight summary of which chunks the LLM saw (count + ids).
  _extract_hsc_metadata is defensive: clamps out-of-range confidences,
  defaults a missing/garbage verdict, caps the snippet at 500 chars,
  returns _empty_hsc_metadata() on every bail-out path.
- run_hybrid_smart_chunked_fields: opt…
* Merged in bugfix/confidence-flagged-missing-field-col (pull request #1006)

Fix KeyError in compute_flagged when a *_CONF column has no value sibling

* Fix KeyError in compute_flagged when a *_CONF column has no value sibling

Production hit a hard crash at the end of every run:

    KeyError: "['DYNAMIC_PRIMARY_ENTITIES'] not in index"
    src/qc_qa/confidence/summary.py:143

compute_flagged was iterating over *_CONF columns and unconditionally
indexing the dataframe with both the FILE_NAME column and the stripped
value column. That assumed every <FIELD>_CONF column has a sibling
<FIELD> value column in final_df. That isn't always true: dynamic-primary
features carry only the _CONF side (their value side is dropped by
reorder_columns since it isn't in FIELD_FORMAT_MAPPING but its _CONF
suffix matches the explicit _CONF carve-out). When the model produced a
below-threshold score for one of these and the value column was absent,
pandas .loc raised KeyError and the runner crashed.

Fix:
  - Build the .loc colu…
* Merged in feature/active-rates (pull request #1007)

Feature/active rates

* amendment intent tag

* prompt update

* Merge remote-tracking branch 'origin/dev' into feature/active-rates

* Merge remote-tracking branch 'origin/dev' into feature/active-rates

* Merge branch 'dev' into feature/active-rates

* exhibit standardization updates

* use AARETE_DERIVED_EXHIBIT_TITLE for intent

* active rates stuff

* Merge branch 'dev' into feature/active-rates

* prompt update

* amendment intent types

* active rates logic update

* unit tests

* Merged dev into feature/active-rates

* logic updates

* Merge remote-tracking branch 'origin/dev' into feature/active-rates

* Merge remote-tracking branch 'origin/dev' into feature/active-rates

* updated instrumentation cost logs

* cost loging

* caching updates

* Merge remote-tracking branch 'origin/dev' into feature/active-rates

* Merge remote-tracking branch 'origin/dev' into feature/active-rates

* logging fix

* null check issue fixes

* caching fix


Approved-by…
* Merged in bugfix/stg-to-main-prep (pull request #1012)

Bugfix/stg to main prep

* Merged in dev (pull request #1001)

Dev

* Revert premature merge of bugfix/retire_stale_client_file_processing

PR #959 was merged into dev without approval. This reverts commits
5e143c10, 63f32c41, 849aa626, and 927abcae to restore dev to its
pre-merge state. The changes will be re-submitted via a new PR
after proper review.

* Merged in bugfix/retire_stale_client_file_processing (pull request #961)

Return None for dashboard output when dashboard postprocessing is off

* Return None for dashboard output when dashboard postprocessing is off

FINAL_RESULT_DF_DASHBOARD was initialized as an empty DataFrame even
when RUN_DASHBOARD_POSTPROCESSING was False, causing downstream code
to needlessly process it (reorder_columns, etc). Now returns None
when dashboard is not requested, matching the postprocess() contract.

* Merged dev into bugfix/retire_stale_client_file_processing

* Merge dev (with revert) into feature branch

* Re-ap…
* Merged in bugfix/sync-stg-into-dev-20260518 (pull request #1015)

Merged in dev (pull request #1001)

* Merged in dev (pull request #1001)

Dev

* Revert premature merge of bugfix/retire_stale_client_file_processing

PR #959 was merged into dev without approval. This reverts commits
5e143c10, 63f32c41, 849aa626, and 927abcae to restore dev to its
pre-merge state. The changes will be re-submitted via a new PR
after proper review.

* Merged in bugfix/retire_stale_client_file_processing (pull request #961)

Return None for dashboard output when dashboard postprocessing is off

* Return None for dashboard output when dashboard postprocessing is off

FINAL_RESULT_DF_DASHBOARD was initialized as an empty DataFrame even
when RUN_DASHBOARD_POSTPROCESSING was False, causing downstream code
to needlessly process it (reorder_columns, etc). Now returns None
when dashboard is not requested, matching the postprocess() contract.

* Merged dev into bugfix/retire_stale_client_file_processing

* Merge dev (with revert) into fe…
* Merged in dev (pull request #1016)

Dev

* Merged in bugfix/DAIP2-2823-generic-issue-fixes-one-to-one (pull request #1011)

Bugfix/DAIP2-2823 generic issue fixes one to one

* updated auto renewal ind prompt

* updated CONTRACT_AMENDMENT_NUM prompt

* Merged dev into bugfix/DAIP2-2823-generic-issue-fixes-one-to-one


Approved-by: Praneel Panchigar
Approved-by: Siddhant Medar

* Merged in feature/DAIP2-2698-phase-3-program-product-lob-mapping (pull request #1000)

Feature/DAIP2-2698 phase 3 program product lob mapping

* added missing phase 2 modifications

* added phase 3 modifications

* Fixed acronym issues

* black format fix

* Merged dev into feature/DAIP2-2698-phase-3-program-product-lob-mapping

* standardization fixes

* Merged dev into feature/DAIP2-2698-phase-3-program-product-lob-mapping

* added hyphenated suffix names fix

* black format fix

* standardization fix

* dynamic primary fix

* Merge Dev into feature/DAIP2-2698-phase-3-program-product-lob-mapping

* updated Program Product Standardiza…

Approved-by: Praneel Panchigar
2026-05-19 20:20:43 +00:00

9.0 KiB

AARETE Derived Program & Product Canonical Derivation


Objective

Raw contracts reference the same healthcare program or health plan product using many different names. This feature normalizes them all to a single canonical name per entity, stored in two Dashboard columns:

  • AARETE_DERIVED_PROGRAM — canonical program name (e.g., CHIP, MMC, STAR)
  • AARETE_DERIVED_PRODUCT — canonical product/plan name (e.g., HA+, HMO)

The canonical is always one of the actual values seen in the data — nothing is invented.


Pipeline (4 Steps)

Raw PROGRAM / PRODUCT column  (comma-separated strings in Dashboard DF)
  │
  ├─ Step 1  Collect all unique non-empty values across the full DataFrame
  │
  ├─ Step 2  Extract payer context from the DataFrame
  │           (PAYER_NAME, PAYER_STATE, FILE_NAME from the first row)
  │
  ├─ Step 3  ONE global LLM batch call with all unique values + payer context
  │           + valid-values soft anchor → returns {raw: canonical} dict
  │           The LLM groups semantically equivalent variants and assigns one
  │           canonical per group, resolving brand-prefix noise vs. meaningful
  │           qualifiers in a single pass
  │
  └─ Step 4  Apply {raw → canonical} map to every row
              Output is comma-separated string, e.g. "CHIP, MMC"

Inputs & Outputs

PROGRAM PRODUCT
Input column PROGRAM PRODUCT
Output column AARETE_DERIVED_PROGRAM AARETE_DERIVED_PRODUCT
Valid-values list constants.VALID_AARETE_DERIVED_PROGRAMS constants.VALID_AARETE_DERIVED_PRODUCTS
Output format str, comma-separated str, comma-separated

Skip condition: If the valid-values list is non-empty and the derived column is already fully populated, the function returns the DataFrame unchanged.


Exceptions & Edge Cases

Scenario Behaviour
PROGRAM / PRODUCT column missing Warning logged, DataFrame returned unchanged
All values empty / null Info logged, no LLM call made
Sentinel values (N/A, UNKNOWN, NONE, NULL) Filtered out before LLM call
LLM omits a raw value from its response Falls back to raw.upper() for that key
LLM response parse failure Error logged; entire map falls back to raw.upper() for all values
Multi-value cell e.g. "CHIP, MMC" Each value mapped independently, re-joined
Duplicate canonicals in one cell Deduplicated before joining

LLM Prompt Design

Both batch prompts follow the same structure: a static cached instruction (registered in src/prompts/cache_registry.py) plus a dynamic template that injects unique values and payer context at runtime.

The instruction defines 7 rules evaluated in order:

# Rule Description
1 PAYER PREFIX Strip the payer brand name when it appears as a leading prefix (e.g., "Molina Medicaid" → "Medicaid"). For PRODUCT, the base name after stripping must be identical for two values to merge — "Options" ≠ "Options Plus".
2 STATE PREFIX NOT NOISE A state abbreviation prefix is meaningful if the payer contracts span multiple states; keep it unless context confirms it is redundant.
3 DISTINGUISHING QUALIFIERS Qualifiers that change meaning (e.g., "Perinatal", "Enhanced") must keep values as separate canonicals. "Plus" is usually distinguishing unless Rule 5 applies.
4 ABBREVIATION / FULL-NAME Prefer the recognized abbreviation as the canonical (e.g., "CHIP" over "Children's Health Insurance Program").
5 ABBREVIATIONS ABSORB BRAND-PREFIX AND PLUS EXPANSIONS A widely-used abbreviation can absorb brand-prefix or "Plus" expansions if it already encodes them (e.g., MMOP absorbs "Molina Medicaid Options Plus").
6 DO NOT STRIP GENERIC SUFFIXES Generic trailing words like "Plan" or "Program" are part of the canonical and must be preserved — unless another group in the same batch already uses the suffix-stripped form as its canonical (collision avoidance).
7 TIEBREAKER — PREFER SPLINTER When uncertain, keep two values as separate canonicals rather than merging them. Over-splitting is safer than over-merging.

The valid-values list is passed as a soft anchor (canonical selection priority (a)). When a raw value clearly matches a valid value, that exact string is used. Otherwise the LLM selects from the raw input values.


Code Changes

File What changed
src/pipelines/shared/postprocessing/aarete_derived.py Removed clustering helpers (_merge_acronym_clusters, _split_subset_clusters); rewrote add_aarete_derived_program_canonical and add_aarete_derived_product_canonical to use a single batch LLM call
src/prompts/prompt_templates.py Added 4 functions: 2 cached batch instructions + 2 dynamic batch templates
src/pipelines/saas/prompts/prompt_calls.py Added 2 batch wrapper functions: prompt_aarete_derived_program_canonical_batch and prompt_aarete_derived_product_canonical_batch
src/prompts/cache_registry.py Added 2 cache entries for the batch prompt instructions
test_program_product_canonical.py Updated for batch architecture; removed clustering imports and per-cluster validation checks

Function Reference

aarete_derived.py

Function Visibility Purpose
_split_multi_value_separators(text) private Splits a raw cell value on commas, semicolons, and similar delimiters.
_collect_unique_values_from_list_column(column) private Deduplicated list of non-empty, non-sentinel values from a Series. Handles both comma-string and list formats.
_all_empty_in_list_column(column) private Returns True if every cell is empty — used as the skip-derivation guard.
_map_row_list_to_canonicals(raw_val, canonical_map) private Maps one cell to its canonical(s). Splits on ,, looks up each part, deduplicates, returns comma-separated string.
add_aarete_derived_program_canonical(df, valid_programs) public Runs the 4-step batch pipeline for PROGRAMAARETE_DERIVED_PROGRAM.
add_aarete_derived_product_canonical(df, valid_products) public Same pipeline for PRODUCTAARETE_DERIVED_PRODUCT.

prompt_templates.py

Function Purpose
AARETE_DERIVED_PROGRAM_CANONICAL_BATCH_INSTRUCTION() Static cached system instruction for batch PROGRAM canonicalization (7 rules). Registered in cache_registry.py under label AARETE_DERIVED_PROGRAM_CANONICAL_BATCH.
AARETE_DERIVED_PROGRAM_CANONICAL_BATCH(unique_values, valid_programs, payer_name, payer_state) Dynamic template. Returns (context_text, prompt_text, parser) tuple.
AARETE_DERIVED_PRODUCT_CANONICAL_BATCH_INSTRUCTION() Static cached system instruction for batch PRODUCT canonicalization (7 rules). Registered under label AARETE_DERIVED_PRODUCT_CANONICAL_BATCH.
AARETE_DERIVED_PRODUCT_CANONICAL_BATCH(unique_values, valid_products, payer_name, payer_state) Dynamic template. Returns (context_text, prompt_text, parser) tuple.

prompt_calls.py

Function Purpose
prompt_aarete_derived_program_canonical_batch(unique_values, valid_programs, filename, payer_name, payer_state) Single LLM call for all PROGRAM values. Returns dict[str, str] mapping each raw value to its canonical. Falls back to raw.upper() for any key the LLM omits; falls back to {v: v.upper()} for all values on parse failure.
prompt_aarete_derived_product_canonical_batch(unique_values, valid_products, filename, payer_name, payer_state) Same for PRODUCT.

Example

Raw PROGRAM values (all unique values across the DataFrame): CHIP, Children's Health Insurance Program, Children's Health Insurance Program (CHIP), MMC, Medicaid Managed Care, Medicaid

Batch LLM call (Step 3):

All 6 values are sent in one call with payer context. The LLM determines:

  • CHIP, Children's Health Insurance Program, Children's Health Insurance Program (CHIP) → same program (Rule 4: prefer abbreviation) → canonical: CHIP
  • MMC, Medicaid Managed Care → same program (Rule 4: prefer abbreviation) → canonical: MMC
  • Medicaid → distinct broader program; no other group uses MEDICAID → canonical: MEDICAID

Resulting {raw → canonical} map:

Raw value Canonical
CHIP CHIP
Children's Health Insurance Program CHIP
Children's Health Insurance Program (CHIP) CHIP
Medicaid Managed Care MMC
MMC MMC
Medicaid MEDICAID

Final output applied to rows:

PROGRAM (raw) AARETE_DERIVED_PROGRAM
CHIP CHIP
Children's Health Insurance Program CHIP
Children's Health Insurance Program (CHIP) CHIP
Medicaid Managed Care MMC
Medicaid MEDICAID
CHIP, Medicaid Managed Care CHIP, MMC