8c9060e425
Feature/multithreading * fixes * Merge main into feature/multithreading Resolved conflicts: - Kept timing instrumentation in file_processing.py - Kept new 3-step Exhibit-based approach for one-to-n processing - Maintained parallelization improvements (20 workers) Changes include: - Timing utils integration for performance monitoring - Increased max_workers from 5 to 20 across all components - Parallelized code_breakout and grouper_breakout - Fixed tin_npi_funcs function call parameters * Fix max_workers error for empty documents - Add check to skip parallel processing when no pages exist - Use min(len(all_page_tasks), 20) to prevent max_workers=0 - Handles edge case of documents with no exhibits or pages * Fix max_workers=0 errors in one_to_n_funcs - Add checks before all ThreadPoolExecutor creations - Prevents errors when processing empty lists: - carveout_and_special_case - breakout - special_case_breakout - filter_services_without_reimbursements - run_lob_relationship - Ensures executor only created when there are items to process * Reduce code processing parallelism to prevent API throttling Lower max_workers from 20 to 10 for code_breakout and grouper_breakout to prevent overwhelming Bedrock API with concurrent requests * fixed * Merged main into feature/multithreading * Move documentation into folder * Fix logging statements * Merge branch 'main' into feature/multithreading * Refactor for clarity * Update previous exhibit passing logic * properly simplify exhibits * Merged main into feature/multithreading * update conditional for None * exhibit multithreading changed * exhibit multithreading changed * Merge remote-tracking branch 'origin/main' into feature/multithreading * synced with main * Parallelize dynamic assignment, refactor HSC field worker, add timing/exhibit unit tests, tidy imports/ignore helpers * analyze_regression.py edited online with Bitbucket * count_pages.py edited online with Bitbucket * compare_regressed_with_baseline.py edited online with Bitbucket * simple_testbed_compare.py edited online with Bitbucket * run_testbed_metrics_regressed.py edited online with Bitbucket Approved-by: Katon Minhas
30 lines
1.2 KiB
Python
30 lines
1.2 KiB
Python
from src.investment.exhibit_funcs import Exhibit, get_exhibit_list
|
|
from src.prompts.fieldset import FieldSet
|
|
|
|
|
|
def test_add_reimbursement_rows_sets_flag():
|
|
exhibit = Exhibit(exhibit_page="1", exhibit_page_nums=["1"], exhibit_header="H", exhibit_text="text")
|
|
exhibit.add_reimbursement_rows([{"a": 1}], [])
|
|
assert exhibit.has_reimbursements is True
|
|
assert len(exhibit.reimbursement_rows) == 1
|
|
|
|
|
|
def test_set_exhibit_level_and_previous_dynamic_fields():
|
|
prev = Exhibit(exhibit_page="1", exhibit_page_nums=["1"], exhibit_header="H1", exhibit_text="text1")
|
|
dyn_fields = FieldSet()
|
|
prev.dynamic_primary_fields = dyn_fields
|
|
current = Exhibit(exhibit_page="2", exhibit_page_nums=["2"], exhibit_header="H2", exhibit_text="text2", prev_exhibit=prev)
|
|
|
|
current.set_exhibit_level_data({"X": "Y"}, FieldSet())
|
|
assert current.exhibit_level_answers["X"] == "Y"
|
|
assert current.get_previous_exhibit_dynamic_fields() is dyn_fields
|
|
|
|
|
|
def test_get_exhibit_list_links_previous():
|
|
text_dict = {"1": "A", "2": "B"}
|
|
mapping = {"1": ["1"], "2": ["2"]}
|
|
headers = {"1": "H1", "2": "H2"}
|
|
exhibits = get_exhibit_list(text_dict, mapping, headers)
|
|
assert len(exhibits) == 2
|
|
assert exhibits[1].prev_exhibit is exhibits[0]
|