Bugfix/mcs issue fixes
* debugging and initial prompt change
* Remove redundant reimb info
* Merged main into bugfix/mcs-issue-fixes
* Prep for merge
* update unit tests
Approved-by: Alex Galarce
Feature/provider state
* Add prompt for extracting provider state information in investment prompts
* Add state normalization function and integrate into context field extraction
* black format string_utils.py; add tests for normalize_state_to_abbreviation function
* black format string_utils_test.py
* Add PROVIDER_STATE to investment column order
* Merged main into feature/provider-state
Approved-by: Katon Minhas
Bugfix/mcs issue fixes
* Update code primary prompt
* Merge branch 'main' into bugfix/mcs-issue-fixes
* No longer treat non-reimbursement tables as 'tables' for the purpose of preprocessing
* No longer extract Reimb Effective Date from footer of the page
* Merge branch 'main' into bugfix/mcs-issue-fixes
* Update unit tests
* Add postprocessing step to remove Reimb Date values when the 1:1 date values are identical
* Update unit tests
* prep for merge
* Remove excess prints
* Merged main into bugfix/mcs-issue-fixes
* Merge branch 'main' into bugfix/mcs-issue-fixes
* Prep for merge
* remove print
Approved-by: Alex Galarce
Bugfix/fix overzealous duplication
* Refactor duplicate processing logic to scope seen pairs by base page, enhancing clarity and accuracy in deduplication.
* Enhance deduplication tests to include base page identifiers in seen pairs
* Merged main into bugfix/fix-overzealous-duplication
Approved-by: Katon Minhas
Feature/split reimbursements
* Update reimbursement test cases to reflect accurate reimbursement terms and improve deduplication logic
* move VALIDATE_REIMBURSEMENTS_PROMPT to investment_prompts.py
* black, isort formatting
* Remove IDENTIFY_REIMBURSEMENT_EXHIBITS_PROMPT function (unused)
* Add MODEL_ID_CLAUDE4_SONNET for new model integration (doesn't work right now)
* pipe delimited-output and parsing
* Merge remote-tracking branch 'origin/main' into feature/split-and-filter-reimbursements
* Merge remote-tracking branch 'origin/main' into feature/split-reimbursements
* Add functionality to split compound reimbursement terms into distinct methodologies
* splitting logic for reimbursement terms
* put in split compound reimbursement step
* Fix JSON output formatting in reimbursement terms example
* reimbursement splitting: move prompt to investment_prompts.py, move keywords to investment_values.py
* Merge remote-tracking branch 'origin/main' into feature/split-reimbursements
* Merge remote-tracking branch 'origin/main' into feature/split-reimbursements
* Change order (split -> filter over filter -> split), improve logging
* Refactor reimbursement filtering logic to use a two-stage approach with pattern-based negative filtering and LLM validation
* Replace clear reimbursement indicators with obvious non-reimbursement indicators for improved clarity in reimbursement filtering
* Simplify reimbursement filtering by removing pattern-based checks and relying solely on LLM validation
* Remove obvious non-reimbursement indicators from reimbursement filtering keywords
* Enhance reimbursement level test by adding LLM validation and updating mock return values for service terms
* Merge remote-tracking branch 'origin/main' into feature/split-reimbursements
* Fix failing test
* Merge remote-tracking branch 'origin/main' into feature/split-reimbursements
* update mock return values to fix failed tests
Approved-by: Siddhant Medar
Feature/filter reimbursements
* First-pass implementation for reimbursement filtering
* Enhance reimbursement level processing by filtering out services without reimbursement terms and logging when no valid pairs are found.
* Update mock return values in reimbursement level tests to reflect new reimbursement terms
* Merge remote-tracking branch 'origin/main' into feature/split-and-filter-reimbursements
* Refine reimbursement filtering by adding 'billed' to indicators and ensuring REIMB_TERM is a string before processing.
* Refactor is_empty function signature to support multiple input types and clarify return values
* adding LLM review for ambiguous reimbursement methodology cases
* more stringent keywords for stage 1 of reimbursement filtering
* refining rate and cost patterns for reimbursement filtering
* Enhance LLM response handling in reimbursement methodology check to improve accuracy and logging for ambiguous cases.
* Merge remote-tracking branch 'origin/main' into feature/split-and-filter-reimbursements
* Update reimbursement test cases to reflect accurate reimbursement terms and improve deduplication logic
* move VALIDATE_REIMBURSEMENTS_PROMPT to investment_prompts.py
* black, isort formatting
* Remove IDENTIFY_REIMBURSEMENT_EXHIBITS_PROMPT function (unused)
* Add MODEL_ID_CLAUDE4_SONNET for new model integration (doesn't work right now)
* pipe delimited-output and parsing
* Merge remote-tracking branch 'origin/main' into feature/split-and-filter-reimbursements
* Refactor reimbursement filtering by moving clear indicators to constants
Approved-by: Katon Minhas
Feature/prov group tin
* don't prompt for INTRO_PAGE
* test complete
* test 2
* Update test script
* Merge branch 'main' into feature/prov-group-tin
* Merged main into feature/prov-group-tin
* refactor
* cleaning
* Update save
* Fix return type
* Update for reimbursement_tin_npi
* Comment out reimbursement-tin-npi
* Pass unit tests
* Merge branch 'main' into feature/prov-group-tin
* Merged main into feature/prov-group-tin
* Remove length validation
* Merge branch 'main' into feature/prov-group-tin
* Remove ON_INTRO from prompt
* Restore
* Remove test
* Move reimb prov info to dynamic reimb info
* e2e test
* Fix for unit test
* remove prints
* update docstrings and typehints
Approved-by: Alex Galarce
Feature/rework table handling
* test: add cases for configuration boundary
* test: add table post-processing tests (cases to ensure preservation of post-table text and metadata during table operations)
* test: enhance table handling tests for small and large tables, ensuring proper suffixing and splitting behavior
* Fix and clarify intent in comments
* fix: enhance handling of post-table text when combining tables
* refactor: remove commented-out code in clean_tables function for clarity
* test: update post-table text preservation assertions and add debug tracing
* refactor: remove unused remove_table_from_page function and its test
* refactor: enhance combine method documentation and clarify split_by_size_with_smart_tables logic
* Merge remote-tracking branch 'origin/main' into feature/rework-table-handling
* refactor: reduce large table threshold from 1000 to 50 rows for better table splitting
* Add simple splitting
* refactor: uncomment dynamic answer retrieval in run_one_to_n_prompts for improved functionality
* Merge remote-tracking branch 'origin/main' into feature/rework-table-handling
* refactor: remove unused table handling functions and related constants for cleaner code
* refactor: remove unused test functions in preparation for new tests
* refactor: streamline continuation table handling by utilizing existing flags
* refactor: update table combination logic to include all tables from the main page
* refactor: enhance table handling to automatically resolve column mismatches, respect row limits, and improve unit/integration test coverage
* rename clean_tables_simple to clean_tables
* Update docstrings
* refactor: rename table handling functions
* Fix indentation and add tests
* isort, black
* Merge remote-tracking branch 'origin/main' into feature/rework-table-handling
Approved-by: Katon Minhas
Bugfix/dynamic primary issues
* Merge branch 'main' into bugfix/dynamic-primary-issues
* Merged main into bugfix/dynamic-primary-issues
* Crosswalk updates
* Merge branch 'main' into bugfix/dynamic-primary-issues
* dont include section information in Exhibit Header
* Add Molina Health Benefit Exchange to Product-LOB mapping
* PPO-POS separate mapping
* Prevent inference of LOB from Program or Product
* Add SoonerSelect
* Update mapping rules to allow for direct values
* update prompts
* Prep for run
* Merged main into bugfix/dynamic-primary-issues
* Update unit tests
* Update Moda Product mappings
* Remove Care1st as Product for Centene
* testing
* Mapping update
* Mapping updates
* prep for merge
* Merged main into bugfix/dynamic-primary-issues
* Merged main into bugfix/dynamic-primary-issues
* docstrings and formatting
* Get group provider for > 1
* Merged main into bugfix/dynamic-primary-issues
Approved-by: Alex Galarce
Add (service, methodology) deduplication with accumulator
* Add deduplication with accumulator
* Merge remote-tracking branch 'origin/main' into feature/service-reimb-deduplication
* Replace print statement with logging.debug for cross-exhibit duplicate detection
* Fix and add tests for one_to_n_funcs
* Merge remote-tracking branch 'origin/main' into feature/service-reimb-deduplication
* Add seen_pairs parameter to reimbursement_level function in tests
Approved-by: Katon Minhas
Feature/exhibit linking
* Initialize new exhibit header prompt
* Save
* Merge branch 'main' into feature/exhibit-linking
* update prompt
* cleanup return type
* remove test
* Test fixes
* Fix test
* Add tests
* Merge branch 'main' into feature/exhibit-linking
* Merge branch 'main' into feature/exhibit-linking
* Remove test notebooks
* Delete test notebook
* Save
* Merge branch 'main' into feature/exhibit-linking
* try except for exhbit page finding
* Prep for merge
* update test
* Remove test
* Merge branch 'main' into feature/exhibit-linking
* update unit test
* Fix unit test
* replace try except
* Update page sort
* Add docstring
Approved-by: Alex Galarce
Don't consider table-subpages each individually for reimbursement exhibit identification
* Refine EXHIBIT_CHECK prompt
* test: rework `get_exhibit_pages` for table subpages
* Fix indentation
* Experimental table de-chunking
* Increase max_tokens for Claude model in exhibit page processing
* fix get_pages_to_process docstring
* refactor: improved docstring for get_exhibit_pages
* move `get_pages_to_process` to `preprocess.py`
* test: add unit tests for get_pages_to_process function
* Comment out debugging lines
* Changed prompt that was breading universal_json_load parsing
* fix: handle "no results" case in get_reimbursement_primary function
* fixed docstring
Approved-by: Katon Minhas
find_overlap function added to handle inconsistent chunk size
* find_overlap function added to handle inconsistent chunk size
* set min chunk overlap needed to remove overlap to 3
* Merge branch 'main' into bugfix/smart-chunking
* pipeline error fixed
* Merge branch 'main' into bugfix/smart-chunking
* spelling fix
Approved-by: Katon Minhas
DRAFT: Bugfix/effective date fix
* Merged in bugfix/extra-fields (pull request #500)
Do not run Exhibit level prompt if there are no exhibit level fields
* Do not run Exhibit level prompt if there are no exhibit level fields
* Update dynamic test
* Remove test from one-to-n test module
* Update test
* Fix unit test
Approved-by: Alex Galarce
* Merged in bugfix/tin-npi-prompt-fixes (pull request #499)
Bugfix/tin npi prompt fixes
* Change delimiter for other provider information from comma to pipe
* filter out invalid providers
* Update extraction and formatting instructions for provider entities in TIN_NPI_TEMPLATE
* Merged main into bugfix/tin-npi-small-bugfixes
* Refactor deduplication logic to use string_utils for checking empty provider fields
* Merge branch 'bugfix/tin-npi-small-bugfixes' of https://bitbucket.org/aarete/doczy.ai into bugfix/tin-npi-small-bugfixes
* Update deduplication logic to exclude providers with 'UNKNOWN' TIN, NPI, or NAME
* Merge remote-tracking branch 'origin/main' into bugfix/tin-npi-small-bugfixes
* Fix deduplication logic to not drop subsequent names
* Fix formatting of provider names in deduplication logic
* add payer name filtering
* reorder execution order to allow to find payer_name first
* Add handling for previously unfound group providers in identify_group_provider function
* Merge remote-trackin…
* Merged in feature/update-testbed-script (pull request #501)
Feature/update testbed script
* Update preprocessing
* Update preprocessing
* remove quit
* Update date for filename
* Merged main into feature/update-testbed-script
Approved-by: Alex Galarce
* Merged in bugfix/smart-chunks (pull request #503)
missing of text in chunks fixed
* missing of text in chunks fixed
Approved-by: Katon Minhas
* Merged in feature/rework-reimb-tin-npi (pull request #502)
Feature/rework reimb tin npi
* use LLM for reimb_[tin/npi/name]
* Merge remote-tracking branch 'origin/main' into feature/rework-reimb-tin-npi
* Merge remote-tracking branch 'origin/main' into feature/rework-reimb-tin-npi
* deduplicate tin/name/npi before adding to prompt
* Pull out prompt and llm_response in methodology_breakout_single_row for easier debugging
* Refactor reimbursement_tin_npi to handle multiple providers and update valid TIN, NPI, and NAME fields accordingly
* Reorder REIMB_PROV_NAME field in investment_prompts.json
* Update REIMB_PROV_NAME prompt to specify valid values for provider names
* Remove valid_values constraint from prov_name to raise accuracy
* Remove debugging print statements from reimbursement_tin_npi and related functions
* Remove debugging print statements from run_provider_info_fields function
* Merge remote-tracking branch 'origin/main' into feature/rework-reimb-tin-npi
* Remove debugging print statem…
* Merged in feature/preformat-single-quotes-input (pull request #506)
Replace double quotes with single quotes in context strings for parsability
* Replace double quotes with single quotes in context strings for parsability
* Merged main into feature/preformat-single-quotes-input
Approved-by: Katon Minhas
* updated sig pages fnxn
* Merge branch 'main' into bugfix/effective_date_fix
* poetry add pymupdf
* created vision funcs
* Merged main into bugfix/effective_date_fix
* fixed pipeline issue
* updated string_utils_test
* updated test values
* updated extract signature page fxn
* updated test values
* updated test values
* Merged main into bugfix/effective_date_fix
* Prompt template in all-caps
* remove debugging prints
* Merged main into bugfix/effective_date_fix
* Change vision to False by default
* Change vision to False by default
* fix tests
* fix tests
Approved-by: Katon Minhas
Do not run Exhibit level prompt if there are no exhibit level fields
* Do not run Exhibit level prompt if there are no exhibit level fields
* Update dynamic test
* Remove test from one-to-n test module
* Update test
* Fix unit test
Approved-by: Alex Galarce
Collect TIN/NPI/Name page-by-page and run a second pass for IS_GROUP
* configure first pass: no providers marked IS_GROUP. Allows simplification of prompt
* Add group provider identification logic in second pass
* group identification: don't take 20% of pages. Take first 3 and last 3
* Refine group provider identification logic to use pipes for JSON formatting and improve error handling
* fix mypy error for `identify_group_provider`
* Merge remote-tracking branch 'origin/main' into feature/tin-npi-prompt-formatting
* Add "PROV_INFO_JSON" to investment column order
* Refactor group provider identification to use signature page extraction and fix JSON format
* move signature page extraction to `string_utils.py` and add unit tests for it
* Update identify_group_provider to return provider name along with TIN and NPI in JSON format
* Update identify_group_provider to always use group-provided name when available
* move out GROUP_TIN_NPI_TEMPLATE to `investment_prompts.py`
* Remove deprecated PROV_GROUP_TIN_CHECK and PROV_GROUP_NPI_CHECK prompts
* Move out provider info fields from `one_to_one_funcs` to `tin_npi_funcs`
* isort
* remove debug print statements
* Remove debug print statements from provider info functions
* add note to regex_utils.py about where functions are used
* Add page_key_sort function and corresponding tests for sorting page keys
* Refactor identify_group_provider to sort pages and avoid duplicates in extracted sections
* Remove debug print statements from identify_group_provider function
* Merge remote-tracking branch 'origin/main' into feature/tin-npi-prompt-formatting
* extract_signature_page function: prioritize Textract marker for signatures and fallback to keyword search
* fix mypy error
* Update test cases to reflect changes in signature extraction logic
Approved-by: Katon Minhas
Field/initialize patient age
* Update split_text so that it works on single-page contracts (discovered when testing)
* Add PATIENT_AGE_RANGE intermediate column
* Remove PATIENT_AGE_RANGE in postprocessing
* Dont convert columns to int type
* Add unit tests
* Remove test
* Update unit test
* Update unit test
* Update unit test
* Update unit test
Approved-by: Alex Galarce
Field/dynamic codes
* remove test
* Merge branch 'main' into field/dynamic-codes
* Merged main into field/dynamic-codes
* Update test
* Fix unit test
* Merged main into field/dynamic-codes
* Merged main into field/dynamic-codes
* Merged main into field/dynamic-codes
* Update fields with crosswalks
* Update mapping to return empty string if no mapping
* Update dynamic_funcs
* Restructure
* Update file_processing
* Update dynamic primary
* Fix dynamic primary
* Update base fields
* Update one-to-n process
* Merged main into field/dynamic-codes
* genericized get_dynamic_answers
* Update tests
* Remove test file
* dynamic_funcs cleanup
* remove prints
* add exhibit_header
* Fix imports
Approved-by: Alex Galarce
Bugfix/flatten singleton list ints problem
* Fix flatten_singleton_string_list to handle non-list int strings and improve test coverage
* Fix test_flatten_singleton_string_list to handle integer input correctly
Approved-by: Katon Minhas
Feature/align tin npi
* Add merge_provider_info function to consolidate provider data into one-to-one results
* Merge remote-tracking branch 'origin/main' into feature/align-tin-npi
* Remove debug print statement from run_one_to_one_prompts function
* turn run_regex_fields into an orchestrator function by breaking parts of it into separate functions
* Remove unused import of get_matches function and add docstring to chunk_on_matches for clarity
* Add docstring to run_regex_fields for improved clarity and documentation
* fix provider name keys in merge_provider_info
* Remove unused functions get_matches, tin_npi_prompt, and clean_tin_npi from tin_npi_funcs.py
* Set provider name to "UNKNOWN" if cleaned name is empty in clean_provider_info
* Add unit tests for clean_provider_info function to validate data processing
* Enhance docstring for get_all_matches function
* Add tests for get_all_matches and chunk_on_matches functions
* Merge remote-tracking branch 'origin/main' into feature/align-tin-npi
* Implement merge_provider_info function to consolidate provider data into one-to-one results
* Refactor merge_provider_info to simplify list conversion for group and other provider information
* Remove unnecessary TIN and NPI postprocessing step from postprocess function
* Remove clean_tin_npi_other function and its invocation from postprocess function
* Use json prompts for TIN/NPI/Name and inject them into TIN_NPI_TEMPLATE
* Rename run_regex_fields to run_provider_info_fields
* Enhance TIN_NPI_TEMPLATE to include detailed instructions for extracting provider entities and their identifying information, including IS_GROUP logic for main contracting parties.
* moved TIN/NPI regexes to regex_patterns.py
* Remove unnecessary blank line in investment_values.py
* Remove debug print statements from run_provider_info_fields and get_provider_info functions
* Enhance clean_provider_info to handle various representations of IS_GROUP as boolean
* Enhance get_provider_info to include error handling for JSON parsing and ensure consistent return format
Approved-by: Katon Minhas
Hotfix/flatten list to string
* updated
* rate fields fix
* updated the singleton_list
* updated
* added
* Checked_files
* added
* Merged main into hotfix/flatten-list-to-string
* format_rate_fields_with_commas
* Update test
* flatten_singleton_string_list - new cases
* Update
* Removed unneccessary case
* Remove case
* Updated the test cases
* Merge remote-tracking branch 'origin/main' into hotfix/flatten-list-to-string
* Refactor flatten_singleton_string_list to handle empty input and return a comma-separated string for multiple elements
Approved-by: Katon Minhas
Hotfix/reimb issue fixes
* updated DYNAMIC_CHECK prompt
* updated date_fix_prompt to YYYYMMDD format
* filled N/A values for REIMB_PROV_TIN
* changed YYYYMMDD to YYYY/MM/DD
* extracting TIN, NPI,NAME as a tuple
* updated TIN, NPI,NAME as a tuple prompt
* Merge remote-tracking branch 'origin/main' into hotfix/reimb_issue_fixes
* fix: correct spelling of 'memorize' to 'memoize' in cache decorator comment
* combined postprocess date format check functions
* Merged main into hotfix/reimb_issue_fixes
* updated postprocessing
* rename and slight refactor
* remove unused imports, re-sort
* remove unused imports from postprocess.py
* Fix unit tests
* re-add imports which are needed for tests to pass (?)
* Fix column name in auto-renewal termination date handling
* Fix formatting in postprocess function
Approved-by: Katon Minhas
Feature/dynamic one to one
* Pass N/A to one-to-one
* E2E test passed
* Merged main into feature/dynamic-one-to-one
* Pass ANY non N/A answer to reimbursement-level
* Update test
* Remove test
* Merge branch 'main' into feature/dynamic-one-to-one
* Wrap clean_dates
* Revert postprocess
* Merged main into feature/dynamic-one-to-one
* Merged main into feature/dynamic-one-to-one
* Remove prints
Approved-by: Alex Galarce
Bugfix/table error fix
* handle multipage table data correctly
* same metadata for all tables on the page
* Multiple tables on page, non-continuous, split
* Add back align_and_format_tables
* update unit tests
* Merged main into bugfix/table-error-fix
* improved align and format table prompt
* Successful E2E test on San Joaquin - without running align and format table prompt
* San Joaquin e2e test passed
* Merged main into bugfix/table-error-fix
* Merge branch 'main' into bugfix/table-error-fix
* Add same number of columns condition
* Merged main into bugfix/table-error-fix
* Clean investment_values
* Remove import
* docstrings
* Remove test doc
* Add back original ALIGN_AND_FORMAT_TABLES prompt for client work
* Remove import
* Add config reference
* Merged main into bugfix/table-error-fix
* Move table_utils to investment.table_funcs
* restore table_utils.py
* Update tests
* update preprocess
Approved-by: Alex Galarce
Update date formatting in prompts to YYYY/MM/DD
* Update date formatting in prompts to YYYY/MM/DD
* fix tests for new date formatting
Approved-by: Katon Minhas
Bugfix/dynamic assignment issue
* Merged main into feature/proc-desc
* Uncomment test code
* Bugfix
* script optimisation for codes
* modified is empty function in string utils
* Remove test.py
* bug fix to handle list of proc_codes
* Merged main into bugfix/proc-code-lists
* changed any method to all method in list handling
* code mapping moved outside to run once
* Merge branch 'feature/daip2-97' into bugfix/proc-code-lists
* merge conflict fixed
* merge conflicts fixed
* incorporated latest changes from main
* Merged main into bugfix/proc-code-lists
* Move load all dataset to io_utils
* Fix list is_empty
* Updated poetry.lock
* Bugfix
* Docstring
* Docstrings
* Merge branch 'main' into bugfix/proc-code-lists
* Merge branch 'bugfix/proc-code-lists' into bugfix/dynamic-assignment-issue
* Update test for no default exhibit_level_answer_dict
* Merged main into bugfix/dynamic-assignment-issue
Approved-by: Alex Galarce
Document dates optimization
* simplified file processing for testing
* modified derived_term_date_prompt
* fixed day values with spaces or leading zeros
* modified contract termination date prompt
* updated termination date prompt
* reranking of chunks for better context
* exclude last date while calculating term date
* clean up
* clean up 2
* Merge branch 'main' into document_dates_optimization
pulled latest changes from main
* Merged main into document_dates_optimization
* fix unit tests
* Merged main into document_dates_optimization
* refactor: update termination date prompt for clarity and accuracy
* Merged main into document_dates_optimization
* fix pipeline
* Merge remote-tracking branch 'origin/main' into document_dates_optimization
* Merge branch 'document_dates_optimization' of https://bitbucket.org/aarete/doczy.ai into document_dates_optimization
Approved-by: Katon Minhas
Test/check doc dates against test bed
* first commit
* Enhance prompts and processing for contract termination dates
- Updated prompt to include 'year to year' as a valid example output.
- Implemented case-insensitive processing for contract text and field keywords for smart chunking.
- Clarified prompt for derived termination date to specify it is exclusive of renewals
* Merge remote-tracking branch 'origin/main' into test/check-doc-dates-against-test-bed
* Update termination date prompt tests
* Refine keywords for contract termination date prompt in investment prompts
Approved-by: Katon Minhas
Initialize Reimbursement-Level TINs
* Initialized Reimbursement-Level TIN and NPI
* Update unit tests
* Docstrings and cleanup
* Update test
Approved-by: Alex Galarce
Fields/document dates further iteration
* add concise date fix prompt
* Refactor and prepare for postprocessing formatting
* Merged main into fields/document-dates-further-iteration
* Merged main into fields/document-dates-further-iteration
* Implement memoization for date postprocessing and add derived termination date calculation
* Merge branch 'fields/document-dates-further-iteration' of https://bitbucket.org/aarete/doczy.ai into fields/document-dates-further-iteration
* fix mypy error
* tweaks to derived termination date
* Add docstring to derived_term_date_prompt and create unit tests for prompt generation
* fix memoization in postprocess.py
* Merged main into fields/document-dates-further-iteration
Approved-by: Katon Minhas
Feature/doc date iterations
* add parse_chunk()
* add stitch_chunks() and tests
* create new constant
* put in new stitch_chunks() function
* resolve output validation TODO
* Merged main into feature/clean-up-smart-chunking
* fix Mayank's callout of noncontiguous chunks, fix typos, add tests
* add comment on test
* Merge remote-tracking branch 'origin/main' into feature/doc-date-iterations
* add 'signature page' to effective date keywords. loosen instructions on termination date.
* enhance prompts with examples and clarify instructions for termination date
* fix for group_fields()
* clean up test imports
Approved-by: Katon Minhas
Feature/clean up smart chunking
* add parse_chunk()
* add stitch_chunks() and tests
* create new constant
* put in new stitch_chunks() function
* resolve output validation TODO
* Merged main into feature/clean-up-smart-chunking
* fix Mayank's callout of noncontiguous chunks, fix typos, add tests
* add comment on test
* Merged main into feature/clean-up-smart-chunking
Approved-by: Katon Minhas
feature/dynamic and generalized branch
* included txt files for nltk_data
* move nltk_data to src
* Fix last upload count
* Last upload count bugfix
* Fixed B processing
* Remove client-specific postprocessing
* split consolidate_output
* Run by file
* Fix output
* Reconfigure smart_chunk fields
* Add Full Context
* Fix merge conflicts
* Regex, Smart-Chunked, and Full working - not adding smart-chunked-->full when necessary
* Modernized run_full_context_fields()
* Switched set to list in field_context
* Move fields from smart_chunked to full_context as part of 'field_context' function
* Working version with placeholders
* Update poetry and pyproject
* Update s3 output
* Remove deprecated unit test
* Updated error messages
* Updated smart chunk ac name to one to one
* Update dependencies - end-to-end test for write s3 functional
* Add basic multithreading
* Send individual output to s3/local
Approved-by: Alex Galarce
Feature/integrate standard and complex
* last edits to merge_tables_draft_v2 before moving to table_funcs
* migrate to table_utils.py
* fix typehint that caused mypy error
* fix mypy errors
* black and isort
* fix error with multi-table pages
* update poetry.lock
* add tests, fix bug when table start marker or table end marker is missing
* added more tests
* tests for get_str_dictionaries_from_text()
* fix edge case in get_str_dictionaries_from_text
* fix mypy error
* more tests
* add tests, add handling for invalid dictionaries
* add tests
* add tests for insert_column_headers()
* update requirements
* add new simple/complex logic to preprocess.py
* remove files used solely for testing
* split pages into sub-pages by tables
* add table split on end marker. Also add tests
* add docstrings, change control flow
* remove unused tests, add TODO for new tests
* get in prompt changes and intermediate decisions
* TODO for future enhancement for get_exhibit_pages
Approved-by: Katon Minhas
Task/unit tests investment and client
* move io_utils.py
* add __init__.py
* refactor: move string_funcs and update import paths to use utils namespace
* rename string_funcs to string_utils
* move and rename table_utils.py
* refactor: rename claude_funcs to llm_utils in multiple files
* update unit tests
* added unit tests for src/string_utils.py
* added unit tests for src/table_utils.py
* added unit tests for src/llm_utils.py
* added unit tests for src/io_utils.py
* Merge branch 'main' into task/unit-tests-investment-and-client; TODO: fix unit tests
* Unit tests to fix
* fix unit tests
* fix invalid references
* add unit tests for preprocessing_funcs.py
* move pytest-mock to test dependencies
* Merge remote-tracking branch 'origin/main' into task/unit-tests-investment-and-client
* add tests/string_utils_test.py
* removed filename parameters for a few functions
* save postprocessing_funcs_test midway
* save postprocessing_funcs_test
* postprocessing_funcs unit tests complete
* Merge branch 'main' into task/unit-tests-investment-and-client
Approved-by: Katon Minhas