9db0fb328e
Don't consider table-subpages each individually for reimbursement exhibit identification * Refine EXHIBIT_CHECK prompt * test: rework `get_exhibit_pages` for table subpages * Fix indentation * Experimental table de-chunking * Increase max_tokens for Claude model in exhibit page processing * fix get_pages_to_process docstring * refactor: improved docstring for get_exhibit_pages * move `get_pages_to_process` to `preprocess.py` * test: add unit tests for get_pages_to_process function * Comment out debugging lines * Changed prompt that was breading universal_json_load parsing * fix: handle "no results" case in get_reimbursement_primary function * fixed docstring Approved-by: Katon Minhas
Field Extraction
Running the Code
This project uses Python's module system for imports. To run the code properly:
-
DO NOT run files directly like this:
python src/client/main.py # ❌ This won't work -
INSTEAD, use Python's module flag (
-m) from the project root:python -m src.client.main # ✅ This is correct
Why?
The code uses absolute imports (e.g., from src.utils import llm_utils) to maintain a clear and consistent package structure. Running with python -m ensures Python can properly resolve these imports.
Development
- All imports should use the
src.prefix (e.g.,from src.utils import llm_utils) - Always run code from the project root directory using the
-mflag - Tests are configured to handle these imports automatically via pytest settings in pyproject.toml
Configuration
The project uses poetry for dependency management and pytest for testing. Key configurations in pyproject.toml:
[tool.pytest.ini_options]
pythonpath=["."]
testpaths = [
"tests"
]
Dependencies
To install dependencies:
poetry install
To run tests:
poetry run pytest
To run type checking:
poetry run mypy .