Files
doczyai-pipelines/fieldExtraction
Alex Galarce c0fa621aed Merged in bugfix/multi-tin-npi-fix (pull request #699)
Bugfix/multi tin npi fix

* Add debug logging for OCR candidates and provider info extraction

* Update TIN and NPI extraction instructions to handle multiple values in a single record

* Add validation layer to cross-check TINs and NPIs with regex findings in provider info extraction

* Enhance exact match extraction by normalizing identifiers and validating lengths for TIN and NPI


Approved-by: Katon Minhas
2025-09-05 21:21:54 +00:00
..

Field Extraction

Running the Code

This project uses Python's module system for imports. To run the code properly:

  1. DO NOT run files directly like this:

    python src/client/main.py  # ❌ This won't work
    
  2. INSTEAD, use Python's module flag (-m) from the project root:

    python -m src.client.main  # ✅ This is correct
    

Why?

The code uses absolute imports (e.g., from src.utils import llm_utils) to maintain a clear and consistent package structure. Running with python -m ensures Python can properly resolve these imports.

Development

  • All imports should use the src. prefix (e.g., from src.utils import llm_utils)
  • Always run code from the project root directory using the -m flag
  • Tests are configured to handle these imports automatically via pytest settings in pyproject.toml

Configuration

The project uses poetry for dependency management and pytest for testing. Key configurations in pyproject.toml:

[tool.pytest.ini_options]
pythonpath=["."]
testpaths = [
    "tests"
]

Dependencies

To install dependencies:

poetry install

To run tests:

poetry run pytest

To run type checking:

poetry run mypy .