Files
doczyai-pipelines/fieldExtraction
Katon Minhas 27e09b5fbc Merged in feature/row-count-by-pages (pull request #585)
Feature/row count by pages

* Add page-level row count comparison to analyze discrepancies between testbed and results

* Fix key naming in page mismatch comparison for consistency

* Merged main into feature/row-count-by-pages

* Implement page-level grouping and comparison for row counts in testbed and results

* Fix row count difference calculation in page comparison


Approved-by: Alex Galarce
2025-06-24 17:42:45 +00:00
..

Field Extraction

Running the Code

This project uses Python's module system for imports. To run the code properly:

  1. DO NOT run files directly like this:

    python src/client/main.py  # ❌ This won't work
    
  2. INSTEAD, use Python's module flag (-m) from the project root:

    python -m src.client.main  # ✅ This is correct
    

Why?

The code uses absolute imports (e.g., from src.utils import llm_utils) to maintain a clear and consistent package structure. Running with python -m ensures Python can properly resolve these imports.

Development

  • All imports should use the src. prefix (e.g., from src.utils import llm_utils)
  • Always run code from the project root directory using the -m flag
  • Tests are configured to handle these imports automatically via pytest settings in pyproject.toml

Configuration

The project uses poetry for dependency management and pytest for testing. Key configurations in pyproject.toml:

[tool.pytest.ini_options]
pythonpath=["."]
testpaths = [
    "tests"
]

Dependencies

To install dependencies:

poetry install

To run tests:

poetry run pytest

To run type checking:

poetry run mypy .