Files
doczyai-pipelines/fieldExtraction
Alex Galarce 485859e954 Merged in bugfix/chroma-too-many-files (pull request #679)
Bugfix/chroma too many files

* Add emergency cleanup and monitoring for vectorstore cache

* Refactor vectorstore handling and enhance file descriptor monitoring in field_context

* Refactor vectorstore cleanup logic to ensure safe reference counting and improve error handling

* Enhance monitoring and cleanup mechanisms to prevent file descriptor leaks in smart chunking functions

* Add garbage collection to safe_process_file for improved memory management

* Refactor docstring in safe_process_file to clarify error handling and resource cleanup details

* isort

* Enable Chroma reset functionality for improved resource management and logging


Approved-by: Katon Minhas
2025-08-25 20:56:04 +00:00
..

Field Extraction

Running the Code

This project uses Python's module system for imports. To run the code properly:

  1. DO NOT run files directly like this:

    python src/client/main.py  # ❌ This won't work
    
  2. INSTEAD, use Python's module flag (-m) from the project root:

    python -m src.client.main  # ✅ This is correct
    

Why?

The code uses absolute imports (e.g., from src.utils import llm_utils) to maintain a clear and consistent package structure. Running with python -m ensures Python can properly resolve these imports.

Development

  • All imports should use the src. prefix (e.g., from src.utils import llm_utils)
  • Always run code from the project root directory using the -m flag
  • Tests are configured to handle these imports automatically via pytest settings in pyproject.toml

Configuration

The project uses poetry for dependency management and pytest for testing. Key configurations in pyproject.toml:

[tool.pytest.ini_options]
pythonpath=["."]
testpaths = [
    "tests"
]

Dependencies

To install dependencies:

poetry install

To run tests:

poetry run pytest

To run type checking:

poetry run mypy .