Merged in feature/textract-fix (pull request #178)

fix textract accuracy issue

* fix textract
This commit is contained in:
Jay Brown
2025-09-03 20:43:27 +00:00
parent 720a84be92
commit 0ad41de26b
179 changed files with 405 additions and 29635 deletions
@@ -0,0 +1,140 @@
# Query System Explanation
## Overview
The query system in this application is designed to extract structured data from processed documents. Queries are configurable data extraction rules that can depend on each other, creating a pipeline of data transformations.
## Query Types
The system currently supports two query types:
### 1. CONTEXT_FULL Query
- **Purpose**: Originally intended to provide the full document text/context
- **Current Implementation**: Returns a hardcoded JSON placeholder `{"keyone":"valueone","keytwo":"valuetwo"}`
- **No Dependencies**: This query type doesn't require any other queries to run first
- **Note**: The current implementation appears to be a mock/placeholder - in production, this would likely return the actual extracted text from the document
### 2. JSON_EXTRACTOR Query
- **Purpose**: Extracts specific values from JSON data using path expressions
- **Dependencies**: Requires exactly one other query result as input (typically a CONTEXT_FULL query)
- **Configuration**: Takes a JSON config with a `path` field specifying what to extract
- **Uses gjson**: Leverages the `tidwall/gjson` library for JSON path extraction
## How Queries Work
### Query Definition Structure
```go
type Query {
ID uuid
Type queryType // CONTEXT_FULL or JSON_EXTRACTOR
ActiveVersion int32 // Current version being used
RequiredQueryIDs []uuid // Dependencies - other queries that must run first
Config string // JSON configuration (e.g., {"path":"keyone"})
}
```
### Example Query Chain
1. **First Query (CONTEXT_FULL)**:
- Type: `CONTEXT_FULL`
- Config: None needed
- Dependencies: None
- Returns: `{"keyone":"valueone","keytwo":"valuetwo"}`
2. **Second Query (JSON_EXTRACTOR)**:
- Type: `JSON_EXTRACTOR`
- Config: `{"path":"keyone"}`
- Dependencies: [First Query ID]
- Input: Result from First Query
- Returns: `"valueone"`
3. **Third Query (JSON_EXTRACTOR)** - could extract from second:
- Type: `JSON_EXTRACTOR`
- Config: `{"path":"some.nested.path"}`
- Dependencies: [Second Query ID]
- Would extract from the result of the second query
## Query Execution Flow
### 1. querySyncRunner Phase
- After document text extraction completes, querySyncRunner is triggered
- It queries the database for all queries that:
- Are configured for this client (via collector configuration)
- Have NO dependencies, OR
- Have all dependencies already satisfied (results exist in database)
- For each eligible query, sends a message to the QUERY queue
### 2. queryRunner Phase
- Receives {document_id, query_id} from QUERY queue
- Retrieves the query definition and active version
- If query has dependencies:
- Fetches results from all required queries
- Validates all dependencies are satisfied
- Executes the appropriate processor based on query type:
- CONTEXT_FULL: Returns the mock JSON (should return document text)
- JSON_EXTRACTOR: Extracts specified path from dependency result
- Stores the result in database with version tracking
- **Triggers dependent queries**: Finds all queries that depend on this one and sends them to QUERY queue
### 3. Recursive Processing
- As each query completes, it triggers its dependents
- This creates a cascade effect through the dependency graph
- Continues until all queries in the chain have been executed
## Collector Configuration
Collectors define which queries should run for each client:
```sql
-- collectorQueries table maps queries to clients
CREATE TABLE collectorQueries (
clientId varchar(255), -- Which client
name varchar(255), -- Named reference (e.g., "JSON_QUERY")
queryId uuid, -- Which query to run
...
)
```
When documents are processed, only queries configured in the client's collector will be executed.
## Current Implementation Status
### What's Working
- Query dependency resolution and scheduling
- Recursive execution of dependent queries
- Version tracking for queries and results
- JSON path extraction from query results
- Client-specific query configuration via collectors
### What Appears Incomplete
- **CONTEXT_FULL Implementation**: Currently returns mock data instead of actual document text
- The system stores extracted text in S3 (via docTextRunner) but CONTEXT_FULL doesn't retrieve it
- There's infrastructure for text storage (`documentTextExtractions` table) but it's not connected to the query processor
## Real-World Use Case
In a production scenario, this system would work like:
1. **Document uploaded** → Text extracted via OCR
2. **Base Query** (CONTEXT_FULL) → Returns full document text as JSON
3. **Field Extraction Queries** (JSON_EXTRACTOR) → Extract specific fields like:
- Customer name
- Invoice number
- Date
- Amount
4. **Derived Queries** → Further process extracted data:
- Calculate totals
- Validate formats
- Cross-reference with other data
## Key Design Benefits
1. **Flexibility**: New extraction rules can be added without code changes
2. **Reusability**: Queries can be shared across clients or documents
3. **Dependency Management**: Complex multi-step extractions are handled automatically
4. **Versioning**: Changes to queries don't break existing results
5. **Client Isolation**: Each client can have different extraction rules for the same document types
## Summary
The query system is a sophisticated, dependency-aware data extraction pipeline. While the current implementation has placeholder logic for the base CONTEXT_FULL query, the infrastructure for chaining queries, managing dependencies, and extracting structured data from documents is fully functional. The system allows for complex, multi-step data extraction workflows to be defined declaratively and executed automatically as documents are processed.