Files
query-orchestration/docs/README.md
T
Michael McGuinness 1211697f61 Merged in feature/autodocs (pull request #94)
Auto Docs for all cmds

* docs
2025-03-06 19:26:01 +00:00

132 lines
3.6 KiB
Markdown

# Query Orchestration
## docCleanRunner
**Listens For**: Body Containing Document ID.
**Action**.
If the document is not already clean according to the job collector standards, a clean is triggered.
An event with the document id is pushed to the DOCTEXT queue.
## docInitRunner
**Listens For**: S3 Event notifications - limited to PUT notifications.
**Action**.
Checks if the uploaded document is a duplicate.
Currently whether the document content is a duplicate or not is determines using the document ETag, and check if it already exists for the job.
Find reference entity.
If NOT a duplicate - Create a document reference entity.
Create a document entry to log the successful upload of a document.
Add an event to the DOC SYNC queue.
## docSyncRunner
**Listens For**: Body Containing Document ID.
**Action**.
Checks if the document is eligible for syncing.
By validating: the client can_sync parameter, the job can_sync parameter.
If eligible, send an event to the DOCCLEAN queue.
## docTextRunner
**Listens For**: Body Containing Document ID
**Action**.
If the document text is not already extracted according to the job collector standards, the text is extracted.
An event with the document id is pushed to the QUERYSYNC queue.
## jobSyncRunner
**Listens For**: Body Containing Job ID
**Action**.
Adds an event for each document into the DOCSYNC queue for the given Job.
## metricsExample\_test
This main.go serves as an example of how to use the custom Prometheus metrics package.
It demonstrates:
1. Setting up metrics with proper context handling
2. Configuring different types of metrics (Counter, Gauge, Histogram, Summary)
3. Recording metrics safely in a concurrent environment
4. Implementing graceful shutdown
5. Proper error handling
## queryRunner
**Listens For**: Body Containing Document ID and Query ID
**Action**.
Run the query for the given document and store the result.
Events with the document id and downstream query ids are pushed to the QUERY queue.
## queryService
This API currently manages all user entities to be managed. This allows users of the service to change required entities and processes.
The entities in question:
- Client - A client is a single client of the Doczy project, all configurations and jobs for a client will be in reference to the respective client.
- Job - A job is a single instance of requirements for a client. This maintains a set of queries and processes in relation to itself.
- Collector - A collector specifies the required queries for export for a given job.
- Sample - A sample is a subset of documents to be used for testing. When it is active a job will only process documents in the sample.
- Query - A query is a block of processing of data which takes an input and produces an output.
- Export - An export is a manually started process which outputs the results for a job to a given location.
## querySyncRunner
**Listens For**: Body Containing Document ID
**Action**.
Find the queries with no dependencies that do not have a valid result for a document.
The reasoning for only extracting these is to start the query extraction process only for necessary queries, and the query runner will manage updating downstream dependencies.
There are 2 situations that cause a query to show up in this list:
The query has no required dependencies.
All required dependencies have a valid result.
An event with the document id is pushed to the DOCTEXT queue.
## queryVersionSyncRunner
**Listens For**: Body Containing Query ID
**Action**.
Adds an event for each job into the JOBSYNC queue for the given query, i.e. each job that contains the query in its dependency tree.
Generated by Doczy