# Server API Package This package contains the HTTP API server implementation for the Query Orchestration system, including REST endpoints, request handlers, and background processing workers. ## Overview The `internal/server/api` package provides: - HTTP server setup and configuration - REST API endpoint handlers for document management - Batch upload processing capabilities - Background task processing for asynchronous operations - Integration with AWS services (S3, SQS) - Document lifecycle management ## Key Components ### 1. Server Setup (`listener.go`) - Initializes the Echo HTTP server with middleware - Configures authentication and authorization - Sets up route handlers for API endpoints - Initializes background task runners - Manages server lifecycle and graceful shutdown ### 2. Document Handlers (`document_handler.go`) - Handles individual document upload requests - Validates document metadata - Manages S3 storage operations - Triggers document processing pipeline ### 3. Batch Processing (`batch_handler.go`) - Manages ZIP file uploads containing multiple documents - Creates batch records for tracking - Queues batches for background processing - Provides batch status and progress endpoints ### 4. Background Batch Worker (`batch_worker.go`) The batch worker is a background task processor that handles asynchronous batch document processing. ## Batch Worker Architecture ### Purpose The batch worker processes ZIP archives containing multiple documents that were uploaded through the batch upload API. It runs as a background task to avoid blocking the API response and to handle potentially long-running batch operations. ### How It Works #### 1. **Task Scheduling** - Runs periodically (configurable interval, default 30 seconds) - Queries the database for unprocessed batches (status = "processing") - Processes batches sequentially to manage resource usage #### 2. **Batch Processing Flow** ``` ZIP Upload → S3 Storage → Batch Record → Background Worker → Document Processing ``` Detailed steps: 1. **Batch Discovery** - Worker queries `batch_uploads` table for batches with status "processing" - Retrieves batch metadata including S3 location of ZIP file 2. **ZIP Download** - Downloads the ZIP archive from S3 using batch's `archive_key` - Loads entire ZIP into memory for processing 3. **File Extraction & Processing** - Iterates through each file in the ZIP archive - For each file: - Skips directories and system files (`.DS_Store`, `__MACOSX/`) - Sanitizes filename to prevent path traversal attacks - Extracts file content - Uploads extracted file to S3 temporary directory - Validates file type (currently only PDFs supported) - Calls document upload handler to create document record - Document enters standard processing pipeline 4. **Progress Tracking** - Maintains counters for: - `processedCount`: Successfully processed documents - `failedCount`: Documents that failed processing - `invalidCount`: Documents with invalid file types - Updates batch progress in database after processing all files - Records failed filenames for troubleshooting 5. **Status Management** - Batch starts with "processing" status when created - Updates to "completed" if at least one document processed successfully - Updates to "failed" if no documents processed successfully - Can be "canceled" by user intervention #### 3. **Configuration** The worker receives configuration through a map containing: - `ConfigKeyBatchService`: Service for batch database operations - `ConfigKeyS3Client`: AWS S3 client for file operations - `ConfigKeyBucket`: S3 bucket name for storage - `ConfigKeyUploadHandler`: Function to handle individual document uploads #### 4. **Error Handling** - Individual file failures don't stop batch processing - Failed files are logged and recorded in `failed_filenames` - Batch marked as "failed" only if ZIP can't be downloaded or opened - All errors logged with context for debugging #### 5. **Security Considerations** - Filename sanitization using `path.Base()` to prevent directory traversal - File type validation (PDF only) - Proper resource cleanup (file readers closed after use) - Safe integer conversion with bounds checking ### Batch States | Status | Description | |--------|-------------| | `processing` | Default state when batch created, being processed by worker | | `completed` | At least one document processed successfully | | `failed` | No documents processed successfully or critical error | | `canceled` | Manually canceled by user | ### Integration Points 1. **Database**: - Reads from `batch_uploads` table - Updates batch progress and status - Creates document records 2. **S3 Storage**: - Downloads ZIP archives - Uploads extracted files to temporary directories - Files organized by batch ID for traceability 3. **Document Pipeline**: - Processed documents enter standard pipeline - Associated with batch via `batch_id` foreign key - Follows normal document lifecycle (init → sync → clean → text extraction) ### Performance Considerations - Processes one batch at a time to control resource usage - ZIP files loaded entirely into memory (consider streaming for large files) - Background processing prevents API blocking - Configurable processing interval for load management ### Monitoring The worker logs: - Number of unprocessed batches found - Processing start/completion for each batch - Individual file processing results - Error details for troubleshooting - Final counts (processed/failed/invalid) ## API Endpoints ### Document Management - `POST /clients/:clientId/documents` - Upload single document - `GET /clients/:clientId/documents` - List client documents - `GET /documents/:documentId` - Get document details ### Batch Operations - `POST /clients/:clientId/documents/batches` - Upload ZIP batch - `GET /clients/:clientId/documents/batches` - List client batches - `GET /clients/:clientId/documents/batches/:batchId` - Get batch status - `DELETE /clients/:clientId/documents/batches/:batchId` - Cancel batch ## Testing The package includes comprehensive tests: - Unit tests for individual handlers - Integration tests with test containers - Batch worker processing tests - S3 operation mocking for isolated testing ## Dependencies - Echo framework for HTTP server - AWS SDK v2 for S3 operations - PostgreSQL for data persistence - Background task runner for async processing