Files
query-orchestration/plans/mutable.metadata.plan.codex.md
T
Jay Brown 17fc813823 Merged in feature/mutable-metadata1 (pull request #221)
M1, M2 and M3 complete

* M1, M2 and M3 complete

* review changes

* docs

* docs
2026-04-16 23:11:26 +00:00

17 KiB

Mutable Metadata Plan

What I Studied

The current field extraction design is rigid end to end:

  • serviceAPIs/queryAPI.yaml hard-codes FieldExtractionRequest, SingleFields, and ArrayFieldItem.
  • internal/database/migrations_collapsed/00000000000008_field_extractions.up.sql hard-codes 19 single-value columns and 112 array-value columns.
  • internal/database/queries/fieldextractions.sql and internal/fieldextraction/service.go implement document-scoped versioning by inserting immutable rows and a separate versions table.
  • internal/database/migrations_collapsed/00000000000006_documents.up.sql shows that documents currently has no schema pointer.
  • internal/document/get.go, internal/folder/service.go, and serviceAPIs/queryAPI.yaml expose field extraction data as textRecord / hasTextRecord, even though it is really structured metadata.
  • internal/server/api/listener.go uses the OpenAPI middleware for request-shape validation, but dynamic per-client schema validation does not exist today.
  • internal/database/README.md mentions json_matches_schema(), but I did not find an active migration or function definition for it, so this plan assumes application-side JSON Schema validation.
  • cmd/auth_related/permit.setup/permit_policies.yaml only maps the existing field extraction endpoints, so any new routes here will require Permit mapping updates.

Recommendation

Use an additive hybrid design.

  • Keep the current fixed-column field extraction model for legacy documents where no custom schema is assigned.
  • Add a new client-scoped JSON Schema registry.
  • Add a new JSONB-backed metadata record store for documents that opt into a custom schema.
  • Pin each document to exactly one schema revision through a nullable document column.
  • Allow schema assignment changes only until the first metadata record exists for that document.
  • Do not try to generate physical columns per client schema. The current codebase is too strongly coupled to generated OpenAPI types, SQLC queries, and static SQL for that to scale cleanly.

This gives the team a safe path:

  • Existing POST/GET /field-extractions flows keep working for legacy data.
  • New clients can use fully custom schemas without changing the database schema every time they add fields.
  • The system can later deprecate the legacy fixed-field model without blocking this first implementation.

Core Design

1. Document-level schema pinning

Add a nullable schema reference to documents.

Recommended column name:

  • metadataSchemaId uuid NULL

Meaning:

  • NULL: this document uses the legacy fixed SingleFields + ArrayFieldItem model.
  • non-NULL: this document uses the custom JSON Schema-backed metadata model.

Rule:

  • A document may change metadataSchemaId only while it has no field extraction record of any kind.
  • Once the first legacy extraction or custom metadata record exists, schema assignment is locked forever.

2. Client-managed schema registry

Store each schema revision as its own immutable row.

Recommended table:

CREATE TYPE metadataSchemaStatus AS ENUM ('active', 'superseded', 'retired');

CREATE TABLE documentMetadataSchemas (
    id uuid PRIMARY KEY DEFAULT uuid_generate_v7(),
    clientId varchar(255) NOT NULL REFERENCES clients(clientId),
    schemaKey varchar(255) NOT NULL,
    version int NOT NULL,
    title text NOT NULL,
    description text,
    documentClass text,
    status metadataSchemaStatus NOT NULL DEFAULT 'active',
    supersedesSchemaId uuid REFERENCES documentMetadataSchemas(id),
    schema jsonb NOT NULL,
    schemaHash text NOT NULL,
    createdAt timestamp NOT NULL DEFAULT NOW(),
    createdBy varchar(255) NOT NULL,
    CHECK (jsonb_typeof(schema) = 'object'),
    UNIQUE (clientId, schemaKey, version)
);

CREATE INDEX idx_documentmetadataschemas_client
    ON documentMetadataSchemas(clientId);

CREATE INDEX idx_documentmetadataschemas_client_key_status
    ON documentMetadataSchemas(clientId, schemaKey, status);

Why a single table instead of family + version tables:

  • every revision already needs a new system-assigned ID
  • schemaKey + version is enough to group revisions
  • it matches the current repo preference for straightforward tables and version rows

Suggested semantics:

  • schemaKey: stable logical name inside a client, for example aircraft-engineering
  • version: starts at 1, increments on each revision
  • id: exact revision identifier that documents reference
  • status:
    • active: assignable
    • superseded: replaced by a newer revision
    • retired: manually disabled for future assignment

3. Custom metadata record storage

Use JSONB for the extracted metadata itself.

Recommended tables:

CREATE TABLE documentMetadataExtractions (
    id uuid PRIMARY KEY DEFAULT uuid_generate_v7(),
    documentId uuid NOT NULL REFERENCES documents(id),
    metadataSchemaId uuid NOT NULL REFERENCES documentMetadataSchemas(id),
    metadata jsonb NOT NULL,
    createdAt timestamp NOT NULL DEFAULT NOW(),
    createdBy varchar(255) NOT NULL,
    CHECK (jsonb_typeof(metadata) = 'object')
);

CREATE INDEX idx_documentmetadataextractions_document
    ON documentMetadataExtractions(documentId);

CREATE INDEX idx_documentmetadataextractions_schema
    ON documentMetadataExtractions(metadataSchemaId);

CREATE TABLE documentMetadataExtractionVersions (
    id uuid PRIMARY KEY DEFAULT uuid_generate_v7(),
    metadataExtractionId uuid NOT NULL REFERENCES documentMetadataExtractions(id),
    documentId uuid NOT NULL REFERENCES documents(id),
    metadataSchemaId uuid NOT NULL REFERENCES documentMetadataSchemas(id),
    version bigint NOT NULL,
    createdAt timestamp NOT NULL DEFAULT NOW(),
    createdBy varchar(255) NOT NULL,
    UNIQUE (documentId, version)
);

CREATE INDEX idx_documentmetadataextractionversions_document
    ON documentMetadataExtractionVersions(documentId);

Recommended view:

CREATE VIEW currentDocumentMetadataExtractions AS
SELECT DISTINCT ON (dme.documentId)
    dme.id,
    dme.documentId,
    dme.metadataSchemaId,
    dme.metadata,
    dmev.version,
    dmev.createdAt,
    dmev.createdBy
FROM documentMetadataExtractions dme
JOIN documentMetadataExtractionVersions dmev
  ON dmev.metadataExtractionId = dme.id
ORDER BY dme.documentId, dmev.version DESC, dmev.id DESC;

Important choice:

  • do not add a GIN index on metadata initially
  • index only documentId, metadataSchemaId, and version paths for now
  • add GIN or projected search columns later only if querying custom fields becomes a real requirement

API Design

Routing recommendation

To stay aligned with the existing FieldExtractionService tag and current Permit resource naming, use the /field-extractions/... prefix for the new endpoints.

Recommended routes:

Method Path Purpose
POST /field-extractions/schemas Create schema revision v1
GET /field-extractions/schemas List client schemas
GET /field-extractions/schemas/{schemaId} Get one schema revision
POST /field-extractions/schemas/{schemaId}/revisions Create new revision with new ID
DELETE /field-extractions/schemas/{schemaId} Delete unused schema revision
PUT /field-extractions/documents/{documentId}/schema Assign or clear a document schema before first extraction
POST /field-extractions/metadata Create custom metadata record version
GET /field-extractions/metadata Get current custom metadata record
GET /field-extractions/metadata/version Get custom metadata record by version
GET /field-extractions/metadata/history Get custom metadata version history

Cleaner but slightly broader routing alternative:

  • /document/{id}/metadata-schema instead of /field-extractions/documents/{documentId}/schema

That route is more document-centric, but it will require matching Permit and middleware updates outside the current field-extractions path family.

Schema registry payloads

Create request:

{
  "clientId": "acme",
  "schemaKey": "aircraft-engineering",
  "title": "Aircraft Engineering Document",
  "description": "ACME aircraft engineering extraction schema",
  "documentClass": "aircraft engineering",
  "jsonSchema": {
    "$schema": "https://json-schema.org/draft/2020-12/schema",
    "type": "object",
    "additionalProperties": false,
    "properties": {
      "manufacturer": { "type": "string", "maxLength": 200 },
      "engine_hours": { "type": "number", "minimum": 0 }
    }
  },
  "createdBy": "admin@acme.com"
}

Schema detail response:

{
  "id": "019....",
  "clientId": "acme",
  "schemaKey": "aircraft-engineering",
  "version": 3,
  "title": "Aircraft Engineering Document",
  "description": "ACME aircraft engineering extraction schema",
  "documentClass": "aircraft engineering",
  "status": "active",
  "supersedesSchemaId": "019....",
  "createdAt": "2026-03-23T12:00:00Z",
  "createdBy": "admin@acme.com",
  "usage": {
    "assignedDocumentCount": 42,
    "metadataRecordCount": 37,
    "canDelete": false
  },
  "jsonSchema": { "...": "..." }
}

Revision create request:

{
  "title": "Aircraft Engineering Document",
  "description": "Added section for turbine notes",
  "jsonSchema": { "...": "..." },
  "createdBy": "admin@acme.com",
  "changeSummary": "Added turbine_notes and tightened serial number validation"
}

Revision semantics:

  • creates a brand new row with a new id
  • keeps the same schemaKey
  • increments version
  • marks the old schema row as superseded
  • does not mutate any document already pinned to the older schema ID

Document schema assignment payload

Request:

{
  "schemaId": "019....",
  "assignedBy": "admin@acme.com"
}

To clear before first extraction:

{
  "schemaId": null,
  "assignedBy": "admin@acme.com"
}

Response:

{
  "documentId": "019....",
  "metadataSchemaId": "019....",
  "locked": false
}

Custom metadata record payloads

Create request:

{
  "documentId": "019....",
  "createdBy": "user@acme.com",
  "metadata": {
    "manufacturer": "GE",
    "engine_hours": 481.25,
    "inspection_notes": [
      { "section": "fuel", "status": "pass" }
    ]
  }
}

Current record response:

{
  "id": "019....",
  "documentId": "019....",
  "metadataSchemaId": "019....",
  "schemaKey": "aircraft-engineering",
  "schemaVersion": 3,
  "version": 2,
  "metadata": {
    "manufacturer": "GE",
    "engine_hours": 481.25
  },
  "createdAt": "2026-03-23T12:34:56Z",
  "createdBy": "user@acme.com"
}

History response:

{
  "versions": [
    {
      "id": "019....",
      "documentId": "019....",
      "metadataSchemaId": "019....",
      "version": 2,
      "createdAt": "2026-03-23T12:34:56Z",
      "createdBy": "user@acme.com"
    }
  ]
}

Document API Changes

Extend the document responses so clients can see schema assignment without fetching the schema registry separately.

Add to DocumentEnriched and DocumentInFolder:

  • metadataSchemaId: uuid | null
  • hasMetadataRecord: boolean
  • metadataRecord optional, only when requested

Recommended query parameter:

  • metadataRecord=true

Compatibility note:

  • keep the existing hasTextRecord / textRecord fields for legacy fixed-field records
  • do not overload textRecord with custom JSON metadata
  • treat metadataRecord as the new preferred field for schema-bound documents

This avoids breaking current consumers that deserialize FieldExtractionResponse.

Validation and Invariants

JSON Schema rules

The server should reject schema creation unless all of the following pass:

  • valid JSON object
  • compiles successfully as JSON Schema
  • root type is object
  • root additionalProperties is explicitly present
  • size stays under a practical guardrail, for example 256 KB

Recommended draft:

  • Draft 2020-12

Recommended validation location:

  • application layer in the new schema service
  • do not rely on database-side JSON Schema validation for the first implementation

Reason:

  • the repo already uses OpenAPI middleware for static contracts, but dynamic schema compilation is not available there
  • I did not find a live DB implementation of json_matches_schema()

Custom metadata rules

When POST /field-extractions/metadata is called:

  1. load the document
  2. confirm metadataSchemaId is non-null
  3. load the exact schema row referenced by the document
  4. confirm schema and document belong to the same client
  5. validate metadata against that schema
  6. lock the version rows for that document
  7. insert immutable extraction row plus new version row

Important: the request body should not carry schemaId. The server should derive it from the document assignment to avoid client/server drift.

Legacy/custom exclusivity rules

  • if documents.metadataSchemaId IS NULL, only the existing legacy /field-extractions endpoints are allowed
  • if documents.metadataSchemaId IS NOT NULL, only the new custom /field-extractions/metadata endpoints are allowed
  • assigning a schema to a document that already has a legacy extraction is rejected with 409 Conflict
  • changing or clearing a schema after the first custom metadata record exists is rejected with 409 Conflict

Delete rules

DELETE /field-extractions/schemas/{schemaId} should succeed only when all are true:

  • no document points to that schema in documents.metadataSchemaId
  • no metadata extraction row references it
  • no metadata extraction version row references it

Otherwise return 409 Conflict.

DB-enforced protections

I recommend two database triggers in addition to service checks.

1. Prevent schema reassignment after first extraction

On documents, before update of metadataSchemaId:

  • if the value changes
  • and the document has any row in documentFieldExtractionVersions
  • or any row in documentMetadataExtractionVersions
  • then raise an error

This enforces the "locked after first extraction" rule even if another code path updates the row later.

2. Validate client/schema match

On documents, before insert/update of metadataSchemaId:

  • if non-null, ensure documents.clientId = documentMetadataSchemas.clientId

This can also be enforced in the assignment service, but a trigger makes the invariant durable.

Why JSONB is the right tradeoff here

This codebase currently pays a large change cost whenever the extraction shape changes:

  • OpenAPI changes
  • generated Go type changes
  • conversion helper changes
  • SQLC query changes
  • migration changes
  • view changes

That is acceptable for one fixed 19/112-field model. It is not acceptable for per-client schemas that can change frequently.

JSONB plus JSON Schema avoids that churn:

  • per-client field definitions become data, not code
  • extraction validation becomes runtime validation against the stored schema
  • version history stays consistent with the current service style

Tradeoff:

  • JSONB is not ideal for ad hoc SQL reporting across arbitrary custom fields

Mitigation:

  • keep the initial implementation retrieval-focused
  • if analytics/search later require it, add projected indexes or a derived flattened table keyed by JSON path

Auth and middleware impact

The current Permit mapping only includes:

  • /field-extractions
  • /field-extractions/version
  • /field-extractions/history

This design requires new entries for:

  • /field-extractions/schemas
  • /field-extractions/schemas/{schemaId}
  • /field-extractions/schemas/{schemaId}/revisions
  • /field-extractions/documents/{documentId}/schema
  • /field-extractions/metadata
  • /field-extractions/metadata/version
  • /field-extractions/metadata/history

If the team chooses the cleaner /document/{id}/metadata-schema route instead, the document Permit resource mapping also needs to be extended.

Rollout Plan

Phase 1

  • add documents.metadataSchemaId
  • add documentMetadataSchemas
  • add documentMetadataExtractions
  • add documentMetadataExtractionVersions
  • add current-view and usage queries

Phase 2

  • add schema registry endpoints
  • add assignment endpoint
  • add custom metadata endpoints
  • add JSON Schema validation service

Phase 3

  • extend document responses with metadataSchemaId, hasMetadataRecord, and optional metadataRecord
  • keep legacy textRecord behavior unchanged

Phase 4

  • optionally add deprecation notices around the legacy fixed-field model
  • evaluate whether new clients should be prohibited from using the legacy model at all

Do not include these in the first pass:

  • migrating existing legacy field extraction rows into JSONB
  • replacing the current legacy /field-extractions response shape
  • building SQL search/filtering across arbitrary custom metadata fields
  • building a flattened field catalog table unless a UI requirement forces it immediately

Final Recommendation

Implement custom schemas as a new schema registry plus JSONB-backed metadata version store, while preserving the current fixed-column field extraction path for legacy documents.

That approach fits the current repo architecture best:

  • it respects the existing version-per-document pattern
  • it avoids forcing a breaking rewrite of the static FieldExtractionResponse
  • it gives each client independent schema control
  • it keeps future migration toward a fully generic metadata model open