M1, M2 and M3 complete * M1, M2 and M3 complete * review changes * docs * docs
16 KiB
Mutable Metadata v3 Simplification Findings (Codex)
Purpose
This document reviews plans/mutable.metadata.plan.combo.v3.md and plans/mutable.metadata.plan.combo.v3.tracking.md with one goal: preserve the real requirements while reducing implementation risk, regression surface, and rollout complexity.
The current v3 plan is trying to land too many things at once:
- a new schema system
- a new metadata storage path
- new admin routes
- new user routes
- three database triggers
- document read-path changes
- delete-path changes
- bulk folder assignment
- destructive schema-reset workflow
- a large auth and test matrix
The core requirement set is smaller than that. The simplest safe plan is to reduce redundant state, avoid touching stable paths until necessary, and ship the feature in vertical slices instead of infrastructure layers.
Executive Summary
The biggest simplification is this:
- Keep
documents.custom_schema_id. - Store versioned metadata by
document_idonly. - Derive the schema from the document binding at write/read time.
- Keep only the invariants that are truly needed for v3.
- Defer optional surface area until the core path is stable.
That change alone removes one trigger, one FK, one index, one view, several queries, several tests, and a large amount of reasoning complexity.
The second major simplification is rollout shape: stop treating this as twelve horizontal phases. Implement it as a few end-to-end vertical milestones so the system is usable and testable much earlier.
What Should Stay
These are the actual requirements and should remain intact:
- client-scoped custom schemas stored as data, not DDL
- schema definitions validated as JSON Schema
- documents gain an optional
custom_schema_id - metadata writes must validate against the document's assigned schema
- once custom metadata exists, the document's schema binding becomes immutable
- legacy field extractions and custom metadata remain mutually exclusive per document
- schema updates create a new schema version/new schema ID
- schema administration remains
super_admin-only
Everything else should be judged by whether it is necessary to satisfy those requirements.
Recommended Simplifications
1. Remove schema_id from document_custom_metadata
This is the highest-value simplification.
The current plan stores both:
documents.custom_schema_iddocument_custom_metadata.schema_id
That is redundant in v3.
Why it is redundant:
POST /custom-metadataalready derives schema fromdocuments.custom_schema_id.- the plan already forbids schema reassignment after metadata exists
- reset-metadata deletes all metadata rows before rebinding to a different schema
Because of those three facts, every surviving metadata row for a document necessarily belongs to the document's current schema binding. A second schema pointer on the metadata row adds no new capability for v3.
Simpler table
CREATE TABLE document_custom_metadata (
id uuid PRIMARY KEY DEFAULT uuid_generate_v7(),
document_id uuid NOT NULL REFERENCES documents(id) ON DELETE CASCADE,
metadata jsonb NOT NULL,
version int NOT NULL DEFAULT 1,
created_at timestamptz NOT NULL DEFAULT NOW(),
created_by varchar(255) NOT NULL,
CONSTRAINT uq_doc_custom_metadata_version UNIQUE (document_id, version)
);
CREATE INDEX idx_dcm_doc_version_desc
ON document_custom_metadata(document_id, version DESC);
What this removes
document_custom_metadata.schema_ididx_dcm_schema_id- Trigger 3 (
trg_enforce_consistent_schema_id) - the
GetSchemaMetadataRecordCountquery - the extra "schema consistency across metadata versions" app-layer checks
- the need to reason about metadata rows carrying a schema separate from the document row
Why this does not violate requirements
It still preserves:
- schema validation on write
- immutable schema binding after first metadata write
- destructive reset before rebinding
- schema version history at the schema table level
- metadata version history at the document level
Do not pre-pay v4 complexity for a non-requirement. If a future version needs non-destructive in-place schema migration, then add row-level schema_id later with a deliberate migration.
2. Replace Trigger 2 with a composite foreign key
The current plan uses a trigger to ensure the schema belongs to the same client as the document.
That is better expressed as a database constraint.
Simpler constraint
Keep:
client_metadata_schemas(id, client_id)documents(id, clientId, custom_schema_id)
Add:
ALTER TABLE client_metadata_schemas
ADD CONSTRAINT uq_cms_id_client UNIQUE (id, client_id);
ALTER TABLE documents
ADD CONSTRAINT fk_documents_custom_schema_same_client
FOREIGN KEY (custom_schema_id, "clientId")
REFERENCES client_metadata_schemas(id, client_id);
Why this is better
- declarative integrity instead of PL/pgSQL
- no trigger code
- no trigger-specific unit tests
- no divergence risk between service logic and DB logic
With recommendation 1 applied, the only remaining trigger worth keeping is Trigger 1: "once metadata or legacy extraction rows exist, custom_schema_id cannot change."
That shrinks the trigger surface from three triggers to one.
3. Drop the current_document_custom_metadata view
The new metadata table is simple enough that a view is unnecessary.
Use:
SELECT ...
FROM document_custom_metadata
WHERE document_id = $1
ORDER BY version DESC
LIMIT 1;
With idx_dcm_doc_version_desc, this is straightforward and efficient.
Why remove the view
- one less schema object
- one less migration artifact
- one less thing to test
- less drift between query logic and view logic
The legacy field-extraction system benefits from a view because its state is spread across multiple tables. The new custom-metadata design does not need the same indirection.
4. Use FK cascades so delete code does not need feature-specific edits
The current plan adds delete-path work in internal/document/delete.go and internal/client/delete.go.
That is avoidable.
Because these are new tables, design their FKs so existing delete flows continue to work without feature-specific code.
Recommended FK behavior
document_custom_metadata.document_id -> documents.id ON DELETE CASCADEclient_metadata_schemas.client_id -> clients.clientId ON DELETE CASCADE
Result
- document hard-delete automatically removes custom metadata rows
- client hard-delete automatically removes client schema rows after client deletion
DeleteDocumentCascadedoes not need custom-metadata-specific stepsclient.HardDeletedoes not needDeleteClientMetadataSchemas
Important correction
documents.custom_schema_id does not need to be nulled before deleting the document row.
The current plan says that delete must null custom_schema_id first because the FK to client_metadata_schemas would otherwise block the delete. That is incorrect. Deleting the document removes the referencing row; it does not require clearing the reference first.
Keep NullifyDocumentCustomSchemaId only for the reset-metadata flow, where the document row survives.
This is a meaningful regression-risk reduction because it avoids touching existing explicit delete code paths at all.
5. Remove convenience-only fields from the first release
The current plan adds several useful but non-essential response decorations:
documentCountcanDeletepreviousVersionIdcustomSchemaNameon unrelated responses- embedded
customMetadatainsideGET /document/{id}
These fields are not required to satisfy the business rules.
Simpler rule
For the first release:
- schema list returns basic schema summary only
- delete attempts return
409if the schema is still in use - full metadata stays on
/custom-metadata, notGET /document/{id}
If usage hints are still wanted later, add them as a dedicated opt-in field or includeUsage=true expansion after the core system is stable.
6. Do not modify GET /document/{id} in the core milestone
The current code in internal/document/get.go and api/queryAPI/documents.go is intentionally narrow: it composes document basics, labels, and optionally legacy text-record data.
Do not expand that path in the first pass.
Reasons:
- it is an existing stable read path
- the feature already has dedicated metadata endpoints
- adding full metadata enrichment multiplies controller, service, OpenAPI, and test work
- it increases regression risk in a widely used endpoint
Recommendation
For the core release:
- leave
GET /document/{id}unchanged
Optional later milestone:
- add
customSchemaId - add
hasCustomMetadata - only add full
customMetadataif there is a concrete consumer that needs it
This keeps the first delivery focused on the new feature itself instead of modifying a stable cross-cutting read path.
7. Defer bulk folder assignment until after the core feature is stable
POST /super-admin/folders/{folderId}/assign-schema adds a large amount of complexity:
- recursive tree traversal
- large-transaction behavior
- 10K-document guardrail
- skip-reason reporting
- mixed-state subtree handling
- many more integration tests
It is not required to prove the mutable-metadata model.
Recommendation
Move bulk folder assignment out of the core implementation and into a later milestone.
Core assignment should be:
- single-document only
Then, if bulk assignment is still wanted after the single-document path is stable, add it separately with real operational evidence about folder sizes and expected admin workflows.
This is one of the easiest ways to cut failure risk without weakening the main feature.
8. Defer reset-metadata to a later additive milestone
The reset endpoint is useful, but it is not part of the minimum path needed to deliver custom schemas and validated custom metadata.
It is also currently coupled to unrelated work in the tracker through delete-cascade query reuse.
Recommendation
Treat reset-metadata as an additive milestone after the core path is working:
- schema create/list/get
- single-document assign
- metadata write/read current
- legacy guard
- schema update/version
Only then add:
POST /super-admin/documents/{id}/reset-metadata
If Q wants reset kept within the same overall initiative, that is still fine. The key simplification is sequencing: do not let reset block the first stable end-to-end slice.
Also, once recommendation 4 is applied, reset no longer needs to depend on delete-path work. It stands on its own.
9. Remove createdBy from all new request bodies
The plan already says actor identity must come from auth context, not the request body.
Because these are new endpoints, there is no backward-compatibility reason to keep misleading createdBy request fields in the contract.
Recommendation
For all new mutable-metadata endpoints:
- do not accept
createdByin the request schema - derive actor from JWT / auth middleware
- return audit actor in responses if needed
This simplifies:
- OpenAPI
- handler code
- audit rules
- tests
- client expectations
The only place createdBy should remain is existing legacy endpoints that already use it today.
10. Replace schema-version PUT with an explicit version-creation POST
The current plan uses:
PUT /super-admin/custom-schemas/{schemaId}
But that endpoint does not replace the resource in place. It creates a new schema row with a new ID and supersedes the old one.
That is creation semantics, not update-in-place semantics.
Simpler route
Use:
POST /super-admin/custom-schemas/{schemaId}/versions
Why this is better
- matches actual behavior
- avoids the Permit
PUTaction-mapping question entirely - reduces auth configuration ambiguity
- makes code review and API docs clearer
This preserves the requirement that schema updates create a new ID, while making the API easier to reason about.
Revised Minimal Design
If I were simplifying the v3 design directly, I would target this shape:
Database
client_metadata_schemas- keep immutable schema rows
- keep version numbers
- keep lifecycle state if desired
documents.custom_schema_id- add optional binding
- enforce same-client binding with composite FK
document_custom_metadatadocument_id,metadata,version, audit fields- no
schema_id ON DELETE CASCADE
- DB trigger surface:
- keep only the "schema cannot change after metadata or legacy extraction exists" trigger
API in the first stable slice
POST /super-admin/custom-schemasGET /super-admin/custom-schemasGET /super-admin/custom-schemas/{schemaId}PATCH /super-admin/documents/{id}/schemaPOST /custom-metadataGET /custom-metadata
API deferred to later slices
- schema version creation
- schema retirement
- metadata history/version endpoints
- reset-metadata
- bulk folder assignment
GET /document/{id}enrichment
This still reaches a real, usable feature early.
Revised Implementation Milestones
The current tracking doc is organized as horizontal layers. That is the wrong shape for a risky feature. It encourages long periods where nothing is fully integrated.
Use vertical slices instead.
Milestone 1: Schema CRUD foundation
Goal: prove schema storage and validation end to end.
Includes:
- migrations for
client_metadata_schemas documents.custom_schema_id- composite FK for same-client binding
- schema validator
POST/GET/LIST /super-admin/custom-schemas- Permit
super-adminresource update
Exit proof:
- super_admin can create and list schemas
- invalid schema rejected
- cross-client document binding cannot be represented in DB
Milestone 2: Core document flow
Goal: prove the actual business workflow end to end.
Includes:
document_custom_metadatatable withoutschema_id- single-document schema assignment
POST /custom-metadataGET /custom-metadata- legacy field-extraction guard plus controller
409mapping
Exit proof:
- create schema
- assign schema to document
- write valid metadata
- reject invalid metadata
- reject legacy field extraction write afterward
At this point the feature already satisfies the central requirement.
Milestone 3: Administrative completion
Goal: complete the admin lifecycle without expanding into optional convenience features.
Includes:
- schema version creation endpoint
- schema retirement/delete-if-unused behavior
- optional metadata history endpoints if they are required immediately
Exit proof:
- create v1
- create v2
- v1 no longer assignable
- v2 assignable
- in-use schema cannot be retired/deleted incorrectly
Milestone 4: Additive extras
Only add these after Milestones 1-3 are green:
- reset-metadata
- bulk folder assignment
GET /document/{id}enrichment- usage/count decorations in schema list responses
These should not block the core mutable-metadata rollout.
How the Tracking Doc Should Change
The tracking doc should be simplified in these specific ways:
Remove whole workstreams from the core path
- remove Phase 5b if FK cascades are used correctly
- move Phase 9 behind the core milestones
- move Phase 10 behind the core milestones
- move Phase 10b behind the core milestones
Collapse redundant DB work
- remove view-related tasks
- remove Trigger 3 tasks
- remove
schema_id-related metadata tasks - remove
GetSchemaMetadataRecordCounttasks - remove document-delete nullification tasks
Change phase ordering
Do not wait until Phase 8 for the first real end-to-end flow.
Instead, each milestone should contain:
- migration/query work
- service work
- controller work
- at least one end-to-end integration test
That is the right unit of verification for a change this invasive.
Bottom Line
The current v3 plan is trying to solve the core problem plus several convenience and escape-hatch problems at the same time. That is where most of the risk is coming from.
The safest simplification is:
- remove redundant metadata-side schema state
- replace trigger logic with declarative constraints where possible
- let new FKs handle delete cleanup
- keep stable document read/delete paths unchanged in the first pass
- ship the feature in vertical slices
- defer bulk assignment, reset, and read-path enrichment until after the core path is proven
If Q wants the shortest version of this advice:
Make the document row the single source of truth for schema binding, make the metadata table just versioned JSON for that document, and do not ship bulk/reset/document-enrichment in the first milestone.