Files
Jay Brown 6dccf494f8 Merged in feature/permit-policy-cleanup (pull request #210)
cleanup permit policies and correct documentation

* docs

* policy cleanup

* refactor

* Merge remote-tracking branch 'origin/feature/permit-policy-cleanup' into feature/permit-policy-cleanup

* docs and edits
2026-02-19 20:22:59 +00:00

17 KiB

Appendix

Table of Contents

  1. S3 Storage Guide

S3 Storage Guide

Overview

Every document uploaded through the platform is stored in AWS S3 using a structured key (path) that encodes enough metadata to identify and recover the document without a database lookup. The key contains the client ID, processing stage, date, partition number, timestamp, entity ID, and optionally a batch ID and file extension.

The implementation lives in:

File Purpose
internal/serviceconfig/objectstore/bucketkey.go BucketKey struct, String(), ParseBucketKey()
internal/serviceconfig/objectstore/bucketkey_test.go Unit tests with concrete path examples
internal/serviceconfig/objectstore/config.go Bucket config, partition logic (GetDirectoryPart)
internal/document/upload/get.go Upload flow that constructs keys
internal/document/batch/storage/service.go Batch (ZIP) upload storage

Bucket Configuration

The bucket name is configured via environment variable:

Variable Required Description
BUCKET Yes S3 bucket name
AWS_ENDPOINT_URL_S3 No Override for local dev (default: http://localhost:4566 via localstack)
AWS_S3_USE_PATH_STYLE No Path-style addressing (default: false)
AWS_REGION No AWS region (default: us-east-1)

BucketKey Path Format

A full S3 object key follows this structure:

{ClientID}/{Location}/{Date}/{Part}/{Timestamp}~{ClientID}~{Location}~{EntityID}[~{BatchID}][.FileType]
|_______________ prefix _________________|  |______________________ filename __________________________|

For the export location, the Part segment is omitted:

{ClientID}/export/{Date}/{Timestamp}~{ClientID}~export~{EntityID}[~{BatchID}][.FileType]

Path Segment Reference

The prefix portion determines the directory structure in S3:

Segment Format Example Description
ClientID [a-zA-Z0-9_#-]+ acme_corp Client identifier. Same value repeated in the filename.
Location import, text, or export import Processing stage (see Location Types below).
Date YYYYMMDD 20250331 Date the document was uploaded, in UTC.
Part Integer (0-based) 0 Partition shard within a date directory. Omitted for export.

Location Types

Location Meaning Has Part?
import Original uploaded document Yes
text Text extraction output Yes
export Exported results (CSV, etc.) No

Filename Segment Reference

The filename portion is delimited by the tilde character (~). It comes in two variants depending on how the document was uploaded:

4-segment filename (single upload via POST /client/{id}/document):

{Timestamp}~{ClientID}~{Location}~{EntityID}[.FileType]

5-segment filename (batch upload via POST /client/{id}/document/batch):

{Timestamp}~{ClientID}~{Location}~{EntityID}~{BatchID}[.FileType]

The only difference is the 5th ~-delimited segment containing the BatchID. Both formats may optionally end with a .{FileType} extension, though in practice the current upload code does not set FileType for import documents -- the file extension is typically absent.

Segment Position Format Example Description
Timestamp 1 (always) YYYY-MM-DDTHHMMSSZ (ISO 8601 UTC) 2025-03-31T040816Z Exact creation time. Uses Go format 2006-01-02T150405Z.
ClientID 2 (always) [a-zA-Z0-9_#-]+ acme_corp Repeated from prefix for self-describing keys.
Location 3 (always) import, text, or export import Repeated from prefix for self-describing keys.
EntityID 4 (always) UUID b745910f-f529-43c5-87e1-25629d46ef40 The upload ID (for import) or document/entity ID.
BatchID 5 (optional) UUID 12345678-1234-5678-9012-123456789012 Present only for documents uploaded as part of a batch ZIP.
FileType suffix Extension without dot pdf Appended after a . separator on the last segment. Often absent for imports.

Real-world examples:

Single upload (4 segments, no extension):

2026-02-07T002850Z~demo-data-loader-1770424122~import~438b605f-1b80-4c9f-bfb9-3f6bfbf853d8

Batch upload (5 segments, no extension):

2026-02-06T183937Z~int-test-1770403169~import~13234476-2881-4b22-924d-e3ebaf63e7c0~019c3440-a4c1-7da0-a3da-3c97f6287a24

Client ID Constraints

Client IDs must match the regex pattern:

[a-zA-Z0-9_#-]+

Valid characters: letters, numbers, underscore (_), hash (#), hyphen (-).

The tilde (~) character is explicitly banned from client IDs because it serves as the delimiter between segments in the filename portion of the key. See internal/client/create.go:

// NOTE: Make sure the char ~ never makes way in as a valid client id character
// Reference bucket key filename
const CLIENT_ID_REGEX = `[a-zA-Z0-9\_#-]+`

Directory Partitioning

Within each {ClientID}/{Location}/{Date}/ prefix, documents are distributed across numbered partitions to avoid placing too many objects in a single virtual directory.

  • The default partition limit is 1000 objects per partition.
  • Partition numbering starts at 0.
  • When partition N reaches 1000 objects for a given date, new uploads go to partition N+1.
  • The current part and count are tracked in the database via GetDocumentUploadCurrentPart.

Logic from ObjectStoreConfig.GetDirectoryPart():

MaxDirSize = 1000 (default)

if count >= MaxDirSize:
    return part + 1
else:
    return part

Example progression for a busy client on a single day:

acme_corp/import/20250331/0/...    (objects 1-1000)
acme_corp/import/20250331/1/...    (objects 1001-2000)
acme_corp/import/20250331/2/...    (objects 2001-3000)

The export location does not use partitions. Export keys skip the part segment entirely.

S3 Object Metadata

In addition to the structured key, each uploaded object carries S3 metadata headers:

Metadata Key Value Example
original-path Full path from the upload request (includes folders) contracts/2025/Q1/agreement.pdf
original-filename Just the filename agreement.pdf
Content-Type MIME type detected from file extension application/pdf

These metadata values preserve the user's original file naming, which is not encoded in the S3 key itself.

Batch Upload Paths

ZIP files uploaded via the batch endpoint (POST /client/{id}/document/batch) use a separate, simpler path scheme:

batches/{ClientID}/{YYYY}/{MM}/{DD}/{HH}/{BatchID}-{OriginalFilename}

Example:

batches/acme_corp/2025/03/15/14/550e8400-e29b-41d4-a716-446655440000-uploads.zip
Segment Description
batches/ Fixed prefix for all batch uploads
ClientID Client identifier
YYYY/MM/DD/HH Timestamp directory (year/month/day/hour)
BatchID UUID assigned to the batch
OriginalFilename The user's original ZIP filename

Once the batch ZIP is extracted, individual documents from the ZIP are stored using the standard BucketKey format with the import location and the BatchID field populated.

Concrete Examples

1. Single PDF upload for client acme_corp:

s3://my-bucket/acme_corp/import/20250331/0/2025-03-31T040816Z~acme_corp~import~b745910f-f529-43c5-87e1-25629d46ef40.pdf

Decoded:

  • Client: acme_corp
  • Stage: import (original upload)
  • Date: March 31, 2025
  • Partition: 0
  • Created at: 2025-03-31 04:08:16 UTC
  • Upload/Entity ID: b745910f-f529-43c5-87e1-25629d46ef40
  • File type: pdf

2. Document from a batch upload:

s3://my-bucket/acme_corp/import/20250331/0/2025-03-31T040816Z~acme_corp~import~b745910f-f529-43c5-87e1-25629d46ef40~12345678-1234-5678-9012-123456789012.pdf

Same as above, plus:

  • Batch ID: 12345678-1234-5678-9012-123456789012

3. Text extraction output:

s3://my-bucket/acme_corp/text/20250331/3/2025-03-31T040816Z~acme_corp~text~b745910f-f529-43c5-87e1-25629d46ef40
  • Stage: text (extraction result)
  • Partition: 3 (high volume, rolled past partitions 0-2)
  • No file extension

4. Export result:

s3://my-bucket/acme_corp/export/20250331/2025-03-31T040816Z~acme_corp~export~b745910f-f529-43c5-87e1-25629d46ef40.csv
  • Stage: export
  • No partition number (exports never have one)
  • File type: csv

5. Client IDs with special characters:

s3://my-bucket/A_A/import/20250331/0/2025-03-31T040816Z~A_A~import~b745910f-f529-43c5-87e1-25629d46ef40
s3://my-bucket/A-A/import/20250331/0/2025-03-31T040816Z~A-A~import~b745910f-f529-43c5-87e1-25629d46ef40
s3://my-bucket/A#A/import/20250331/0/2025-03-31T040816Z~A#A~import~b745910f-f529-43c5-87e1-25629d46ef40

Parsing a Key Back Into Its Components

The ParseBucketKey() function in bucketkey.go reverses the process. It uses regex to extract the prefix components and then parses the tilde-delimited filename.

Step 1: Split the prefix from the filename.

Parse regex for the prefix:

^([a-zA-Z0-9_#-]+)/([a-z]+/?[a-z]+?)/([0-9]{8})/(.+)$
  |___ ClientID ___|  |__ Location __|  |_ Date _| |_ rest (part/filename or just filename) _|

The "rest" portion is then checked for a leading integer partition number. If present, it splits into {Part}/{Filename}. If not, the entire rest is the filename.

Step 2: Parse the filename.

The filename is a ~-delimited string with either 4 or 5 segments, plus an optional .{extension} suffix on the last segment. The parsing algorithm is:

  1. Split on . first. If there are exactly 2 parts, the second part is the FileType. If there is 1 part, there is no FileType. More than 2 .-separated parts is invalid.
  2. Split the base (part before .) on ~. This yields either 4 or 5 segments:
    • 4 segments: [Timestamp, ClientID, Location, EntityID] -- single upload, no batch
    • 5 segments: [Timestamp, ClientID, Location, EntityID, BatchID] -- batch upload
  3. The segment count is the only way to determine whether a BatchID is present. Count the ~ delimiters: 3 tildes = 4 segments (no batch), 4 tildes = 5 segments (has batch).

Decision table for segment count:

Tilde count Segment count Has BatchID? Upload source
3 4 No Single document upload API
4 5 Yes Batch (ZIP) upload API
Other N/A N/A Invalid -- reject

Guide: Watching an S3 Bucket for Client Document Imports

This section shows how to monitor an S3 bucket for newly imported PDF documents belonging to a specific client.

Prerequisites

  • AWS CLI configured with credentials that have s3:ListBucket and s3:GetObject permissions on the target bucket.
  • The bucket name (from the BUCKET environment variable).
  • The client ID you want to monitor.

Step 1: List All Imports for a Client

To see all imported documents for client acme_corp:

aws s3 ls s3://my-bucket/acme_corp/import/ --recursive

This lists every object under the client's import location across all dates and partitions.

Step 2: Filter by Date

To narrow down to a specific date (e.g., March 31, 2025):

aws s3 ls s3://my-bucket/acme_corp/import/20250331/

This shows the partition directories for that date:

PRE 0/
PRE 1/

Then list a specific partition:

aws s3 ls s3://my-bucket/acme_corp/import/20250331/0/

Step 3: Filter for PDFs Only

aws s3 ls s3://my-bucket/acme_corp/import/ --recursive | grep '\.pdf$'

Step 4: Watch for New Imports in Real Time

Use a polling loop to detect new documents as they arrive:

BUCKET="my-bucket"
CLIENT="acme_corp"
INTERVAL=30  # seconds between checks

LAST_CHECK=$(date -u +"%Y-%m-%dT%H:%M:%SZ")

while true; do
    echo "Checking for new imports since $LAST_CHECK ..."

    # List objects and filter by those newer than last check
    aws s3api list-objects-v2 \
        --bucket "$BUCKET" \
        --prefix "$CLIENT/import/" \
        --query "Contents[?LastModified>=\`$LAST_CHECK\`].[Key,Size,LastModified]" \
        --output table

    LAST_CHECK=$(date -u +"%Y-%m-%dT%H:%M:%SZ")
    sleep "$INTERVAL"
done

Step 5: Extract Document Identity from a Key

Given a key, you can extract the client ID, document/upload ID, batch ID, and timestamp without a database query. The filename portion has either 4 or 5 ~-delimited segments (see Filename Segment Reference).

Example 1 -- single upload (4 segments, no batch, no extension):

acme_corp/import/20250331/0/2025-03-31T040816Z~acme_corp~import~b745910f-f529-43c5-87e1-25629d46ef40

Example 2 -- batch upload (5 segments, has batch ID, no extension):

acme_corp/import/20250331/0/2025-03-31T040816Z~acme_corp~import~b745910f-f529-43c5-87e1-25629d46ef40~12345678-1234-5678-9012-123456789012

Using bash to parse either format:

KEY="$1"  # pass the S3 key as argument

# Extract the filename (everything after the last /)
FILENAME="${KEY##*/}"

# Strip file extension if present (e.g. ".pdf")
# Note: import documents typically have NO extension
BASE="${FILENAME%%.*}"
if [[ "$FILENAME" == *"."* ]]; then
    FILE_EXT="${FILENAME##*.}"
else
    FILE_EXT=""
fi

# Split on ~ into an array
IFS='~' read -ra SEGMENTS <<< "$BASE"
SEGMENT_COUNT="${#SEGMENTS[@]}"

# First 4 segments are always present
TIMESTAMP="${SEGMENTS[0]}"
CLIENT_ID="${SEGMENTS[1]}"
LOCATION="${SEGMENTS[2]}"
ENTITY_ID="${SEGMENTS[3]}"

# 5th segment is the BatchID (only present for batch uploads)
if [[ "$SEGMENT_COUNT" -eq 5 ]]; then
    BATCH_ID="${SEGMENTS[4]}"
else
    BATCH_ID=""
fi

echo "Segments:   $SEGMENT_COUNT"
echo "Client ID:  $CLIENT_ID"
echo "Location:   $LOCATION"
echo "Timestamp:  $TIMESTAMP"
echo "Entity ID:  $ENTITY_ID"
echo "Batch ID:   ${BATCH_ID:-none}"
echo "File Type:  ${FILE_EXT:-none}"

Output for example 1 (single upload):

Segments:   4
Client ID:  acme_corp
Location:   import
Timestamp:  2025-03-31T040816Z
Entity ID:  b745910f-f529-43c5-87e1-25629d46ef40
Batch ID:   none
File Type:  none

Output for example 2 (batch upload):

Segments:   5
Client ID:  acme_corp
Location:   import
Timestamp:  2025-03-31T040816Z
Entity ID:  b745910f-f529-43c5-87e1-25629d46ef40
Batch ID:   12345678-1234-5678-9012-123456789012
File Type:  none

Step 6: Download a Document by Key

aws s3 cp \
    "s3://my-bucket/acme_corp/import/20250331/0/2025-03-31T040816Z~acme_corp~import~b745910f-f529-43c5-87e1-25629d46ef40.pdf" \
    ./recovered-document.pdf

Step 7: Retrieve the Original Filename

The original filename from the user's upload is stored in S3 object metadata, not in the key:

aws s3api head-object \
    --bucket my-bucket \
    --key "acme_corp/import/20250331/0/2025-03-31T040816Z~acme_corp~import~b745910f-f529-43c5-87e1-25629d46ef40.pdf" \
    --query "Metadata"

Output:

{
    "original-path": "contracts/2025/Q1/agreement.pdf",
    "original-filename": "agreement.pdf"
}

Step 8: Bulk Recovery Script

To recover all PDFs for a client on a given date and rename them using their original filenames:

BUCKET="my-bucket"
CLIENT="acme_corp"
DATE="20250331"
OUTPUT_DIR="./recovered"

mkdir -p "$OUTPUT_DIR"

aws s3api list-objects-v2 \
    --bucket "$BUCKET" \
    --prefix "$CLIENT/import/$DATE/" \
    --query "Contents[].Key" \
    --output text | tr '\t' '\n' | grep '\.pdf$' | while read -r KEY; do

    # Get original filename from metadata
    ORIGINAL=$(aws s3api head-object \
        --bucket "$BUCKET" \
        --key "$KEY" \
        --query "Metadata.\"original-filename\"" \
        --output text)

    # Extract entity ID for uniqueness (4th segment, works for both 4 and 5 segment filenames)
    FILENAME="${KEY##*/}"
    BASE="${FILENAME%%.*}"
    IFS='~' read -ra SEGS <<< "$BASE"
    ENTITY_ID="${SEGS[3]}"

    # Use original filename if available, otherwise use entity ID
    if [ "$ORIGINAL" != "None" ] && [ -n "$ORIGINAL" ]; then
        DEST="$OUTPUT_DIR/${ENTITY_ID}_${ORIGINAL}"
    else
        DEST="$OUTPUT_DIR/${ENTITY_ID}.pdf"
    fi

    echo "Downloading: $KEY -> $DEST"
    aws s3 cp "s3://$BUCKET/$KEY" "$DEST"
done

Source code reference: internal/serviceconfig/objectstore/bucketkey.go