# Underlay - AI Integration Guide Underlay is a versioned, content-addressed registry for structured knowledge. Apps push snapshots of their data; Underlay preserves them and serves them via HTTPS API. Built by Knowledge Futures (501c3): https://www.knowledgefutures.org Base URL: https://underlay.org/api --- ## Authentication There are two auth methods: 1. API Key (for programmatic access): Header: Authorization: Bearer Keys are prefixed with "ul_" (e.g. ul_abc123...) when created through the UI. Keys have scopes: read, write, admin. Keys can optionally be scoped to specific collections via metadata. Create keys at https://underlay.org/settings/keys (personal) or /:owner/settings/keys (organization). Keys are managed via better-auth's apiKey plugin at /api/auth/api-key/*. 2. Session cookie (for browser use): Users sign in via KF Auth SSO (OAuth2/PKCE) at https://underlay.org/login. Accounts are created automatically on first sign-in, along with a default organization. GET /api/accounts/me returns the current user and their organization memberships. All GET requests are public — no auth required to read public data. All write requests (POST, PATCH, PUT, DELETE) require authentication. If a Bearer token is provided but invalid, the request is rejected immediately (401). ## Rate Limits All API requests are rate-limited per IP (unauthenticated) or per account (authenticated). | Auth status | Limit | |-------------------|---------------| | Unauthenticated | 60 req/min | | Authenticated | 5,000 req/min | Rate limit headers are included on every response: - X-RateLimit-Limit: max requests in the window - X-RateLimit-Remaining: requests left - X-RateLimit-Reset: seconds until the window resets When exceeded, you'll get a 429 response with a Retry-After header. To get the higher limit, authenticate with an API key (recommended for any automated access). --- ## Core Concepts - Organization: an entity that owns collections. Every user gets a default organization on signup. Identified by :slug. Managed via better-auth's organization plugin at /api/auth/organization/*. - Collection: a named, versioned body of data owned by an organization. Identified by :owner/:slug. - Version: an immutable snapshot containing a JSON Schema, records, and file references. Identified by semver (e.g. "v1.0.0"). Major = schema change, Minor = records/files change, Patch = metadata-only change. - Record: a flat JSON object with { id, type, data }. Records are content-addressed: the SHA-256 hash of the canonical JSON `{"id":...,"type":...,"data":...}` is the record's identity. Records are stored globally and deduplicated — the same record in ten collections is stored once. Records reference other records by id and files by hash. Wire format is JSONL (one record per line). - File: a binary blob stored by SHA-256 hash, referenced in record data as {"$file": "sha256:"}. - Schema: a JSON Schema document for a single record type, stored as a global, immutable, content-addressed entity. Each type gets its own schema. Schema changes trigger a major version bump. - Schema labeling: schemas can be labeled post-hoc with URIs or names (e.g. "schema.org/Person") for cross-collection discovery. --- ## Web URLs (for linking humans to a view) The API paths above are for fetching data. When you need to point a person at something in the browser, use these page URLs. The rule: a version is a path prefix, a view is a path segment, and omitting the version prefix means "latest". Semver appears WITHOUT the leading "v" in page URLs (/v/1.2.0), unlike API paths (/versions/v1.2.0). Both forms resolve, but the bare form is canonical. https://underlay.org/:owner → account/organization profile https://underlay.org/:owner/:slug → collection overview (latest): README, record types, stats https://underlay.org/:owner/:slug/records → browse records (latest) — add ?type=TypeName to pick a type https://underlay.org/:owner/:slug/schemas → schemas (latest) https://underlay.org/:owner/:slug/files → files (latest) https://underlay.org/:owner/:slug/versions → version history https://underlay.org/:owner/:slug/versions/compare?from=v1.0.0&to=v1.1.0 → diff two versions https://underlay.org/:owner/:slug/v/1.2.0 → overview pinned to v1.2.0 https://underlay.org/:owner/:slug/v/1.2.0/records → records in v1.2.0 (?type=TypeName) https://underlay.org/:owner/:slug/v/1.2.0/schemas → schemas as pinned in v1.2.0 https://underlay.org/:owner/:slug/v/1.2.0/files → files in v1.2.0 https://underlay.org/records/:hash → a record's content-addressed permalink + provenance https://underlay.org/schemas/:id → a schema's detail page + which collections use it Prefer a version-pinned URL when citing data, since /records (no prefix) follows the latest version and its contents will change on the next push. ARK identifiers (ark:NAAN/id.v1.2.0) also resolve to these pages and are the most durable option. --- ## Reading Data ### Browse collections GET /api/collections → list public collections (?q=search&limit=50&offset=0) GET /api/collections/:owner/:slug → collection metadata + latest version summary ### Read versions GET /api/collections/:owner/:slug/versions → list versions (newest first, ?limit=50&offset=0) GET /api/collections/:owner/:slug/versions/latest → latest version with full metadata GET /api/collections/:owner/:slug/versions/:semver → specific version by semver (e.g. /versions/v1.2.0) ### Read records and files GET /api/collections/:owner/:slug/versions/:semver/records.ndjson → ALL records, streamed as NDJSON in one request (?type=TypeName&after=recordId) GET /api/collections/:owner/:slug/versions/:semver/records → records for a version (?type=TypeName&limit=100&after=) GET /api/collections/:owner/:slug/versions/:semver/manifest → manifest: record ids/types/hashes + file hashes + schema hashes (?since=v1.0.0 for delta) GET /api/collections/:owner/:slug/versions/:semver/files → list files for a version (hash, size, content type) GET /api/collections/:owner/:slug/files/:hash → download a file: access-checked, then 302-redirects to a short-lived presigned URL (follow it, e.g. curl -L). Public files are anonymous; private needs a session or a Bearer share/agent token. This API path is the durable locator — the redirect target is ephemeral, never persist it. HEAD /api/collections/:owner/:slug/files/:hash → check if a file exists/is accessible (returns Content-Length, Content-Type) POST /api/collections/:owner/:slug/files/presign → presign many files at once: body {"hashes":[...]} → {hash: presignedUrl | null}. One round trip for a page full of files. ### Records (global, content-addressed) GET /api/records/:hash/provenance → find all collections/versions containing this record hash POST /api/records/batch → fetch records by hash: {"hashes": ["abc..."]} → JSONL stream ### Diff GET /api/collections/:owner/:slug/versions/:semver/diff?from=:semver → diff between two versions (added, updated, removed records) Returns full record bodies, keyset-paginated: limit defaults to 500, maxes at 5000, and ?cursor= walks the rest. When you only need hashes, prefer .../manifest?since= — its entries are ~120 bytes instead of whole records. ### Export GET /api/collections/:owner/:slug/export → download .tar.gz archive (manifest.json + records/*.ndjson + files/*) GET /api/collections/:owner/:slug/export?version=v2.0.0 → export a specific version Archive layout: manifest.json, records/.ndjson per record type, and files/. Types too large for a single archive entry are split into numbered parts — records/.0000.ndjson, records/.0001.ndjson, ... — at 25,000 records per part. A type that fits in one part keeps the unnumbered records/.ndjson name, so archives of ordinary collections are unchanged. Read every records/*.ndjson entry and you have the version, whichever form it took. The archive streams: the download starts immediately and is produced as you read it, so pipe it to disk or straight into tar rather than buffering it. Exporting the same version again gives a byte-identical archive (entries carry the version's creation time), so a checksum of one export holds for the next. manifest.json is the last entry; its files_missing lists any file that couldn't be fetched. A failure partway through cuts the response off instead of finishing it, so an archive that gunzips and untars without error is complete. Versions over 2,000,000 records return 413 — use records.ndjson for those. Export is capped at 2,000,000 records and returns 413 above that — a guard against handing back a multi-gigabyte tarball from a single GET, not a memory limit. For bulk reads prefer records.ndjson (below): it streams, resumes, and needs no unpacking. The SQL explorer (/api/query/...) has a lower limit of 250,000 records, because unlike export it genuinely does hold the whole version in memory to build a SQLite copy. It is a UI feature rather than a documented API; on a large collection use records.ndjson, or Hot, which hydrates a collection into a queryable database built for the purpose. ### Fork POST /api/collections/:owner/:slug/fork → fork collection into caller's org (requires write auth; a collection-scoped key is refused). 403 if you are NOT a member of the source org and the source's latest version holds any private record, private type, or private field — a fork copies the full record bodies, and there is no redacted-fork path. Body: { "targetOrg": "my-org", "slug": "optional-new-slug" } Creates a new collection under targetOrg with the source's latest version. Records, schemas, and files are referenced (not copied) — zero additional storage. Response includes { id, owner, slug, forkedFrom: { owner, slug, version } }. --- ## Writing Data: The Push Flow This is the canonical workflow for syncing an app's data to Underlay: ### Step 1: Get current state GET /api/collections/:owner/:slug/versions/latest Response includes { semver, hash, recordCount, fileCount }. If 404, no versions exist yet — your first push should use base_version: null. ### Step 2: Fetch the manifest (optional, for diffing) GET /api/collections/:owner/:slug/versions/:semver/manifest?limit=10000 Response: { "semver": "v1.2.0", "hash": "abc123...", "schemas": {"Article": "schema-hash...", ...}, "records": [{"id": "rec-1", "type": "Article", "hash": "record-hash..."}, ...], "files": ["deadbeef...", ...], "pagination": {"limit": 10000, "hasMore": true, "nextCursor": "eyJhZGRlZCI6..."} } The manifest is paginated: limit defaults to 10000 and maxes at 100000. Pass ?cursor= to continue; repeat until hasMore is false. The cursor is opaque — pass back exactly what you were given, don't construct or parse it. This is by far the cheapest way to learn what a version contains: manifest entries are ~120 bytes each, so a million records is one request-order-of-magnitude smaller than fetching the records themselves. Prefer it over walking /records when you only need hashes. For delta manifests, add ?since= to get only the changes between two versions: GET /api/collections/:owner/:slug/versions/:semver/manifest?since=:semver Response: { ..., "delta": { "added": [{"id": "rec-9", "type": "Article", "hash": "..."}], "updated": [{"id": "rec-1", "type": "Article", "hash": "...", "previousHash": "..."}], "removed": [{"id": "rec-4", "type": "Article", "hash": "..."}] }, "pagination": {"limit": 10000, "hasMore": false, "nextCursor": null}, "truncated": false } Deltas are keyset-paginated the same way, so a delta of any size can be walked to completion with ?cursor=. The three lists drain independently, so a page late in the walk may contain only "updated" entries. "truncated" is legacy: it now just mirrors pagination.hasMore, and older clients treated it as "give up and rebuild from the full manifest". If you understand the cursor, page instead of rebuilding. Compare this against your local data to determine what changed. ### Step 3: Upload new files (if any) For each file your app has that Underlay doesn't: PUT /api/collections/:owner/:slug/files/sha256: Content-Type: application/octet-stream Body: raw file bytes The server verifies the SHA-256 hash matches the body. Existing hashes are idempotent (200 OK). Check existence first with HEAD if you want to skip uploads. ### Step 4: Push the version (negotiate protocol) All pushes use the negotiate protocol — a three-step flow similar to git's pack negotiation. For very large collections two of those steps can be broken up (the manifest uploads in chunks, the commit runs in the background) — see 4a-chunked and 4c-async below; the shape is the same. The client sends a manifest of record hashes; the server says which it needs; the client sends only those records; then commits. #### Step 4a: Negotiate POST /api/collections/:owner/:slug/versions/negotiate Content-Type: application/json Authorization: Bearer ul_ { "base_version": "v1.2.0", "message": "Daily archive 2026-04-27", "app_id": "my-app", "actor_id": "my-app:cron-job", "schemas": { "Article": { "type": "object", "properties": { "title": {"type": "string"}, "body": {"type": "string"}, "publishedAt": {"type": "string", "format": "date-time"}, "authorId": {"type": "string", "x-ref-type": "Author"} } }, "Author": { "type": "object", "properties": { "name": {"type": "string"}, "email": {"type": "string", "private": true} } } }, "manifest": [ {"id": "article-42", "type": "Article", "hash": "abc123..."}, {"id": "article-10", "type": "Article", "hash": "def456..."} ], "files": ["7a8b9c..."], "metadata": { "description": "Daily archive of publications" } } Each record hash is SHA-256 of the canonical JSON — see "Record Hashing" section below for the exact algorithm. Clients MUST canonicalize data (sort object keys recursively) before hashing, or the server will reject the records. Field reference: - base_version: the semver string of the version you diffed against (e.g. "v1.2.0"). null for first push. Used for optimistic locking. - schemas: per-type JSON Schema map. Required on every push. - manifest: array of {id, type, hash, private?} for every record in the new version. Capped at 500,000 entries — above that, upload it in chunks instead (see 4a-chunked below). Omit when using manifest_expected. - private (per manifest entry): optional boolean. true hides that record from non-owners in THIS version. OMITTING IT MEANS PUBLIC — it is not inherited from the base version and must be re-sent on every push. Applies identically to inline manifests and JSONL manifest chunks. See "Private Records" below. - manifest_expected: number of distinct record hashes you will upload in chunks. Mutually exclusive with manifest; sending both returns 400. - files: array of file hashes (SHA-256 hex strings) referenced by records. - metadata: optional JSON object for version metadata (description, readme, license, etc.). Merged with previous version's metadata. - message: human-readable commit message (optional). - app_id: identifier for the pushing application (optional). - actor_id: identifier for the user or process that triggered the push (optional). - strip_unknown_fields: if true, records with fields not defined in the schema will have those fields silently stripped. Default: false (returns 422 with extra field list). Response: { "session_id": "uuid", "needed_records": ["def456..."], "needed_files": [], "total_records": 2, "total_files": 1, "already_have_records": 1, "already_have_files": 1 } #### Step 4a-chunked: Large collections — upload the manifest in chunks The manifest above is one JSON body. That is fine up to 500,000 entries; past that it would be hundreds of megabytes parsed in one go, so upload it in chunks instead. Omit "manifest" and declare the count: POST /api/collections/:owner/:slug/versions/negotiate { "base_version": null, "schemas": {...}, "manifest_expected": 3110000, "message": "arXiv metadata" } Response — nothing here is proportional to the collection: { "session_id": "uuid", "manifest_expected": 3110000, "manifest_received": 0, "needed_files": [], "total_files": 0, "already_have_files": 0, "next": "POST .../versions/negotiate/uuid/manifest" } Then send the manifest as JSONL, up to 50,000 entries per request: POST /api/collections/:owner/:slug/versions/negotiate/:sessionId/manifest Content-Type: application/x-ndjson Authorization: Bearer ul_ {"id":"article-42","type":"Article","hash":"abc123..."} {"id":"article-10","type":"Article","hash":"def456..."} Response: { "received": 50000, "needed_records": ["def456..."], "manifest_received": 150000, "manifest_expected": 3110000 } Each response tells you which records from THAT chunk the server needs, so you can start sending record bodies (step 4b) before the whole manifest is uploaded. Chunks are idempotent: entries are keyed by hash, so re-sending a chunk after a timeout is safe and manifest_received will not move. Commit refuses to build a version until manifest_received equals manifest_expected, so a client that dies partway through cannot silently produce a version that dropped records. #### Step 4b: Send needed records POST /api/collections/:owner/:slug/versions/negotiate/:sessionId/records Content-Type: application/x-ndjson Authorization: Bearer ul_ {"id":"article-10","type":"Article","data":{"title":"Updated Title","body":"..."}} Each line is one JSON record. Only send records whose hashes appear in needed_records. Call this endpoint multiple times for large datasets (up to 10,000 records per batch). If needed_records was empty, skip this step entirely. Response: { "received": 1, "remaining": 0 } When remaining reaches 0, all needed records have been received. #### Step 4c: Commit POST /api/collections/:owner/:slug/versions/negotiate/:sessionId/commit Authorization: Bearer ul_ No request body needed. The server validates all records against schemas, computes version hashes, and creates the new immutable version. Response (201): { "semver": "v1.3.0", "hash": "def456...", "recordCount": 2, "fileCount": 1 } #### Step 4c-async: Large collections — commit in the background Commit work is proportional to collection size, so on a very large collection it can run for minutes — longer than a proxy or client will hold a request open. Add ?async=true (or send {"async": true}) and the server accepts the commit and builds the version in the background: POST /api/collections/:owner/:slug/versions/negotiate/:sessionId/commit?async=true Response (202): { "session_id": "uuid", "status": "committing", "message": "Commit accepted. Poll GET .../versions/negotiate/uuid until status is \"committed\" or \"failed\"." } Then poll GET .../versions/negotiate/:sessionId until status is "committed" or "failed": { "session_id": "uuid", "status": "committed", "finalize_started_at": "2026-07-31T12:00:00.000Z", "result": {"semver": "v1.3.0", "hash": "private:def456...", "recordCount": 3110000, "fileCount": 0}, "error": null } On success "result" holds exactly what the synchronous 201 would have returned. On failure "status" is "failed" and "error" holds the rejection body the synchronous path would have returned (same statusCode and shape), so the two paths are interchangeable apart from timing. The version is invisible to readers until the finalize completes — there is no window in which a half-built version can be read. The finalize is server-side work and does not depend on your connection staying open: a client that disconnects while polling can reconnect and read the result. A finalize whose server process dies is swept and marked "failed". ### Step 5: Handle errors Conflict (409 — someone pushed while you were diffing): { "error": "Version conflict", "currentVersion": "v1.3.0", "statusCode": 409 } → Re-negotiate with the new base_version. Missing records (400 — commit called before all records submitted): { "error": "Missing records", "missing_hashes": ["def456..."], "statusCode": 400 } → Send the remaining records via the /records endpoint, then retry commit. Missing files (422 — records reference files not yet uploaded): { "error": "Missing files", "filesNeeded": ["sha256:abc..."], "statusCode": 422 } → Upload the listed files, then retry commit. Extra fields (422 — records contain fields not in the schema): { "error": "Records contain fields not defined in schema", "extraFields": [...], "totalRecords": 12, "statusCode": 422 } → Either fix the records, or re-negotiate with "strip_unknown_fields": true. extraFields lists at most the first 100; totalRecords is how many were affected. Schema validation failures (422) are reported the same way, with "totalErrors". Manifest incomplete (400 — chunked upload, commit called before every chunk arrived): { "error": "Manifest incomplete", "manifest_expected": 3110000, "manifest_received": 3050000, "statusCode": 400 } → Upload the remaining chunks, then retry commit. Manifest too large (413 — inline manifest over 500,000 entries): { "error": "Inline manifests are limited to 500000 entries...", "statusCode": 413 } → Re-negotiate with manifest_expected and upload the manifest in chunks. Sessions expire after 10 minutes of INACTIVITY. Every manifest chunk and record batch pushes the expiry back, so a push that legitimately runs for an hour will not expire underneath you. If a session does expire, re-negotiate. ### Session management GET /api/collections/:owner/:slug/versions/negotiate/:sessionId → check session status DELETE /api/collections/:owner/:slug/versions/negotiate/:sessionId → cancel session (204) Session status is one of: - open — accepting manifest chunks and records - committing — async commit accepted, finalize running in the background - committed — done; "result" holds the version - failed — finalize rejected or died; "error" holds why - expired — timed out or cancelled ### First push (no existing versions) Set base_version to null. Include all records in the manifest. Include schemas for all types. The first version will be v1.0.0. --- ## Pagination (Records Endpoint) The records endpoint uses cursor-based pagination for efficient traversal of large collections. GET /api/collections/:owner/:slug/versions/:semver/records?limit=100&after= Response: { "records": [ ...up to `limit` records... ], "pagination": { "limit": 100, "hasMore": true, "nextCursor": "eyJyIjpbInB1Yi0wMDIiLCJkZWY0NTYiXX0", "total": 2000000 } } Parameters: - limit: max records per page (default 100, max 2000) - after: opaque keyset cursor — pass pagination.nextCursor back unchanged. Records are ordered by (record id, record hash), so records sharing an id across types are never skipped at a page boundary. This is the canonical, scalable method: it is an index seek and stays fast at any depth. `cursor` is accepted as an alias for `after`. A bare record ID is still accepted (records with IDs strictly after it), but it skips other records that share the last ID. - offset: legacy offset-based pagination. Cost grows with the offset, so it is capped: an offset greater than 10000 returns 400. Use `after` to page deeper. - type: filter by record type Notes: - `pagination.total` respects the `type` filter and excludes private types. On collections that mark individual records private it is an upper bound for anonymous callers, since those records are hidden but still counted — use `hasMore` for an exact end-of-set signal. - A query that exceeds the server's statement timeout returns 503 with a Retry-After header. If you hit this on `offset`, switch to `after`. To paginate through all records (works at any collection size): 1. First request: GET .../records?limit=2000 2. If pagination.hasMore is true, use pagination.nextCursor for the next request: GET .../records?limit=2000&after= 3. Repeat until hasMore is false. BUT: if you want the whole collection, do not page it. Use the NDJSON stream below — paging costs one round trip per page purely to re-establish a cursor the server just had. Paging is for browsing a slice; streaming is for reading everything. Ask for the largest page you can handle. Walking a whole collection by paging is bounded by request count, not bytes — 60 requests/minute unauthenticated, 5,000 authenticated — so a 3-million-record collection is 6,200 requests at 500/page and 1,550 at 2,000/page. Authenticate for any full-collection walk. Do NOT paginate large collections with ?offset=; it is capped at 10000 and will 400 beyond that. Use ?after= keyset pagination instead. --- ## Bulk Read (the fast way to get a whole collection) GET /api/collections/:owner/:slug/versions/:semver/records.ndjson Streams every record in the version as newline-delimited JSON in a single response. One request regardless of size: a 3.1M-record collection is one call here versus 1,556 paged ones. The server reads through a database cursor and writes as it goes, so memory is constant on both ends — you can start processing the first line before the last is sent. Content-Type: application/x-ndjson X-Underlay-Record-Count: 3113504 ← how many lines to expect One JSON object per line: {"id":"arxiv:0704.0001","type":"Preprint","data":{...},"hash":"abc123..."} Parameters: - type: restrict to one record type - after: resume — return records with ids strictly after this value Guarantees you can rely on: - Records are ordered by id, ascending. This is what makes `after` work. - `hash` is the same content-address the records endpoint serves: the full record hash for owners, the public hash for everyone else. - Privacy filtering is identical to /records — private types and private records are absent, private fields are stripped. Verify completeness yourself. A stream that dies halfway cannot report an error: the 200 status and headers were already sent. Count the lines and compare against X-Underlay-Record-Count. If they differ, resume with ?after= rather than starting over. X-Underlay-Record-Count is the count for THIS request — privacy-filtered for your access level and scoped to ?type= if you passed one — so the comparison is exact for every caller. Do NOT compare against the version's `recordCount`: that is the full total and includes private records and private types you may not be receiving. That resume behaviour makes this strictly better than paging for bulk reads: the same recovery from a dropped connection, at a fraction of the requests. One edge: a record id is not guaranteed unique within a version — the same id can appear under more than one hash. `after` resumes strictly past the id, so if a stream broke between two lines sharing an id, resuming from it skips the second. Rare, and identical to how ?after= behaves on the paged endpoint, but if you need exactness compare the final line count against X-Underlay-Record-Count and re-read from an earlier id if it falls short. ### Compression All /api/ responses are compressed when you send Accept-Encoding: gzip — about 3x on record data. Most HTTP clients do this automatically. It applies to the NDJSON stream too. ### When to use which - Whole collection, one pass → records.ndjson - A page, or browsing → records?limit=2000&after= - Only ids/types/hashes → manifest (≈120 bytes/record, far smaller than bodies) - What changed since a version → manifest?since= (delta) - Specific records you know hashes for → POST /api/records/batch (up to 10,000 per call) - An archive to keep → export (.tar.gz) --- ## Record Format { "id": "unique-stable-id", "type": "TypeName", "data": { "title": "Some value", "authorId": "author-123", "attachment": {"$file": "sha256:abc123def456..."} } } - id: stable, unique within the collection. Use your app's primary key or generate a deterministic one. - type: groups records by kind (e.g. "Article", "Author", "Grant"). - data: flat JSON object. Reference other records by id. Reference files with {"$file": "sha256:"}. Records are flat — no nested joins. Relationships are expressed by storing the id of the related record. The schema declares which fields are references, so tools can resolve them at read time. --- ## Record Hashing (required for push) Records are content-addressed. The hash is the record's identity across the system — it determines deduplication, manifest membership, and version integrity. Any client that pushes data must compute hashes exactly as the server does, or the push will fail with "Unexpected record hash". ### Algorithm 1. Build the canonical object: `{"id": , "type": , "data": }` - The top-level keys MUST appear in this exact order: id, type, data. - The `data` value MUST be canonicalized (see below). - The `private` flag is NOT part of the hash. Two records with identical id/type/data but different privacy flags produce the same hash. 2. Serialize to JSON with no extra whitespace (standard JSON.stringify behavior). 3. Compute SHA-256 over the UTF-8 bytes of that JSON string. 4. Encode as lowercase hex (64 characters). ### Canonicalization Canonicalization recursively sorts all object keys alphabetically. This ensures that `{"b":1,"a":2}` and `{"a":2,"b":1}` produce the same hash. Rules: - Objects: sort keys lexicographically (by Unicode code point), recurse into values. - Arrays: preserve order, recurse into elements. - Primitives (strings, numbers, booleans, null): unchanged. Reference implementation (JavaScript): ```javascript function canonicalize(value) { if (value === null || typeof value !== 'object') return value; if (Array.isArray(value)) return value.map(canonicalize); const sorted = {}; for (const key of Object.keys(value).sort()) { sorted[key] = canonicalize(value[key]); } return sorted; } function hashRecord(record) { const canonical = JSON.stringify({ id: record.id, type: record.type, data: canonicalize(record.data), }); // Node.js: // const hash = createHash('sha256').update(canonical).digest('hex'); // Browser: // const buf = await crypto.subtle.digest('SHA-256', new TextEncoder().encode(canonical)); // const hash = Array.from(new Uint8Array(buf)).map(b => b.toString(16).padStart(2, '0')).join(''); return { hash, canonical }; } ``` ### Example Input record: { "id": "article-1", "type": "Article", "data": { "title": "Hello", "body": "World" } } Canonical JSON (data keys already sorted): {"id":"article-1","type":"Article","data":{"body":"World","title":"Hello"}} SHA-256 hash: e3b7a... (64 hex characters) Note: if data keys were in a different order (e.g. `{"title":"Hello","body":"World"}`), canonicalization sorts them to `{"body":"World","title":"Hello"}` — producing the same hash. ### Schema hashing Schemas are also content-addressed. The hash is SHA-256 of `JSON.stringify(canonicalize(schemaBody))`. The same canonicalization rules apply. --- ## Schema Discovery Schemas are globally deduplicated by content hash. If two collections use the same type shape, they share the same schema row. This enables cross-collection discovery. GET /api/schemas → search schemas (?q=text&slug=TypeName&label=uri&schema_hash=sha256:...&limit=50&offset=0) GET /api/schemas/:id → single schema with labels and usage info GET /api/collections/:owner/:slug/schemas → schemas for latest version (?version=v1.2.0 for specific, ?raw=true to skip label enrichment) POST /api/schemas/:id/labels → add label {"label": "schema.org/Person"} (requires write scope) DELETE /api/schemas/:id/labels/:label → remove label (requires admin scope) When schemas are returned via the collection schemas endpoint, known labels are injected as "x-underlay-labels" on the schema body (opt-out with ?raw=true). --- ## Versioning - Versions are identified solely by semver (e.g. "v1.0.0", "v1.2.3"). - Semver semantics: major bump = schema change (types added/removed or schema_id changed), minor bump = records/files change, patch bump = metadata-only change. - Each version has a content-addressed hash computed from sorted schema hashes + sorted record hashes + sorted file hashes + metadata. - Each version has a `metadata` field: a JSON object that can contain `readme`, `license`, and other arbitrary metadata. - Records are content-addressed: SHA-256 of canonical JSON (see "Record Hashing" section). Records are globally deduplicated. - Versions are immutable once created. - The provenance endpoint (GET /api/records/:hash/provenance) shows every collection and version that includes a given record. --- ## Organization Management Organizations are managed via better-auth's organization plugin at /api/auth/organization/*. Every user gets a default organization on signup. Users can create additional organizations. POST /api/auth/organization/create → create org {"name", "slug"} GET /api/auth/organization/list → list user's organizations PATCH /api/auth/organization/update → update org DELETE /api/auth/organization/delete → delete org Member management (invite, remove, update roles) is also under /api/auth/organization/*. ## Collection Management POST /api/accounts/:owner/collections → create collection {"slug", "name", "public"} PATCH /api/collections/:owner/:slug → update {"name", "slug", "public"} PATCH /api/collections/:owner/:slug/metadata → update version metadata {"description", "readme", "license", ...} — creates a patch version (requires write scope) DELETE /api/collections/:owner/:slug → delete collection (requires admin scope) GET /api/accounts/:owner/collections → list collections for an organization --- ## Privacy & Visibility Underlay supports privacy at four levels that compose — a reader sees content only if it passes all of them: 1. Collection (collections.public) — a private collection 404s entirely; nothing below is evaluated. 2. Type — "private": true on the type's schema. Per-version. 3. Field — "private": true on the property. Per-version. 4. Record — "private": true on the negotiate MANIFEST ENTRY. Per-version. Levels 2-4 are bound to the version, so the same content can be public in one collection's version and hidden in another's. Private data is stored alongside public data in the same version and is visible only to members of the owning organization. ### Private Types Mark an entire type as private in its schema. All records of that type are hidden from public readers. "schemas": { "Article": { "type": "object", "properties": { "title": {"type": "string"} } }, "InternalNote": { "type": "object", "private": true, "properties": { "note": {"type": "string"}, "articleId": {"type": "string"} } } } Public readers see only the Article type. InternalNote is completely hidden (including from the schema response). ### Private Fields Mark individual fields as private within a type's schema. The type itself is visible, but those fields are stripped for public readers. "Author": { "type": "object", "properties": { "name": {"type": "string"}, "email": {"type": "string", "private": true}, "phone": {"type": "string", "private": true} } } Public readers see Author records with only "name". The owner sees all fields. ### Private Records Mark individual records as private on their MANIFEST ENTRY in the negotiate body. The type and schema stay visible; that specific record is hidden from non-owners. "manifest": [ {"id": "article-1", "type": "Article", "hash": ""}, {"id": "article-2", "type": "Article", "hash": "", "private": true} ] article-2 is only visible to members of the owning org. Public readers see article-1 only. Two rules you must design around: 1. PRIVACY IS PER-VERSION AND MUST BE RE-DECLARED ON EVERY PUSH. The flag is stored on the (version, record) edge, not on the record itself. Omitting `private` on a manifest entry means PUBLIC — it does NOT inherit the previous version's value. A push that sends a full manifest without the flags publishes everything it omits. (Dropping the flag is therefore also how you deliberately un-hide a record.) Read the current flags back from GET .../versions/:semver/manifest, which echoes `private: true` on the entries that have it. 2. REDACTION IS FORWARD-ONLY. Marking a record private in v2 hides it in v2 only. Versions are immutable, so v1 still serves that record at /versions/v1.0.0/records. There is no retroactive purge. Files are looser still: file access resolves across ALL ready versions, so a file referenced publicly in v1 stays downloadable after the referencing record is redacted in v2. `private` belongs ONLY on the manifest entry. A `private` key on a record BODY sent in step 4b is parsed and silently ignored — no error, and the record ships public. The same content can be private in one collection and public in another: the flag lives on the version edge, not on the globally deduplicated record body. ### How it works - Public hash: the digest of the public projection — private types omitted, private records omitted, private fields stripped from the records that remain. Version metadata and the file list are NOT filtered: both digests are computed over the same metadata and the same files. - Version identity is BOTH digests. A push is a duplicate (409 "No changes detected") only if `hash` AND `public_hash` match an existing version. Privacy is deliberately not folded into `hash`, so `hash` stays independently verifiable from the content; privacy moves `public_hash` instead. That is why re-pushing byte-identical content with a record newly marked private is a legitimate new version — it is what makes redaction-in-place possible. Converse: flagging a record private whose TYPE is already private changes neither digest and is correctly rejected as a duplicate, because the flag has no observable effect. - A privacy-only push is a PATCH bump (schema change → major, record-set change → minor, everything else → patch). - Public record hash: a record of a type with private fields is listed in public manifests under the hash of its filtered projection ({"id", "type", "data"} with private fields stripped). Record endpoints resolve either address; hashing the document you receive always reproduces the address you requested. - Private hash: the full digest over ALL schema and record hashes, plus files and metadata. Served only to org members; everyone else receives public_hash in the `hash` field. Both are prefixed — "private:<64hex>" and "public:<64hex>" — so the two forms are never confusable. - Schema filtering: the schema returned to public readers omits private types and private fields - Record filtering: queries by non-owners automatically exclude private records and strip private fields --- ## API Key Management API keys are managed via better-auth's apiKey plugin. All endpoints are under /api/auth/api-key/*. POST /api/auth/api-key/create → create key {"name": "my-app", "metadata": {"scope": "write"}, "prefix": "ul"} The scope in metadata is translated to permissions server-side. Response includes the key once: {"key": "ul_abc123...", "id": "..."} GET /api/auth/api-key/list → list keys (id, name, start, permissions, metadata, createdAt, expiresAt) POST /api/auth/api-key/delete → revoke a key {"keyId": "..."} --- ## Error Codes 400 — Validation error (bad request body) 401 — Authentication required 403 — Insufficient scope (e.g. read key used for write) 404 — Not found (or private collection you can't access) 409 — Version conflict (re-fetch and retry with new base_version) 413 — Payload too large (file upload exceeds size limit) 422 — Missing files, schema validation failed, or records contain extra fields not in the schema 429 — Rate limited (wait and retry) --- Full documentation: https://underlay.org/docs Protocol specification: https://underlay.org/protocol Integration guide: https://underlay.org/docs/integration Source code: https://github.com/knowledgefutures/underlay