Integration Guide
Everything a developer or LLM needs to push data to the registry. No SDK required. HTTPS and JSON. For a machine-readable version, see llms.txt.
What is Underlay?
Underlay is a versioned registry for structured knowledge. Apps publish versions of their data; Underlay keeps every version in content-addressed trees, so versions share what they have in common, and serves them via a stable API. Think npm for data, or Docker Hub for structured content.
Core Concepts
- Collection: A named, versioned body of structured data. Identified by
:owner/:slug. - Version: An immutable snapshot: a JSON Schema per type, records, file references and metadata. Identified by semver (e.g.
v1.0.0) or by itsulv2:hash. - Record: An
id, atypeand adatapayload conforming to the type’s schema. Content-addressed by SHA-256 hash. - File: A binary blob (PDF, image, etc.) stored by SHA-256 hash. Referenced in records via
{"$file": "sha256:<hex>"}.
Authentication
Create an API key at /settings/keys or via the API. Pass it as:
Authorization: Bearer ul_your_key_hereKeys are scoped read, write or admin; use write for pushing data. A key never exceeds its holder’s role: a write key acts as a member, and only an owner’s or admin’s admin key can change visibility, delete or manage webhooks. A key confined to specific collections acts as a member whatever its scope. To let an AI agent push to one collection, create an agent link from the collection’s Share panel: an instructions page at https://www.underlay.org/agent/<key> carrying a write key for that collection only, which expires after an hour.
The Push Flow
Every push is a delta push: you send what changed since a base version, and the server builds the new version.
- Open a session with
baseset to the semver you diffed against (nullfor the first push;nullmeans no conflict check), plus any schemas, metadata and files to declare.schemas, when sent, is the full type set: a type left out is removed.metadatareplaces the metadata,metadata_patchmerges into it. The response listsneeded_filesand the server’slimits. - Upload files listed in
needed_files, by hash. - Upload records that are new or changed, as NDJSON, in batches within
limits.batch_linesandlimits.batch_bytes. - Upload deletes,
{"type", "id"}per line, for records that are gone. - Commit. On
409 Conflict, someone else published first: diff against the new head and push again.
# 1. Open a session against the current version (its semver, e.g. "v1.2.0")
curl -X POST https://www.underlay.org/api/collections/:owner/:slug/push \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $KEY" \
-d '{
"base": "v1.2.0",
"message": "Daily sync",
"files": {"add": ["9f86d0..."]}
}'
# → {"session_id":"...","base":"v1.2.0","needed_files":["9f86d0..."],"limits":{...},...}
# 2. Upload the files the server doesn't have yet
curl -X PUT "https://www.underlay.org/api/collections/:owner/:slug/files/sha256:9f86d0..." \
-H "Authorization: Bearer $KEY" \
-H "Content-Type: application/pdf" \
--data-binary @paper.pdf
# 3. Upload new and changed records (NDJSON, repeatable)
curl -X POST .../push/SESSION_ID/records \
-H "Content-Type: application/x-ndjson" \
-H "Authorization: Bearer $KEY" \
--data-binary '{"id":"record-2","type":"Article","data":{...}}'
# 4. Delete records that are gone (NDJSON, repeatable)
curl -X POST .../push/SESSION_ID/deletes \
-H "Content-Type: application/x-ndjson" \
-H "Authorization: Bearer $KEY" \
--data-binary '{"type":"Article","id":"record-9"}'
# 5. Commit
curl -X POST .../push/SESSION_ID/commit \
-H "Authorization: Bearer $KEY"
# → 201 {"semver":"v1.3.0","hash":"ulv2:...","recordCount":...,"fileCount":...,"changes":{...}}Within a session, the later upload of a (type, id) wins, whether a record or a delete. Above 100,000 uploaded records (or when a schema change revalidates more than 100,000 existing records) the commit runs in the background: it answers 202, and you poll GET .../push/SESSION_ID until its status is committed or failed. Add ?async=true to ask for that at any size.
Pushing a Full Export
If your app exports its whole dataset each time rather than tracking changes, diff the export against the current version’s manifest, then push the differences. The upload is the size of the changes, whatever the size of the collection. See clients without a copy.
// hashRecord(record) → hex, as defined under "Record Hashing" below.
const key = (r) => JSON.stringify([r.type, r.id])
// 1. Read the current version's manifest, every page.
const have = new Map() // key → { id, type, hash, private? }
let base = null
let cursor = null
do {
const query = cursor ? `?cursor=${encodeURIComponent(cursor)}` : ''
const res = await fetch(`${api}/versions/latest/manifest${query}`, { headers: auth })
if (res.status === 404) break // no versions yet
const page = await res.json()
base = page.semver
for (const m of page.records) have.set(key(m), m)
cursor = page.pagination.hasMore ? page.pagination.nextCursor : null
} while (cursor)
// 2. Diff your full export against it. A record of a private type is private
// whatever its own flag says (privateTypes: the types whose schema has "private": true).
const isPrivate = (r) => !!r.private || privateTypes.has(r.type)
const upserts = records.filter((r) => {
const m = have.get(key(r))
return !m || m.hash !== hashRecord(r) || !!m.private !== isPrivate(r)
})
const keep = new Set(records.map(key))
const deletes = [...have.values()]
.filter((m) => !keep.has(key(m)))
.map((m) => ({ type: m.type, id: m.id }))
// 3. Open a session with "base": base, upload upserts and deletes in batches, commit.A record’s set is part of the diff: a record that should become private or public is uploaded again with its new private flag.
Record Hashing
The server hashes the records you upload, so a push needs no hashing. You need the hash to diff against a manifest. It is the SHA-256 of a fixed { id, type, data } envelope with data in canonical JSON (RFC 8785), so any implementation produces the same hash for the same content. In JavaScript:
import { createHash } from 'node:crypto'
// RFC 8785 (JCS), written out as a string: a sorted object passed to
// JSON.stringify would put integer-like keys ("9", "10") first.
function jcs(value) {
if (value === null || typeof value !== 'object') return JSON.stringify(value)
if (Array.isArray(value)) return '[' + value.map(jcs).join(',') + ']'
const keys = Object.keys(value).sort()
return '{' + keys.map((k) => JSON.stringify(k) + ':' + jcs(value[k])).join(',') + '}'
}
function hashRecord(record) {
const canonical =
'{"id":' + JSON.stringify(record.id) + ',"type":' + JSON.stringify(record.type) +
',"data":' + jcs(record.data) + '}'
return createHash('sha256').update(canonical).digest('hex')
}See Records and schemas for the full rules, including the input rules every record line must pass.
Record Format
Every record has three fields: id (stable string), type (matches schema), and data (the payload).
- Relationships are plain ID strings (e.g.
"authorId": "author-1") - Files are referenced as
{"$file": "sha256:<hex>"} - No joins. Prefer flat records; nesting is allowed up to 64 levels
Metadata
Each version carries a metadata object that can include description, readme, license, and any other key-value pairs. Metadata lives on the version, not the collection; it's versioned alongside your data. Set it on your first push and update it via subsequent pushes or the metadata endpoint.
To update metadata without changing records or schemas (e.g. editing the readme), POST /api/collections/:owner/:slug/metadata with the fields to change (null clears one). This creates a patch version over the same trees, so it is quick at any collection size.
First Push Example
The body that opens the first push. Include schemas (a per-type JSON Schema map) and metadata:
{
"base": null,
"message": "Initial import",
"app_id": "my-app",
"metadata": {
"description": "Articles and authors from my app",
"readme": "# My App Data\nExported from the app database."
},
"schemas": {
"Article": {
"type": "object",
"properties": {
"title": {"type": "string"},
"body": {"type": "string"},
"authorId": {"type": "string"},
"publishedAt": {"type": "string", "format": "date-time"}
}
},
"Author": {
"type": "object",
"properties": {
"name": {"type": "string"},
"email": {"type": "string"}
}
}
}
}Then upload the records as NDJSON and commit. See the Quickstart for the complete curl walkthrough.
Mapping a SQL Database
Most apps store data in SQL. Here's how to map it to Underlay records:
-- For each table, generate a JSON Schema type:
-- table name → type name
-- column name → property name
-- column type → JSON Schema type (text→string, integer→integer, etc.)
-- foreign keys → note as ID references in the schema description
-- Example: a "publications" table with columns (id, title, doi, author_id)
-- becomes a "Publication" type with properties {title: string, doi: string, authorId: string}
-- The record id is the primary key value.General rules:
- Each table becomes a record type
- Each row becomes a record (primary key → record
id) - Foreign keys become string ID references
- Binary columns (BLOBs) → upload as files, replace with
$filereferences - Generate a JSON Schema from your column types
Versioning
Versions are identified by semver (e.g. v1.0.0). The semver is derived automatically from what changed:
- Schema changes → major bump
- Record changes, including a record becoming private or public → minor bump
- Metadata or file changes only → patch bump
The first version of a collection is always v1.0.0. A push that changes nothing makes no version. The base when opening a push is a semver string (or null for the first push).
Privacy
You can control what's publicly visible at two levels:
- Private types: Add
"private": trueat the root of a type’s schema. All records of that type are hidden from public readers. - Private records: Add
"private": trueto a record line when uploading it. That record is hidden from public readers.
"private": true on a field inside a schema is refused. Put private fields in a private type, or push the whole record as private.
Private content is stored in the same version; members of the owning organization see everything, and public readers see the public set only. The version hash covers the private set through a salted commitment, so public readers can verify the public set without learning anything about the private one.
API Reference
Full API docs are at /docs. The key endpoints:
POST .../push | Open a push session against a base version |
POST .../push/:id/records | Upload new or changed records (NDJSON, repeatable) |
POST .../push/:id/deletes | Delete records by type and id (NDJSON, repeatable) |
POST .../push/:id/commit | Build the version. Add ?async=true to get a 202 and poll instead of holding the request open |
GET .../push/:id | Session status, and the result or error of an async commit |
DELETE .../push/:id | Abandon a push session |
GET .../versions/latest | Get latest version |
GET .../versions/:semver/records | Get records (paginated) |
GET .../versions/:semver/records.ndjson | Stream every record in one request (NDJSON). The bulk read path — use this instead of paging when you want the whole collection |
GET .../versions/:semver/manifest | Record ids, types and hashes, paged (supports delta via ?since=) |
GET .../versions/:semver/diff?from= | Diff two versions |
PUT .../files/:hash | Upload a file |
POST /api/records/batch | Fetch up to 100 records by hash (NDJSON response) |
GET /api/records/:hash/provenance | Find which collections you can read contain a record |
GET /api/collections | Browse public collections |
Unknown Fields
If a record has top-level fields its schema’s properties doesn’t list, the records upload answers 422 with the extra fields per line. To strip those fields instead, set "strip_unknown_fields": true when opening the push.
When stripping is enabled, the server removes the extra fields before hashing, and stores only the schema-conformant data.
Error Handling
409 Conflict: Another version was published since yourbase(the answer namescurrentVersion, at open or at commit), or the push changes nothing. Diff against the new head and push again.413 Payload Too Large: A batch or file is over the session’slimits. Split it.422 Unprocessable: Records or deletes fail the input rules or their schema (validationErrors, by line), a schema is refused when the session opens, or the commit references files that haven’t been uploaded (filesNeeded).429 Too Many Requests: Too many push sessions open at once (limits.open_sessions), or a rate limit. Wait and retry.503 Service Unavailable: Storage cleanup ran while the push was committing. Push again.
Pushing from Scripts
The most common pattern for pushing data from a script, cron job, or CI pipeline:
- Query your source (database, API, filesystem) and build an array of records in
{id, type, data}format. - Diff against the current version’s manifest, unless your app already knows what changed. See Pushing a Full Export above.
- Open a push with
baseset to that version. - Upload the new and changed records and the deletes as NDJSON, in batches within the session’s
limits. - Commit to create the version.
A minimal Node.js or Python script typically takes 30-50 lines: query your data, map rows to records, diff, push. No SDK needed. See the Quickstart for a curl-based walkthrough.
Source Code
Underlay is open source: github.com/knowledgefutures/underlay
Built by Knowledge Futures, a 501(c)(3) public charity. Contact: team@knowledgefutures.org