Versions API
Versions are the core of Underlay. Each version is an immutable snapshot of a collection: schemas, records and file references. You publish one with a delta push: open a session against the version you started from, upload the records you add or change and the ids you delete, and commit. The same exchange works against any Underlay node (Push and pull).
Version hashes are ulv2:<sha256>, the hash of the version’s root, and are the same for every reader: the private set is in the root only as a salted commitment (see Trees and versions). Wherever a version is named in a path (:n), it can be a semver (v1.1.0), a version hash, or latest. Record, schema and file hashes are bare hex with no prefix.
A version never changes, so responses under .../versions/:n are cached when :n is a semver or hash: public, max-age=600, stale-while-revalidate=60 for anonymous readers, private, max-age=3600 for signed-in ones. latest moves, so it isn’t cached. Only 200s are cached.
Delta push (open → upload → commit)
A push costs work in proportion to what changed, not to the size of the collection. Only the records you send are validated and hashed; unchanged records are never re-sent.
POST /api/collections/:owner/:slug/push
Auth: write scope
Open a session. Every field is optional.
Request
{
"base": "v1.0.0",
"schemas": {
"Publication": {
"type": "object",
"properties": {
"title": {"type": "string"},
"pdf": {"type": "object"}
}
}
},
"metadata_patch": {"description": "PubPub archive"},
"files": {"add": ["7a8b9c..."]},
"message": "Add new publications"
}Fields
base | The semver you started from. If it isn’t the collection’s head, the answer is 409 with currentVersion. null or absent applies the changes to whatever the head is (use null for the first version). |
schemas | The full type set, {"TypeName": schema}. It replaces the base’s: a type left out is removed with its records. Absent keeps the base’s types. Every schema of the resulting type set, kept ones included, is checked in full when the session opens (the draft-07 meta-schema, patterns, $refs that resolve, and the schema rules); one that fails is 422. |
metadata / metadata_patch | metadata replaces the version metadata (description, readme, license, …) and must be an object or null; metadata_patch, an object, merges its top-level members into the base’s. Anything else is 400. With neither, the metadata is kept. |
files | {"add": [hash, …], "remove": [hash, …]}: files to declare or drop beyond those your records reference. Each hash is bare 64-character lowercase hex. Any other string, including one with a sha256: prefix, is ignored without an error. |
message, app_id, actor_id | Strings recorded in the version’s signed log entry. |
strip_unknown_fields | If true, top-level fields a record’s schema doesn’t list in properties are dropped before the record is hashed, instead of refusing the record. |
Response 200
{
"session_id": "uuid",
"base": "v1.0.0",
"needed_files": ["7a8b9c..."],
"expires_at": "2026-10-04T13:00:00.000Z",
"limits": {
"open_bytes": 8388608,
"batch_bytes": 16777216,
"batch_lines": 10000,
"session_idle_seconds": 3600,
"open_sessions": 20,
"file_bytes": 33554432
}
}needed_files are the declared files this collection doesn’t hold yet (files already uploaded and verified for it are held): upload them before you commit. limits are this node’s; size your batches by them. The session expires after session_idle_seconds without an upload. Each records or deletes batch pushes expires_at back; file uploads don’t.
POST .../push/:sid/records
Auth: write scope
Upserts as NDJSON (Content-Type: application/x-ndjson), one record per line: {"id", "type", "data", "private"?}. Call it as many times as you need, up to batch_lines lines and batch_bytes bytes each.
Request
{"id":"pub-002","type":"Publication","data":{"title":"New Paper"}}
{"id":"pub-003","type":"Publication","data":{"title":"Draft"},"private":true}Response 200
{ "received": 2 }Each line passes the input rules and its type’s schema, and its type must be in the session’s type set. If any line fails, the answer is 422 and nothing from the batch is stored. validationErrors lists the first 100 failing lines, each with its 1-based line number; totalErrors is the full count.
{
"error": "Invalid records",
"validationErrors": [
{"line": 2, "recordId": "pub-003", "type": "Publication",
"errors": ["..."]}
],
"totalErrors": 1,
"statusCode": 422
}A batch with no lines is 400. A session that is no longer open is 409 ("Session is not open").
POST .../push/:sid/deletes
Auth: write scope
Deletes as NDJSON, one {"type", "id"} per line; the answer is { "received": n }. Deleting an id the base doesn’t hold is not an error. Within a session the later upload of a (type, id) wins, whether it is a record or a delete.
{"type":"Publication","id":"pub-old"}Each line passes the input rules, as a record line does, and its id and type the same checks. A line that fails them, isn’t {"type", "id"}, or names a type not in the session’s type set fails the batch: 422, with validationErrors (the first 100 failing lines) and totalErrors as for records, and nothing from the batch is stored. An empty batch is 400; a session that is no longer open is 409.
PUT .../files/:hash
Auth: write scope
Upload a file’s bytes under its SHA-256 hash: 201, or 400 if the bytes don’t match the hash. Files over file_bytes are a 413; upload those through the Files API. Every file a new record references must be uploaded before the commit.
POST .../push/:sid/commit
Auth: write scope
Build the version. No body is needed; {"async": true} is the one field read.
Response 201
{
"semver": "v1.1.0",
"hash": "ulv2:a1b2c3d4...",
"recordCount": 3,
"fileCount": 1,
"changes": {"added": 2, "removed": 1, "updated": 0}
}recordCount here is every record in the version, private ones included: the pusher is a member. Reads of the collection and its versions count what the caller may see, so a public reader’s count leaves private records out.
The semver follows from what changed against the base: major when a type was added or removed or a schema changed, otherwise minor when any record was added, removed or changed (moving between public and private counts), otherwise patch (metadata or files only). A push that changes nothing makes no version and answers 409.
Committing a session that is already committed answers its 201 result again. A session in any other state that isn’t open is 409 ("Session is <status>").
With ?async=true, ?async=1 or a body of {"async": true}, the answer is 202 and the version is built in the background. The commit runs this way without being asked when the session uploaded more than 100,000 records, or when a schema change means more than 100,000 of the base’s records must be revalidated.
{
"session_id": "uuid",
"status": "committing",
"poll": "GET /api/collections/:owner/:slug/push/uuid"
}A synchronous commit can also answer 202, with {"session_id", "status": "committing"} and no poll, when the server splits a large commit into parallel jobs. Treat both the same way: poll GET .../push/:sid until status is committed (result is the 201 body) or failed (error is the rejection). The version isn’t visible to readers until it is committed.
A commit can be refused:
409{"error": "Version conflict", "currentVersion"}: the head moved after the session opened;currentVersionis the head now. Read it and push again.409"No changes detected", with the head’shash.422"Missing files":filesNeededlists up to 100 files not uploaded, as bare hex.422"Schema validation failed", withvalidationErrorsandtotalErrors: a schema change makes records of the base invalid.503"Storage cleanup ran while this push was committing. Push again."
Clients that keep no copy
A client that exports its whole dataset each time, rather than tracking changes, reads the head’s manifest, diffs against it, and pushes only the differences. The upload is the size of the changes, whatever the size of the collection.
# 1. What the head holds: page the manifest (members also see private records)
GET .../versions/latest/manifest?limit=25000 # then ?cursor=<nextCursor>
# 2. Compare by (type, id): hash your current records (canonical form, SHA-256)
# new, changed hash, or changed privacy -> upsert
# in the manifest but not in your data -> delete
# 3. Push only those, against the manifest's semver
POST .../push {"base": "<manifest semver>"}
POST .../push/:sid/records ...
POST .../push/:sid/deletes ...
POST .../push/:sid/commitPrivacy
- Type-level:
"private": trueat a schema’s root hides every record of that type from people outside the owning organization. - Record-level:
"private": trueon a record line puts that record in the version’s private set. It belongs to this version’s reference to the record: a record keeps its flag until a later upload of the same id changes it.
"private": true on a property inside a schema (field-level privacy) is refused. Put private fields in a private type, or push the whole record as private. Privacy is per version: marking a record private in v1.1.0 doesn’t change v1.0.0, which still serves it.
Errors
400 | A body that isn’t a JSON object, metadata that isn’t an object or null, metadata_patch that isn’t an object, an empty records or deletes batch, or file bytes that don’t match their hash. |
401 | You aren’t signed in. |
403 | You can read the collection but not write to it, or the session is another user’s. |
404 | Collection or session not found, or not visible to you. |
409 | base isn’t the head, or the head moved before the commit (both with currentVersion, the head now); the session isn’t open; or the push changes nothing ("No changes detected"). |
413 | A body over open_bytes or batch_bytes, a batch over batch_lines, or a file over file_bytes. |
422 | Records or deletes that fail (validationErrors), a line whose type isn’t in the session’s type set, a schema that is refused (at open), a commit with files not uploaded (filesNeeded), or a schema change that base records fail ("Schema validation failed"). |
429 | You have open_sessions sessions in progress (commit or abandon one), or a rate limit (Retry-After). |
503 | Storage cleanup ran while the push was committing. Push again. |
GET /api/collections/:owner/:slug/push/:sid
Auth: write scope; your own sessions only
A push session’s status. Poll it after a commit answers 202.
Response 200
{
"session_id": "uuid",
"status": "committed",
"records_received": 3110000,
"expires_at": "2026-10-04T13:00:00.000Z",
"created_at": "2026-10-04T12:00:00.000Z",
"finalize_started_at": "2026-10-04T12:20:00.000Z",
"result": {
"semver": "v1.1.0",
"hash": "ulv2:a1b2c3d4...",
"recordCount": 3110000,
"fileCount": 0,
"changes": {"added": 3110000, "removed": 0, "updated": 0}
},
"error": null
}status | open (taking uploads), committing, committed, failed, or expired (abandoned, or idle too long). |
records_received | Record lines accepted so far, across batches. |
finalize_started_at | When the commit began, or null. |
result | Once committed, the commit’s 201 body; otherwise null. |
error | Once failed, the rejection body (for example {"error": "Version conflict", "statusCode": 409}); otherwise null. |
DELETE /api/collections/:owner/:slug/push/:sid
Auth: write scope; your own sessions only
Abandon an open session: it becomes expired and stops counting toward open_sessions. The answer is {"ok": true}. A session that isn’t open is left as it is.
GET /api/collections/:owner/:slug/versions
No auth for public collections
List versions, newest first.
Query parameters
limit | Max results (default 50, max 100) |
offset | Pagination offset |
Response 200
[
{
"semver": "v1.1.0",
"major": 1,
"minor": 1,
"patch": 0,
"hash": "ulv2:a1b2c3d4...",
"baseSemver": "v1.0.0",
"message": "Add new publications",
"appId": "pubpub-sync",
"pushedBy": "user-42",
"pushedByName": "Ada Lovelace",
"pushedBySlug": "ada",
"actorId": "user-42",
"recordCount": 150,
"fileCount": 12,
"totalBytes": 52428800,
"typeCounts": {"Publication": 150},
"createdAt": "2026-04-01T00:00:00.000Z",
"ark": "https://www.underlay.org/ark:12345/ulb9bq4n5gmv3k0.v1.1.0"
}
]recordCount, fileCount, totalBytes and typeCounts count the sets the caller may read: for anyone outside the collection’s members, public records only. baseSemver is the version the push started from (null for the first). pushedBy, pushedByName, pushedBySlug (the pusher’s personal account) and actorId are for the collection’s members only. ark is null when the collection’s ARK is off.
GET /api/collections/:owner/:slug/versions/latest
No auth for public collections
Get the most recent version. Returns the same object as .../versions/:n. Not cached.
GET /api/collections/:owner/:slug/versions/:n
No auth for public collections
Get a specific version by semver (e.g. v1.1.0) or version hash. Returns the version object as listed above, plus metadata and the schemas of the types the caller may read.
With ?records=<type> (empty for the first type), it also returns recordsPage, {"type", "records", "total"}: a page of that type’s records (offset, limit as on /records) and the type’s total. schemas then holds only that type’s schema; typeCounts still lists every type. The records page uses this to need only one call.
GET /api/collections/:owner/:slug/versions/:n/records
No auth for public collections
Get records for a specific version, in (type, id) order: types in slug order (UTF-8 byte order), then ids within a type.
Query parameters
type | Filter by record type |
limit | Max results (default 100, max 2000) |
after | Opaque keyset cursor from pagination.nextCursor; it names the (type, id) of the last record returned. Ids are unique within a type, so nothing is skipped at a page boundary. cursor is an alias. A bare record id is accepted only with ?type=; without it, the id is ignored and paging starts at the beginning. |
offset | Skip this many records. It is a tree seek, O(tree height), and works at any depth. Ignored when after is given. |
Walking a whole collection is bounded by request count, not bytes: 60 rate units a minute anonymous, 5,000 signed in, and a page costs 1. Ask for the largest page you can handle: a 3-million-record collection is 6,000 requests at 500 per page and 1,500 at 2,000 per page. records.ndjson reads it in one request for 10 units.
Response 200
{
"records": [
{
"id": "pub-001",
"type": "Publication",
"data": {
"title": "Example Paper",
"doi": "10.1234/example"
},
"hash": "def456..."
},
{
"id": "pub-002",
"type": "Publication",
"data": {"title": "Draft"},
"hash": "789abc...",
"private": true
}
],
"pagination": {
"limit": 2,
"hasMore": true,
"nextCursor": "eyJ0IjoiUHVibGljYXRpb24iLCJrIjoicHViLTAwMiJ9",
"total": 150
}
}Each record carries its hash. A member’s view includes private records, marked "private": true; anyone else sees public records only.
Use pagination.nextCursor as the after parameter in the next request, unchanged — treat it as opaque. When hasMore is false, you’ve reached the end.
pagination.total is the exact count of records in the sets the caller may read (public only for non-members), within type if given. The one exception: records withheld from serving are counted but skipped.
GET /api/collections/:owner/:slug/versions/:n/records/:type/:id
No auth for public collections
One record at a version. A member’s private record is marked "private": true. A record that doesn’t exist, or that the caller may not read, is 404.
Response 200
{
"id": "pub-001",
"type": "Publication",
"data": {"title": "Example Paper"},
"hash": "def456...",
"semver": "v1.1.0"
}GET /api/collections/:owner/:slug/records/:type/:id/history
No auth for public collections
A record’s history: each version where it was added, updated or removed, oldest first. hash is the record’s hash in that version, null when removed. Only the sets the caller may read are consulted, so for a non-member a record made private reads as removed. It looks back over the newest 500 versions; truncated is true when there may be older ones. A record never seen is 404. It costs 5 rate units.
Response 200
{
"type": "Publication",
"id": "pub-001",
"changes": [
{"seq": 1, "semver": "v1.0.0", "createdAt": "2026-03-01T00:00:00.000Z",
"change": "added", "hash": "abc123..."},
{"seq": 2, "semver": "v1.1.0", "createdAt": "2026-04-01T00:00:00.000Z",
"change": "updated", "hash": "def456..."}
],
"truncated": false
}GET /api/collections/:owner/:slug/versions/:n/records.ndjson
No auth for public collections
Every record in the version, streamed as newline-delimited JSON in a single response. This is the bulk read path: one request, costing 10 rate units, against 1,500 pages for a 3-million-record collection. The server streams from the version’s record trees as it reads them, so memory stays constant on both ends and you can process the first line before the last is sent.
Lines come in the same order as /records: types in slug order, then ids. A member gets private records too, each line ending "private":true; anyone else gets public records only.
Query parameters
type | Restrict to a single record type |
after | Resume after the record with this id. With after_type, the stream continues after (after_type, after) through the later types; with type, it stays within that type. With neither, it is 400. |
after_type | The type of the after record. Without after it is 400. |
Response 200
HTTP/1.1 200 OK
Content-Type: application/x-ndjson
X-Underlay-Record-Count: 3110000
{"id":"pub-001","type":"Publication","data":{"title":"..."},"hash":"def456..."}
{"id":"pub-002","type":"Publication","data":{"title":"..."},"hash":"789abc...","private":true}
{"id":"pub-003","type":"Publication","data":{"title":"..."},"hash":"a1b2c3..."}hash is the same content address /records serves.
Check completeness yourself. A stream that fails partway cannot report it: the 200 and headers were sent before anything went wrong. X-Underlay-Record-Count tells you how many lines to expect: the count for this request, for the sets you may read and within ?type= if you passed one, and after the resume point if you gave one. It can exceed the lines sent only by records withheld from serving. (Don’t compare against the version’s recordCount: it covers every type.)
If you receive fewer, resume rather than start over: take the type and id of the last complete line you parsed and request ?after_type=<type>&after=<id>. The stream picks up after that record and runs to the end of the version.
GET /api/collections/:owner/:slug/versions/:n/records.ndjson.gz
No auth for public collections
Public records as a gzip file (Content-Type: application/gzip, sent as an attachment). With ?type=, one type’s; without it, every type’s in slug order. The server sends the records as stored, so the file is several gzip members concatenated. Most gzip readers handle that; some, such as browsers’ DecompressionStream, stop after the first member.
Lines are canonical records, {"id", "type", "data"}, without hash (it is the SHA-256 of the line). Private records are never included, even for members: use records.ndjson for those. A type with no public records is 404. It costs 10 rate units.
GET /api/collections/:owner/:slug/versions/:n/manifest
No auth for public collections
Get the manifest: every record’s id, type and content hash, without the bodies, in (type, id) order. Members also see private records, marked "private": true. This is the cheapest way to learn what a version contains — at roughly 120 bytes per entry, a million records is one order of magnitude smaller than fetching them — and what a client that keeps no copy diffs against before it pushes.
Query parameters
limit | Entries per page (default 10000, max 25000). The first page also lists the version’s files, at most 25,000, with filesTruncated: true past that; later pages have files: []. |
cursor | Opaque keyset cursor from pagination.nextCursor. Do not construct or parse it — pass back exactly what you were given. |
since | A semver or ulv2: version hash. Return a delta against that version instead of the full manifest: which records were added, updated and removed between the two. |
Response 200
{
"semver": "v1.1.0",
"hash": "ulv2:a1b2c3d4...",
"schemas": {"Publication": "abc123..."},
"records": [
{"id": "pub-001", "type": "Publication", "hash": "def456..."},
{"id": "pub-002", "type": "Publication", "hash": "789abc...", "private": true}
],
"files": ["a1b2c3...", "d4e5f6..."],
"pagination": {
"limit": 2,
"hasMore": true,
"nextCursor": "eyJ0IjoiUHVibGljYXRpb24iLCJrIjoicHViLTAwMiJ9"
}
}Response with ?since= 200
{
"semver": "v1.1.0",
"hash": "ulv2:a1b2c3d4...",
"since": "v1.0.0",
"schemas": {"Publication": "abc123..."},
"delta": {
"added": [{"id": "pub-003", "type": "Publication", "hash": "def456..."}],
"updated": [{"id": "pub-001", "type": "Publication", "hash": "def456...",
"previousHash": "abc123..."},
{"id": "pub-002", "type": "Publication", "hash": "789abc...",
"previousHash": "789abc...", "private": true, "previousPrivate": false}],
"removed": [{"id": "pub-old", "type": "Publication", "hash": "def456..."}]
},
"files": ["a1b2c3..."],
"pagination": {
"limit": 10000,
"hasMore": false,
"nextCursor": null
}
}A delta is one walk over both versions in (type, id) order, with a single cursor; limit counts entries across the three lists. Keep re-requesting with cursor=pagination.nextCursor until hasMore is false.
For a member, an entry in the private set carries "private": true (for a removal, the set it left). A record that moved between the sets is listed under updated, with previousPrivate naming its former set; previousPrivate is present only when the set changed, and a move that kept the record’s hash has previousHash equal to hash. A non-member sees the public set only: a record made private reads as removed, and one made public as added. files in a delta is the full file list of :n that the caller may read (first page only), not a file delta.
GET /api/collections/:owner/:slug/versions/:n/diff
No auth for public collections
Diff two versions, with full record bodies. Without from, the diff is against the version before :n, which the response’s from names. The first version has none before it, so its diff is against an empty version: every record is added, from is null and schemaChanged is true. It costs 5 rate units.
Query parameters
from | Semver or version hash to diff from (e.g. v1.0.0). Default: the version before :n. |
limit | Entries per page, across the three lists (default 500, max 5000) |
cursor | Opaque keyset cursor from pagination.nextCursor, as on the manifest endpoint. Diff returns full record bodies, so pages are much larger than manifest pages — prefer manifest?since= when you only need the hashes. |
Response 200
{
"from": "v1.0.0",
"to": "v1.1.0",
"added": [
{"id": "pub-003", "type": "Publication", "data": {...}}
],
"updated": [
{"id": "pub-001", "type": "Publication", "data": {...}}
],
"removed": [{"id": "pub-old", "type": "Publication"}],
"pagination": {
"limit": 500,
"hasMore": false,
"nextCursor": null
},
"meta": {
"schemaChanged": false,
"metadataChanged": false,
"filesAdded": 0,
"filesRemoved": 0
}
}removed entries are {"id", "type"}, without bodies. filesAdded and filesRemoved are computed on the first page only; later pages report 0.
GET /api/collections/:owner/:slug/versions/:n/files
No auth for public collections
The files of the version that the caller may read, sorted by hash, at most 10,000. referenceCount is how many records in those sets reference the file. references is always empty: which records reference a file isn’t indexed.
Response 200
[
{
"hash": "a1b2c3...",
"size": 1048576,
"mimeType": "application/pdf",
"createdAt": "2026-04-01T00:00:00.000Z",
"referenceCount": 2,
"references": []
}
]