Skip to content

The JSONL Wire Format

This page is the contract for the file the reference adapter's Jsonl writes: a documented, versioned, stable interchange format. The format belongs to that adapter rather than to harmos — harmos publishes one order, one fold contract, and the encode boundary, and never chooses a format — so what is pinned here is one worked implementation, pinned because an interchange format nobody can read twice is not one. It exists so that another backend — a database-backed source, an analysis script, a tool in another language — can read and write the same history without reading the Rust. Everything below is held by the golden-file tests in examples/jsonl-adapter/tests/format.rs; the format does not move without this document and a new version number moving with it.

An application that copies examples/jsonl-adapter owns the copy and may change every byte described here. What it gives up by doing so is exactly what this page is: the guarantee that some other reader already knows how to meet the file.

The current format version is 3. Version 2 files remain readable. Their first append validates the complete existing history, writes a version 3 replacement beside it, syncs it, and atomically renames it before appending. Inspection and recovery never rewrite a version 2 file. Version 1 remains retired and is refused.

A journal copy is UTF-8 JSON Lines: one JSON value per line, each line terminated by \n alone. The wire is LF-only — CRLF is not the wire, and a reader meeting a line that ends in a carriage return must refuse the file rather than strip the byte. Redelivery forwards stored bytes, so a reader that quietly normalized line endings would rewrite history byte by byte.

The first line is the header. Every other line is a complete logical commit group or a legacy standalone entry, in position order. The one exception is the sidecar reset described in chapter 8, which replaces the whole file with a fresh header.

The member sets below are closed. A line or header carrying a member this document does not name is refused by name, never tolerated — even beside otherwise valid members — because an extra member is a new format version, not an extension.

{"schema":3,"generation":0,"origin":0}
MemberEncodingMeaning
schemaunsigned integerThe format version marker. It names this whole document: this reader accepts 2 and 3, migrates 2 on append, and refuses other values at open.
generationunsigned integerThe document save generation the entries below were folded from. 0 for an infrastructure journal that stands beside no document.
originunsigned integerThe position of the origin state this copy stands on. Entries at or beneath it are already folded into that origin, so recovery hands back only what lies above.

A live transaction and its zero to 256 emitted records occupy one physical line:

{"group":1,"first":1,"frontier":2,"entries":[{"position":1,"kind":"Transaction","id":"ledger.deposit","version":1,"metadata":{"principal":null,"correlation":1,"causation":null,"request":null,"link":null,"at_us":0},"payload":{"amount":40}},{"position":2,"kind":"Record","id":"ledger.noted","version":1,"metadata":{"principal":null,"correlation":1,"causation":1,"request":null,"link":null,"at_us":0},"payload":"deposited"}]}

group is the frame version, currently 1. first is the transaction position; frontier is the complete last position. entries is nonempty, contiguous, starts with exactly one transaction, and contains records after it. Record witnesses share the transaction's principal, correlation, and commit time; causation names its position, while request and reversal links are absent. Each member retains its own durable definition identity and version.

The writer runs Transaction::records(&pre_state, &mut Records) after check and before apply, then reifies every emitted leaf through the current record catalog. Encoding, membership, count, and aggregate byte failures are sticky and refuse the whole group before mutation, even when a caller ignores a Records::push error. Original prepared leaf data reaches storage without invoking custom serialization again. The payload budget is 4 MiB including the transaction; principal and request strings are each limited to 1,024 UTF-8 bytes. A physical frame is bounded at 8 MiB including metadata and framing.

LF terminates a committed group. An unterminated canonical group at the end of a v3 file is outside the recovered prefix and is removed only before an append. Every byte cut therefore exposes the previous frontier or the complete new group. A terminated malformed group, a gap, or corruption before later history is refused. Legacy v2 complete final entry lines without LF remain readable. A v2 file never contains group frames.

A receipt's position identifies its transaction; frontier identifies the whole logical commit. Admission does not promise synchronous disk durability: wait on wait_until_stored(receipt.frontier) when needed. Request-key dedupe remains process-local. Occurrence times belong in record payloads; at_us is commit metadata. Replay invokes only the originally stored transaction version's apply and exposes stored records without re-running the records hook. Undo and redo are live transactions and collect their own records.

Tail eviction keeps complete groups. Queries can resume between member positions of a verified group; state reconstruction and recovery origins must name complete boundaries. Snapshots checkpoint only complete groups and restore historical apply from original prepared payloads. Pruning inside a group rounds down to the previous complete boundary. Unknown application records remain skippable during replay and remain present in forwarded group bytes.

A canonical member, also accepted as a standalone legacy entry:

{"position":1,"kind":"Transaction","id":"ledger.deposit","version":1,"metadata":{"principal":"alice","correlation":1,"causation":null,"request":null,"link":null,"at_us":1700000000000000},"payload":{"amount":40}}
MemberEncodingMeaning
positionunsigned integerThe entry's permanent identity in the one total order. Strictly ascending, contiguous, never reused.
kind"Transaction" or "Record"Which half of the order the payload belongs to, decidable without a catalog: replay must apply an unknown change and may skip an unknown fact.
idstringThe permanent wire name of the payload's definition — "ledger.deposit" — stamped from the authoring attributes. Never a Rust type or catalog variant name.
versionunsigned integerThe definition version the payload was written with. The stored (id, version) pair is the whole of what a payload is read by.
metadataobjectWhat the runtime witnessed when the entry was sealed, below.
payloadany JSON valueFor application definitions, the definition's own serde shape under that (id, version) — an object, a bare string, whatever the definition derives. For the harmos-owned fact namespaces below, an array of byte values. Nothing of any catalog appears here.

Application definition ids remain the permanent names their authoring attributes declare, such as ledger.deposit. Two prefixes are reserved for harmos-owned records whose structure must remain readable without application or work service code:

ID shapeKindMeaning
harmos:work/<service>/<name>"Record"A fact authored by harmos's resident work layer for the named service. The first / after harmos:work/ separates the service from its local fact name.

The namespace uses the existing entry frame and the existing "Record" discriminator; there is no third kind and no schema-version change. Their version is the source-local payload version, and payload is the admitted opaque byte string encoded as a JSON array of unsigned byte values. Recovery recognizes these id namespaces before consulting the application record catalog, reconstructs the typed system fact, and skips it during state replay because it is a record.

Direct guests use the ordinary metadata.principal string with the durable spelling guest:<id>. That is a namespaced value, not a new metadata member.

Every member is always present; an absent value is null.

MemberEncodingMeaning
principalstring or nullWho committed, from the requesting handle's scope.
correlationunsigned integerThe family this entry belongs to, named by the position that opened it. An entry that starts a family names its own position.
causationunsigned integer or nullThe position of the entry that directly caused this one.
requeststring or nullThe idempotency key the entry was admitted under.
linknull, {"UndoOf": p}, or {"RedoOf": p}The earlier entry this one undoes or redoes, by position.
at_ussigned 64-bit integerWhen the writer witnessed the entry: whole microseconds since the Unix epoch, UTC. Negative for an instant before the epoch.

at_us is the one place the wire diverges from what a serde derive would have written. A timestamp is stored as a single explicit signed integer — no SystemTime structure, no nested seconds-and-nanoseconds object appears anywhere in the format. Sub-microsecond precision is not stored; the wire's resolution is the microsecond, deliberately and permanently. The wire's range is the i64 range — some ±292,000 years around the epoch — which every witnessed clock lies comfortably inside.

A durable copy that lagged behind the frontier is redelivered each stored entry or group's own bytes, and writes them verbatim: a stored historical entry redelivers byte-identical. A group is forwarded once, tagged with its complete frontier. History is forwarded, never re-encoded. Two consequences are part of the contract:

  • A writer always emits the canonical member order and spacing shown above, but a reader must accept any valid JSON member order and spacing. A forwarded line keeps the exact bytes some earlier writer chose, and those bytes are as much the format as the canonical ones.
  • A record whose (id, version) this build cannot decode still travels whole. The fold skips it; the byte lane forwards it; a copy is never silently shorn of what only a later build can read.

A stream may declare a track: the typed tuple that names one partition of its rows. This section is the canonical text form of such a tuple, so a reader in another language can order and select partitions the way the adapter that wrote them does.

It is not part of the entry frame above, and it is not part of this adapter at all. No journal line carries a track, and nothing here moves the schema marker: this is the form a series adapter stores partitions under, written down here because it is the one place a cross-language reader would look, and such an adapter versions its own files. Series implementations are an application's own (Serve Beneath the Window); this form is the one that makes the two laws below hold.

A key is each component's label, outermost first, with the unit separator U+001F after every one of them — including the last:

("run_7", "probe_1") -> run_7␟probe_1␟
("run_7",) -> run_7␟
() ->

A label is text carrying no control character, and never empty. That is the one rule, and it is a rule about correctness rather than tidiness: the separator is what a component ends at, so a label containing one would name a partition that a selection reads as two.

Two properties follow, and both are the point:

  • Byte order is tuple order. Every character a label may carry sorts above U+001F, so comparing two keys byte by byte answers what comparing the two tuples component by component would. A shorter tuple sorts immediately before everything beneath it.
  • A prefix is a selection. The keys beginning with run_7␟ are exactly the partitions whose tuple begins with ("run_7",). The trailing separator is what makes that exact: without it, run_7 would also begin run_70.

The empty key is every partition, which is what an unnarrowed read asks for and what a stream declaring no track stores under.

The schema marker names the whole format. Any change to a member, an encoding, or the header is a new version and a new revision of this document — the two are versioned together.

Version 1, whose metadata carried a serde-default SystemTime structure under at, is retired. A version-2 implementation refuses it at open, naming the version it met; it never dual-reads and never guesses. A metadata object carrying at under the current marker — with or without a valid at_us beside it — is likewise refused by member name, so the retired shape can never be silently misread as, or smuggled into, the current one.

A trusted host transaction can call Records::push_system(SystemFact) alongside its typed records. Both lanes share the count and byte budgets and refuse before mutation on any ignored collection error. FactSource::Sidecar(owner) persists as harmos:sidecar/{owner}/{name}. The owner starts with a lowercase ASCII letter and contains only lowercase letters, digits, dots, and hyphens; the local name is nonempty and preserved without normalization. Recovery preserves the admitted opaque bytes and version without consulting the application record catalog.

This API validates identity shape, not caller authenticity. The host must derive provenance from the installed supervised handle and static declaration, validate dynamic catalog membership and payload before collection, and never accept a source supplied by the payload. Existing Work and Guest historical identities retain their recovery semantics. Typed record admission refuses all three reserved system prefixes, preventing a typed record from changing into a system fact during recovery.

A state read returns At { value, position }. Use journal.query_through::<Recorded<Application>>(after, position) to read every recognized transaction, typed record, and system fact through that same position, even when later commits have arrived. Inspection supports the same finite query. query delegates using its captured applied frontier, and watch::<Recorded<Application>> provides one ordered live feed. Individual query cursors may lie inside an already verified group; state and snapshot boundaries still require complete commits.

A future frontier returns FrontierAhead. Evicted history requires a recovery source; lagging or pruned storage produces an explicit refusal, never a successful partial result. Callers must propagate that failure instead of substituting empty history. Queries do not promise synchronous durability or retain an extra cache.

Stokker Technologies markDesigned and built by Stokker Technologies