The JSONL Wire Format
This page is the contract for the file the reference adapter's Jsonl writes:
a documented, versioned, stable interchange format. The format belongs to that
adapter rather than to harmos — harmos publishes one order, one fold contract,
and the encode boundary, and never chooses a format — so what is pinned here is
one worked implementation, pinned because an interchange format nobody can read
twice is not one. It exists so that another backend — a database-backed source,
an analysis script, a tool in another language — can read and write the same
history without reading the Rust. Everything below is held by the golden-file
tests in examples/jsonl-adapter/tests/format.rs; the format does not move
without this document and a new version number moving with it.
An application that copies examples/jsonl-adapter owns the copy and may change
every byte described here. What it gives up by doing so is exactly what this
page is: the guarantee that some other reader already knows how to meet the
file.
The current format version is 3. Version 2 files remain readable. Their first append validates the complete existing history, writes a version 3 replacement beside it, syncs it, and atomically renames it before appending. Inspection and recovery never rewrite a version 2 file. Version 1 remains retired and is refused.
The File
Section titled “The File”A journal copy is UTF-8 JSON Lines: one JSON value
per line, each line terminated by \n alone. The wire is LF-only — CRLF is
not the wire, and a reader meeting a line that ends in a carriage return must
refuse the file rather than strip the byte. Redelivery forwards stored bytes,
so a reader that quietly normalized line endings would rewrite history byte
by byte.
The first line is the header. Every other line is a complete logical commit group or a legacy standalone entry, in position order. The one exception is the sidecar reset described in chapter 8, which replaces the whole file with a fresh header.
The member sets below are closed. A line or header carrying a member this document does not name is refused by name, never tolerated — even beside otherwise valid members — because an extra member is a new format version, not an extension.
The Header
Section titled “The Header”{"schema":3,"generation":0,"origin":0}| Member | Encoding | Meaning |
|---|---|---|
schema | unsigned integer | The format version marker. It names this whole document: this reader accepts 2 and 3, migrates 2 on append, and refuses other values at open. |
generation | unsigned integer | The document save generation the entries below were folded from. 0 for an infrastructure journal that stands beside no document. |
origin | unsigned integer | The position of the origin state this copy stands on. Entries at or beneath it are already folded into that origin, so recovery hands back only what lies above. |
Logical Commit Groups
Section titled “Logical Commit Groups”A live transaction and its zero to 256 emitted records occupy one physical line:
{"group":1,"first":1,"frontier":2,"entries":[{"position":1,"kind":"Transaction","id":"ledger.deposit","version":1,"metadata":{"principal":null,"correlation":1,"causation":null,"request":null,"link":null,"at_us":0},"payload":{"amount":40}},{"position":2,"kind":"Record","id":"ledger.noted","version":1,"metadata":{"principal":null,"correlation":1,"causation":1,"request":null,"link":null,"at_us":0},"payload":"deposited"}]}group is the frame version, currently 1. first is the transaction position;
frontier is the complete last position. entries is nonempty, contiguous,
starts with exactly one transaction, and contains records after it. Record
witnesses share the transaction's principal, correlation, and commit time;
causation names its position, while request and reversal links are absent.
Each member retains its own durable definition identity and version.
The writer runs Transaction::records(&pre_state, &mut Records) after check
and before apply, then reifies every emitted leaf through the current record
catalog. Encoding, membership, count, and aggregate byte failures are sticky
and refuse the whole group before mutation, even when a caller ignores a
Records::push error. Original prepared leaf data reaches storage without
invoking custom serialization again. The payload budget is 4 MiB including the
transaction; principal and request strings are each limited to 1,024 UTF-8
bytes. A physical frame is bounded at 8 MiB including metadata and framing.
LF terminates a committed group. An unterminated canonical group at the end of a v3 file is outside the recovered prefix and is removed only before an append. Every byte cut therefore exposes the previous frontier or the complete new group. A terminated malformed group, a gap, or corruption before later history is refused. Legacy v2 complete final entry lines without LF remain readable. A v2 file never contains group frames.
A receipt's position identifies its transaction; frontier identifies the
whole logical commit. Admission does not promise synchronous disk durability:
wait on wait_until_stored(receipt.frontier) when needed. Request-key dedupe
remains process-local. Occurrence times belong in record payloads; at_us is
commit metadata. Replay invokes only the originally stored transaction version's
apply and exposes stored records without re-running the records hook. Undo and
redo are live transactions and collect their own records.
Tail eviction keeps complete groups. Queries can resume between member positions of a verified group; state reconstruction and recovery origins must name complete boundaries. Snapshots checkpoint only complete groups and restore historical apply from original prepared payloads. Pruning inside a group rounds down to the previous complete boundary. Unknown application records remain skippable during replay and remain present in forwarded group bytes.
The Entry Line
Section titled “The Entry Line”A canonical member, also accepted as a standalone legacy entry:
{"position":1,"kind":"Transaction","id":"ledger.deposit","version":1,"metadata":{"principal":"alice","correlation":1,"causation":null,"request":null,"link":null,"at_us":1700000000000000},"payload":{"amount":40}}| Member | Encoding | Meaning |
|---|---|---|
position | unsigned integer | The entry's permanent identity in the one total order. Strictly ascending, contiguous, never reused. |
kind | "Transaction" or "Record" | Which half of the order the payload belongs to, decidable without a catalog: replay must apply an unknown change and may skip an unknown fact. |
id | string | The permanent wire name of the payload's definition — "ledger.deposit" — stamped from the authoring attributes. Never a Rust type or catalog variant name. |
version | unsigned integer | The definition version the payload was written with. The stored (id, version) pair is the whole of what a payload is read by. |
metadata | object | What the runtime witnessed when the entry was sealed, below. |
payload | any JSON value | For application definitions, the definition's own serde shape under that (id, version) — an object, a bare string, whatever the definition derives. For the harmos-owned fact namespaces below, an array of byte values. Nothing of any catalog appears here. |
Definition ID Namespaces
Section titled “Definition ID Namespaces”Application definition ids remain the permanent names their authoring attributes
declare, such as ledger.deposit. Two prefixes are reserved for harmos-owned
records whose structure must remain readable without application or work
service code:
| ID shape | Kind | Meaning |
|---|---|---|
harmos:work/<service>/<name> | "Record" | A fact authored by harmos's resident work layer for the named service. The first / after harmos:work/ separates the service from its local fact name. |
The namespace uses the existing entry frame and the existing "Record" discriminator;
there is no third kind and no schema-version change. Their version is the
source-local payload version, and payload is the admitted opaque byte string
encoded as a JSON array of unsigned byte values. Recovery recognizes these id
namespaces before consulting the application record catalog, reconstructs the
typed system fact, and skips it during state replay because it is a record.
Direct guests use the ordinary metadata.principal string with the durable
spelling guest:<id>. That is a namespaced value, not a new metadata member.
The Metadata Members
Section titled “The Metadata Members”Every member is always present; an absent value is null.
| Member | Encoding | Meaning |
|---|---|---|
principal | string or null | Who committed, from the requesting handle's scope. |
correlation | unsigned integer | The family this entry belongs to, named by the position that opened it. An entry that starts a family names its own position. |
causation | unsigned integer or null | The position of the entry that directly caused this one. |
request | string or null | The idempotency key the entry was admitted under. |
link | null, {"UndoOf": p}, or {"RedoOf": p} | The earlier entry this one undoes or redoes, by position. |
at_us | signed 64-bit integer | When the writer witnessed the entry: whole microseconds since the Unix epoch, UTC. Negative for an instant before the epoch. |
at_us is the one place the wire diverges from what a serde derive would have
written. A timestamp is stored as a single explicit signed integer — no
SystemTime structure, no nested seconds-and-nanoseconds object appears
anywhere in the format. Sub-microsecond precision is not stored; the wire's
resolution is the microsecond, deliberately and permanently. The wire's range
is the i64 range — some ±292,000 years around the epoch — which every
witnessed clock lies comfortably inside.
Byte-Forward Redelivery
Section titled “Byte-Forward Redelivery”A durable copy that lagged behind the frontier is redelivered each stored entry or group's own bytes, and writes them verbatim: a stored historical entry redelivers byte-identical. A group is forwarded once, tagged with its complete frontier. History is forwarded, never re-encoded. Two consequences are part of the contract:
- A writer always emits the canonical member order and spacing shown above, but a reader must accept any valid JSON member order and spacing. A forwarded line keeps the exact bytes some earlier writer chose, and those bytes are as much the format as the canonical ones.
- A record whose
(id, version)this build cannot decode still travels whole. The fold skips it; the byte lane forwards it; a copy is never silently shorn of what only a later build can read.
The Track Key
Section titled “The Track Key”A stream may declare a track: the typed tuple that names one partition of its rows. This section is the canonical text form of such a tuple, so a reader in another language can order and select partitions the way the adapter that wrote them does.
It is not part of the entry frame above, and it is not part of this adapter at
all. No journal line carries a track, and nothing here moves the schema
marker: this is the form a series adapter stores partitions under, written
down here because it is the one place a cross-language reader would look, and
such an adapter versions its own files. Series implementations are an application's
own (Serve Beneath the Window); this form
is the one that makes the two laws below hold.
A key is each component's label, outermost first, with the unit separator
U+001F after every one of them — including the last:
("run_7", "probe_1") -> run_7␟probe_1␟("run_7",) -> run_7␟() ->A label is text carrying no control character, and never empty. That is the one rule, and it is a rule about correctness rather than tidiness: the separator is what a component ends at, so a label containing one would name a partition that a selection reads as two.
Two properties follow, and both are the point:
- Byte order is tuple order. Every character a label may carry sorts above
U+001F, so comparing two keys byte by byte answers what comparing the two tuples component by component would. A shorter tuple sorts immediately before everything beneath it. - A prefix is a selection. The keys beginning with
run_7␟are exactly the partitions whose tuple begins with("run_7",). The trailing separator is what makes that exact: without it,run_7would also beginrun_70.
The empty key is every partition, which is what an unnarrowed read asks for and what a stream declaring no track stores under.
Versioning
Section titled “Versioning”The schema marker names the whole format. Any change to a member, an
encoding, or the header is a new version and a new revision of this document —
the two are versioned together.
Version 1, whose metadata carried a serde-default SystemTime structure
under at, is retired. A version-2 implementation refuses it at open, naming
the version it met; it never dual-reads and never guesses. A metadata object
carrying at under the current marker — with or without a valid at_us
beside it — is likewise refused by member name, so the retired shape can
never be silently misread as, or smuggled into, the current one.
Atomic sidecar facts
Section titled “Atomic sidecar facts”A trusted host transaction can call Records::push_system(SystemFact) alongside
its typed records. Both lanes share the count and byte budgets and refuse before
mutation on any ignored collection error. FactSource::Sidecar(owner) persists as
harmos:sidecar/{owner}/{name}. The owner starts with a lowercase ASCII letter and
contains only lowercase letters, digits, dots, and hyphens; the local name is
nonempty and preserved without normalization. Recovery preserves the admitted
opaque bytes and version without consulting the application record catalog.
This API validates identity shape, not caller authenticity. The host must derive provenance from the installed supervised handle and static declaration, validate dynamic catalog membership and payload before collection, and never accept a source supplied by the payload. Existing Work and Guest historical identities retain their recovery semantics. Typed record admission refuses all three reserved system prefixes, preventing a typed record from changing into a system fact during recovery.
Reading one captured frontier
Section titled “Reading one captured frontier”A state read returns At { value, position }. Use
journal.query_through::<Recorded<Application>>(after, position) to read every
recognized transaction, typed record, and system fact through that same position,
even when later commits have arrived. Inspection supports the same finite query.
query delegates using its captured applied frontier, and watch::<Recorded<Application>>
provides one ordered live feed. Individual query cursors may lie inside an already
verified group; state and snapshot boundaries still require complete commits.
A future frontier returns FrontierAhead. Evicted history requires a recovery
source; lagging or pruned storage produces an explicit refusal, never a successful
partial result. Callers must propagate that failure instead of substituting empty
history. Queries do not promise synchronous durability or retain an extra cache.