Skip to content

Persist

Everything up to here has run in memory. This chapter attaches durability without changing a line above it: one storage declaration, one contract every downstream fold shares, and two frontiers whose distance is exactly what a crash would cost. You leave knowing where the one line that waits for durability belongs, and why it belongs nowhere else.

Everything through the previous chapter runs entirely in memory. The state is in memory, the journal's tail is in memory, and if the process exits it is all gone. That was the correct default, not a gap: a test, a prototype, or a tool that produces one output file needs no durability at all, and the ones that do should say so deliberately.

This chapter is where you say so. None of the code you have already written changes.

Storage is a value you build and hand to the builder before runtime():

let runtime = Runtime::builder(Schematic::default(), ())
.register(Transactions::catalog())
.register(Records::catalog())
.storage(storage)
.runtime()
.await?;

That is chapter 3's assembly with one argument filled in. What is inside storage is a list of attached consumers, of which at most one is designated as the place to recover from:

let storage = Storage::new()
.recover(Jsonl::open(&journal_file)?) // the complete copy the next boot loads
.attach(Jsonl::open(&mirror)?) // an off-machine copy, folded the same way
.snapshot(Snapshot::folding(Schematic::default(), Position::ORIGIN, write_snapshot)
.every(1_000))
.retain(Horizon::BeneathSnapshot)
.tail(65_536);
Runtime
└── Storage ──▶ JSONL journal copy RECOVERY SOURCE — the next boot loads this
──▶ JSONL mirror the same fold, somewhere else
──▶ snapshot producer folds a replica of Schematic, writes on a policy
──▶ net-index projection derived read model, in memory or on disk

Jsonl and Snapshot are the reference adapter's, and harmos ships neither: what it publishes is the contract they implement. A redb file, an S3 bucket, or a Postgres table is that same contract with a different handler, and adding one teaches the runtime nothing. Copy the worked implementation or write your own — jsonl-adapter exists to be read end to end, not to be named in a manifest.

Designating one adapter as the recovery source means one thing: use this complete journal copy at the next boot. It does not make commits synchronous, and it does not change what commit returns.

There may be zero or one recovery sources, never two. The reason is worth following, because "two copies for redundancy" is such a natural instinct. Parallel copies lag independently — the local file is a few entries ahead of the off-machine upload at almost every instant. At boot, a runtime with two designated copies would have to either guess which one is authoritative or merge them, and both of those are ways of inventing history. So harmos never does either: a second recover(..) call demotes the first to an ordinary mirror rather than leaving the choice ambiguous. The demoted copy remains genuinely useful; it is simply a mirror, not a claimant.

A stream recording is declared in the same place, through attach_streams. The schematic keeps both of its storage shapes in one storage.rs, so the composition is a declaration the entry verbs call rather than something a shell assembles (mise run run:minischematic):

/// Records the selected feed and recovers nothing.
pub fn recordings(config: &Config) -> Result<Storage<SchematicEditor>, io::Error> {
Ok(Storage::new().attach_streams(recording(config)?))
}
/// Recovers `document`'s sidecar, and records the selected feed beside it.
pub fn attach(document: &Document, config: &Config) -> Result<Storage<SchematicEditor>, io::Error> {
Ok(document.storage()?.attach_streams(recording(config)?))
}
/// The one feed this build records, opened in the configured directory.
fn recording(config: &Config) -> Result<JsonlStreams<SchematicEditor>, io::Error> {
Ok(JsonlStreams::<SchematicEditor>::open(config.recordings())?.select::<NetVoltage>())
}

and here is the same run stopping the runtime and reopening that capture through its typed row, which is what makes the recording readable rather than merely written:

let schematic = whole(&runtime).await;
drop(handle);
runtime.stop().await.expect("runtime capture drained");
capture.await.expect("the capture task stops without panic");
let columns = analysis::read(config.recordings());
println!(" demand 0 -> capture stopped; the recording is durable");
println!(" recorded position -> {:?}", columns.position);
println!(" recorded gap_missed -> {:?}\n", columns.gap_missed);

Here is the idea that keeps this layer small. Every single thing that consumes the journal is the same object: a checkpointed consumer — a fold over the entry order with a persisted resume Position.

Only the handler differs.

Attached foldIts handler does
storage adapterencodes the envelope and persists it
snapshot producerfolds a replica of state and serializes it on a policy
durable listenerperforms an external effect and acknowledges it
projectionfolds a derived read model
per-file historyfilters the order to one document and writes its log
collaborative clientfolds a replica for display

Six rows, one contract. There are not two supervision stacks to keep straight, no separate "event handler" system beside the storage system, and no third thing for projections. Storage itself reduces to three jobs: supervise the consumers, designate the recovery source, and publish the stored frontier.

The entry-fed contract is named by its source: JournalConsumer<A>. Its two positions answer different questions. resume() says what the fold has taken in, so delivery starts after it; drain(..) answers what the fold made durable. For the designated recovery source, that durable answer becomes stored. The distance between the two is the fold's own risk window — never collapse them into one checkpoint just because a consumer normally advances both together.

Consumers read immutable, ordered entries, so many of them run in parallel over the same order and lag independently — a slow off-machine upload never holds up the local file, and neither holds up the writer. And by the fold test, every consumer's output is a cache: it may be deleted and rebuilt at any time. (There is exactly one exception, and it appears under snapshots below.)

This is also where the previous chapter's BeyondTail stops being a problem. A consumer that folds forward as entries arrive and keeps its own checkpoint never asks about a position below the in-memory floor, because it is never that far behind.

pub fn applied(&self) -> Position;
pub fn stored(&self) -> Position;
pub async fn wait_until_stored(&self, at: Position);
  • applied is the position of the last entry applied in memory.
  • stored is the position acknowledged by the recovery source.

The gap between them, applied - stored, is the crash-risk window: the entries that exist, are visible, and would not come back if the machine lost power right now.

crash-risk window

commit

apply in memory

applied frontier

recovery-source fold

stored frontier

That window exists because commit does not await storage. The commit boundary is state mutation plus in-process publication; persistence follows behind it. This is a deliberate decision with a consequence you should hear stated once, plainly: storage failure cannot roll back an already committed state transition. There is no un-apply. A write that fails is a durability problem to be handled by retrying and by watching the frontier — never by pretending the change did not happen.

So how do you get durability where you actually need it? You gate on it, at the places that need it and nowhere else. The rule is short:

Anything leaving the process gates on stored.

External acknowledgements, HTTP replies, emails, payments, machine control — anything a crash could not take back. In-process readers do not gate on anything and should read the volatile tail freely, because a crash rewinds the whole node together: the reader and the entry it read disappear at the same instant, so there is nobody left holding a belief the journal has forgotten.

With no recovery source attached, stored tracks applied and wait_until_stored resolves immediately. Chapter 3 mentioned this; here is what it buys. A handler that waits for durability before acting is correct in a persisted deployment and in a memory-only test, with no #[cfg], no mode flag, and no branch in your code. You write the gate once, and it is a no-op exactly when there is nothing to wait for.

Read it as a convention, not a durability claim. In memory-only mode the risk window is nominally zero because nothing is durable at all — the entire process is the window.

/// Export the schematic and tell the operator it is done — without ever
/// claiming something a crash could take back.
async fn export(
runtime: &Runtime,
format: ExportFormat,
) -> Result<(), Error> {
// All I/O lives outside the writer: this is ordinary async code.
let symbols = write_export_file(runtime, format).await?;
// The fact enters history. It is sealed, published, and visible to every
// watcher the moment this returns — but not yet durable.
let receipt = runtime.journal
.record(ExportCompleted { format, symbols })
.await?;
// The acknowledgement leaves the process, so it waits for the frontier.
runtime.journal.wait_until_stored(receipt.position).await;
notify_operator(receipt.position);
Ok(())
}

Three lines, three different notions of "done", in the right order: the file is written, the fact is in history, the fact is durable. Only the third one is allowed to be told to the outside world. Move the wait_until_stored and you have chosen a different risk profile — which is the point of it being one visible line rather than a hidden default.

Replaying a million entries at every boot is not a plan, so state is normally reconstructed as a snapshot plus the entries after it. Treat that as the normal form rather than an optimisation you switch on later.

A snapshot identifies the journal position it represents, and that position is the entire interface between the two halves: load the snapshot taken at p, then apply every stored entry after p.

A snapshot producer is an ordinary checkpointed consumer. It folds its own replica of the state from the entry stream and serializes that on a policy you declare at assembly. It never locks the live state, never pauses the writer, and never reads through the state lock — so the naive design's stop-the-world moment does not exist here. It does not need to exist, because a fold over the order arrives at the same value the live state holds.

A snapshot then has two lives, and which one it is in is decided by retention, not by you:

  • Cache. While the entries beneath it are still retained, the snapshot is disposable. A schema mismatch is not an error — discard it and replay. Nothing versioned is required of it at all.
  • Truth. Once retention prunes the entries beneath it, the snapshot becomes the only surviving record of that prefix. From that moment it carries the same versioned-evolution obligation as any other durable application data, and the chapter on evolution is what it owes.

That is the fold test applied to itself: a snapshot is a cache exactly as long as it is rebuildable, and becomes truth precisely when it stops being.

A desktop editor has a problem the server shape does not. Its document file — schematic.sch — is an artifact: application-shaped, versioned, meant for humans and other tools, and exported deliberately when the user presses Save. It is not a journal, and it should not become one. So what happens to twenty minutes of edits between two saves when the process dies?

The recovery equation answers it — the next chapter takes it apart properly, but its shape is simply state = origin + stored entries since it. Editor mode plugs both halves in:

  • the artifact file is the origin — decoded before the runtime is built, and handed to the builder as origin;
  • the sidecar journal supplies the entries since.

Together they satisfy the equation, so unsaved work survives a crash. In the example the document owns that pair — the state it decoded, the generation it was saved at, the position that state reflects, and the sidecar beside it — and answering with its own storage is one call plus a designation:

/// Opens the sidecar this document's history is recovered from.
///
/// The generation and the position travel with it: a sidecar left behind
/// by an older save no longer matches, and the adapter resets it rather
/// than replaying entries the artifact already folds.
pub fn storage(&self) -> Result<Storage<SchematicEditor>, io::Error> {
Ok(Storage::new().recover(Jsonl::sidecar(
&self.sidecar,
self.generation,
self.position,
)?))
}

generation is the save the artifact came from and at is where history stood when it was written — the two numbers that let the next boot tell a sidecar that belongs to this document from one that does not.

Two rules make it trustworthy, and both are worth putting in your test suite:

  1. Save is atomic in effect. It exports the artifact and resets the sidecar as one operation. Writing the artifact without resetting the sidecar replays old edits on top of a file that already contains them; resetting the sidecar without writing the artifact loses them entirely. Neither is a mode. Both are defects. The JSONL adapter hands the application that operation rather than leaving it to be assembled: copy.sidecars() is a handle whose save takes your export closure, gives it the generation and position to stamp into the artifact, and resets the file only once the export returned.
  2. The sidecar records the save generation it was folded from. A sidecar whose generation does not match the artifact beside it is not applied — which is what stops a stale sidecar, left behind by a different save, from corrupting a good file.

Naming the file and deciding whether to prompt "recover unsaved changes?" are application choices. Harmos supplies the mechanism, not the policy.

That split has a seam worth putting in the right crate. The format — what an artifact contains, how it decodes, and which sidecar stands beside it — belongs to the application, next to the state it is a picture of; the schematic keeps it in document.rs. When to save is the shell's: it is a user gesture, an idle timer, or a window closing, and none of those are things the application crate can know. A shell that has both in hand calls Sidecar::save with its export closure at a settled position, which is the rule above stated as code.

Two consequences worth pulling out. Artifact codecs work standalone, before any runtime exists — they have to, because editor-mode boot decodes the document first and only then constructs the runtime around it. Chapter 3's mandatory origin is what forces that, and this is where it pays. And there is no public save() on the journal: durability is a frontier you observe, not a boundary a caller crosses. What the user calls Save is an artifact export.

A session with many documents open decouples recovery from user files entirely: one infrastructure snapshot plus one sidecar restores the whole session, and the per-file histories are filtered projections — derived, never load-bearing.

Entries have two shapes, and which one you are holding tells you where you are.

In memory a consumer reads Entry<X> — position, metadata, and your typed value, from the previous chapter. At the storage boundary, and only there, the same fact materializes as an envelope:

pub struct Envelope<P> {
/// The entry's permanent identity in the one total order.
pub position: Position,
/// Whether the payload is a change or a fact.
pub kind: Kind,
/// The permanent wire name of the payload's definition.
pub id: DefinitionId,
/// The definition version the payload was written with.
pub version: u32,
/// What the runtime witnessed when the entry was sealed.
pub metadata: Metadata,
/// The encoded payload, in the adapter's own representation.
pub payload: P,
}
pub enum Kind {
Transaction,
Record,
}

Envelope<P> is not held in memory and never appears in ordinary application code. If you are looking at one, you are writing an adapter.

Three of those fields are stamped at encode time, from the attribute macros you wrote in Chapter 2:

  • id — the DefinitionId, "symbol.move". The permanent wire name.
  • version — the number beside it, dormant until the evolution chapter.
  • kind — transaction or record.

kind earns its place on the envelope by being decidable without a catalog. A runtime that meets a definition id it has never heard of can still tell whether it is looking at a fact or a change, without asking anything. An unknown fact can be skipped safely, since it mutates nothing; an unknown change cannot, since everything after it depends on its mutation. The evolution chapter cashes that in; the reason it is possible is one enum stored beside the payload.

An adapter needs three sentences from you, and the trait that asks for them is Wire: the permanent name a value is stored under, how it is written, and how a stored payload reads back as the version that wrote it. Every one of them is one line over the catalogs, because the attributes already closed those:

impl Wire for SchematicEditor {
fn identify(value: &Recorded<Self>) -> (DefinitionId, u32) {
match value {
Recorded::Transaction(change) => change.identify(),
Recorded::Record(fact) => fact.identify(),
Recorded::System(_) => unreachable!("system facts are encoded by the adapter"),
}
}
fn encode<S: Serializer>(value: &Recorded<Self>, serializer: S) -> Result<S::Ok, S::Error> {
match value {
Recorded::Transaction(change) => change.encode_with(serializer),
Recorded::Record(fact) => fact.encode_with(serializer),
Recorded::System(_) => unreachable!("system facts are encoded by the adapter"),
}
}
fn decode<'de, D: Deserializer<'de> + Copy>(
kind: Kind,
id: &DefinitionId,
version: u32,
payload: D,
) -> Result<Option<Payload<Self>>, D::Error> {
Ok(match kind {
Kind::Transaction => {
Transactions::restore(id, version, payload)?.map(|(change, apply)| Payload {
prepared: None,
id: id.clone(),
version,
value: Recorded::Transaction(Arc::new(change)),
apply,
})
}
Kind::Record => Records::restore(id, version, payload)?.map(|fact| Payload {
prepared: None,
id: id.clone(),
version,
value: Recorded::Record(fact),
apply: None,
}),
})
}
}

There is no per-definition list anywhere in it. identify, encode_with, and restore are what #[harmos::transactions] and #[harmos::records] generated, so a new definition joins the wire the moment it joins the catalog. Attaching the adapter is then two lines:

let copy = Jsonl::<SchematicEditor>::open(&journal_file)?;
let storage = Storage::new().recover(copy);

Notice the P on Envelope<P>. The envelope is generic over payload representation, so each adapter owns its own:

Envelope<Recorded<A>> the in-process value the tail holds
Envelope<serde_json::Value> the JSONL adapter's representation
Envelope<Vec<u8>> a byte-oriented backend's representation

Harmos itself depends on the format-neutral Serde traits and nothing more: what it derives are the witnessed fields — Position, Metadata, Kind, DefinitionId — while the payload's shape is the definition's own. serde_json is one adapter's codec and one adapter's dependency. An application that copies no adapter compiles no concrete format at all.

One rule adapters follow: encode the stable id and version plus the inner payload, and never the Rust enum variant name. The catalog enum is an in-memory dispatch type, not a wire format. Here is one real line the JSONL adapter wrote for the schematic example, reformatted for the page:

{
"position": 1,
"kind": "Transaction",
"id": "symbol.place",
"version": 1,
"metadata": {
"principal": "alice",
"correlation": 1,
"causation": null,
"request": null,
"link": null,
"at_us": 1785598323058904
},
"payload": { "at": { "x": 10, "y": 20 }, "reference": "R1", "rotation": "Zero", "symbol": 1 }
}

Read the payload object: it is PlaceSymbol's own serde shape, and there is no "PlaceSymbol" tag anywhere on the line. Rename Transaction::PlaceSymbol to Transaction::Drop next year and every stored file still reads — for exactly the reason renaming the struct was safe in Chapter 2. What names the payload is the (id, version) pair beside it, which is what chapter 10 reads it back by.

The two ats on that line are different things, and the spelling keeps them apart. metadata.at_us is the witnessed clock — whole microseconds since the epoch, as one signed integer — and the at inside payload is a field PlaceSymbol declared, meaning a point on the canvas. Witness and declare, one more time. The JSONL wire format is the full contract for this line, pinned by golden files.

1. A colleague wants commit to await storage: “then a receipt means it is on disk, the risk window disappears, and nobody has to remember to call wait_until_stored.” The throughput argument is not the strongest one. What is?

Ask what commit would do when the write fails. By the time storage is involved, check has passed and apply has run — the state has already changed, and there is no un-apply. Harmos never requires Clone on your state, so there is no pre-image to restore, and even if there were, the entry has been published and other observers have already seen it. So a storage-awaiting commit has exactly two options on failure: return an error for a change that definitely happened, or hang. The first is a lie, the second is an outage. The obligation is impossible, so the API must not take it on.

And the window does not actually disappear. It moves and shrinks. There is still a gap between the adapter's acknowledgement and the bytes being genuinely safe on the device, and there is still a machine that can lose power in it. What awaiting inside commit really buys is that every commit pays full durability latency, including the thousands that nobody outside the process will ever hear about.

The design harmos chose instead is to make the window visible and gateable: applied - stored is a number you can read, and wait_until_stored is a line you place at the one boundary that matters — where something leaves the process. Fewer waits, at the only places where waiting means anything.

2. For redundancy, a teammate attaches two journal adapters — a local file and an off-machine copy — and calls recover(..) on both. What did they actually build, and what should they do instead?

They built one recovery source and one mirror, and probably not the ones they meant. Storage::recover holds a single designation: naming a second source demotes the first to an ordinary attached fold, still written and still folded, simply no longer the copy boot reads. The state they were trying to express is not representable, which is the point.

The reason is that the two copies are never at the same position. Consumers lag independently; that is a feature, and it is why the off-machine upload cannot stall the local write. A boot handed two complete journal copies that disagree about how much history exists would have to either pick one (guessing) or combine them (inventing an order nobody committed). Harmos does neither.

A second breakage is quieter and worse. stored is defined as the recovery source's checkpoint — one consumer's, by definition. With two claimants there is no single answer, so applied - stored stops naming a real risk window, and every wait_until_stored in the codebase starts gating on something with no clear meaning.

What to do: keep both adapters and choose deliberately — .recover(local).attach(remote). The mirror is still valuable, and its checkpoint is still readable; it is just not the thing the next boot loads from. Redundancy is a property of having two copies, not of both of them claiming to be the original.

3. A user saved at 14:00, kept editing, and the process was killed at 14:07. They reopen the file and their 14:06 edits are there — but the .sch file on disk still has a 14:00 timestamp and 14:00 contents. Where did the edits come from, and which slice of them might still be missing?

From the sidecar, through the recovery equation. State is origin plus stored entries since it. In editor mode the artifact file is the origin: it is decoded by an ordinary codec, before any runtime exists, and handed to the builder as origin. The sidecar journal beside it holds the entries committed since that save, and replaying them onto the decoded document reproduces the 14:06 state. The .sch file was never asked to be a journal, and it did not have to be.

Before applying the sidecar, boot checks that the sidecar's recorded save generation matches the artifact. A sidecar left over from a different save is discarded rather than replayed onto a file it does not belong to.

What may be missing is the last applied - stored slice. The sidecar is a checkpointed consumer, and commit never waited for it, so the last few entries — the ones applied in memory but not yet acknowledged — die with the process. If the user's last keystroke was at 14:06:59.8, it may not be there. That is the accepted operating decision, and it is the number the frontier exposes so an application can decide how much of it to tolerate.

4. After a refactor, the schema of the snapshot no longer matches what the snapshot producer writes, and loading an old snapshot fails. Is that a bug? Careful — the honest answer is “it depends”, and on exactly one thing.

It depends on whether the entries beneath that snapshot are still retained.

If they are, it is not a bug and barely an event. The snapshot is a cache: everything in it can be recomputed by replaying the entries it summarizes. The correct handling of a mismatch is to discard the snapshot and replay — slower to boot, identical result. Nothing versioned is required of a snapshot in this life, which is exactly why a snapshot must never be presented to users as the application's file format.

If retention has pruned the entries beneath it, it is serious. At that moment the snapshot stopped being a summary of something else and became the only surviving record of that prefix — truth. Truth carries the full versioned-evolution obligation, so a snapshot below the retention horizon needs the same discipline as any other durable application data, and a failure to load it is data loss rather than a slow boot.

The general rule underneath both answers is the fold test applied to itself: a thing is a cache exactly as long as it can be rebuilt, and becomes truth precisely when it cannot. Which is also why the retention policy and the snapshot policy are one conversation, not two.

Stokker Technologies markDesigned and built by Stokker Technologies