Persist
Everything up to here has run in memory. This chapter attaches durability without changing a line above it: one storage declaration, one contract every downstream fold shares, and two frontiers whose distance is exactly what a crash would cost. You leave knowing where the one line that waits for durability belongs, and why it belongs nowhere else.
Nothing So Far Has Touched a Disk
Section titled “Nothing So Far Has Touched a Disk”Everything through the previous chapter runs entirely in memory. The state is in memory, the journal's tail is in memory, and if the process exits it is all gone. That was the correct default, not a gap: a test, a prototype, or a tool that produces one output file needs no durability at all, and the ones that do should say so deliberately.
This chapter is where you say so. None of the code you have already written changes.
Attaching Storage
Section titled “Attaching Storage”Storage is a value you build and hand to the builder before runtime():
let runtime = Runtime::builder(Schematic::default(), ()) .register(Transactions::catalog()) .register(Records::catalog()) .storage(storage) .runtime() .await?;That is chapter 3's assembly with one argument
filled in. What is inside storage is a list of attached consumers, of which at
most one is designated as the place to recover from:
let storage = Storage::new() .recover(Jsonl::open(&journal_file)?) // the complete copy the next boot loads .attach(Jsonl::open(&mirror)?) // an off-machine copy, folded the same way .snapshot(Snapshot::folding(Schematic::default(), Position::ORIGIN, write_snapshot) .every(1_000)) .retain(Horizon::BeneathSnapshot) .tail(65_536);Runtime └── Storage ──▶ JSONL journal copy RECOVERY SOURCE — the next boot loads this ──▶ JSONL mirror the same fold, somewhere else ──▶ snapshot producer folds a replica of Schematic, writes on a policy ──▶ net-index projection derived read model, in memory or on diskJsonl and Snapshot are the reference adapter's, and harmos ships neither:
what it publishes is the contract they implement. A redb file, an S3 bucket, or
a Postgres table is that same contract with a different handler, and adding one
teaches the runtime nothing. Copy the worked implementation or write your own —
jsonl-adapter exists to be read end to end, not
to be named in a manifest.
Designating one adapter as the recovery source means one thing: use this complete
journal copy at the next boot. It does not make commits synchronous, and it
does not change what commit returns.
There may be zero or one recovery sources, never two. The reason is worth
following, because "two copies for redundancy" is such a natural instinct.
Parallel copies lag independently — the local file is a few entries ahead of the
off-machine upload at almost every instant. At boot, a runtime with two
designated copies would have to either guess which one is authoritative or merge
them, and both of those are ways of inventing history. So harmos never does
either: a second recover(..) call demotes the first to an ordinary mirror
rather than leaving the choice ambiguous. The demoted copy remains genuinely
useful; it is simply a mirror, not a claimant.
A stream recording is declared in the same place, through attach_streams.
The schematic keeps both of its storage shapes in one storage.rs, so the
composition is a declaration the entry verbs call rather than something a shell
assembles (mise run run:minischematic):
/// Records the selected feed and recovers nothing.pub fn recordings(config: &Config) -> Result<Storage<SchematicEditor>, io::Error> { Ok(Storage::new().attach_streams(recording(config)?))}
/// Recovers `document`'s sidecar, and records the selected feed beside it.pub fn attach(document: &Document, config: &Config) -> Result<Storage<SchematicEditor>, io::Error> { Ok(document.storage()?.attach_streams(recording(config)?))}
/// The one feed this build records, opened in the configured directory.fn recording(config: &Config) -> Result<JsonlStreams<SchematicEditor>, io::Error> { Ok(JsonlStreams::<SchematicEditor>::open(config.recordings())?.select::<NetVoltage>())}and here is the same run stopping the runtime and reopening that capture through its typed row, which is what makes the recording readable rather than merely written:
let schematic = whole(&runtime).await; drop(handle); runtime.stop().await.expect("runtime capture drained"); capture.await.expect("the capture task stops without panic"); let columns = analysis::read(config.recordings()); println!(" demand 0 -> capture stopped; the recording is durable"); println!(" recorded position -> {:?}", columns.position); println!(" recorded gap_missed -> {:?}\n", columns.gap_missed);Everything Downstream Is One Mechanism
Section titled “Everything Downstream Is One Mechanism”Here is the idea that keeps this layer small. Every single thing that consumes
the journal is the same object: a checkpointed consumer — a fold over the
entry order with a persisted resume Position.
Only the handler differs.
| Attached fold | Its handler does |
|---|---|
| storage adapter | encodes the envelope and persists it |
| snapshot producer | folds a replica of state and serializes it on a policy |
| durable listener | performs an external effect and acknowledges it |
| projection | folds a derived read model |
| per-file history | filters the order to one document and writes its log |
| collaborative client | folds a replica for display |
Six rows, one contract. There are not two supervision stacks to keep straight,
no separate "event handler" system beside the storage system, and no third thing
for projections. Storage itself reduces to three jobs: supervise the
consumers, designate the recovery source, and publish the stored frontier.
The entry-fed contract is named by its source: JournalConsumer<A>. Its two
positions answer different questions. resume() says what the fold has taken
in, so delivery starts after it; drain(..) answers what the fold made durable.
For the designated recovery source, that durable answer becomes stored. The
distance between the two is the fold's own risk window — never collapse them
into one checkpoint just because a consumer normally advances both together.
Consumers read immutable, ordered entries, so many of them run in parallel over the same order and lag independently — a slow off-machine upload never holds up the local file, and neither holds up the writer. And by the fold test, every consumer's output is a cache: it may be deleted and rebuilt at any time. (There is exactly one exception, and it appears under snapshots below.)
This is also where the previous chapter's BeyondTail stops being a problem. A
consumer that folds forward as entries arrive and keeps its own checkpoint never
asks about a position below the in-memory floor, because it is never that far
behind.
Two Frontiers: applied and stored
Section titled “Two Frontiers: applied and stored”pub fn applied(&self) -> Position;pub fn stored(&self) -> Position;pub async fn wait_until_stored(&self, at: Position);appliedis the position of the last entry applied in memory.storedis the position acknowledged by the recovery source.
The gap between them, applied - stored, is the crash-risk window: the
entries that exist, are visible, and would not come back if the machine lost
power right now.
That window exists because commit does not await storage. The commit
boundary is state mutation plus in-process publication; persistence follows
behind it. This is a deliberate decision with a consequence you should hear
stated once, plainly: storage failure cannot roll back an already committed
state transition. There is no un-apply. A write that fails is a durability
problem to be handled by retrying and by watching the frontier — never by
pretending the change did not happen.
So how do you get durability where you actually need it? You gate on it, at the places that need it and nowhere else. The rule is short:
Anything leaving the process gates on
stored.
External acknowledgements, HTTP replies, emails, payments, machine control — anything a crash could not take back. In-process readers do not gate on anything and should read the volatile tail freely, because a crash rewinds the whole node together: the reader and the entry it read disappear at the same instant, so there is nobody left holding a belief the journal has forgotten.
Memory-Only Is the Same Code Path
Section titled “Memory-Only Is the Same Code Path”With no recovery source attached, stored tracks applied and
wait_until_stored resolves immediately. Chapter 3 mentioned this; here is what
it buys. A handler that waits for durability before acting is correct in a
persisted deployment and in a memory-only test, with no #[cfg], no mode
flag, and no branch in your code. You write the gate once, and it is a no-op
exactly when there is nothing to wait for.
Read it as a convention, not a durability claim. In memory-only mode the risk window is nominally zero because nothing is durable at all — the entire process is the window.
Gating an External Effect
Section titled “Gating an External Effect”/// Export the schematic and tell the operator it is done — without ever/// claiming something a crash could take back.async fn export( runtime: &Runtime, format: ExportFormat,) -> Result<(), Error> { // All I/O lives outside the writer: this is ordinary async code. let symbols = write_export_file(runtime, format).await?;
// The fact enters history. It is sealed, published, and visible to every // watcher the moment this returns — but not yet durable. let receipt = runtime.journal .record(ExportCompleted { format, symbols }) .await?;
// The acknowledgement leaves the process, so it waits for the frontier. runtime.journal.wait_until_stored(receipt.position).await; notify_operator(receipt.position);
Ok(())}Three lines, three different notions of "done", in the right order: the file is
written, the fact is in history, the fact is durable. Only the third one is
allowed to be told to the outside world. Move the wait_until_stored and you
have chosen a different risk profile — which is the point of it being one
visible line rather than a hidden default.
Snapshots
Section titled “Snapshots”Replaying a million entries at every boot is not a plan, so state is normally reconstructed as a snapshot plus the entries after it. Treat that as the normal form rather than an optimisation you switch on later.
A snapshot identifies the journal position it represents, and that position is
the entire interface between the two halves: load the snapshot taken at p,
then apply every stored entry after p.
A snapshot producer is an ordinary checkpointed consumer. It folds its own replica of the state from the entry stream and serializes that on a policy you declare at assembly. It never locks the live state, never pauses the writer, and never reads through the state lock — so the naive design's stop-the-world moment does not exist here. It does not need to exist, because a fold over the order arrives at the same value the live state holds.
A snapshot then has two lives, and which one it is in is decided by retention, not by you:
- Cache. While the entries beneath it are still retained, the snapshot is disposable. A schema mismatch is not an error — discard it and replay. Nothing versioned is required of it at all.
- Truth. Once retention prunes the entries beneath it, the snapshot becomes the only surviving record of that prefix. From that moment it carries the same versioned-evolution obligation as any other durable application data, and the chapter on evolution is what it owes.
That is the fold test applied to itself: a snapshot is a cache exactly as long as it is rebuildable, and becomes truth precisely when it stops being.
The Sidecar Pattern
Section titled “The Sidecar Pattern”A desktop editor has a problem the server shape does not. Its document file —
schematic.sch — is an artifact: application-shaped, versioned, meant for
humans and other tools, and exported deliberately when the user presses Save. It
is not a journal, and it should not become one. So what happens to twenty
minutes of edits between two saves when the process dies?
The recovery equation answers it — the next chapter takes it apart properly, but its shape is simply state = origin + stored entries since it. Editor mode plugs both halves in:
- the artifact file is the origin — decoded before the runtime is built, and
handed to the builder as
origin; - the sidecar journal supplies the entries since.
Together they satisfy the equation, so unsaved work survives a crash. In the example the document owns that pair — the state it decoded, the generation it was saved at, the position that state reflects, and the sidecar beside it — and answering with its own storage is one call plus a designation:
/// Opens the sidecar this document's history is recovered from. /// /// The generation and the position travel with it: a sidecar left behind /// by an older save no longer matches, and the adapter resets it rather /// than replaying entries the artifact already folds. pub fn storage(&self) -> Result<Storage<SchematicEditor>, io::Error> { Ok(Storage::new().recover(Jsonl::sidecar( &self.sidecar, self.generation, self.position, )?)) }generation is the save the artifact came from and at is where history stood
when it was written — the two numbers that let the next boot tell a sidecar that
belongs to this document from one that does not.
Two rules make it trustworthy, and both are worth putting in your test suite:
- Save is atomic in effect. It exports the artifact and resets the sidecar
as one operation. Writing the artifact without resetting the sidecar replays
old edits on top of a file that already contains them; resetting the sidecar
without writing the artifact loses them entirely. Neither is a mode. Both are
defects. The JSONL adapter hands the application that operation rather than
leaving it to be assembled:
copy.sidecars()is a handle whosesavetakes your export closure, gives it the generation and position to stamp into the artifact, and resets the file only once the export returned. - The sidecar records the save generation it was folded from. A sidecar whose generation does not match the artifact beside it is not applied — which is what stops a stale sidecar, left behind by a different save, from corrupting a good file.
Naming the file and deciding whether to prompt "recover unsaved changes?" are application choices. Harmos supplies the mechanism, not the policy.
That split has a seam worth putting in the right crate. The format — what an
artifact contains, how it decodes, and which sidecar stands beside it — belongs
to the application, next to the state it is a picture of; the schematic keeps it
in document.rs. When to save is the shell's: it is a user gesture, an idle
timer, or a window closing, and none of those are things the application crate
can know. A shell that has both in hand calls Sidecar::save with its export
closure at a settled position, which is the rule above stated as code.
Two consequences worth pulling out. Artifact codecs work standalone, before
any runtime exists — they have to, because editor-mode boot decodes the document
first and only then constructs the runtime around it. Chapter 3's mandatory
origin is what forces that, and this is where it pays. And there is no
public save() on the journal: durability is a frontier you observe, not a
boundary a caller crosses. What the user calls Save is an artifact export.
A session with many documents open decouples recovery from user files entirely: one infrastructure snapshot plus one sidecar restores the whole session, and the per-file histories are filtered projections — derived, never load-bearing.
Under the Hood: The Encode Boundary
Section titled “Under the Hood: The Encode Boundary”Entries have two shapes, and which one you are holding tells you where you are.
In memory a consumer reads Entry<X> — position, metadata, and your typed
value, from the previous chapter. At the storage boundary, and only there, the
same fact materializes as an envelope:
pub struct Envelope<P> { /// The entry's permanent identity in the one total order. pub position: Position, /// Whether the payload is a change or a fact. pub kind: Kind, /// The permanent wire name of the payload's definition. pub id: DefinitionId, /// The definition version the payload was written with. pub version: u32, /// What the runtime witnessed when the entry was sealed. pub metadata: Metadata, /// The encoded payload, in the adapter's own representation. pub payload: P,}
pub enum Kind { Transaction, Record,}Envelope<P> is not held in memory and never appears in ordinary application code.
If you are looking at one, you are writing an adapter.
Three of those fields are stamped at encode time, from the attribute macros you wrote in Chapter 2:
id— theDefinitionId,"symbol.move". The permanent wire name.version— the number beside it, dormant until the evolution chapter.kind— transaction or record.
kind earns its place on the envelope by being decidable without a catalog.
A runtime that meets a definition id it has never heard of can still tell
whether it is looking at a fact or a change, without asking anything. An unknown
fact can be skipped safely, since it mutates nothing; an unknown change cannot,
since everything after it depends on its mutation. The evolution chapter cashes
that in; the reason it is possible is one enum stored beside the payload.
What an Adapter Asks of Your Application
Section titled “What an Adapter Asks of Your Application”An adapter needs three sentences from you, and the trait that asks for them is
Wire: the permanent name a value is stored under, how it is written, and how
a stored payload reads back as the version that wrote it. Every one of them is
one line over the catalogs, because the attributes already closed those:
impl Wire for SchematicEditor { fn identify(value: &Recorded<Self>) -> (DefinitionId, u32) { match value { Recorded::Transaction(change) => change.identify(), Recorded::Record(fact) => fact.identify(), Recorded::System(_) => unreachable!("system facts are encoded by the adapter"), } }
fn encode<S: Serializer>(value: &Recorded<Self>, serializer: S) -> Result<S::Ok, S::Error> { match value { Recorded::Transaction(change) => change.encode_with(serializer), Recorded::Record(fact) => fact.encode_with(serializer), Recorded::System(_) => unreachable!("system facts are encoded by the adapter"), } }
fn decode<'de, D: Deserializer<'de> + Copy>( kind: Kind, id: &DefinitionId, version: u32, payload: D, ) -> Result<Option<Payload<Self>>, D::Error> { Ok(match kind { Kind::Transaction => { Transactions::restore(id, version, payload)?.map(|(change, apply)| Payload { prepared: None, id: id.clone(), version, value: Recorded::Transaction(Arc::new(change)), apply, }) } Kind::Record => Records::restore(id, version, payload)?.map(|fact| Payload { prepared: None, id: id.clone(), version, value: Recorded::Record(fact), apply: None, }), }) }}There is no per-definition list anywhere in it. identify, encode_with, and
restore are what #[harmos::transactions] and #[harmos::records] generated,
so a new definition joins the wire the moment it joins the catalog. Attaching
the adapter is then two lines:
let copy = Jsonl::<SchematicEditor>::open(&journal_file)?;let storage = Storage::new().recover(copy);The Journal Never Picks a Format
Section titled “The Journal Never Picks a Format”Notice the P on Envelope<P>. The envelope is generic over payload
representation, so each adapter owns its own:
Envelope<Recorded<A>> the in-process value the tail holdsEnvelope<serde_json::Value> the JSONL adapter's representationEnvelope<Vec<u8>> a byte-oriented backend's representationHarmos itself depends on the format-neutral Serde traits and nothing more: what
it derives are the witnessed fields — Position, Metadata, Kind,
DefinitionId — while the payload's shape is the definition's own. serde_json
is one adapter's codec and one adapter's dependency. An application that copies
no adapter compiles no concrete format at all.
One rule adapters follow: encode the stable id and version plus the inner
payload, and never the Rust enum variant name. The catalog enum is an
in-memory dispatch type, not a wire format. Here is one real line the JSONL
adapter wrote for the schematic example, reformatted for the page:
{ "position": 1, "kind": "Transaction", "id": "symbol.place", "version": 1, "metadata": { "principal": "alice", "correlation": 1, "causation": null, "request": null, "link": null, "at_us": 1785598323058904 }, "payload": { "at": { "x": 10, "y": 20 }, "reference": "R1", "rotation": "Zero", "symbol": 1 }}Read the payload object: it is PlaceSymbol's own serde shape, and there is
no "PlaceSymbol" tag anywhere on the line. Rename Transaction::PlaceSymbol
to Transaction::Drop next year and every stored file still reads — for exactly
the reason renaming the struct was safe in Chapter 2. What names the payload is
the (id, version) pair beside it, which is what
chapter 10 reads it back by.
The two ats on that line are different things, and the spelling keeps them
apart. metadata.at_us is the witnessed clock — whole microseconds since the
epoch, as one signed integer — and the at inside payload is a field
PlaceSymbol declared, meaning a point on the canvas. Witness and declare, one
more time. The JSONL wire format is the full
contract for this line, pinned by golden files.
Test Your Knowledge
Section titled “Test Your Knowledge”
1. A colleague wants commit to await storage: “then a receipt means it is on disk, the risk window disappears, and nobody has to remember to call wait_until_stored.” The throughput argument is not the strongest one. What is?
commit to await storage: “then a receipt means it is on disk, the risk window disappears, and nobody has to remember to call wait_until_stored.” The throughput argument is not the strongest one. What is?Ask what commit would do when the write fails. By the time storage is
involved, check has passed and apply has run — the state has already
changed, and there is no un-apply. Harmos never requires Clone on your
state, so there is no pre-image to restore, and even if there were, the
entry has been published and other observers have already seen it. So a
storage-awaiting commit has exactly two options on failure: return an
error for a change that definitely happened, or hang. The first is a lie,
the second is an outage. The obligation is impossible, so the API must not
take it on.
And the window does not actually disappear. It moves and shrinks. There
is still a gap between the adapter's acknowledgement and the bytes being
genuinely safe on the device, and there is still a machine that can lose
power in it. What awaiting inside commit really buys is that every
commit pays full durability latency, including the thousands that nobody
outside the process will ever hear about.
The design harmos chose instead is to make the window visible and
gateable: applied - stored is a number you can read, and
wait_until_stored is a line you place at the one boundary that matters —
where something leaves the process. Fewer waits, at the only places where
waiting means anything.
2. For redundancy, a teammate attaches two journal adapters — a local file and an off-machine copy — and calls recover(..) on both. What did they actually build, and what should they do instead?
recover(..) on both. What did they actually build, and what should they do instead?They built one recovery source and one mirror, and probably not the ones
they meant. Storage::recover holds a single designation: naming a second
source demotes the first to an ordinary attached fold, still written and
still folded, simply no longer the copy boot reads. The state they were
trying to express is not representable, which is the point.
The reason is that the two copies are never at the same position. Consumers lag independently; that is a feature, and it is why the off-machine upload cannot stall the local write. A boot handed two complete journal copies that disagree about how much history exists would have to either pick one (guessing) or combine them (inventing an order nobody committed). Harmos does neither.
A second breakage is quieter and worse. stored is defined as the
recovery source's checkpoint — one consumer's, by definition. With two
claimants there is no single answer, so applied - stored stops naming a
real risk window, and every wait_until_stored in the codebase starts
gating on something with no clear meaning.
What to do: keep both adapters and choose deliberately —
.recover(local).attach(remote). The mirror is still valuable, and its
checkpoint is still readable; it is just not the thing the next boot loads
from. Redundancy is a property of having two copies, not of both of them
claiming to be the original.
3. A user saved at 14:00, kept editing, and the process was killed at 14:07. They reopen the file and their 14:06 edits are there — but the .sch file on disk still has a 14:00 timestamp and 14:00 contents. Where did the edits come from, and which slice of them might still be missing?
From the sidecar, through the recovery equation. State is origin plus
stored entries since it. In editor mode the artifact file is the origin: it
is decoded by an ordinary codec, before any runtime exists, and handed to
the builder as origin. The sidecar journal beside it holds the entries
committed since that save, and replaying them onto the decoded document
reproduces the 14:06 state. The .sch file was never asked to be a journal,
and it did not have to be.
Before applying the sidecar, boot checks that the sidecar's recorded save generation matches the artifact. A sidecar left over from a different save is discarded rather than replayed onto a file it does not belong to.
What may be missing is the last applied - stored slice. The sidecar is
a checkpointed consumer, and commit never waited for it, so the last few
entries — the ones applied in memory but not yet acknowledged — die with the
process. If the user's last keystroke was at 14:06:59.8, it may not be
there. That is the accepted operating decision, and it is the number the
frontier exposes so an application can decide how much of it to tolerate.
4. After a refactor, the schema of the snapshot no longer matches what the snapshot producer writes, and loading an old snapshot fails. Is that a bug? Careful — the honest answer is “it depends”, and on exactly one thing.
It depends on whether the entries beneath that snapshot are still retained.
If they are, it is not a bug and barely an event. The snapshot is a cache: everything in it can be recomputed by replaying the entries it summarizes. The correct handling of a mismatch is to discard the snapshot and replay — slower to boot, identical result. Nothing versioned is required of a snapshot in this life, which is exactly why a snapshot must never be presented to users as the application's file format.
If retention has pruned the entries beneath it, it is serious. At that moment the snapshot stopped being a summary of something else and became the only surviving record of that prefix — truth. Truth carries the full versioned-evolution obligation, so a snapshot below the retention horizon needs the same discipline as any other durable application data, and a failure to load it is data loss rather than a slow boot.
The general rule underneath both answers is the fold test applied to itself: a thing is a cache exactly as long as it can be rebuilt, and becomes truth precisely when it cannot. Which is also why the retention policy and the snapshot policy are one conversation, not two.