Recover and Survive
The machine lost power mid-edit, or the app was killed, or a transaction panicked. This chapter is what comes back, answered with one equation and no repair procedure: the three origins that satisfy it, the fixed boot order that makes a healthy start and a post-crash start the same operation, and the exact slice of work a hard crash costs.
What Is Left When the Process Is Gone
Section titled “What Is Left When the Process Is Gone”Chapter 8 attached storage, and entries began landing on disk behind the
stored frontier. This chapter answers the question that makes all of it
worth doing: the machine lost power mid-edit, or the app was killed, or a
transaction panicked. What comes back when the user opens the document again,
and how much of their work is in it?
Harmos answers with one equation and no repair procedure.
The Recovery Equation
Section titled “The Recovery Equation”state = origin + stored entries since itThat is the whole of recovery. An origin is a state value the runtime can stand on, at a known position in history. The stored entries since it are the entries a recovery source acknowledged after that position. Replay applies them in order, and the result is the state. Nothing else participates: no repair pass, no consistency check, no partial-write scan.
Because the equation is the only mechanism, a healthy start and a post-crash start are the same operation. There is no "recovering" mode to enter and no degraded path to test separately.
The Three Origins
Section titled “The Three Origins”An origin is one of exactly three things.
1. Initial state. A fresh boot from Schematic::default(), or a full
replay when the complete journal is stored. The origin is the state before any
entry existed, and every entry ever committed replays on top of it.
2. An infrastructure snapshot. The runtime's own fold of state, taken at a
known position, written by a snapshot producer. Reconstruction is the snapshot
at position p, then every stored entry after p.
3. An application document file. The document your editor opened — the
.schematic file the user double-clicked.
The third one carries a boundary worth stating plainly, because it is the difference between a file format and a database. A document exists for people and other tools; it is not the runtime's recovery format. A document file serves as an origin only in the degenerate single-document case, where the session happens to contain exactly one file. A session with six documents open does not recover from six files. It recovers from one infrastructure snapshot plus one sidecar, and treats every user file as an export only.
Journal and snapshot combine freely — journal only, snapshot only, both, or neither. All four are valid modes of the same equation. None of them is an error and none is a degraded path.
Snapshot Plus Tail Is the Normal Form
Section titled “Snapshot Plus Tail Is the Normal Form”It is tempting to read snapshots as an optimization — the thing you add when replay gets slow. They are not. Snapshot-plus-tail is the ordinary shape of a long-lived runtime, and full replay from initial state is the special case that happens to be available while history is still short.
A snapshot has two lives, and retention decides which one it is in. While the entries beneath it are still retained, a snapshot is a cache: it can be thrown away and refolded, and a schema mismatch after an upgrade is not an error — discard it and replay. Once retention prunes the entries beneath it, the snapshot becomes the only surviving record of that prefix, and from that moment it is truth, with the same versioned obligations any durable application data carries. That transition is also what governs when an old transaction version may be retired, which is chapter 10's business.
Opening a File: Artifact Plus Sidecar
Section titled “Opening a File: Artifact Plus Sidecar”Here is the arrangement that makes an editor crash-safe without turning its file format into a database.
The document file is the origin. A sidecar journal sits beside it and holds the entries committed since the last save. Together they satisfy the recovery equation, so unsaved work survives a crash:
// 1. Decode the document with your own codec. No runtime exists yet —// document codecs work standalone, before harmos is involved.let opened = schematic_file::read(&path)?;
// 2. Open the sidecar beside it, stamped with the save generation it was// folded from and the position history stood at when that save was written.let copy = Jsonl::<SchematicEditor>::sidecar(&sidecar_of(&path), opened.generation, opened.at)?;
// 3. Boot. start() returns a runtime that has already recovered: the// document is the origin, the sidecar's entries replay on top of it.let runtime = Runtime::builder(opened.schematic, ()) .register(Transactions::catalog()) .register(Records::catalog()) .storage(Storage::new().recover(copy)) .runtime() .await?;Two rules keep this honest.
Save is atomic in effect. Saving exports the document and resets the sidecar, as one operation. A save that wrote the document without resetting the sidecar would replay already-saved edits on top of a file that already contains them; a save that reset the sidecar without writing the document would throw the work away. Either half alone is a defect, not a partial success.
The sidecar records the save generation it was folded from. If a user
restores an older copy of the .schematic file over the current one, the
sidecar beside it now describes entries since a save that this file is not. A
sidecar whose generation does not match its document is not applied. The
mechanism stops there — whether your app then discards it silently, keeps it
for inspection, or asks the user, is an application decision. Harmos supplies the
mechanism, not the policy.
Under the Hood: Boot Ordering, and What Replay Refuses to Do
Section titled “Under the Hood: Boot Ordering, and What Replay Refuses to Do”Boot runs in a fixed order, every time:
freeze the catalogs -> invoke the recovery source's finite load -> replay -> open the writer -> wire the stream windows -> start live consumers -> start resident services -> report readyRead that sequence for what it excludes. The writer opens after replay, so
no commit can interleave with reconstruction. The windows are wired before
the consumers, so an attached recorder cannot miss a row published in the gap
between them. Live consumers start after the writer, so no projection sees
a half-rebuilt world, and services start after those, so a resident never
publishes into a fold that has not begun listening. And the whole sequence
happens inside the builder call: runtime() returns a runtime that has already
recovered. There is no readiness gate to poll, because a half-restored runtime
is not representable — the same reason origin is mandatory back in chapter 3.
Replay itself is deliberately weaker than the commit path:
- It applies transactions and never calls
check. Chapter 2's rule, now load-bearing: the stored entry was already accepted, under the rules that existed then. A tightened validation rule cannot invalidate history. - It retains records and never applies them. A record changed nothing when it was written; it changes nothing now. It stays in the order, available to consumers, and contributes nothing to state.
- It performs no I/O. Not a guideline — a conformance test (chapter 18). Replay that read a file or called a service would make recovery depend on the world being the same as it was, which it never is.
- It notifies no live consumer. Consumers resume from their own persisted checkpoints. Nothing gets a torrent of ten thousand doorbell rings for history it already processed.
The Risk Window
Section titled “The Risk Window”commit does not wait for storage. State mutation and in-process publication
are the commit boundary; persistence lags behind it. That lag has a name and a
measurement:
risk window = applied - storedThose are the two watch frontiers from chapter 8. applied is the last
entry applied in memory; stored is the last position the recovery source
acknowledged. Everything between them is real, visible, and not yet durable —
and a hard crash loses exactly that much. Storage failure cannot roll back a
committed state transition, so the window is an accepted operating decision,
not a bug to design away.
What it buys you is a rule that is easy to apply and hard to get wrong:
Anything leaving the process gates on
stored. External acknowledgements, HTTP replies, durable effect dispatch — all of it waits onwait_until_stored(position). In-process readers may read the volatile tail freely, because a crash rewinds the whole node together.
A reader that sees an entry the disk has not got is not observing an inconsistency; it is a passenger on the same node. A client that was told about that entry over the network is holding a claim the node may not be able to honor. That is the entire distinction, and it is why the gate lives at the process boundary rather than at every read.
Memory-only runtimes take the same path: with no recovery source attached,
stored tracks applied and wait_until_stored resolves immediately. Code
that gates correctly in production gates correctly in a unit test, with no
#[cfg] and no mode branch.
Ending a Session on Purpose
Section titled “Ending a Session on Purpose”A clean shutdown is the same machinery, run deliberately:
drop(handle); // and every other Journal clonelet settled = runtime.stop().await; // the position every fold reachedstop consumes the runtime, drops the journal inside it, and waits. Dropping
the last handle closes the one request channel, so the writer finishes what it
already admitted and stops; the attached folds then drain the intact tail and
stop reports the frontier they all reached. In the schematic example that
number is what the transcript prints as "the process ends at 4 without a save",
and it is the same number the next boot recovers to.
The order matters, and it is the one lifecycle rule worth memorising: drop
every Journal clone before you call stop(). A live clone is a live sender,
a live sender keeps the channel open, and stop waits for a writer that has
been given no reason to finish. Nothing is lost — it simply does not return.
When apply Panics
Section titled “When apply Panics”apply is infallible by contract (chapter 2). A panic inside it is therefore
not a domain rejection and not a recoverable error — it is an implementation
bug that has already run partway through mutating your state. Harmos treats it
as the one catastrophe, and the response is a shutdown, never a restart:
1. apply panics state is poisoned2. writer stops admitting in-flight and later calls fail: Unavailable3. nothing was sealed seal follows apply — no poison in history4. state reads are refused runtime-wide; nobody sees half-mutated state5. consumers keep draining until stored == applied6. the runtime stops7. next boot origin + stored entries since itStep 3 is the one that makes the rest possible. The pipeline seals after applying, so a transaction that panicked mid-mutation never got a position, never reached the tail, and cannot be in the journal. Recovery can never replay the thing that killed you.
Step 5 is the one people misread. "Reads are refused" is about the state — the poisoned value nobody may observe. The entry tail is intact, because no poison ever reached it, and consumers go on draining it to storage until durability catches up with what was applied. The runtime spends its last moments making sure the work that survives is written down. Graceful shutdown uses the identical drain; evacuation is that drain with the door already locked.
The consequence, stated exactly: a panic loses no committed data. Only a hard crash loses the risk window.
And the writer is never restarted in place. Supervised restart is the right policy for a service that failed to do its job; it is the wrong policy for a writer holding a state value that is now wrong in an unknown way. Recovery from the stored prefix is the only correct response, and it is the response the runtime already knows how to perform.
Test Your Knowledge
Section titled “Test Your Knowledge”
1. Support forwards a bug report: “the app threw a panic and closed itself.” The customer wants to know whether they lost work. Reconstruct the timeline and answer them.
They lost nothing that was committed, and the panic is not the thing that would have cost them anything.
In order: a transaction's apply panicked partway through mutating
state. The writer stopped admitting immediately, so every in-flight and
subsequent commit failed with Unavailable rather than touching the
poisoned value. No entry was sealed for the panicking transaction — seal
follows apply — so it never got a position and never entered history.
State reads were refused runtime-wide, which is why the window went dead
rather than showing something wrong. Meanwhile the entry tail was
intact, and consumers kept draining it to storage until
stored == applied. Then the runtime stopped.
On reopening, the equation runs as always: origin plus stored entries since it. Everything the customer committed before the panic is there. The transaction that panicked is not, and cannot be — it was never in history to replay.
The one honest caveat is unrelated to the panic. Had this been a hard
crash, power loss rather than a panic, the risk window applied - stored
would have been lost, because there would have been no opportunity to
drain. The panic path gets that drain; the power cord does not.
2. A teammate wires the writer into the supervision tree: “if it panics, restart it — that is what supervisors are for. Restarting beats killing the whole runtime.”
Supervised restart assumes the failed component can come back to a known state. The writer cannot, because the thing that failed is not the writer — it is the state the writer owns.
apply panicked partway through a mutation. That state value is now
wrong in a way nobody has characterized: some fields updated, some not,
invariants your check methods rely on possibly violated. A restarted
writer would resume admitting commits against that value, and every
subsequent entry would be sealed on top of it. Worse, those entries
would be perfectly legitimate as far as history is concerned — they
would replay faithfully at the next boot and produce something else
entirely, because the corrupt starting point was never a fact in the
journal.
The correct response is the one harmos already implements: refuse reads,
drain the intact tail to stored == applied, stop, and rebuild from the
stored prefix. That prefix is known-good by construction, because seal
follows apply and no poison ever reached it. Restart discards a working
recovery mechanism in favor of continuing on a value with no known
meaning.
There is also a structural reason the writer cannot be a supervised item at all: work depends on the journal, so supervising the writer would create an ownership cycle, and the writer must exist before the first service starts and still be draining after the last one stops.
3. A user reports “the app opened my file and threw away an hour of work.” You find they restored yesterday's .schematic from a backup over today's file, and today's sidecar was still sitting beside it. What did the runtime do, and was it right?
.schematic from a backup over today's file, and today's sidecar was still sitting beside it. What did the runtime do, and was it right?It refused to apply the sidecar, and yes.
The sidecar records the save generation it was folded from. Its entries are meaningful only as "the changes since that save" — they are a delta against one specific origin. The restored file is a different origin: an earlier generation, whose contents the sidecar's entries were never recorded against. Applying them anyway would replay edits like "move symbol R7" onto a document where R7 may not exist, or may sit somewhere else entirely, producing a state that corresponds to no history at all. The recovery equation holds only when the origin is the origin the entries were recorded after.
So the generation stamp is checked, it does not match, and the sidecar is not applied. What happens next is an application decision rather than a harmos one — discard it, keep it aside for inspection, or offer the user the choice. Harmos supplies the mechanism; sidecar naming and whether autorecovery prompts are yours.
The hour of work is not gone from the disk, incidentally. It is in the sidecar. It is simply not applicable to the file the user put there.
4. The server-shaped deployment replies 200 OK the moment commit resolves Ok(receipt). Chapter 4 established that publish precedes reply, so the receipt is genuinely later than the change. What is still wrong, and what is the smallest fix?
200 OK the moment commit resolves Ok(receipt). Chapter 4 established that publish precedes reply, so the receipt is genuinely later than the change. What is still wrong, and what is the smallest fix?The receipt proves the entry is applied, not that it is stored. Those are two different frontiers, and the gap between them is the risk window.
When commit resolves, the transaction has been checked, applied,
sealed at a position, and published to in-process observers. It has not
necessarily been written anywhere. commit does not await a storage
acknowledgement — persistence lags deliberately, and the lag is
applied - stored. If the node loses power in that interval the entry
is gone, but the client has already been told it happened and has no way
to find out otherwise.
The rule the receipt does not satisfy: anything leaving the process gates
on stored. The fix is one line before the reply.
let receipt = journal.commit(tx).await?;journal.wait_until_stored(receipt.position).await;// now the 200 is a claim this node can honorIn-process readers need none of this — a watcher may observe the volatile tail freely, because a crash rewinds the whole node together and the reader goes down with it. The gate belongs at the process boundary, where a claim outlives the node that made it. And in a memory-only test the same line resolves immediately, so the gated code is the code you ship.