Skip to content

Run Work

The runtime can change, observe, persist, and recover its truth; it still has things to do. This chapter gives finite work and resident work two different nouns, because their callers need two different answers, and then works through typed lane admission, what one attempt owes the next, composing an attempt out of its own children, where serialization actually belongs, the two failure regimes a resident can exit under, the fixed stop choreography, and the diagnostic for a shutdown that will not finish.

The editor can now change, observe, persist, and recover its truth. It still has things to do: export a netlist once, keep an instrument adapter alive, or run analysis without making any of those activities part of state. Harmos gives those two different lifetimes two different nouns.

A Job is finite. It starts, produces one typed outcome, and ends. A Service is resident. It belongs to the runtime's lifetime and reports why an attempt exited. A declared Lane owns job concurrency; the service's own exit value chooses whether a resident failure heals or surfaces.

That split is mechanical rather than stylistic. If callers await the result, it is a Job. If runtime shutdown must stop it, it is a Service.

The authoring shape keeps the policy beside the operation: a struct whose fields are the inputs, #[harmos::job(...)] on the struct, and a plain impl Errand naming the outcome, the error, and the one attempt:

#[harmos::job(lane = Lanes::Export, retries = 2, timeout_ms = 30_000)]
pub struct ExportNetlist {
pub request: ExportRequest,
}
impl Errand for ExportNetlist {
type Outcome = Exported;
type Error = ExportError;
async fn attempt(&mut self, _scope: Scope<'_>) -> Result<Exported, ExportError> {
write_netlist(&self.request).await
}
}

attempt's second parameter is the attempt's own Scope, which the runtime constructs fresh per attempt. This one starts no children, so it names the scope _scope and ignores it; Compose an Attempt is where the name earns itself. The struct's fields are the job's arguments — construction is ExportNetlist { request }, not a positional call, so the moment a job takes two same-typed parameters they read as named fields rather than an order a caller has to remember.

Each attempt gets thirty seconds. A refusal retries twice under the default backoff, so the job runs at most three attempts. The deadline bounds an attempt, never the whole sequence; the wait between attempts is the job's own to declare, which Retry Needs No Hooks comes to.

Outcome and Error are ordinary associated types on Errand, not something the attribute infers, so a one-parameter Result alias out of your own error.rs names the error exactly as the spelled form does:

impl Errand for ExportNetlist {
type Outcome = Exported;
type Error = crate::error::Error;
async fn attempt(&mut self, _scope: Scope<'_>) -> crate::error::Result<Exported> {
write_netlist(&self.request).await
}
}

Concurrency belongs to the lane vocabulary, declared once and frozen at boot. A lane catalog is the one authority for the set, closed the way a change catalog is: one enum, one variant per lane, and a catalog() the assembly consumes verbatim, so the lanes named on job attributes and the lanes frozen at boot cannot drift. It lives in app.rs, beside the declaration that freezes it:

#[harmos::lanes]
pub enum Lanes {
/// Fully concurrent; every job that names no specific lane lands here.
#[default]
Open,
/// Netlist exports, one at a time.
#[lane(serial)]
Export,
/// Up to four analyses at once.
#[lane(concurrent(4))]
Analysis,
/// Detached feed mirroring, which must never queue behind an export.
#[lane(unbounded)]
Mirror,
}
let runtime = Runtime::builder(origin, resources)
.register(Transactions::catalog())
.register(Lanes::catalog())
.storage(storage)
.runtime()
.await?;

The enum's name is yours; Lanes is house style. A variant is spelled for the work it admits and not suffixed — Lanes::Export already says lane once, and the reading at a job's own attribute is what the name is for.

There are exactly three capacities. serial admits one attempt at a time; concurrent(n) admits at most n; unbounded admits everything immediately. Every catalog's #[default] variant is unbounded by construction — the compiler enforces it, and combining #[default] with an explicit #[lane(...)], or declaring zero or two of them, refuses to compile — so a job that lands there because it named its catalog rather than a variant never queues: it starts now, or the catalog itself refused to exist. A catalog can still spell a named unbounded lane beside the default, and the reason is legibility rather than capacity: Lanes::Mirror above reads as "detached feed mirroring" at the job's attribute, where landing on the anonymous default would say only "nothing else claimed this." Reach for a named unbounded lane when a submitter should see which detached work it is; reach for the default when it should not matter. Nothing structural bounds an unbounded lane either way, and that is worth being honest about — a bounded lane's queue is the runtime's back-pressure, and declaring unbounded is you saying you will supply your own. A zero-width lane, duplicate name, or empty name refuses boot. An undeclared lane no longer exists at runtime at all — every job names its lane or its lane catalog, so a lane nobody declared is a variant that does not exist and a name that does not compile.

Submitting routes the job by type and returns the outcome itself as an awaitable handle. One statement is the norm. The lane already lives on the job's own attribute, so there is nothing left to route by hand — construct the value and submit it in the same expression:

let exported = runtime.work.submit(ExportNetlist { request }).await?;

Split into a let job = …; followed by submit(job) only once an override — .under(…), .caused_by(…) — would otherwise stretch the line past what a reader can take in at a glance; One Declaration, Two Budgets is where that split first earns itself.

Cancellation is explicit and cooperative. Call handle.cancel(). A queued job is removed without running; a running attempt is dropped at its next await and the handle answers Refusal::Cancelled. Cooperative is worth saying plainly: whatever the attempt was mid-way through — a half-sent device command, a partially written file — stays however that operation left it, exactly as it would if the attempt's own deadline had expired there. Shutdown uses the same rule for queued work.

Job failure remains the submitter's fact. Exhausted retries answer the last error the job itself refused with, handed back as itself; a final timeout answers Refusal::TimedOut. Harmos does not write either one into history. If finishing the export belongs in the application's story, application code records ExportCompleted deliberately, as in Persist.

There is no before_retry, no on_failure, and no cleanup callback, because each attempt already owns two responsibilities that leave nothing for a hook to do.

An attempt cleans up after itself. Whatever it opened — a connection, a file, a lock, a temporary directory — is dropped when its future ends, and equally when the deadline drops that future mid-flight. That is ordinary RAII rather than a work-layer rule, and it is why a timed-out attempt leaves nothing behind for the next one to sweep.

An attempt also enters defensively. It assumes nothing about how far an earlier attempt got: it re-reads what it needs and commits under a RequestKey, so a retry that repeats a commit the previous attempt already landed answers the original receipt rather than doubling it — the idempotency from Commit, doing exactly the job it was built for.

A cleanup hook would run after the values it must clean have already dropped, and a setup hook would state what a defensive attempt establishes for itself anyway. Neither one exists.

What a hook would really have been reached for is pacing, and pacing is a declaration rather than a callback. Between two attempts harmos waits a capped exponential backoff: a base doubled after each failure, held under a ceiling, plus a jitter read off the attempt index. Undeclared, that base is 10 ms and that ceiling is 5 s — the two numbers an undeclared ServicePolicy already restarts a resident under — so ten retries wait 12, 24, 41, 100, 198, 335, 753, 1410, 2707, and 5164 ms: 10.74 s in total, with the ceiling binding from the tenth wait onward, where 10 · 2⁹ = 5120 ms first exceeds it.

A rate-limited vendor endpoint needs its own numbers, and says so where the retries and the deadline are already written:

#[harmos::job(
lanes = Lanes,
retries = 5,
timeout_ms = 10_000,
backoff_base_ms = 1_000,
backoff_cap_ms = 60_000,
)]
pub struct PushReading {
pub reading: Reading,
}
impl Errand for PushReading {
type Outcome = Accepted;
type Error = VendorError;
async fn attempt(&mut self, _scope: Scope<'_>) -> Result<Accepted, VendorError> {
vendor::push(self.reading.clone()).await
}
}

The two words are ServicePolicy's own — backoff_base and backoff_cap — so a resident and an errand describe their pacing in one vocabulary; the _ms suffix belongs to the attribute layer, exactly as in timeout_ms. They travel together or not at all. A base with no ceiling of its own would be clamped by the default one — a 10 s base never reaches its second wait under a 5 s cap — and the attribute refuses that at compile time rather than accept a knob that reads configurable while doing nothing. A cap beneath its base, or a cap of zero, is refused by the same rule.

Both attribute knobs bottom out in two constructors every macro expansion calls: Job::declared(retries, timeout_ms, lane, errand) for the flat knobs above, and Job::declared_with(policy, lane, errand) once a whole JobPolicy and the resolved lane name are in hand — the shape .under(policy) reaches for at the call site instead of the declaration. Nothing about this expansion is code you write in the ordinary case; a leaf declares knobs, .under replaces them, and the constructors stay a fact about the macro rather than a second dialect to choose. A policy whose cap is not a ceiling over its base is refused at submission with Refusal::UndeclarablePacing, for the reason a zero-width lane is refused at boot: the declaration is inert, and that is the first place a caller is listening.

The jitter is a function of the attempt index and nothing else. No random source and no per-job seed enters it, which is why a replayed retry sequence waits exactly what the original one waited — and why Simulate can put the whole backoff on virtual time.

The declaration is the default, not a straitjacket. One operation sometimes needs two budgets — an interactive attempt at the desk, a patient background retry overnight — and that is one declaration with a call-site replacement, never two declarations:

let job = ExportNetlist { request }.under(JobPolicy {
retries: 5,
timeout: Duration::from_secs(120),
..JobPolicy::default()
});
runtime.work.submit(job).await?;

under replaces the whole policy for that one submission and stays on the construction line, where everything that shapes a job belongs. This is the split the one-statement norm above named: the override makes the value long enough that giving it a name earns its keep. The scratch name is always job; the meaningful names go to handles and outcomes, and rebinding job for the next submission is deliberate — a reader never wonders whether an old job value is still live, because the name always means "the one I am about to submit." A replacement is held to the same rules a declaration is: a pacing that is not a range refuses the submission with UndeclarablePacing.

Once a second job wants that same patience, the four numbers stop being a policy and start being a duplication. Give the budget a name where the application keeps its vocabulary:

/// The budget for work a person is not waiting on.
pub fn patient() -> JobPolicy {
JobPolicy {
retries: 5,
timeout: Duration::from_secs(120),
backoff_base: Duration::from_secs(1),
backoff_cap: Duration::from_secs(60),
}
}

The attribute takes it as the declared default, and the call site takes it as a replacement — the same word in both places, because it is the same value:

#[harmos::job(lanes = Lanes, under = patient())]
pub struct PostProcess {
pub report: PathBuf,
}
impl Errand for PostProcess {
type Outcome = Report;
type Error = ReportError;
async fn attempt(&mut self, _scope: Scope<'_>) -> Result<Report, ReportError> {
render(&self.report).await
}
}
// And where some other job wants that budget for one submission:
let job = ExportNetlist { request }.under(patient());
runtime.work.submit(job).await?;

under = … accepts any expression that resolves to a JobPolicy — a call, a path to a constant, a struct literal — and it is evaluated where the job value is built, so a preset read from configuration is read afresh for each job.

The two forms are exclusive. under answers the whole policy, so writing it beside retries, timeout_ms, backoff_base_ms, or backoff_cap_ms is a compile error that names both spellings and asks which one you meant — the mistake is nearly always a job that grew a named preset and kept one leftover knob, and a knob that silently loses to a preset is the outcome a declaration must never have.

Reach for the flat knobs first. A leaf whose budget is nobody else's business says so best in four words on its own attribute; a name earns itself the moment a second job wants the same budget, or the moment the budget has a reason worth writing down.

A job's struct is not just named arguments; its fields can genuinely adapt across attempts — rotate an endpoint, remember a cursor, widen a page. impl Errand borrows &mut self for exactly one attempt at a time, so state that must survive a retry lives in a plain field rather than anything shared:

#[harmos::job(lane = Lanes::Export, retries = 2, timeout_ms = 30_000)]
pub struct Export {
endpoints: VecDeque<Endpoint>,
cursor: u64,
}
impl Errand for Export {
type Outcome = Exported;
type Error = ExportError;
async fn attempt(&mut self, _scope: Scope<'_>) -> Result<Exported, ExportError> {
let endpoint = self.endpoints.pop_front().ok_or(ExportError::NoEndpoint)?;
let page = endpoint.export_from(self.cursor).await?;
self.cursor = page.next;
Ok(page.exported)
}
}
let exported = runtime.work.submit(Export { endpoints, cursor: 0 }).await?;

Retries are sequential — attempt three cannot begin until attempt two has settled — so the retry loop borrows that one value &mut for exactly one attempt at a time. Nothing is shared, so nothing needs guarding: the endpoint queue and the cursor are plain fields, not an Arc<Mutex<_>>. And the moment a job takes two same-typed parameters, naming them as struct fields is the only spelling there is — there is no positional call for a caller to get wrong.

That is the same contract a Service's resident has, on the finite lifetime. Errand and Resident differ only in what an attempt is handed and what it answers with, which the Resident Services section comes to.

A workflow-shaped job is not one operation; it is an orchestration. It starts observers — a dialog watcher, a log stream — and runs mutators one at a time while they watch. The trap in that shape is ownership: an observer started with a detached work.submit outlives the attempt that started it, and every ? between the start and the hand-written cancel at the bottom is an exit that leaks a running observer onto hardware the attempt no longer controls.

Every attempt receives its own submit seam through Errand::attempt's scope: Scope<'_> parameter, constructed by the runtime per attempt. There is no flag to remember and no second shape to migrate to: a leaf writes _scope, and the day it grows a child it deletes one underscore. Children started through the scope are owned by the attempt. The focused work tests exercise the complete teardown and retry contract; application examples keep their own stories small instead of embedding a second orchestration tutorial.

Three rules do all the work:

  • Any exit tears the children down. Return, error, deadline, cancellation from outside — the runtime cancels every outstanding child and drains it to termination before the pass's own handle answers. There is no cleanup callback to forget, because there is no exit path without the teardown.
  • A retry drains first. The next attempt is admitted only after the previous attempt's children terminated; the backoff sleep may overlap that drain, admission may not. Attempt two never fights attempt one's observer for the instrument.
  • The scope carries only what the runtime alone knows. Its lifetime, its admission lane, its cause, the clock its deadline is measured on, and the one Resources value the assembly declared. Everything narrower than that — this attempt's input, anything the submitter computed — arrives through the job's own fields. There is no service locator here: one type, named where the application is declared, and read-only.

Detached submission stays for work meant to outlive the attempt, and the two spellings are the visible difference: scope. is owned, work. is not. Explicit handle.cancel() stays too, for tearing an observer down early — before the scope would.

A scope starts a child one way — scope.submit(job) — and the choice that used to split submit from a separate spawn verb is now just the lane the child's own attribute names: a specific variant, queued for that variant's capacity, or the catalog's #[default] variant, which is always unbounded and never queues.

// Routed onto a declared lane, and queued for its capacity.
let reading = scope.submit(TakeMeasurement { tool: tool.clone(), what: Measure::fast() }).await?;
// On the catalog's default, unbounded lane: it starts now, or it refuses now.
let dialogs = scope.submit(WatchDialogs { tool: tool.clone() });

WatchDialogs above is the observer shape: its attribute names lanes = Lanes rather than a specific variant, so it lands on Lanes' own #[default], which every catalog declares exactly once and which the compiler holds to unbounded. Landing there is never admission in the queued sense — the job starts immediately or the seam refuses it immediately — precisely because a watcher that is queued is not watching, and by the time it is admitted the thing it existed to see has happened.

Everything else is the submit rule verbatim, whichever lane a job names. A submitted child is owned by the scope, cancelled and drained on every attempt exit, and it inherits the attempt's cause. Its handle is awaitable and cancellable like any other; an observer simply is not meant to be awaited for a value, because the value of an observer is that it ran. let _watch = scope.submit(job); is the ordinary shape, and dropping the handle entirely still leaves the child owned.

A child landing on the default cannot hit the self-lane refusal below: the default is never the parent's own serial admission, because the default is always unbounded.

Two residuals are worth knowing about submit. First, a submission from an attempt to its own serial lane is refused immediately as Refusal::SelfLane: awaited, the child would deadlock on the one slot its own parent holds; unawaited, it would queue behind a parent whose scope close cancels it before it ever runs. A concurrent(n) lane is not refused — the deadlock there needs all n holders to be waiting on their own children at once, which no single submission can know — so keep the rule in view: children go on a different lane than their parent, and a lane deep enough for its parents is a capacity you declare, not a property the runtime can check. Second, a refused child whose handle nobody awaits does not vanish: a scope that closes over an unobserved refusal fails the attempt with it, because an observer that never existed is not an attempt that succeeded.

The full shape is a table, not a sentence:

Parent laneChild laneBehaviour
serialsamerefused Refusal::SelfLane at the seam
serialother serialqueues for that slot while parent holds its own; cycles are undetectable, deadline is the backstop
concurrent(n)sameadmitted, queues if full; deadlock needs n holders, not refused
anydefaultadmitted at once, never queues (the former spawn)
defaultanythe orchestrator shape

Read the last two rows as the orchestration guidance they are: an orchestrator names no specific lane — it declares lanes = Lanes and lands on the default, which is why it can start any mix of children without first asking whether they fit behind its own admission; a resource-bound leaf names the lane guarding that resource; and a serial-lane job should not orchestrate — a job holding a serial slot that goes on to submit and await several children of its own is exactly the shape the self-lane rule and the concurrent-lane residual above exist to make you notice.

Awaiting a child answers Result<Outcome, ChildError> — the child's own error type, not a wrapper over it. Give that type one derived variant and every child await in the application is a bare ?:

#[derive(Debug, thiserror::Error)]
pub enum Refusal {
#[error("the instrument service is not accepting commands")]
Unreachable,
// …the application's own vocabulary…
#[error(transparent)]
Refused(#[from] harmos::Refusal),
}
let captured = scope.submit(Sample { adapter: adapter.clone(), millivolts }).await?;

Two facts meet at that seam, and the variant is there because they are two. A job's own error is handed back as itself: a child that refused on its own terms already said why, and "a child job failed" laid over the top of it says strictly less than the child did, so harmos wraps nothing. The runtime's own refusals keep their type: a deadline, a cancellation from outside, or a submission the seam could not admit are the runtime ending the child rather than the child ending itself — a different fact, carried by harmos::Refusal, which is not generic over anything and therefore lands in an application's enum with #[from] writing the conversion.

AbsorbsRefusal is the bound both submit seams state, and an error type without the conversion is told so where it is submitted:

error[E0277]: a job's error type absorbs the runtime's own refusals
= note: add the one derived variant to it:
`#[error(transparent)] Refused(#[from] harmos::Refusal)`

Reach for .map_err(|error| error.to_string()) on a child await and you have thrown away a typed answer to get out of one function. There is nothing to write instead: the variant is one line, and children whose error types differ from the parent's convert between themselves the way any two error types in an application do.

An attempt that wants to know how long something took reaches for Instant::now(), and that is the one reach the scope exists to intercept. The call compiles inside an attempt, runs inside an attempt, and reports wall time — so the moment the same job runs under Simulate, where an hour of scripted history passes in a millisecond of real time, the number it produces is microseconds wearing the name of an hour. Nothing fails. The measurement is simply about a different clock than the deadline beside it.

scope.now() is the clock the runtime actually runs the attempt on — the one its deadline is measured against and its backoff waits on:

let demanded = scope.now();
let reading = scope.submit(job).await?;
let latency = scope.now() - demanded;

Elapsed time is ordinary subtraction, and the two answers are in one currency: a latency a job reports cannot exceed the deadline the same job ran under, under simulation and under wall time alike.

This is the third thing the scope carries, and it is the same kind of thing as the other two: which clock is in force is a fact about the attempt that only the runtime knows. It takes no type parameter because it is a reading rather than something the application declared.

The fourth thing the scope carries is the one value your application named as Resources and the assembly passed at boot:

impl Application for Tooling {
type State = Rig;
type Transaction = Transactions;
type Record = Records;
type Streams = Streams;
type Resources = Instruments;
}
let runtime = Runtime::builder(origin, Instruments::connect(&settings).await?)
.register(Transactions::catalog())
.register(Lanes::catalog())
.runtime()
.await?;

Inside any attempt, at any depth, it is one borrow:

async fn attempt(&mut self, scope: Scope<'_>) -> Result<Exported, ExportError> {
let instruments = scope.resources::<Tooling>();
instruments.client.push(&self.rows).await
}

Named the same way scope.journal::<Tooling>() is, and for the same reason: A is the runtime's own application, so an application writes its alias once and naming another application's is a wiring bug answered loudly.

Declare once at boot, borrow in the attempt. That is the whole pattern. A client, a connection pool, a resolved configuration — the things that are expensive to build, identical for every attempt, and wrong to copy into each job value — are built before boot, named once, and borrowed here. What stays in the job's own fields is what differs per attempt: which rows, which channel, which request.

Two things it deliberately is not. It is not a second state: the writer never sees it, replay never folds it, nothing in it is recovered, and a simulated reboot re-declares it rather than restoring it. And it is not a registry: there is exactly one value of exactly one type, so there is no key to get wrong and no "did someone register this?" to answer at runtime.

An attempt also knows which fact it answers. Name that fact on the construction line, and read the journal handle inside the attempt:

let job = DrainBacklog { tool: tool.clone() }.caused_by(receipt.position);
runtime.work.submit(job).await?;
// Inside the attempt: a handle already derived on that cause.
let journal = scope.journal::<Tooling>();
journal.record(BacklogDrained { rows }).await?;

Every fact the attempt records through that handle joins the causal family of the position the submission named — the same caused_by from Commit, applied where the cause is actually known instead of threaded by hand through the job's arguments.

And it does not stop at the attempt that was told. The scope's cause flows down until someone re-roots it. Every child a scope submits inherits the cause its parent's attempt carries, and because each child's own attempt opens a scope carrying what it inherited, the cause reaches grandchildren and below without anyone threading it:

let job = RunPass { tool: tool.clone() }.caused_by(receipt.position);
runtime.work.submit(job).await?;
// Inside `RunPass`'s attempt, and inside everything it starts, and inside
// everything *those* start: `scope.journal()` is already in that family.

To leave the family, say so. .caused_by(other) on a child re-roots that child and everything beneath it:

// This branch answers a different fact, and its own children follow it there.
let job = ArchiveBatch { rows }.caused_by(batch.position);
scope.submit(job).await?;

Two things do not inherit, both for the same reason. The detached work.submit has no scope, so there is nothing to inherit from — which is precisely what "detached" means, and it is why a background job started from inside an attempt carries no cause unless you name one. And a job submitted before any of this, from ordinary application code, is a root: it starts a family rather than joining one.

One honesty note, restated from the cancellation rule above: teardown is cooperative. A mutator cancelled mid-flight — parent deadline, external cancel — stops at its next await, and the device is in whatever state that operation left it. That is the same exposure a deadline expiry always had; the scope changes who remembers to cancel, not what cancellation is. And no job, observer included, runs without a declared deadline: an observer declares its parent's budget as its own, and the scope bounds its effective lifetime tighter on every ordinary exit. An unbounded deadline is a lie the first hung tool exposes.

Once an application has both lanes and residents, one question arrives: should this be a serial lane, or a resident?

Residents serialize resource access; lanes bound work admission. If a serial lane exists only because a device cannot take two commands at once, the serialization belongs to the resident that owns the device; the lane is then a throughput choice, not a correctness one.

There is a one-question test for it, and it is the question a reviewer should ask out loud: widening the lane must corrupt nothing. Change serial to concurrent(4) and read what breaks.

  • "More work runs at once, some of it waits longer, nothing is wrong" — the lane was a throughput choice. The declaration is honest and the number is yours to tune.
  • "Two commands reach the instrument" — the lane was carrying an invariant it cannot enforce. Move that invariant into the resident that owns the device, where a second command has to pass through one place, and let the lane go back to being a number.

A lane cannot hold that invariant, for two reasons worth knowing rather than memorizing. It admits attempts, not operations: an attempt that hands work to something else and awaits the answer is still holding its slot without holding the device — and an attempt that returns while a child it started is mid-command has released the slot without releasing it. And a lane is declared once at boot, far from the code that touches the hardware, while the resident is right there. A resident owns the channel; a #[lane(serial)] Device variant is admission policy, not the lock.

A Service is one mutable resident value owned by the supervisor, borrowed per attempt exactly as an errand is. The difference is what crosses that borrow: an attempt is handed a context and answers with a regime rather than an outcome:

impl Resident<Tooling> for Instrument {
type Error = String;
fn attempt(
&mut self,
context: Service<Tooling>,
) -> impl Future<Output = Result<(), Fault<Self::Error>>> + Send {
self.attempts += 1;
run_instrument(context, self.attempts)
}
}

Assembly names it once, and the runtime owns it from there:

let runtime = Runtime::builder(origin, resources)
.service("instrument", Instrument::default())
.runtime()
.await?;

The context hands each attempt journal() for deliberate application work, streams() for typed high-rate rows, and the stop signal tied to runtime shutdown. A plain async function or stateless closure still satisfies the resident contract directly.

A demand-driven capture loop from Publish Streams is a Service now: it owns a live resource and must settle before the journal's final drain. That makes demand something a resident reads, not something the runtime acts on:

let mut demand = context.streams().demand::<Reading>();
while context.until_stopping(demand.changed()).await.is_some() {
if *demand.borrow() > 0 {
instrument.capture().await?;
}
}

An attempt exits in one of three ways:

  • Ok(()) means it completed by design.
  • Fault::Environmental(e) means the world outside application logic failed: a bus reset, a disconnected device, a transient busy response.
  • Fault::Application(e) means the program's own rule or invariant failed.

Environmental exits restart under an exponential backoff and spend a per-service restart credit. Sustained health refills that credit. The default preserves the standard policy; assembly can declare another one explicitly:

let policy = ServicePolicy {
credit: 5,
sustained_health: Duration::from_secs(30),
backoff_base: Duration::from_millis(100),
backoff_cap: Duration::from_secs(10),
};
let builder = builder.service_with("instrument", policy, instrument);

Application exits and panics never restart; they deactivate immediately. Persistent environmental failure eventually exhausts credit and follows the same surfaced path.

The status is observable without inventing another manager:

let mut services = runtime.work.services();
let status = services
.wait_for(|all| {
matches!(
all.get("instrument"),
Some(ServiceStatus::Running { credit: 5 })
)
})
.await?;

Starting, Running { credit }, BackingOff, Stopped, and Inactive { cause, history } say exactly where the resident stands. Credit is updated when sustained health refills it, so an observer waits on the fact instead of sleeping past a private timer. Deactivation also appends a harmos-owned SystemFact under harmos:work/<service>/...; recovery can read that fact without loading the service code.

attempt begins with credit

environmental exit

credit remains

credit exhausted

application exit or panic

completes or shutdown wins

Starting

Running

BackingOff

Inactive

Stopped

The Service context owns two lifecycle races. context.rest(duration) sleeps through the work clock — which simulation virtualizes — and returns true when shutdown woke it early. context.until_stopping(future) runs any future until it completes, returning None when stop won:

if context.rest(Duration::from_millis(50)).await {
return Ok(());
}
let Some(command) = context.until_stopping(device.next_command()).await else {
return Ok(());
};

Everything domain-shaped stays on its owning primitive. Demand and service status are Tokio watch channels, so use Receiver::wait_for for predicates:

let watchers = demand.wait_for(|count| *count > 0).await?;
let running = services
.wait_for(|all| matches!(all.get("instrument"), Some(ServiceStatus::Running { .. })))
.await?;

Row conditions belong to ordinary stream combinators such as filter, find, and take_until, not to the Service context. The context deliberately stops at time and shutdown: adding event-wait or stream-condition helpers there would duplicate watch and stream semantics while hiding which source owns the wait.

Runtime::stop consumes the composition in one fixed order:

cancel queued jobs
-> let running attempts finish within their deadlines
-> signal and join resident services
-> signal and await registered host participants
-> drain the journal -> flush durability -> join
-> close stream delivery

Framework adapters and other external hosts join through runtime.stop_participant(). A participant receives the stop signal, closes admission, finishes bounded in-flight work, drops its journal handles, and calls finish(). Chapter 18 shows the complete host snippet.

If the writer join crosses its grace window, runtime.notices() names the number of live Journal clones. That is an ownership diagnostic: some host activity kept the writer-channel capability alive. Lengthening the grace window does not fix it.

1. A report generator retries twice and then fails. Should harmos append a work fact automatically?

No. It is finite work with a caller waiting for one answer, so the last typed error belongs to that caller. Record an outcome only when the application decides the outcome belongs in its durable story.

2. A USB adapter loses its device for three seconds. Environmental or application — and what happens?

Environmental. The service enters shared backoff, spends restart credit, and tries again. Persistent exits eventually exhaust credit and surface.

3. A capture loop follows stream demand and owns an open device. Job or Service?

Service. Its lifetime follows the runtime and shutdown must signal it before durability drains. Its samples still belong on Streams, not state.

4. A serial lane exists because the instrument takes one command at a time. Is that the right home for the rule?

No. Widen the lane to concurrent(4) and ask what breaks: if the answer is "two commands reach the instrument", the lane was carrying an invariant a queue cannot enforce. The serialization belongs to the resident that owns the channel; the lane stays as a throughput number beside it.

5. A workflow starts a dialog watcher, then one of its fifteen ? sites returns early. Who stops the watcher?

The runtime, if the watcher was started through the workflow's scope — every attempt exit cancels and drains the attempt's children before the handle answers, and a watcher wants scope.submit on the catalog's default lane, which never queues it behind anything. If it was started with the detached work.submit, nobody does until shutdown: detached is exactly the spelling for work meant to outlive the attempt, which a watcher is not.

Stokker Technologies markDesigned and built by Stokker Technologies