Skip to content

The Sidecar Facility

The sidecar facility isolates native code without inventing another application runtime. A sidecar is one executable supervised as resident work.

harmos-sidecar owns:

  • static executable declaration parsing;
  • Executable attachment and private stdio pipes;
  • the Tonic/prost gRPC protocol;
  • route and publication envelopes;
  • the author-side serve loop and local test double.

The root harmos crate owns runtime supervision, validation of the embedded restart policy, service status, stop choreography, and publication delivery into application code.

One bidirectional streaming RPC carries calls, answers, refusals, pushes, individual publications, and publication batches. HTTP/2 and protobuf supply framing and multiplexing over an already-open stdio connection.

A binary transport over a readable line protocol rests on three points:

  • Publications and telemetry are the volume on the wire; they need framing that sustains throughput and coalesces consecutive samples, not a line-per-frame text protocol sized for control traffic. Bulk data is unaffected — video, captures, and files still go to provided paths and never ride the wire; only the high-rate structured stream does.
  • Calls, lifecycle, refusals, pushes, and publications all cross the one bidirectional stream. A readable control protocol beside a separate high-throughput channel would mean two framings, two failure modes, and an ordering question nothing answers.
  • Wire readability is not a goal: authors touch only the Rust declaration surface — attributes and route types — so nobody hand-reads or hand-writes a frame.

Author-owned Rust values use #[harmos::message]; the attributes generate the sidecar declarations from those same types. Protobuf therefore owns both the fixed envelope and every typed payload without a second .proto model beside the Rust source.

Frames are bounded at 16 MiB. Frame::request_fits and Frame::reply_fits budget typed routes through the complete protobuf envelope with the exact maximum-width call id. The host refuses an oversized request before queueing it to Tonic, and the sidecar replaces an oversized answer with a bounded refusal under the stable Frame::OVERSIZED code, leaving the exchange usable.

Chunking an application value does not itself satisfy the frame bound: a batch of individually valid chunks is still one route reply and must fit as a whole. This is a control-plane escape hatch, not a bulk-data facility; video, captures, and files still belong on provided paths. Publication uses a bounded channel and non-blocking Supervisor::publish; monotonically increasing sequence numbers make loss visible.

Run mise run bench:sidecar:capacity for a real child-process stdio sweep. The default uses 16 KiB and 64 KiB samples, 100/300/500/700 MiB/s requested rates for two nominal seconds each, a 256 MiB uncapped transfer, and three repetitions of each case. It usually takes about a minute after the release build; slow or saturated cases take longer. Each sample size gets a 64 MiB uncapped warmup, printed but excluded from summaries. Uncapped transfers have fixed useful work, not fixed duration; brief runs are sensitive to startup and scheduler effects. Increase work and repetitions when confirming a result.

Terminal window
HARMOS_BENCH_SIZES=16384,65536 HARMOS_BENCH_RATES=400,500,600 HARMOS_BENCH_SECONDS=5 HARMOS_BENCH_REPETITIONS=3 mise run bench:sidecar:capacity
Environment VariableDefaultAccepted Values
HARMOS_BENCH_SIZES16384,655361–1048576 useful bytes/sample, up to 8 comma-separated integers
HARMOS_BENCH_RATES100,300,500,7001–4096 useful MiB/s, up to 8 comma-separated integers; uncapped is added automatically
HARMOS_BENCH_SECONDS21–30 nominal seconds per paced case
HARMOS_BENCH_REPETITIONS31–10 repetitions per case
HARMOS_BENCH_BURST_MIB2561–4096 useful MiB per uncapped repetition

Lists are sorted numerically and duplicates rejected. Empty, malformed, zero, and overflowing values fail with the setting's name before any child starts. Combinations must also fit 2,000,000 samples and 8 GiB per case, 256 recorded cases, 600 nominal sustained seconds, and 96 GiB total useful work including warmup. Very small samples may need reduced rates/durations to fit the sample bound. Every receive has a five-second stall timeout. A case allows four times the larger of its nominal duration and its transfer time at 100 MiB/s, plus ten seconds, capped at 180 seconds. The whole run has a 30-minute limit. Timeout, stream failure, or failed integrity assertions stop and reap the child before reporting the failed benchmark.

The receiver decodes every protobuf Sample, checks its complete byte pattern and length, and verifies count, stream/schema, and strictly increasing sequence. Sequence accounting carries across warmups and cases; the case's total sequence delta must equal attempted publications, so gaps match rejected enqueues. Both parent and child retain their current-thread Tokio runtimes. This measures the stdio/gRPC publication path with decoding and full payload validation; it excludes application work, full runtime stream retention, storage, hardware sampling, and a network link. Results are not directly interchangeable with a decode-and-length-only benchmark because pattern checking scans the full buffer.

Each repetition reports useful payload MiB/s, encoded protobuf payload MiB/s (excluding the gRPC envelope), samples/s, elapsed time, and queue rejected attempts. Summaries show median and minimum/maximum for these values and the achieved/target ratio. The producer retries rejected enqueues until the requested number is accepted and paces by accepted bytes. Rejections therefore count retry attempts, not permanent sample losses. A producer can also miss its requested rate with zero rejections.

A target is classified as keeping pace only when achieved useful throughput is at least 98% of the request and no enqueue was rejected. The summary names the highest requested rate meeting both conditions in every repetition, and separately the peak observed rate among uncapped or missed-target trials. Neither is a precise mathematical ceiling or a production guarantee. Saturation is reported, not a CI timing failure; integrity and transport failures do fail.

mise run bench:sidecar still runs the lightweight process and in-memory profiles (without capacity receiver decoding/pattern checks). The capacity task filters the actual child test; it does not silently skip an in-memory run. mise run test:sidecar:bench checks configuration and summary rules without running expensive measurements; mise run check:sidecar:bench lints the harness and child fixture.

The transport coalesces only immediately queued publications, up to 64 rows and at most 16 MiB for the complete encoded protobuf frame (including all nested wrappers). A row that would exceed the batch budget is deferred to the next frame, preserving order. A single available row is sent immediately. The benchmark supports 256 KiB and 1 MiB chunks:

Terminal window
HARMOS_BENCH_SIZES=262144,1048576 mise run bench:sidecar:capacity

Supervisor::publish returns Result<(), PublicationError>. Full means the bounded outgoing queue has no room; Closed means the connection ended; Oversized { encoded, limit } is a permanent refusal of a single row, before payload encoding. Only Full is retryable. Every attempted publication takes a stream sequence, including a refusal, so the next accepted row exposes the source gap. Oversized rows never enter the transport and do not kill the connection. Concurrent producers enqueue in assigned sequence order.

#[harmos::sidecar] declares identity, restart policy, and the catalogs naming what a sidecar contributes, on its struct: literal data a host reads before anything runs. Everything typed is on the author's own impl Sidecar — Config, Error, Resources, initialize, resources, optional interrupt and optional cleanup — which the attribute never reads, because a type belongs in an associated type. The executable starts only after static metadata validation and dependency planning. Registration runs before initialization and is stable for the loaded artifact even while an attempt is restarting or backing off.

Declarations are static: the inventory is read out of the embedded bytes, so nothing may be registered in initialize and initialization opens handles and nothing else.

A sampler's instance is a second lifecycle inside that one. It starts on the first lease of its stream and stops after the last release, both of which arrive as reserved pushes the host drives off its own consumer count. Finalization cancels and joins every running instance before the sidecar's cleanup, so nothing a sampler reads is released while it is still reading it.

An announced exit is application behavior. A broken process is an environmental fault and spends restart credit. A registration or initialization refusal is application behavior and is not retried. Runtime stop requests interruption and graceful cleanup — which quiesces every sampler, then runs the sidecar’s own cleanup — then ends a child that outlives the bound.

Scanning metadata executes nothing, but a native binary can still lie in its declaration. Installation must establish provenance and integrity. Runtime enforcement then limits the artifact to declared routes, bounded messages, and the OS permissions under which the process runs.

Runtime statistical grouping is independent of publication transport packing. See Statistical Stream Groups for scalar declarations, observations, source gaps, numeric validity and retention.

Stokker Technologies markDesigned and built by Stokker Technologies