What shipped since 0.12, and what it cost
A deep engineering update on outl: the elected endpoint lease that lets a headless daemon hold P2P sync without starving the GUI, a pairing handshake where anyone who knew your address could take the slot, page history read straight out of an append-only op log that turned out to be lying about 65,141 of its own Move operations, a 24.7 second boot caused by 5,656 fsyncs, client parity turned into an exhaustive match, and the one that cost most: 233 pages holding 1,426 lines of my own notes that existed in no operation, deleted by a repair command that printed 708 fixed while it ran. Every decision here comes with the alternatives I rejected and why.
This is an engineering update, written the way I’d want to read one. Every number came off a real workspace, every decision comes with the alternatives I rejected, and every fix is followed by what it broke. Most of them broke something.
Two workspaces show up repeatedly. One is mine: 2,560 pages, 213,859 ops, edited daily since June. The other is a 64k-block graph. Nothing here was measured on a benchmark I wrote to make outl look good.
the endpoint is elected, not assigned
Two devices converge when both are reachable. A laptop closed at 18:00 never overlaps a phone edited at 22:00, and that’s the whole reason people think P2P sync is unreliable.
Continuous sync was never the missing piece. Once any process on a device holds its iroh endpoint, the catch-up loop and the gossip converge with every paired peer on their own. What was missing was a headless process willing to hold it. outl serve was already the long-lived background process and it never called build_default_transport, so a machine running it under launchd synced with nobody and said nothing about it. That’s the same shape as #220, where a box running only outl mcp serve synced with nothing, in a different command.
outl serve holds the endpoint now (#244). Both halves have an off switch, --no-watch for endpoint only and --no-sync for watcher only.
The design problem isn’t holding an endpoint. It’s that there can be exactly one per device identity, and three processes want it: the desktop GUI, the TUI, and now a daemon that outlives both.
The naive daemon wins that race and never gives it up. That pushes every GUI on the machine permanently into the degraded mode where the sync indicator never turns green and Refresh can’t force a pass. So the supervisor retries instead of competing:
/// How long to wait before asking for the endpoint lease again.
///
/// Only ever paid while something else legitimately holds the endpoint (a GUI,
/// a TUI, `outl mcp serve`) or while there is nothing to sync with, so the cost
/// of a longer wait is "the supervisor takes up to this long to notice the GUI
/// closed". 30s keeps that unnoticeable without polling the lease file hard.
const LEASE_RETRY: Duration = Duration::from_secs(30);
A refusal is a normal state, not a failure. A GUI that’s already running keeps what it has, and the daemon takes over the moment that GUI exits.
It does not hand the endpoint back, and that asymmetry is deliberate rather than lazy. SyncTransport::start moves the lease onto the endpoint’s own thread, which releases it on shutdown or on a peers.json change. Nothing watches for a GUI that wants it. So a GUI opened after the daemon won stays degraded until the daemon stops.
I’d like that to be symmetric. It isn’t, because flock gives no signal that another process is waiting. Yielding needs a design of its own: a second channel where a waiter announces itself, and then a protocol for what happens when the announcing process dies mid-request. Deferring on the way in is cheap and correct. Yielding once held is a distributed handshake wearing a lock’s clothing, and I’d rather ship the half that works than the half that looks complete.
Two more behaviours that exist because of a specific failure:
It stands down entirely when no devices are paired. Holding the endpoint to sync with nobody only denies it to a GUI that could be using it to pair. A daemon that starts on boot on a fresh machine would otherwise make pairing impossible, on the machine that most needs to pair.
It re-reads peers.json on mtime change. PeersStore is read once, at transport build. A device paired from the GUI after the supervisor started would otherwise never be synced with, while the daemon reported itself perfectly healthy. Silent divergence with a green light is worse than a red one.
And the lease releases on SIGTERM and SIGINT. A lease left held by a killed process locks every outl process on that device out of an endpoint, which is a strictly worse bug than the one being fixed. Guards that outlive the thing they authorise are their own failure class.
the cost that couldn’t decide anything
I nearly rejected this whole design, and the reason I nearly rejected it is more useful than the design.
The argument against putting sync into outl serve: a permanently running daemon takes the per-actor write lock, so every later GUI launch loses that race, gets a fresh ephemeral actor, and mints its own ops-<ulid>.jsonl. On one 20-actor workspace I found 16 of those files, holding between 1 and 221 ops each. That’s real, and it’s ugly.
The conclusion I almost drew was “so sync needs its own command.”
That’s wrong, and it took me embarrassingly long to see why. The actor cost comes from running any process permanently, not from what that process does while it runs. A dedicated sync daemon under launchd pays exactly the same cost. The only reason the proposed alternative appeared cheaper is that it dropped the file watcher, and a flag drops the watcher just as well, at the price of one boolean instead of a second daemon, a second launchd job, and a second thing that can die quietly.
So the rule I wrote into the repo’s invariants: a cost only argues against a design if that design causes it. Before letting one rule something out, ask whether the alternative pays it too, whether the cost tracks the feature or just the lifetime, and whether it existed before your change. --no-watch is the mode to leave running next to a GUI you also use, because the sync half carries no actor cost at all.
the always-on peer, and why there’s no healthcheck
There’s a Dockerfile now, a compose file, and a self-hosting doc, so a machine you own can be the peer that never sleeps.
It’s optional, and it’s not a server your notes sync through. No account, no upload, nothing that holds your graph, same as the relay. It’s one more peer doing what your laptop does, holding a real replica: the op log itself, every op, replayable.
Two things the doc leads with, because both are expensive to get wrong. The server joins your graph, it never hosts the pairing, because the joiner adopts the host’s workspace id (RFC 0038) and pairing the wrong way round makes your laptop adopt the empty box’s identity and stop converging with its own notes. And /home/outl is a volume, not a detail: it holds identity.key, which is the node id every peer has stored. Lose it and the container comes back as a stranger that every device lists as offline.
There’s no HEALTHCHECK, on purpose, and the reasoning is in the Dockerfile so nobody helpfully adds one. outl workspace info would lose the per-actor write lock to the running daemon and mint a fresh ops-<ulid>.jsonl every thirty seconds. outl peer status would fight it for the endpoint lease. The probe would have been the outage.
the pairing handshake had no proof of possession
While a host was armed for pairing, roughly two minutes, the first device to connect on the pairing channel got accepted and handed the workspace identity. No PIN, no challenge, no check that it was the device the code was meant for (#159).
“Knew the address” is a low bar. Membership gossip hands every mesh member every peer’s address, so anyone already in your mesh had everything they needed, and so did anyone who’d ever seen a ticket.
Pairing codes carry a one-time secret now, and the joiner proves it holds the code before the host discloses anything.
The detail that matters is what the proof is bound to. A bare hash of the secret is replayable: capture it off the wire, present it yourself, get in. So the proof is keyed to the joiner’s own device id. A captured proof authenticates the device it was minted for and nothing else, which turns a passive wire attack into a no-op.
A refused attempt no longer burns the pairing window, and this half mattered more than the check. The CLI accepted exactly one connection and the GUI disarmed on any inbound dial. Adding a check that can fail, on top of that, would have converted a single packet into a denial of service against pairing itself: dial once, fail the check, and the legitimate device now has nothing to connect to. Both stay armed until a joiner actually passes.
Codes from an older version are refused with a message telling you to update the other device, rather than silently pairing without the check. A downgrade path that works is not a compatibility feature, it’s the attack.
The honest limit, in the docs and in the CLI output: treat a pairing code like a password for its two-minute life. Anyone who photographs it can still use it.
revocation is rotation, not a broadcast
outl peer remove only ever took effect on the machine you ran it on. Fine for retiring a laptop you still have. Useless for one that’s gone, and the command name implied otherwise, which is its own bug.
outl peer revoke-all (#158, RFC 0155) rotates the workspace identity, drops every pairing, and you re-pair what you still own. The device you don’t re-pair keeps the old id, stops sharing a gossip topic with your devices, and gets refused on any direct connection.
The obvious alternative is a propagated “revoke everywhere”: one device broadcasts a removal, every peer applies it. I rejected it for a reason specific to the threat model. Broadcasting a removal means any paired device can evict any other, and in the stolen-laptop case the attacker holds a paired device. Give that primitive to the mesh and the thief gets to use it first, on you. Rotation puts the authority in whoever holds the workspace identity and can re-pair, which is the property you actually want.
Two things it tells you up front, because people would otherwise learn them the hard way: a running GUI or outl serve still holds the old identity until you restart it, and the revoked device keeps the notes it already synced. Rotation stops new edits reaching it. It cannot take back history, and a security feature that implies otherwise is worse than one that admits the limit.
page history, and an append-only log that lied
Op::Edit carries a Yrs delta, not a snapshot, and the log is append-only. Every past state of every block has always been reconstructible. Nothing surfaced any of it.
The only history a user actually had was outl backup: workspace-granular git snapshots, not wired into desktop or mobile at all. On the workspace I built this against, that meant one commit, twelve days old, on a graph edited daily.
So outl page history <slug> and outl block history <id>, both with --limit and --json, plus a clock button in the desktop page header (#241). Read-only everywhere. Restoring stays outl recover’s job, which covers one narrow case with a provably additive rule, and a general restore needs its own safety argument before I write it. A history feature that can also write is two features, and only one of them can lose data.
Two decisions worth knowing about:
A deleted block stays in its page’s history, carrying the text it held when it went. Delete in outl is Move(node, TRASH_ROOT), never physical removal, so the text is right there. A history that omits deletions omits the exact change people open a history to find.
Not every op is an event. Folds, snoozes, page-slug writes, an edit that re-emitted the block’s existing text, a re-emitted Create or Move that changed nothing. A reconcile produces all of those in volume, and a timeline that shows them buries the three changes you’re looking for under four hundred that aren’t changes at all.
Building a reader for the log is also how I found out the log was lying.
Tree::do_op fills the old_parent, old_position and old_value fields that make an op undoable. Workspace::apply was handing do_op a clone and then persisting the caller’s original. Those fields are only as good as the code that built the op, and while block::moves::move_to reads them off the tree and gets them right, the reconcile and import paths pass root and None.
On the 64k-block workspace: 65,141 of 65,703 stored Move ops named the wrong old parent, and all 14,191 SetProp ops carried a null old value.
Sync and convergence were never affected, which is exactly why this survived so long. Every ingest path runs do_op, which overwrites those fields before the one function that reads them ever gets there. The corruption was invisible to every property the test suite checks, because the tests check convergence and convergence was fine. What was affected is anything that reads the log as data, which is the thing I’d just started doing.
apply stores what the tree recorded now. The values already on disk stay wrong, because append-only means append-only. There’s no migration to write, only a rule for readers: derive from the fields that describe an op’s own effect, never from the fields describing what it replaced.
5,656 fsyncs, and why I kept them
outl’s premise is that it opens fast and is ready for input. A CURRENT_PIPELINE_VERSION bump broke that: every sidecar goes stale by pipeline, so the first boot after an upgrade re-reconciles the whole graph.
On a 2,827-file workspace: 24.7 seconds at 8% CPU.
8% is the number that tells you what this is. It isn’t computation. It’s write_atomic doing two fsyncs per sidecar, 5,656 of them back to back, for 44 ops of actual content.
Moving it to a worker thread wasn’t enough, and this is a keystroke should never wait for the disk arriving from the other direction. The pass reacquired the lock immediately after each page and kept the disk saturated for the entire batch. The UI reads that same disk to paint. Thread placement is irrelevant when the contended resource is the device.
Three ways out, two of them wrong:
Drop the sidecar fsync. Takes it to 0.3 seconds. It’s the wrong trade and the failure mode is not “a stale sidecar”. A rename landing before its data leaves a sidecar full of garbage, which reads as a missing sidecar, which mints a fresh ULID per block. The page duplicates and every ((blk-...)) handle pointing into it breaks. That’s a durability bug that presents as a data-model bug three days later (RFC 0129).
Skip the migration for unopened pages. Worse. The parser fix in the next section then never reaches pages nobody opens, which is precisely the population that quietly holds broken content.
Yield in proportion to the work. BackgroundPace::COOPERATIVE sleeps as long as the page took, so the pass holds about half the device. The loop takes the workspace lock with try_lock, so a click never queues behind a migration, and it sleeps outside the lock, because sleeping while holding it is the same stall with extra steps.
The property I wanted is that a slow disk makes it yield more rather than stutter more. A fixed delay gives you the opposite: tune it for an SSD and it’s invisible there and useless on a spinning disk, tune it for the spinning disk and you’ve slowed everyone down. Making the ratio the constant means the pass adapts to hardware I’ll never test on.
Nobody needed the migration to be fast. It needed to be invisible.
parity that a build can check
Press a shortcut your client doesn’t implement and it used to do nothing at all. The desktop logged a line to a browser console you never open. From the keyboard, a missing feature and a broken one are the same event.
outl_shortcuts::support is now the single owner of which client performs which action, and it carries the sentence you see when one can’t. It’s an exhaustive match, so a new Action variant doesn’t compile until all three clients declare what they do with it. The gap gets recorded by whoever creates it, instead of found later by you pressing a key.
Writing it down immediately falsified three claims docs/shortcuts.md had been making. The desktop bound y r and : with no handler behind either, so both were dead keys. Mobile undo and redo were listed as “toolbar” when mobile had neither. Three hand-maintained copies of one fact, each stale in a different direction, and nothing in the build that could fail.
Two subtleties that took a second pass, and both are about where a fact lives:
The reason text belongs in the catalog, not the client. A client writing its own “not available here” wording is a fourth copy of the fact. The desktop already had the right shape, one shared message covering ten actions, and it still lived on the wrong side of the boundary.
“Reachable” and “has a handler” are different questions. Backspace on an empty textarea works on desktop and has no outl code behind it, because the platform does it. A boolean forces that row to lie in one direction or the other, so Native is its own state: reachable, no handler, no nudge.
The generated table lives at docs/client-parity, pinned by a test that fails if the doc and the code disagree.
That covers chords. Features with no chord, page history, the plugin marketplace, the calendar, templates, diverged the same way and were tracked nowhere, so there’s a second exhaustive match over a Capability enum reusing the same support states (RFC 0253). Mobile closed 24 of the 27 gaps that catalog enumerated (RFC 0254), and the interesting finding is that most of them weren’t missing features at all, they were features with no way in.
There was also a narrower defect underneath: outl-mobile/src-tauri didn’t depend on outl-shortcuts at all. Mobile’s column in the generated table was declared on mobile’s behalf, by a crate mobile doesn’t build against. Nothing in mobile’s own build broke when mobile diverged, which made its column a claim rather than a constraint. It builds against it now.
Colours got the same treatment for the same reason (RFC 0022). applyPaletteToRoot() mapped --color-ios-bg to the palette’s bg but --color-iosd-bg to bg_elev, so on desktop iosd meant elevated. Mobile’s stylesheet used that prefix for something else entirely: dark. One shared component read both namespaces and reached the iosd set through Tailwind’s dark: variant, which resolves off prefers-color-scheme.
Net effect: the operating system appearance setting changed the elevation of markdown blocks on desktop, with no relationship to the theme you picked. One token name, two meanings, and the wrong one got selected by whatever signal happened to be wired to it.
the hash answered a different question
This is the one that cost most, and I did it to myself.
The invariant is simple. The op log is the source of truth. The materialized tree and the .md files are projections. A page’s sidecar carries last_synced_hash, and if it matches the bytes on disk, outl wrote those bytes.
I read that as something stronger: that the op log holds what’s in the file.
sidecar.last_synced_hash == file_hash(disk)
answers: did outl write these bytes last?
read as: did these bytes come from the op log?
Those are different questions, and a page can answer yes to the first and no to the second. That’s exactly the state a reconcile_md leaves behind when it rewrites the sidecar without emitting ops covering everything it read.
On my own workspace, 2,560 pages and 213,859 ops, 233 pages held 1,426 lines that existed in no op (#210, RFC 0210).
The content wasn’t scratch. Infrastructure learnings with run ids, operational briefings, root-cause notes spanning 2022 to 2026. It read correctly on disk, in the editor, and in any grep.
Three consequences, in increasing severity:
- It never reached another device. Peers exchange ops, not files. All 1,426 lines existed on exactly one machine, belonging to a user who was trusting the sync story. That user was me.
- Nothing surfaced it. Not
doctor, not any consistency check, not one log line, becauselast_synced_hashagreed with the bytes and that agreement is what every downstream check tests. - Re-projection deleted it and reported success.
outl doctor --repairprinted708 fixedwhile removing content from 233 pages, then rebuilt each sidecar from the same render, so afterwards nothing could tell the page had ever held more.
It wasn’t limited to --repair. apply_page_md_with_sidecar_if_stale runs on every GUI open path, open_page_by_slug, open_journal_for, open_today_journal, open_ref, so opening the page in the desktop or mobile app was enough.
And the regression has a lineage. #166 fixed a real bug, tree ahead of .md so the page renders empty, by re-projecting whenever the tree outran a faithful projection, where faithful meant the hash matched. That traded one divergence direction for its mirror image, and the mirror is worse. Tree ahead of .md hides content the op log still holds. .md ahead of tree deletes it.
the guard, and the two times I got it wrong
content_lines_missing_from(disk, sidecar_blocks) -> Vec<String> is the single owner of the verdict “would re-projecting delete something?”. apply_page_md_with_sidecar_if_stale calls it after the hash gate passes and returns ActionError::PageMarkdownAheadOfLog { path, lines, sample } instead of writing.
The reference is the sidecar’s blocks, and the first version got that wrong. It compared against a fresh render of the tree, which answers “do disk and tree disagree?”. Every remote edit answers yes to that, because the pre-edit line is on disk and absent from the render. So does every remote delete and every reorder.
So the guard refused to re-project any page a peer had touched. That’s #166 reintroduced for the most ordinary sync case there is, with the blame moved. Worse, the recovery command the error message recommended wrote the pre-edit text back as ops, reverting the peer permanently, because the log is append-only.
That was caught in review by a five-minute executable probe against a real workspace, after 1,687 tests and a green /check had not caught it. The lesson isn’t “write more tests”. It’s that a guard whose correctness depends on which of two similar questions you asked cannot be validated by a suite that only ever constructs one of them.
The doc comment on the function now carries the whole argument, because the next person to simplify it will read the code before they read an RFC:
/// The content lines in `disk` that **no block the op log knows** can
/// account for.
///
/// `sidecar_blocks` is the reference, not a fresh render of the tree, and
/// the distinction is the whole point. The sidecar's blocks are what the
/// log held when the two last agreed, so comparing against them answers
/// *"does the op log know this line"*. Comparing against a render answers
/// *"do disk and tree disagree"*, which is also yes for every remote
/// edit, every remote delete and every reorder.
Three details in the comparison, each of which is a bug I’d otherwise have shipped:
It compares a multiset of trimmed non-blank lines, not a diff. A line the renderer merely moved is not at risk, and an LCS diff reports it as unique to disk. On the measured workspace that’s the difference between flagging 616 pages and the 233 that genuinely hold unlogged content.
Whitespace-only drift is ignored on purpose. The renderer’s trailing-newline behaviour changed between releases, so a large share of stale pages differ by exactly that. Treating it as content would strand every genuine re-projection behind noise, and a guard that fires constantly gets disabled.
The bullet marker is stripped exactly once. Repeating it would turn - - - x into x, matching a logged block x, which is a false negative: unlogged content walking straight past the guard.
Then there’s the case the guard has to refuse rather than answer. A sidecar written before SidecarBlock::text existed carries text: "" on every entry, so it describes which blocks the page had and nothing about what they said. An empty result from that sidecar doesn’t mean “nothing at risk”, it means “I couldn’t check”, and the two are opposite instructions:
/// - *"I checked, nothing is at risk"* → safe to write;
/// - *"I could not check"* → **not** permission to write.
///
/// Reading the second as the first is how a page holding unlogged
/// content gets overwritten by a caller that did ask the question.
sidecar_can_answer exports that condition so callers can tell the two apart, and it lives in the same file as the guard so the two can’t drift. An empty block list is the opposite case and answers true, because a page with no blocks has nothing to lose, and treating it as unanswerable would freeze every freshly created page.
the producer, which wasn’t where the RFC guessed
A guard stops the bleeding. It doesn’t explain how content got outside the log in the first place, and the RFC’s first guess about that was wrong.
render → parse wasn’t a roundtrip:
input: a block whose text carries a blank line and its own indentation
render: correct, every line emitted
parse: one block, first line only
warnings: 0
The parser’s contract is that nothing gets dropped in silence, and three arms of parse_block_list broke it. An over-indented line was recovered only at depth 0 and skipped mutely below it. A blank line inside a block’s text was read as a separator, when the renderer writes it indented and a real separator is empty, so the indent is the thing that tells them apart. A continuation line’s own indentation pushed it out of reach of the level that could have claimed it. A fourth arm warned about an unplaceable line and dropped it from the AST anyway, reasoning that a guard in another crate would keep the bytes on disk, which is a crate boundary being used as an excuse.
The reconcile that followed then wrote the truncation into the op log as an Op::Edit, so the loss reached the one place the RFC had described as still holding the content.
Fixing it took pages holding unlogged content from 41 to 8, and lines from 387 to 49, on a 2,827-file workspace.
Two numbers in this post describe the same graph and disagree, and that’s on purpose. The 233 pages and 1,426 lines came from the render-based detector. The sidecar-based one that replaced it measured 41 and 387 over the same graph. Both were acted on, and neither should be quoted without the method attached. Final residue after everything: 0 and 0.
The first version of the parser fix traded the bug for a worse one, three times over, and none of 237 green tests saw any of them. render → parse has to be a fixpoint, not merely lossless. The original bug at least converged. Two of the replacements mutated the document on every save:
- A bullet at an irregular indent (
- child) came back as verbatim text carrying its own-, so each pass read one more level of nesting:- parent\n - - child, then- parent\n - - child, and so on. - A whitespace-only line before a sibling appended a trailing
\nto the previous block. Invisible to the reader, a differentcontent_hashto the log, so that page emitted anOp::Editforever.
A parser that loses a line is a bug. A parser that changes the document every time it reads it is a bug that compounds, and the op log records every step of the compounding.
the rule, which isn’t about hashes
When you fix one direction of a .md ↔ tree divergence, state what happens in the opposite direction before you merge.
Reconciliation bugs come in mirrored pairs, and the pair that deletes is never the one being reported. Nobody files “my page silently lost four lines a week ago”. They file “my page is empty”, and the fix for that is what deletes the four lines.
Two commands exist because of this, covering the two halves. outl recover reads the op log to bring back text an Op::Edit truncated, with a provably additive rule. outl reconcile --ahead-of-log handles the other half: content the log never saw, recovered from the .md.
And a refusal has to reach the user. PageMarkdownAheadOfLog freezes the page in both directions until you run that command, so a client that swallows it into a log line ships a page that silently stopped syncing. That’s the same “silence is the defect” failure this whole invariant exists to prevent, moved up one layer. Every client surfaces it now.
The generalisation I keep coming back to: a fix relocates a problem far more often than it removes one. After “did I fix it?” comes “where does the problem live now, and what does that place require that the old one didn’t?” (RFC 0211 is the same lesson in a different domain: moving the write actor out of a synced config file into a device-local store was correct, and it put test runs into my real ~/.config/outl and left 1,166 orphaned records behind, because “how does a test get its own copy” and “what cleans it up” went unanswered.)
what’s next
- Closing the last mobile capability gaps, then aligning the CLI and MCP vocabulary (RFC 0255). Design tokens and the capability catalog were the first two steps of the same effort.
- Android release plumbing. The port already shipped and the APK is signed on every release. The doc claiming it was blocked on a filesystem watcher named the wrong blocker and was wrong for months.
- Reminders that fire when the app is closed.
remind::works today only while something is running, which is the least useful time for a reminder. - Per-page op log shards, when the single-jsonl-per-device layout hits the 10k-page wall. Not before.
The CHANGELOG carries the reasoning for everything above, and every RFC is in docs/rfcs, including the alternatives I rejected and what each change made worse.
Break it and tell me where.