blog · · engineering

Two devices, one Wi-Fi, no internet

A laptop and a phone two metres apart had no local path between them. A peer's address was learnable exactly one way, through a discovery service that runs on the internet, so a new DHCP lease killed the stored address, every dial stalled about ten seconds on a corpse, and the fallback was the relay that is slow or blocked on precisely the networks where this gets reported. mDNS fixes it, and the interesting part is everything that had to be true for it to work on three platforms at once: an attach that happens after the bind because a failed mDNS start would otherwise take the whole endpoint down, a multicast lock on Android because the Wi-Fi driver discards the packets below the socket API, a soft-deprecated Apple API because the modern one hides exactly the thing iroh needs, a 63-byte DNS label that would have broken discovery in one direction only and in silence, and a test pinning all of it that spent its first version being a tautology.

A
20 min read

A laptop and a phone, two metres apart, on the same Wi-Fi, had no local path between them.

They could sync. They just did it by going out to the internet and back, because a peer’s address was learnable exactly one way: the relay and discovery service, which is an internet service. (The relay never reads anything, which I walked through byte by byte a while back. Being unreadable does not make it reachable.)

Most of the time you never notice. The failure is specific, and once you see it you see it every week.

the shape of the bug

Your laptop takes a new DHCP lease. New coffee shop, router reboot, a lease that expired overnight.

The address the phone has stored for it, in peers.json, is now a corpse. The phone dials it, and QUIC does what QUIC does with an address nobody answers at: it waits. About ten seconds per attempt. Then it falls back to the relay.

Which is slow, or blocked outright, because the networks where people report this are hotel Wi-Fi, conference Wi-Fi, a corporate network with an egress policy. The exact conditions that killed the direct path are the conditions that make the fallback bad.

So both devices show each other as offline, the journal does not converge, and the user’s model of what happened is “outl sync is broken”, which is a fair reading of the evidence.

The fix is not a better timeout. Two devices on the same link should not be asking the internet where each other are.

mDNS, and the one character that decides where it goes

mDNS resolves a peer’s current address off the LAN itself. Neither the stale stored entry nor relay reachability is on the path any more.

The obvious wiring is one line:

Builder::address_lookup(MdnsAddressLookup::builder())

It is wrong here, and the reason is a ? in somebody else’s code. iroh runs each registered builder inside bind():

into_address_lookup(&ep)?   // iroh-1.0.3, endpoint.rs:303

That question mark means an mDNS service that cannot start takes the whole bind down with it. And MdnsAddressLookupBuilder::build fails whenever the host allows neither IPv4 nor IPv6 multicast, which is not hypothetical: a locked-down corporate network, a container with multicast disabled, and iOS without the multicast entitlement all land there.

So the obvious wiring trades “same-LAN peers cannot find each other” for “sync does not start at all”. That is strictly worse, and it is undiagnosable from the user’s side. The app stops syncing on a network where it used to work fine over the relay, and nothing on screen connects that to a feature about local discovery.

Attaching after the bind keeps mDNS additive:

pub(crate) async fn bind_ipv4_only(
    relay_url: Option<&str>,
    secret_key: &iroh::SecretKey,
    alpns: Vec<Vec<u8>>,
    advertise: Advertise,
) -> Result<Endpoint> {
    let endpoint = n0_builder_ipv4_only(relay_url)
        .secret_key(secret_key.clone())
        .alpns(alpns)
        .bind()
        .await
        .context("bind iroh endpoint")?;
    attach_mdns(&endpoint, advertise);
    Ok(endpoint)
}

It either helps, or it logs one line and the relay path carries sync exactly as it did before this existed. The worst case is the status quo.

the diagnostic that would have made everything slower

Every endpoint in the crate binds under the same device identity. They all advertise the same node id, at different ephemeral ports. For the long-lived sync endpoint that is correct. For a transient one it is poison.

outl peer status binds an endpoint, probes, and drops it seconds later. If it advertised, every device on the LAN would learn a second address for this node id that is dead the moment the command exits. And iroh 1.0.0 multipath stalls on a dead candidate rather than skipping it, which is a thirty-second hang.

So running the diagnostic, the command whose entire job is telling you whether your peers are reachable, would have degraded sync for every peer that happened to resolve during it.

If that setup sounds familiar, it should: a status probe wearing the device identity is exactly how I broke sync once already, by stealing the relay route from the endpoint doing the real work. Same probe, different mechanism, and I did not see it coming the second time either. The lesson I took from the first one was “never share the identity”. The sharper version is that a short-lived endpoint must not publish anything a peer will cache.

There is a worse case than the probe, and it took a review to notice. The one-shot pairing endpoint accepts only PAIRING_ALPN. A peer that resolved it and then dials SYNC_ALPN does not get a stalled path, it gets CONNECTION_REFUSED.

Reading costs nothing and has none of these effects, so both still resolve. They just do not publish:

pub(crate) enum Advertise {
    /// Publish this endpoint's addresses on the LAN. Only the long-lived sync
    /// endpoint, which is the only one whose address stays true.
    Yes,
    /// Resolve other devices, publish nothing. For a transient endpoint whose
    /// address stops being true when the process exits: the `outl peer status`
    /// probe, and the one-shot pairing endpoint.
    No,
    /// Attach no LAN discovery at all: no advertise, and no browse.
    Off,
}

That third variant is there for a reason I would not have predicted. It is not a micro-optimisation for tests: Advertise::No does not avoid the mDNS machinery at all, because upstream gates only with_addrs on that flag and spawns the discoverer regardless. So every endpoint in the test harness bound port 5353 and browsed. At the endpoint counts the chaos suite reaches, that contention was enough to time out loopback dials that have nothing to do with discovery, and the failures presented as flaky sync rather than as a discovery problem.

Android, where it silently did nothing

The Wi-Fi driver drops multicast frames not addressed to the device. It is a battery optimisation and a sensible one.

The problem is where that filter sits. Below the socket API.

The Kotlin file that fixes it opens by saying what would have happened without it, and I would rather quote it than paraphrase, because the failure mode is the whole point:

/**
 * Android's Wi-Fi stack drops multicast and broadcast frames that are not
 * addressed to this device, as a battery optimisation. That filter sits in the
 * driver, below the socket API — so `swarm-discovery` joins `224.0.0.251`,
 * gets no error, sends its queries fine, and simply never receives an answer.
 * Discovery finds nobody, and nothing in the Rust layer can tell that apart
 * from "there are no peers on this LAN". Wiring mDNS without this would have
 * shipped Android a feature that looks present in the logs and does nothing.
 */

“Looks present in the logs and does nothing” is the category I find hardest to catch, because every instrument you would reach for says the code ran. The join succeeded. The query went out. There is no error anywhere. The only evidence is an absence, and an absence is what an empty network looks like too.

WifiManager.MulticastLock lifts the filter and needs CHANGE_WIFI_MULTICAST_STATE in the manifest. It is taken on resume and released on pause rather than held for the activity’s lifetime, because the lock is a documented drain: the radio stops filtering and the CPU processes every multicast frame on the network. It buys something only while a foreground session wants to find a peer right now. Background sync runs off known peers and the relay, neither of which needs multicast, so holding it while backgrounded is pure cost.

One small thing in there that is easy to get wrong: the lock is deliberately not reference counted, so the pause path releases in one call however many times resume ran.

iOS, where the modern API is the wrong one

Since iOS 14, a socket-level multicast join needs com.apple.developer.networking.multicast. Apple grants it by request. Apple commonly declines, and points applicants at Bonjour.

Building on an entitlement that might be refused after you ship is not a plan, so iOS does not join a group. It asks mDNSResponder, the system’s own mDNS daemon, to advertise and browse.

Then the question is whether Bonjour drags the entitlement back in through a side door. Apple’s Local Network Privacy FAQ names exactly two operations that still need it: working with arbitrary service types, and browsing for advertised service types, which is the _services._dns-sd._udp.local. meta-query. outl does neither. One fixed service type, declared in NSBonjourServices. No entitlement, no request form, nothing to wait on.

The part I did not expect is which Apple API you have to use. NetService is soft-deprecated. Network.framework is the modern answer. And the modern answer cannot do this job, in either direction:

/// `NWListener` opens its own port and advertises *that*. The port we must
/// advertise is the one iroh's QUIC socket already holds, and no
/// `Network.framework` type will advertise a socket it did not create.
/// `NetService(domain:type:name:port:)` publishes a record for an arbitrary
/// port, which is exactly the job.
///
/// Browsing has the same constraint from the other side: `NWBrowser` hands back
/// an opaque `NWEndpoint.service`, by design — the framework wants you to
/// connect through it rather than learn addresses. iroh needs the actual
/// `SocketAddr`s to hand to QUIC, and `NetService.addresses` is what exposes
/// them.

Read the second half again, because it is the more interesting one. NWBrowser withholding the addresses is not an oversight, it is the framework’s opinion: you should connect through the endpoint it hands you rather than learn where the peer lives. That is good advice for an app making its own connection, and useless when the thing making the connection is a QUIC stack that wants a SocketAddr.

So the deprecated API stays, and the comment says why, in the file, because “why are we on the old one” is a question somebody will ask in a review a year from now. Deprecated is not removed, and deprecation does not affect App Review.

one byte over

swarm-discovery publishes plain DNS-SD, RFC 6763. Service type _irohv1._udp.local., instance name is the endpoint id, TXT carrying the relay URL, SRV plus A and AAAA. Nobody reimplements a wire format, which is why a laptop advertising through swarm-discovery and an iPhone browsing through mDNSResponder resolve each other unchanged.

The instance name is where this first shipped broken, and the margin is absurd:

/// **Not `EndpointId::to_string()`.** That is `Display`, which `iroh-base`
/// implements as `HEXLOWER` — 64 characters. RFC 6763 §4.1.1 caps one DNS label
/// at 63 bytes, so a hex label is one byte too long and the publish fails
/// outright.

One byte. Base32 without padding is 52 characters and fits.

That alone is a normal bug: it fails loudly, you fix it. The reason it earns a named function and a test of its own is the next paragraph:

/// The failure this prevents is one-directional, which is what makes it worth a
/// named function: `PublicKey::from_str` accepts **both** encodings (hex when
/// the input is 64 chars, base32 otherwise), so a phone publishing hex still
/// *reads* a laptop's base32 label fine. The phone finds the laptop, the laptop
/// never finds the phone, and neither side logs anything.

A parser that accepts both encodings is a kindness that turns a loud failure into a silent asymmetric one. The phone works. The laptop does not. Each side’s logs look fine. And “sync works one way” is a bug report almost nobody words correctly, because from the user’s seat it reads as intermittent.

the test that was a tautology

All of that rests on two constants matching across three files written in three languages. Get one wrong and nothing errors; the peers never meet, which is indistinguishable from an empty LAN.

So there is a test. And the first version of it was worthless, which the current version says out loud:

/// The wire constants are the entire agreement with every other client, and
/// getting one wrong fails in the one way nothing catches: no error, no
/// log, the peers just never see each other.
///
/// So this reads the other two copies **off disk**. The first version of
/// this test asserted `SERVICE_TYPE == "_irohv1._udp."` against a literal
/// in this same file, which is a tautology: it stayed green no matter what
/// the Swift said, while three docs claimed it pinned the contract.
#[test]
fn the_lan_wire_constants_match_the_platform_bridge() {

I want to sit on that one, because I think it is the most useful thing in the whole change.

The test existed. It was green. Three separate doc comments cited it as the thing keeping the Rust and the Swift in agreement. And it compared a constant to a copy of itself, one line below, in the same file. Change the Swift to a different service type and the test stays green while discovery silently dies.

A test that asserts a constant equals its own literal is not a weak test. It is worse than no test, because a missing test is an open question and a green one is a closed one. Everybody downstream, me included, stops looking.

It reads the Swift bridge off disk now, and the Info.plist with it:

assert!(
    plist.contains(&format!("<string>{bare}</string>")),
    "Info.plist must declare {bare} in NSBonjourServices, or iOS treats it \
     as an arbitrary service type and requires the multicast entitlement",
);

That third file is not decoration. NSBonjourServices is what makes the service type declared rather than arbitrary, and an arbitrary service type is one of the two operations that needs the entitlement Apple declines. Drop the plist entry and the iOS build does not fail to compile. It fails to have permission, at runtime, on somebody’s phone.

There is a wrinkle worth stealing: NSBonjourServices takes the bare type and NetService wants the dotted form. Same fact, two spellings. The test trims the shared constant rather than hardcoding a third literal, because a third literal is how you get back to comparing a copy with a copy.

the crate that would not take the FFI

outl-sync-iroh is #![forbid(unsafe_code)], and Bonjour means FFI.

The lazy read is that the attribute got in the way. It did not. A sync transport should not hold a platform bridge, and the constraint pushed the bridge where it belongs: into outl-mobile, next to the background-sync exports that already cross that boundary. What crosses into the transport crate is plain Rust, strings in and strings out, no pointers.

One FFI detail I would have got wrong on my own: the Swift callback is a function pointer handed over at startup rather than a symbol resolved at link time. A pointer passed at runtime cannot be dropped by dead-code elimination, and it needs no guarantee about which crate survives the link. The alternative works in debug and can vanish in release, which is the worst possible schedule for a bug.

a structural guard, because a comment would not hold

Attaching mDNS inside bind_ipv4_only has a cost. A new endpoint that reaches for the builder directly and calls .bind() itself compiles fine and silently has no LAN discovery. Invisible, because sync still works over the relay. It just goes back to being unreachable on exactly the networks this was built for.

A comment saying “always use bind_ipv4_only” does not survive a new contributor, or a tired me. So the guard scans the crate and names the offenders:

assert!(
    offenders.is_empty(),
    "these files bind an endpoint without LAN discovery (issue #149); \
     call `bind::bind_ipv4_only` instead: {offenders:?}",
);

With one line I have started adding to every test of this shape, after the tautology:

assert!(
    scanned > 1,
    "the scan found no source files to check — it would pass vacuously",
);

A test that scans a directory passes when the directory is empty. That is the same failure as comparing a constant to itself: green and meaningless. One assertion that the scan found something turns a vacuous pass into a failure.

the one thing you can still hit

iOS reports nothing at all to an app whose local-network permission was declined.

There is no error and no empty result you can tell apart from a real one, so a denied iPhone looks exactly like an empty LAN from inside the app, forever.

Nothing in outl can detect it, so it is written down where somebody looking will find it: Settings, Privacy and Security, Local Network. That is in docs/sync.md and on the sync page.

I would rather ship a documented blind spot than a heuristic that guesses wrong and tells a user their network is broken when their phone is working exactly as configured.

Counting it up, the feature is a few hundred lines. What it cost was four platform facts that each fail in silence: a driver filtering below the socket API, an entitlement that gets declined, a parser accepting two encodings, and a modern framework hiding the one value the old one exposes. None of them error. All of them look like an empty network.