blog · · engineering

A fence is not always yours

A ```lua code block in outl had a shell, arbitrary file read and write, and getenv. It took 7ms to prove. The interpreter was built with mlua's default library set, where safe means memory-safe rather than sandboxed: it excludes debug and ffi and includes os, io and package. That is a defect rather than a feature because a fence body is not always written by the person running it, arriving over sync from a paired device, through an import of someone else's graph, or from an LLM agent over MCP. This is the allowlist that closed it and why a denylist could not, the same hole found in the Lisp runtime during review, and then the harder half: a deadline the runtime contract demanded and no interpreter honoured, why stopping a Lua block takes three layers rather than one, why the timeout signal is an atomic the script cannot reach instead of a string it could forge, and what a bounded leak of four worker threads buys over an editor that freezes.

A
22 min read

A lua code block in outl had a shell.

$ outl template run pwn --page host --block 01M2… --json
{"result":{"duration_ms":7,"exit":"Ok","stdout":"os:\ttable\tio:\ttable\nHOME:\t/Users/avelino\n"}}

$ ls /tmp/outl-exec-*.txt
/tmp/outl-exec-escape.txt   # io.open + write
/tmp/outl-exec-shell.txt    # os.execute

Seven milliseconds. Arbitrary file read and write, getenv, and a process spawn, out of a fenced code block in a markdown note.

The cause is one line. The Lua interpreter was built with Lua::new(), which loads mlua’s StdLib::ALL_SAFE. I read “safe” and moved on, and that is the whole mistake in one word: safe there means memory-safe. The set excludes debug and ffi, the two that can corrupt the host process. It includes os, which carries execute. It includes io. It includes package.

That is a perfectly good default for an embedder running its own scripts. outl is not that, and it took me a year to notice.

the part that makes it a defect

If every fence in your workspace is one you typed, this is a curiosity. You already have a shell. The block is a slower way to reach it.

But a fence body arrives three other ways, and I built all three on purpose.

It arrives over iroh from a device you paired. It arrives through outl import, out of somebody else’s graph. And it arrives from an LLM agent driving outl_template_run over the MCP server, which is the path I built most recently and thought about least.

Sit with the first one. Pairing a phone is a decision about sync. You scan a QR, the two devices agree on an identity, and from then on your notes converge. Nowhere in that gesture did anybody decide that the paired device may execute code on your laptop. It just did, as a side effect, invisibly, because a page syncs and a page can hold a fence. The user made one decision and got two, and only one of them was on screen.

The third is worse in a quieter way. An agent reads a template page and runs it. The template came from wherever templates come from. Nobody typed it into a terminal, nobody read it before it ran, and the model has no concept of “this fence is load-bearing for my host’s filesystem.” There is also no human in the loop to be suspicious, which is the part of the loop I had been unconsciously relying on.

Once the input is untrusted, a runtime has to be bounded by construction, because the alternative is a boundary that holds exactly as long as nobody tries.

allowlist, because a denylist fails open

The obvious fix is to take os, io and package away. I did not do that, and the reason is worth more than the fix is.

A denylist is correct against the libraries that exist on the day you write it. The next release of mlua that adds a library grants it to every fence in every workspace, and nobody decides that. Nobody even notices: the code compiles, the tests pass, and the boundary quietly moved while the diff said “bump mlua”.

So the runtime names what it permits, and the comment in the source carries the argument:

/// The Lua standard libraries a code block is allowed to see.
///
/// The language, and nothing that reaches the host: string handling,
/// tables, arithmetic, unicode and coroutines. Deliberately absent:
///
/// - `os` — `execute` is a shell, `getenv` reads the environment.
/// - `io` — arbitrary file read and write, outside the workspace.
/// - `package` — `require` pulls a module off disk.
/// - `debug` — reaches into the interpreter itself.
fn stdlib() -> StdLib {
    StdLib::STRING | StdLib::TABLE | StdLib::MATH | StdLib::UTF8 | StdLib::COROUTINE
}

A library not on that list is absent whether or not I have heard of it. That is the only property that survives a dependency bump.

Then there is a second surface an allowlist expressed in flags cannot reach:

/// Base-library globals removed after the interpreter is built.
///
/// These are not part of any [`StdLib`] flag — Lua's base library is
/// always loaded — and each one re-opens what `stdlib` closed: a
/// loader turns a string into code, so leaving them is a lock with the
/// key next to it.
const LOADERS: &[&str] = &["load", "loadfile", "dofile", "require"];

Lua’s base library is always loaded. You cannot flag it off. So load, loadfile, dofile and require sit there after the allowlist has done its work, and each of them turns a string into code. They come off by hand, after the interpreter is built, by setting each global to nil.

That second step is exactly the kind of thing you skip when you believe the problem is already solved. The allowlist is the satisfying part. The four globals are the part that makes it true.

reviewing the fix found the same hole next door

While reviewing the change I went to check the other in-process runtimes, expecting to confirm they were fine.

Two of them were, by construction rather than by care. Python is built with Interpreter::without_stdlib, so there is no os to import and no open to call: the language is there and the library is not. JavaScript is a plain Context with no host bindings, which is ECMAScript and nothing else. Neither has a file or a socket to reach for, and neither needed a fix.

Lisp was not fine. Engine::new() registers steel/filesystem, steel/process, steel/tcp and steel/http. A ```lisp fence had command, open-output-file and tcp-connect. It was never mentioned in the original report, because nobody had looked.

It builds sandboxed now and shadows the host names that survive that. But the honest description of the result is a denylist held shut by a test, and there is a specific reason it cannot be better: Engine::new_sandboxed() still registers the whole steel/meta module, so a fence can construct an unsandboxed engine inside the sandboxed one and run a shell through that. The process module stays reachable through a primitive path. A defmacro body runs in a kernel engine an embedder cannot touch.

You cannot build a boundary out of that. You can only keep naming things and hoping you named them all, which is the definition of failing open.

So the runtime moved behind a cargo feature. lang-lisp is not in the default build, and the reason lives in the manifest where somebody enabling it will actually read it, rather than in a changelog nobody opens. It can be turned on, and it should be turned on by anybody whose fences are their own. It is no longer on for everybody by accident.

and then the timeout that was never there

The runtime contract said implementations must honour ctx.timeout.

sandbox::with_timeout, the helper that enforces it, had zero callers in production. Three of the four runtimes took _ctx and dropped it on the floor.

I measured it the boring way. while true do end was still running after twenty seconds.

Because execution is in-process and synchronous, that is not a stuck block. It is the TUI event loop frozen. It is the desktop wedged while holding the workspace mutex, so every other page is frozen with it. It is outl mcp serve hung, with an agent on the other end waiting on a call that is never coming back. Force Quit is the only way out of all three.

Here is where I expected to add a timeout and be done, and where the problem turned out to have layers.

stopping Lua takes three mechanisms, and each one has a hole the next covers

The first is an instruction hook. mlua can call you every N VM instructions, and past the deadline you return an error, which unwinds the interpreter exactly the way a script calling error() does. This is a real abort: the work stops, the thread ends, nothing leaks.

N matters more than it looks:

/// How often the deadline hook runs, in VM instructions.
///
/// A wall-clock check costs an `Instant::now`, so this trades deadline
/// precision for interpreter throughput. 10k instructions is well under
/// a millisecond of Lua on any machine outl runs on, which keeps
/// overshoot far below the smallest timeout a caller would set.
const DEADLINE_CHECK_INTERVAL: u32 = 10_000;

Check every instruction and you have a slow interpreter that is punctual about a deadline nobody measures that precisely. Check rarely and the deadline drifts. Ten thousand is under a millisecond of Lua, far below any timeout a caller would actually set.

Now the hole. The hook counts VM instructions, so it cannot fire inside a single long C call. And this is a single instruction:

string.rep('x', 2e9)

One instruction. Two gigabytes. The hook never gets a turn and the machine starts swapping.

So the second mechanism is a memory limit, through Lua::set_memory_limit. The allocator refuses and the block traps. Lua is the only one of the four runtimes that can do this, which is worth knowing when you are choosing a language for a note that might get shared.

Then the third hole, which neither of those covers: a C call that is slow without allocating. The test uses a classic:

return string.find(string.rep('a', 2000), '.-.-.-.-b')

Catastrophic backtracking. No allocation to refuse, no VM instruction to interrupt, and the pattern matcher will be in there for a very long time. So the whole run also happens on a worker thread, and the caller is released shortly after the deadline whatever the interpreter is doing.

“Shortly after” is a named constant, and the reasoning behind it is the kind of thing I would have got wrong by guessing:

/// How long past `ctx.timeout` the worker gets before the caller is
/// released without it.
///
/// The hook is the primary path: it fires at the deadline for any Lua
/// code and the worker returns `Err(Timeout)` through the channel like
/// a normal result. Releasing the caller at the exact same instant
/// would race it, and a worker abandoned a millisecond before it
/// finished on its own would be counted against
/// `sandbox::MAX_RUNAWAY_WORKERS` for that millisecond.
const C_CALL_GRACE: Duration = Duration::from_millis(100);

Release at exactly the deadline and you race your own hook. The worker was about to return cleanly, you abandon it a millisecond early, and now it counts as a runaway until it finishes.

One consequence makes Lua better than the other three: a Lua worker is only ever abandoned for the length of one C call. When that call returns, the hook fires and the thread exits on its own. The leak is bounded by the call, not by the script.

the coroutine that ran with no deadline at all

Two escapes turned up in review, and both would have passed a naive test.

The first is that a coroutine the script created ran unbounded:

coroutine.wrap(function() while true do end end)()

Lua::set_hook installs per Lua thread. mlua’s per-thread hook resolves the callback through a registry table keyed by the running thread. A coroutine created by the script goes through lua_newthread directly, is not in that table, and mlua’s response to a missing entry is to turn the hook off. Not to error. To disable it.

So the deadline was correctly installed, correctly firing, and one coroutine.wrap away from not existing. set_global_hook is inherited by every thread, and that is the one you want.

The second is that pcall ate the deadline:

pcall(function() while true do end end)
print('escaped')

The hook signals by returning an mlua::Error. pcall catches ordinary Lua errors. So the deadline error was caught, execution continued, escaped was written to the page, and the block reported Ok.

The fix is to check expiry on the success path too, not only on the error path:

match lua.load(source).eval::<MultiValue>() {
    // The deadline is checked here too, not only on the error
    // arm: the hook raises an ordinary Lua error, and `pcall`
    // catches ordinary Lua errors.
    _ if expired.load(Ordering::Relaxed) => Err(ExecError::Timeout(timeout)),
    Ok(values) => { /* … */ }

the signal is an atomic, because a string would be forgeable

That expired flag is an AtomicBool, and the choice is the most interesting decision in the whole change.

The natural implementation is a sentinel string. The hook raises error("outl-exec: deadline exceeded") and the caller checks whether the message matches. It works. It is also a hole, and the source comment spells out the shape:

// The hook raises the flag *before* returning the error, and the
// flag — not the error's message — is what `execute` reads. A
// string sentinel would be forgeable: a script calling
// `error("<sentinel>")` would be reported as a timeout, and
// since `ExecError::Timeout` is an infrastructure error the
// orchestrator writes no result block, so the script would
// suppress its own failure from the page. The script cannot
// reach this `AtomicBool`.

Follow that through. A timeout is an infrastructure error, so the orchestrator writes no result block: there is nothing useful to say and the block’s previous result stays put. A script failure is a user error, so it writes a trap into the page where you can see it.

Which means a script able to forge a timeout can suppress its own failure. It runs, it fails, and the page looks exactly like the block never ran. On a page that syncs to your other devices, from a fence you did not write.

An AtomicBool living in Rust is not reachable from Lua. There is nothing to forge. And there is a test whose name is the whole argument:

/// A script must not be able to make its own failure look like ours.
#[test]
fn a_script_cannot_forge_a_timeout() {
    let out = outl_exec::runtimes::lua::LuaRuntime.execute(
        r#"error("outl-exec: deadline exceeded")"#,
        &ExecContext::default(),
    );
    // a script calling error() must be a Trap, never a Timeout

the leak, and why it is bounded rather than absent

For Python, JavaScript and Lisp there is no instruction hook to install, so the deadline releases the caller and the worker thread keeps running until the process exits.

I want to be plain that this is a leak. A while (true) {} fence burns a core until you quit outl.

It is still the right trade, because the alternative is not “no leak”, it is the frozen editor from two sections ago. But I only believe that because the leak is capped, and the cap is the part I would have skipped if I had stopped at “release the caller”:

/// How many abandoned workers may still be running before
/// [`with_timeout`] refuses to start another.
///
/// Four is a budget, not a target. One runaway block is a mistake, two
/// is a pattern, and a fifth would mean the user has been shown four
/// timeouts and kept going. The count only includes work that is still
/// running, so a worker that finishes late frees its slot.
pub const MAX_RUNAWAY_WORKERS: usize = 4;

Without it, one non-terminating fence under auto-run:: on spawns a fresh core-burning thread on every page load. Open the page ten times, lose ten cores. And outl mcp serve under an agent does the same without a page load in sight, because an agent will happily retry a call that timed out.

Once four are outstanding, a new run refuses with an error instead of spawning a fifth, and that refusal lands in the result block where you can read it. The failure mode it replaces is a machine that gets slower over an afternoon for no visible reason. I would rather be told no.

A worker that finishes late hands its slot back, so the budget covers work that is still running rather than work that ever overran.

the three embeddings, and the one that lied

Each weak-form runtime is that way for a specific reason, recorded in the source so a dependency bump can re-check it rather than re-derive it:

  • boa 0.22 for JavaScript. RuntimeLimits caps loop iterations, recursion depth and stack size. Nothing per-instruction, nothing wall-clock, and HostHooks has no interrupt point.
  • rustpython 0.5 for Python. eval_breaker_tripped is pub(crate), and the opcode trace hook only fires when a frame’s f_trace_opcodes is set from inside Python.
  • steel 0.8.3 for Lisp. Engine::with_interrupted takes an Arc<AtomicBool> and the VM never reads it. Every occurrence of that flag in the crate is a write.

Spend a moment on the last one. It is an API shaped exactly like a cancellation point: it takes the type one takes, it is named the way one is named, and nothing on the other side looks at it. My first attempt at this fix used it and the test went green, because the block finished on its own before the assertion ran. An API that looks like a cancellation point and is not is worse than no API, because no API sends you looking.

Steel does have a real one behind a different door, get_thread_state_controller().interrupt(), checked by VmCore::safepoint_or_interrupt. So Lisp can be upgraded to the strong form. It has not been yet, and I would rather write that down than imply the runtime is at its ceiling.

the breaking change I am not going to bury

os.date, os.time and os.clock are gone from Lua blocks, along with os.execute and os.getenv. They ship in the same StdLib::OS flag, and the allowlist takes flags.

A Lua block that formatted a date now traps with attempt to index a nil value (global 'os'). That is real breakage for a note doing something entirely harmless, and there was no deprecation path, because the hole would have been open throughout the deprecation.

Re-exposing the clock functions as a hand-built os table keeps the allowlist honest and is worth doing. It is a feature though, and this change was about closing a hole. A fix that ships late because it grew a companion is a mistake I have made before.

what holds it shut

tests/sandbox.rs asks both questions of every runtime in one place: what can this reach, and does it stop. One file, every language, so adding a runtime cannot answer either question by omission. That is the same move as turning client parity into an exhaustive match: make the suite ask, because I will not remember to.

One detail in there I would not have thought of a year ago. The deadline tests cannot call execute directly:

/// Runs the block on its own thread and waits on a channel, because
/// the failure being tested for is *not returning*. Calling `execute`
/// directly here would hang the whole suite on the very defect the
/// test exists to catch — a red test has to fail, not wedge CI.
fn assert_times_out<F>(language: &str, source: &str, run: F)

A test for “this never returns” that calls the thing directly does not fail when the bug is present. It hangs, CI times out with no useful output, and the next person assumes the runner is flaky. The failure message names the three real consequences rather than saying “expected Timeout”, so whoever hits it in two years knows what is at stake:

{language} did not return within {PATIENCE} running {source} under a {DEADLINE} deadline, ctx.timeout is not honoured, and in a client this is the TUI event loop, the desktop holding the workspace mutex, or outl mcp serve frozen with no way out

The thing I keep taking from this one: the sandbox was not wrong because I misunderstood sandboxing. It was wrong because I read the word “safe” in somebody else’s vocabulary and assumed it meant what it means in mine. The /code page now says which language is in which regime, instead of calling the whole thing sandboxed.