Skip to content

How it works & the hard parts

oma runs on any LLM provider and any sandbox platform. This page explains what actually happens when an agent takes a turn, then walks through the problems that are genuinely hard — the ones that shape the whole design.

If you want the component map first, read Architecture. This page is the narrative version, plus the trade-offs.

  1. A message arrives. A user (or a webhook, schedule, or another agent) sends a user.message to a session. The session is a durable object with an append-only event log — the message is written to that log before anything else happens.

  2. The harness wakes up. The platform hands the harness everything the agent is allowed to use — its tools, skills, memory mounts, and the full event history — and the harness decides how to drive the model: how to build the context, what to cache, how many tool steps to allow.

  3. The model is called. The harness resolves the agent’s model handle to a provider (Anthropic, an OpenAI-compatible gateway, or a model card) and streams the response. Streaming chunks are broadcast live but are not the source of truth — the final event is.

  4. Tools run in a sandbox. When the model calls bash, read, write, or an MCP tool, the call is dispatched into the session’s sandbox — a container, a Kubernetes pod, a micro-VM, or a subprocess, depending on the environment. The agent’s code never holds a credential.

  5. Credentials are injected at the edge. If a tool makes an outbound HTTPS call that matches a vault credential, an outbound proxy adds the auth header on the way out. The token is never visible inside the sandbox.

  6. Every step is persisted, then broadcast. Each event is durably written before it is sent to any listener. If the process dies mid-turn, a fresh harness reads the log and continues — no lost work.

That “persist first, broadcast second” ordering, and the “credentials never enter the sandbox” rule, are not details — they are the two decisions everything else bends around.

Building an agent platform is easy to start and hard to finish. These are the problems that take real work.

The sandbox runs untrusted, model-driven code. It must never see a raw token — otherwise a prompt injection could exfiltrate every secret the agent can touch.

oma solves this with an outbound proxy: the sandbox makes a normal HTTPS request, and a proxy in front of it matches the destination against the vault and injects the auth header. The token lives outside the box the whole time.

The cost is subtle constraints that surface in practice:

  • The proxy has to be reachable from inside the sandbox’s network — a localhost URL that works on the host is meaningless inside a Kubernetes pod, so the proxy is exposed on a cluster Service and referenced by DNS.
  • Only clients that honor HTTPS_PROXY (curl, git, most SDKs) get injection. A client that ignores proxy environment variables bypasses it — a real gotcha for coding agents that ship their own HTTP stack.
  • The proxy terminates TLS, so its CA has to be uploaded into the sandbox and trusted before the first authenticated call.

If a session can crash — and at scale it will — then losing the conversation is not acceptable. oma’s answer is an append-only event log as the single source of truth, with a strict rule: persist before broadcast.

Because state lives only in the log, the harness is stateless. A crash mid-turn is recoverable: a new harness instance replays the log, rebuilds context, and resumes. Streaming chunks, thinking, and tool-input previews are broadcast-only — the persisted agent.* event is the record. Getting that split right is what makes “kill the process and it keeps working” true instead of aspirational.

Model providers cache identical prompt prefixes and bill the cached portion far cheaper. To benefit, the system prompt prefix has to be byte-for-byte deterministic across turns — any drift silently invalidates the cache and quietly multiplies cost.

That means the context-building code can’t casually reorder fields, and dynamic injections (like <system-reminder> blocks) have to sit inside the cached prefix in a stable position. It’s an invisible constraint: nothing breaks when you get it wrong, the bill just goes up. This is why context assembly is treated as a contract, not a convenience.

4. One business logic, two very different runtimes

Section titled “4. One business logic, two very different runtimes”

oma runs the same platform two ways: on Cloudflare (Workers + Durable Objects + Containers) and self-hosted Node (docker compose). A Cloudflare Worker is a single-file V8 isolate — no filesystem, no child_process, no dynamic module loading. Node has all three.

So a sandbox provider that shells out or reads the disk (subprocess, a native micro-VM, local Kubernetes) simply cannot run in a Worker, while a provider that only speaks HTTP (a remote BoxRun, a k8s gateway) runs in both. The platform classifies providers up front and, on Cloudflare, fails loudly and early for the node-only ones rather than half-working. Keeping one codebase honest across two runtimes with different physics is a constant tension — even the test suite is split by execution pool because of it.

There’s a third shape: a host the platform can’t fetch() at all, only relay to over a WebSocket — a paired laptop (subprocess via oma bridge daemon) or, newest, a user’s own browser tab running a WASM VM (browser-vm, v86 by default). browser-vm pushes sandbox compute all the way to zero — no container, no VM, no server cost — at the cost of proxied-only networking, no vault-injected outbound calls from the tab, and no memory/output mounts. Cloudflare only for now; see docs/browser-vm-sandbox.md.

5. Talking to a sandbox is harder than it looks

Section titled “5. Talking to a sandbox is harder than it looks”

Executing a command in a remote sandbox — a Kubernetes pod, say — means opening a streaming connection, wiring stdout/stderr/exit-code back, and tearing it down cleanly, every time. Small ordering mistakes turn into hangs.

The lesson generalized into a rule the platform now follows everywhere: a turn that can’t make progress must surface an error and return to idle, never hang in silence. Silent failure is worse than a loud one.

6. Any model, without special-casing every model

Section titled “6. Any model, without special-casing every model”

“Any LLM provider” means the harness can’t assume Anthropic’s wire format. Providers are addressed through a compatibility layer — Anthropic-style /messages, OpenAI-style /chat/completions, or a custom gateway — resolved from the agent’s model handle or a model card, with env-var fallbacks.

The messy reality lives at the edges: a gateway may reject a large max_tokens even for a tiny reply, or advertise a model it can’t actually serve. So the platform keeps output caps configurable and treats provider errors as first-class events, not surprises. Breadth costs you the comfort of a single well-behaved API.

Portable by construction

Because the model provider and the sandbox are both swappable behind interfaces, “any LLM provider, any sandbox platform” is a property of the architecture — not a feature bolted on later.

Safe by default

Credentials never enter the sandbox, and every turn is durable before it is visible. Safety isn’t a mode you turn on; it’s the default path.

Yours to inspect

Every rule above lives in source you can read and fork. When something behaves unexpectedly, you can see exactly why — and your fix can ship as a pull request.

Fails loud, not silent

Unavailable providers error early, stalled turns surface errors, and misconfigurations are reported — the platform prefers a clear failure over a quiet wrong answer.