zviz vs gVisor vs Firecracker: Sandboxing AI Code

Deny, emulate or virtualise? How zviz's selective-denial container runtime compares with gVisor, Firecracker and WASM for running untrusted AI code.

The question

Should I sandbox untrusted AI-generated code with zviz, gVisor, Firecracker, or WASM?

Sandboxing is the standard defence for code you cannot trust, and LLM agents that write and execute code have made the question routine. This is the comparison we wish we had when we started zviz.

The short version: gVisor emulates, Firecracker virtualises, WASM recompiles, zviz denies. gVisor puts a userspace kernel between the workload and the host. Firecracker gives each workload its own guest kernel inside a KVM microVM. WASM runs only code compiled to a portable bytecode. zviz runs a normal container on the host kernel and refuses the syscalls that matter most for escape. The right choice depends on your threat model, your workload’s syscall needs, and how much infrastructure you are willing to run.

What each option is

zviz is an OCI-compatible container runtime written in Zig that takes a selective-denial approach to isolation. It wraps a standard OCI bundle in a layered enforcement stack: namespaces (user, PID, mount, IPC, UTS), all 41 Linux capabilities dropped, a Landlock LSM ruleset, a seccomp-BPF filter and cgroups v2 limits. Of the syscall surface, 132 syscalls reach the host kernel natively, 24 are denied at seccomp and return EPERM, and socket is argument-filtered so safe address families can be permitted while raw ones such as AF_PACKET are refused. It is a single static binary with no daemon and no userspace kernel. It needs Linux 5.13 or newer (for Landlock) with cgroups v2.

gVisor is an application kernel from Google, written in Go. Its Sentry intercepts the sandboxed process’s system calls and services them in userspace, so most of them never reach the host kernel directly; it has been used in production by services such as Google Cloud Run and GKE Sandbox. Its strength is compatibility: because it emulates rather than refuses, it can support syscalls a deny-list cannot. The cost is an emulation layer on the syscall path, which matters most for syscall- and I/O-heavy workloads.

Firecracker is a virtual machine monitor from AWS, written in Rust, that runs lightweight microVMs on KVM with a minimal device model. Each workload gets its own guest kernel behind a hardware-virtualised boundary. It underpins AWS Lambda and Fargate. The costs are a guest kernel and VM image to maintain, a KVM-capable host, and a boot path on every new sandbox.

WASM runtimes (Wasmtime, WasmEdge and others) execute WebAssembly modules in a memory-safe sandbox with no ambient access to the host; capabilities such as files or sockets are granted explicitly through WASI. The constraint is that the code must be compiled to WASM, which rules out arbitrary native binaries and many native libraries.

The dimensions

DimensionzvizgVisorFirecrackerWASM
ArchitectureSelective-denial container runtimeUserspace application kernelKVM microVMBytecode sandbox
Allowed syscalls run…On the host kernel, nativelyInside the SentryIn a guest kerneln/a (WASI host calls)
Host kernel in the trusted baseYesReduced exposure via SentrySeparated by the hypervisorRuntime-mediated
Runs OCI imagesYesYes (runsc)Via a layer on topNo
Language supportAny Linux binary that avoids denied syscallsAny Linux binaryAny Linux binaryWASM-compiled code only
ptrace / mount / unshare insideDeniedSupported (emulated)Supported (in guest)n/a
Runtime shapeSingle static binary, no daemonSentry process per sandboxVMM process per microVMRuntime library or CLI
Host requirementsLinux ≥ 5.13, cgroups v2LinuxKVM-capable hostAny supported platform
MaturityEarlyProductionProductionProduction
Implementation languageZigGoRustVaries
LicenseMITApache-2.0Apache-2.0Varies (Apache-2.0 common)

We have deliberately left out overhead and start-up columns. zviz publishes no benchmark numbers, and figures for the other three vary so much with workload, configuration and host that a single cell would mislead more than it informed. The architectural difference is real, though: zviz has no emulation layer on the syscall path and no guest kernel to boot. Measure the rest on your own workload.

When to use which

Use zviz when:

  • You run untrusted code — agent tool calls, code-runner products, CI steps, multi-tenant user code — that does not need ptrace, mount or unshare.
  • You want the smallest reachable syscall surface with allowed calls at native speed: deny, don’t emulate.
  • You want a single static binary with no daemon that you can ship inside your agent runtime.
  • Your hosts run Linux 5.13+ with cgroups v2.

Use gVisor when:

  • The workload needs syscalls zviz denies: debuggers and strace, mount, unshare, Docker-in-Docker, or Bazel/Nix internal sandboxing.
  • You want the host kernel kept out of reach of most syscalls, not merely a shrunken set of them.
  • You must support kernels older than 5.13, or want a mature runtime with a long production record.

Use Firecracker when:

  • The threat model includes a host-kernel exploit and you want a hardware-virtualised boundary with a separate guest kernel.
  • You need VM-grade isolation for strict multi-tenancy.
  • You can run KVM and are prepared to maintain guest kernels and images.

Use WASM when:

  • You control the toolchain and can compile the untrusted code, or its interpreter, to WASM.
  • You want capability-based access where nothing is reachable unless granted.
  • You are running in a browser or an edge platform that already speaks WASM.

Where zviz loses

Selective denial keeps the host kernel in the trusted computing base for the 132 syscalls it allows. If an attacker has a kernel exploit reachable through one of those, zviz does not stop it; gVisor’s emulation layer and Firecracker’s hypervisor boundary are both designed for exactly that case. If your threat model is a determined attacker with a kernel zero-day, zviz is the wrong tool.

It also breaks workloads that need what it denies. There is no ptrace, so no debuggers or strace; no mount or unshare, so no nested containers or Docker-in-Docker. gVisor handles those safely. And Landlock is a relatively recent kernel feature: on a host that lacks it, the filesystem layer is not there, so verify the kernel feature set at deployment rather than assuming the headline protection is in place.

zviz vs gVisor: the closer comparison

Both raise the bar well above a plain runc container. The difference is emulate versus deny.

  1. Mechanism. gVisor services syscalls in its userspace kernel. zviz uses kernel primitives directly — namespaces, zero capabilities, Landlock, seccomp-BPF, cgroups v2 — and lets allowed syscalls run on the host.
  2. Surface. zviz’s policy is small enough to read: 132 allow, 24 deny, 1 argument-filter. gVisor implements a much larger share of the Linux syscall ABI, which is why it can run workloads zviz refuses.
  3. Runtime shape. zviz is one static Zig binary with no daemon. gVisor runs a Sentry process that supervises the sandbox.
  4. Kernel requirements. zviz needs Landlock, so Linux 5.13+. gVisor does not depend on Landlock.

If your untrusted code is agent-written and fits inside the allowed surface, zviz is a good fit. If it needs the syscalls zviz denies, or your threat model includes the host kernel itself, use gVisor or Firecracker.

Trying it

From the zviz quickstart: build from source with Zig 0.15+, export any image as an OCI rootfs, and run it under the hostile-tenant profile.

git clone https://github.com/Skelf-Research/zviz.git
cd zviz && zig build -Doptimize=ReleaseSafe

mkdir -p ~/bundle/rootfs
docker create --name x redis:alpine
docker export x | tar -C ~/bundle/rootfs -xf -
docker rm x

# add a minimal config.json to ~/bundle, then:
./zig-out/bin/zviz run my-container ~/bundle --profile hostile-tenant

On Ubuntu 24.04+ load the bundled AppArmor profile first; the quickstart has the exact commands.