National Cyber Warfare Foundation (NCWF)

Inside bpfjailer: how it turns the eBPF LSM into a mandatory access control jailer


0 user ratings
2026-10-07 15:25:53
milo
Red Team (CNA)
"Inside

Facebook Incubator's bpfjailer enforces kernel-level Mandatory Access Control on Linux by placing processes into policy-driven pods via eBPF LSM programs, for defenders hardening hosts they own.








Toolfacebookincubator/bpfjailer — eBPF LSM based mandatory access control system and jailer for Linux, written in C++
CategoryKernel-level host hardening / eBPF security tooling
Primary UseConfining processes into role-based pods on hosts you administer, with TOML policies governing filesystem paths, exec, bpf(2), keyring, IPC, Unix sockets and mounts
Safe UseDeployed by authorized administrators on their own Linux hosts and VMs to enforce least privilege; strictly a defensive confinement framework, not an attack tool
Telemetry NotePurely defensive: denials and lifecycle events are written to pinned ring buffers and rendered by bpfjlog; programs pin under /sys/fs/bpf/bpfj-pins, making jailer presence easily auditable

Facebook Incubator's bpfjailer is a fascinating specimen in the modern eBPF security ecosystem because it is not a monitoring tool that observes and reports — it is an enforcement tool that denies. It uses eBPF LSM programs to push processes into kernel-enforced jails, which the project calls pods, each bound to a role defined in a TOML policy file. The project is explicitly a full rewrite of a closed-source internal predecessor, and the README is refreshingly candid: this is completely experimental, leverages newer kernel features like bpf arena that did not exist when the internal version was written, and issues are expected and not bug-bounty eligible. That honesty matters for anyone considering deploying it on production hosts.


The architectural core sits in bpfj/, which contains the jailer, the enforcers, the policy parser, and libbpf C++ helpers. Pod membership is inherited across fork and exec, which is the property that makes a jailer practical: you cannot escape your confinement by spawning children. Enrollment happens three ways. A binary can claim a role through the user.bpfj.policy.exec extended attribute and be enrolled at exec time; running processes can be enrolled directly by an administrator; and an unprivileged process can enroll itself through the bpfjsrv/bpfjclient pair, gated by the roles the policy marks unpriv-enroll.


The policy surface is where the tool's ambitions show. Optional features include signed binaries via fs-verity and named certificate sets, scoped kill and ptrace targeting, control over which roles' eBPF maps and programs may be opened or whether bpf(2) may be called at all, keyring write restrictions tied to fs-verity certificate management, filesystem read/write path rules, executable and shared-object path restrictions, kernel module and kexec loading denial, ownership-aware System V and POSIX message queues and shared memory with variable-expanded name patterns, Unix socket policy for pathname and abstract names, and mount rules scoped by destination and filesystem type. That is a strikingly complete mediation surface for a single LSM-based implementation.


The policy semantics follow a consistent grammar that the README documents well. Most operation gates default to deny when a role has no corresponding option; *-pod options allow resources owned by the same pod; *-roles adds named owner roles; *-any opens an operation completely; and keyring-own is the role-scoped counterpart because fs-verity keyrings belong to roles rather than pods. A role marked any = true opens everything not more specifically constrained, which the authors note is useful for attribution-only pods — tracking what a process does without restricting it. When a process holds several roles, every role must agree, consulted newest first, and an override-stacked role answers for the roles beneath it. Notably, the target side of kill and ptrace ignores override, so every role a potential victim holds must be listed — a sensible defensive asymmetry.


The component layout tells you this is a real system rather than a research artifact. bpfjctl is the general-purpose control binary for attaching, reloading, inspecting, and detaching the jailer; bpfjcmd is bpfjctl with its arguments and optionally its policy compiled in, ignoring argv so it can be statically linked and fs-verity signed as a single unit; bpfjsrv is a socket-activated daemon enrolling unprivileged callers into permitted roles; bpfjclient is a minimal client with no libbpf dependency; bpfjlog consumes diagnostics; and bpfjtest runs the suite. The examples/signed-attach scenario — a signed bpfjcmd being the only binary on the host allowed to update BpfJailer's own BPF programs — is a textbook demonstration of self-protecting security tooling.


Operational maturity shows in the live replacement machinery. bpfjctl replace loads a complete second jailer beside the active one, migrates pod membership, variables, and tracked resource ownership, then atomically swaps the pin trees under /sys/fs/bpf/bpfj-pins. Both trees stay attached during the handoff, forks and enrollments are coordinated with the migration, ownership changes are journaled and replayed, and replacement fails closed if persisted layout versions are incompatible or state cannot be copied safely. bpfjlog automatically reconnects across the swap, printing human-readable BPF diagnostics to stderr and structured events to stdout. This is the kind of hot-reload engineering that distinguishes tools designed for long-lived hosts.


Observability is ring-buffer based: denials and lifecycle events are written to pinned ring buffers, which bpfjlog renders and follows across policy replacement. For defenders, this is the telemetry backbone — every denial is an audit record, and every enrollment and pod lifecycle transition is an event. Because programs pin under a well-known bpffs path and stay loaded until detach, the jailer's presence and configuration are trivially auditable by host-level security tooling. The variables system (vars = ["vm_uuid"] in policy) allows per-pod scoping with ${NAME} expansion and glob wildcards, so two containers can hold the same role yet receive different filesystem permissions — a design clearly aimed at container-adjacent isolation.


The requirements bar is the main adoption hurdle. You need Linux 6.16 or newer with CONFIG_BPF_LSM=y and bpf in the lsm= boot parameter; older kernels are unsupported. The build needs clang for BPF codegen, bpftool, a C++20 compiler, libbpf, and a checkout of libarena for the arena spin lock the BPF programs use. Signed builds additionally require openssl, fsverity, and setfattr with STATIC=1. A basic build is simply make producing build/bpfjctl, with make test running the suite; the tests must run as root because each creates a mount namespace and mounts a bpffs, and they run serially by default because concurrent BPF LSM detach can panic affected kernels — parallelism is recommended only inside a disposable VM.


Day-to-day operation is centered on a handful of bpfjctl verbs: check parses a policy and reports what it holds, attach loads and pins the jailer, wrap ROLE USER_ID -- CMD runs a command in a new pod, enroll admits a running PID with variables, show and list inspect pods, and detach unpins and unloads. The README includes a sharp operational warning worth internalizing: wrap without --drop-cap and a non-root --uid leaves the command able to remove itself from the jail. Misconfiguration that silently regrants escape is exactly the failure mode a mandatory access control deployment cannot afford.


There are honest caveats an operator should weigh. Filesystem operations are currently performed in systemd's mount namespace to prevent mount manipulation from subverting the matcher, files unresolvable in that namespace are allowed, and configurable namespace selection is listed as coming soon — a fail-open edge that defenders should track. The README points to POLICY.md for the complete option matrix and to bpfj/policy/Policy.h for full stacking semantics, and licensing is clean MIT with the BPF programs dual-licensed MIT/GPL for kernel compatibility. At 44 stars and explicitly experimental, bpfjailer is not a drop-in replacement for SELinux or AppArmor today, but for authorized engineers studying how bpf arena-era eBPF can express ownership-aware mandatory access control directly in the kernel, it is one of the most instructive codebases currently published.



Official project repository for facebookincubator/bpfjailer.

Download Tool

Educational analysis for authorized security professionals. Use only in controlled, authorized environments.






Source: OffensiveSec
Source Link: https://www.offsecblog.com/2026/10/inside-bpfjailer-how-it-turns-ebpf-lsm.html


Comments
new comment
Nobody has commented yet. Will you be the first?
 
Forum
Red Team (CNA)



Copyright 2012 through 2026 - National Cyber Warfare Foundation - All rights reserved worldwide.