Terminus Expanse
CurriculumBlogPricingSign in
Back to dispatches
guideslinuxAug 17, 2026

What eBPF actually is, and why half of modern infrastructure quietly depends on it

Terminus Expanse

The problem with the old way of extending the kernel

For most of Linux's history, if you wanted to observe or change something happening deep inside the kernel — which process is opening which file, which network packet is arriving on which interface, how long a system call actually takes — you had two bad options. You could write a kernel module: real C code that gets loaded directly into the kernel and runs with full, unrestricted access to the entire machine. A single bug in that code — a null pointer, a bad memory access — doesn't crash your program, it crashes the entire operating system, instantly, taking down every other process on the machine with it. Or you could avoid the kernel entirely and observe things from userspace instead, which is safe but slow and often can't see what you actually need to see, because the kernel is exactly where the interesting events happen first.

eBPF (extended Berkeley Packet Filter — the name is a historical artifact; it does far more than packet filtering now) is the third option that emerged to solve this: a way to run small, sandboxed programs inside the kernel, triggered by real kernel events, without the crash-the-whole-machine risk of a kernel module.

How it manages to be both fast and safe

An eBPF program is written in a restricted subset of C, compiled to a special bytecode, and then loaded into the kernel through a dedicated system call. Before that bytecode is allowed to run, the kernel's verifier walks through every possible execution path of the program and rejects it outright if it finds anything dangerous: an unbounded loop that could hang forever, a memory access outside the bounds it's allowed to touch, anything that could crash the kernel or leak memory it shouldn't see. Only code that survives this verification gets compiled just-in-time into native machine instructions and attached to whatever kernel event it's meant to respond to — a network packet arriving, a system call being made, a function being entered. Because it's actually running as native code inside the kernel rather than being interpreted, and because it's attached directly to the event it cares about instead of polling for it, an eBPF program can observe and react to things at a speed and level of detail that was previously only possible with the crash-your-whole-machine kind of kernel code — but the verifier's guarantees mean a buggy eBPF program gets rejected at load time or safely terminated, not allowed to take the kernel down with it.

What people actually use it for

Observability is where most engineers first run into eBPF, even if they never write a line of it themselves. Tools like bpftrace and the entire modern generation of Linux performance-debugging tools use eBPF to answer questions that used to require adding debug logging and restarting a service: which function is actually consuming the CPU right now, which process just opened this specific file, how long did this exact database query really take at the kernel level. Because eBPF programs attach directly to the running kernel, they can answer these questions on a live production system, without restarting anything and without the overhead of traditional debugging tools — a genuinely different tier of "can you look at what's happening right now" than existed before.

Networking is the second major use case, and it's why eBPF is quietly running underneath a huge share of modern cloud infrastructure. Cilium, a Kubernetes networking plugin used at serious scale, replaces the traditional Linux networking stack's iptables-based rules — which get slower as the number of rules grows — with eBPF programs that make routing and filtering decisions directly as packets arrive, scaling far better as a cluster grows to thousands of pods and network policies. When a company says their Kubernetes networking is "eBPF-based," this is what they mean: packet routing, load balancing, and network policy enforcement all happening as compiled code inside the kernel instead of as a slow chain of userspace rule lookups.

Security is the third pillar. Because eBPF programs can watch every system call, every file open, and every network connection as it happens, they're an extremely natural fit for runtime security tools that need to detect "this process just tried to do something it's never done before" in real time, rather than reconstructing what happened afterward from logs. Falco, one of the best-known open-source runtime security tools, is built entirely on this idea — eBPF programs watching kernel events live and flagging genuinely suspicious patterns as they occur.

Why "extended" is the right word

The original BPF, from 1992, was a much narrower tool built for exactly one job: filtering which network packets a program like tcpdump was allowed to see, efficiently, without copying every packet into userspace first just to throw most of them away. eBPF kept that same core mechanism — a tiny, verified virtual machine attached to kernel events — and generalized it to attach to almost any kernel event at all, not just packet arrival, which is why a technology that started as a packet filter now underpins observability tools, security tools, and load balancers that have nothing to do with filtering packets.

Why this is worth understanding even if you never write an eBPF program

You don't need to write eBPF bytecode by hand to benefit from understanding what it is — most engineers who rely on it do so entirely through tools built on top of it, the same way most people who benefit from TCP/IP have never written a socket by hand. What's worth internalizing is the shape of the idea: a way to safely run custom code at the exact moment and place inside the kernel where an event actually happens, instead of either taking the crash-the-whole-machine risk of a kernel module or accepting the blind spots and overhead of watching from userspace. Once you recognize that shape, you start noticing it everywhere — in why a company's Kubernetes networking scales the way it does, in how a production incident got debugged live without a restart, and in why "eBPF-based" has become something vendors advertise rather than a detail they bury in documentation.

This post is about