Dropping Packets Before the Stack Wakes Up: A Practical Look at XDP
Linux XDP lets a verified eBPF program drop, pass, or redirect packets right after the NIC receives them, before the network stack pays any cost. A worked ICMP-drop example shows how to compile, attach, verify, and safely detach.
28 Nov 2025, 12:34 UTC

The packet that costs too much
Every packet arriving at a Linux server normally travels a long road: the NIC DMAs it into memory, the driver allocates a socket buffer (an sk_buff), and the packet climbs through the IP stack, netfilter hooks, and routing before any application sees it. For legitimate traffic that journey is the point. But when a host is absorbing junk — a flood of ICMP echo requests, port scans, or traffic you know you'll drop anyway — you're paying full stack-traversal cost for packets destined for the trash.
The useful takeaway: Linux lets you run a small verified program on the packet before the sk_buff is even allocated, using the XDP (eXpress Data Path) hook. Deciding to drop, pass, or redirect at that point can cut per-packet CPU cost dramatically, because the work happens right after the driver receives the frame.
Where XDP sits and what it can decide
XDP is a hook in the network driver path where an eBPF program — a small sandboxed program checked by the kernel's BPF verifier — runs against the raw packet buffer. The program ends with a return code that tells the kernel what to do:
XDP_DROP— discard the packet immediately. Cheapest possible outcome.XDP_PASS— hand it to the normal network stack as usual.XDP_TX— bounce it back out the same interface (useful for responders or load balancers).XDP_REDIRECT— send it to another interface or a userspace AF_XDP socket.XDP_ABORTED— an error path; drops the packet and is visible in tracepoints, so it should never appear in healthy operation.
Because this runs before the stack allocates metadata structures, a drop at XDP is far cheaper than an iptables DROP rule, which still pays for sk_buff allocation and several netfilter hook traversals.
A worked example: dropping ICMP echo requests
Assume a host with interface eth0, kernel 4.8 or newer (in practice, use 5.x or later for good driver support), and clang plus iproute2 installed. You need root or CAP_NET_ADMIN plus CAP_BPF (or CAP_SYS_ADMIN on older kernels) to load programs.
The program, icmp_drop.c:
#include <linux/bpf.h>
#include <linux/if_ether.h>
#include <linux/ip.h>
#include <linux/in.h>
#include <bpf/bpf_helpers.h>
SEC("xdp")
int icmp_drop(struct xdp_md *ctx)
{
void *data = (void *)(long)ctx->data;
void *data_end = (void *)(long)ctx->data_end;
struct ethhdr *eth = data;
if ((void *)(eth + 1) > data_end)
return XDP_PASS; /* malformed: let stack handle */
if (eth->h_proto != __constant_htons(ETH_P_IP))
return XDP_PASS; /* not IPv4: ignore */
struct iphdr *ip = (void *)(eth + 1);
if ((void *)(ip + 1) > data_end)
return XDP_PASS;
if (ip->protocol == IPPROTO_ICMP)
return XDP_DROP; /* drop all ICMP */
return XDP_PASS;
}
char _license[] SEC("license") = "GPL";
The bounds checks against data_end are mandatory — the verifier rejects any program that could read past the packet buffer. Compile on the host (or any machine with kernel headers):
clang -O2 -target bpf -c icmp_drop.c -o icmp_drop.oThen attach it to the interface. This changes live traffic behavior, so do it on a test host first:
sudo ip link set dev eth0 xdp obj icmp_drop.o sec xdpIf the NIC driver supports native XDP, this runs in the driver. If not, ip falls back to generic (SKB) mode, which still works but loses much of the performance gain — check with ip link show eth0 and look for xdp versus xdpgeneric in the flags.
Verify from another machine: ping -c 3 <host> should show 100% packet loss, while SSH or HTTP sessions to the same host keep working. On the host, ethtool -S eth0 exposes driver XDP counters on supported NICs, and sudo bpftrace -e 'tracepoint:xdp:xdp_exception { @[args->action] = count(); }' confirms you're not silently hitting XDP_ABORTED.
To detach and restore normal behavior:
sudo ip link set dev eth0 xdp offThe trade-offs you sign up for
XDP is not free complexity. First, you are parsing raw bytes yourself — every protocol you care about means manual header walking with verifier-proof bounds checks, and bugs drop or pass the wrong traffic. Second, the verifier is strict: no unbounded loops (on older kernels, no loops at all), limited stack space, limited instruction count. Third, debugging is harder than with iptables; a bad rule in netfilter logs cleanly, while a bad XDP program just makes packets vanish, and you need bpftrace or xdp_monitor from the kernel's bpf tools to see why. Fourth, native-mode support depends on the NIC driver — virtio, ixgbe, i40e, mlx5 and others support it, but always confirm before assuming driver-level performance.
A pragmatic pattern: use XDP for the cheap, high-volume decision (drop obvious junk, count by prefix) and pass everything else to the stack, where iptables/nftables and conntrack still do the nuanced work.
Closing: start with a counter, not a drop
If you want to try this safely, write your first XDP program to only count packets into a BPF map and always return XDP_PASS. Read the map with bpftool map dump, confirm the counts match what tcpdump sees, and only then add a drop rule. That gives you a verified observation point at the earliest moment a packet touches your kernel — and a clean rollback path (ip link set dev eth0 xdp off) the entire time.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.