Packet Path/

Kernel Bypass Basics for HFT Networks

Interrupts, copies, and syscalls make the stock kernel path cost 10–30 µs a packet — the whole tick-to-trade budget, spent before your code runs. DPDK, Solarflare Onload, and XDP explained with real numbers, real config, and the gotchas that eat your first month.

Low Latency · Intermediate · 10 min · October 2, 2026

Illustration of packets taking a fast direct path around a tall layered kernel stack

Why the kernel path costs too much

The Linux kernel network stack is excellent engineering — for the job it was designed for: serving everything from laptops to file servers through one general-purpose code path. High-frequency trading is not that job. In a colocation cage your competition reacts to the same market-data tick in single-digit microseconds, and the stock kernel path can burn that entire budget before your application sees a single byte. Walk an inbound packet through it and count the tolls:

StepWhat happensTypical cost
NIC interruptThe frame lands and the NIC raises an IRQ. Default interrupt coalescing may deliberately wait, batching interrupts to save CPU.~1–3 µs, plus 10s of µs of coalescing delay if enabled
Softirq + protocol processingThe driver builds an sk_buff and pushes it up through Ethernet, IP, and UDP/TCP: route lookup, socket lookup, checksums, filters.~2–5 µs
Queueing + wakeupThe packet waits in the socket receive queue while the kernel wakes your thread and the scheduler finds it a core.~1–5 µs — unbounded under load
Copy to user spacecopy_to_user moves the payload from the kernel buffer into your application's buffer.~0.5–1 µs per small packet, plus cache eviction
Syscall overheadrecvmsg() entry and exit, including modern Spectre-era mitigations.~0.5–1 µs per call

Add it up: a realistic kernel path costs 10–30 µs one-way on a tuned, idle box — and the tail under load is far worse, because nearly every step involves scheduling somebody else. That is the number you have to beat.

The bypass model: three ideas

Kernel bypass removes the kernel from the data path entirely. Three ideas do all the work:

The result is wire-to-application latency around 1–2 µs, repeatably. The price: one core pegged at 100% forever, a NIC the kernel can no longer see (no ping, no tcpdump on that port), and a server that has become a single-purpose appliance.

One packet, two host paths
Kernel path · ~10–30 µs
NIC IRQ + softirq Kernel IP/TCP stack Socket queue copy + recvmsg() Strategy app
Kernel bypass · ~1–2 µs
NIC DMA ring in user memory Busy-poll loop (dedicated core) Strategy app

The three approaches worth knowing

DPDK: you own everything above the wire

The Data Plane Development Kit hands you raw frames and gets out of the way. That freedom is total — including the freedom to reimplement ARP, IP, UDP, and TCP yourself, or to license a commercial stack. DPDK fits UDP multicast market data and custom feed handlers, where TCP was never invited anyway. Budget for real software engineering, not just tuning.

Solarflare Onload and TCP Direct: keep your sockets

Onload accelerates the standard POSIX sockets API in user space on Solarflare (now AMD) NICs: most of your existing TCP/UDP application runs unchanged, just faster. TCP Direct is its leaner sibling — a purpose-built API that skips some sockets generality for lower latency. This is the classic fit for order entry: TCP sessions to the exchange where reliability matters and rewriting the app in DPDK would be madness. The trade is vendor lock-in on the NIC.

XDP: the halfway house

XDP runs an eBPF program inside the NIC driver at the earliest possible moment — before the kernel allocates an sk_buff. Programs can drop, rewrite, or redirect packets at line rate, and AF_XDP sockets deliver frames to user space while the kernel stack keeps handling everything else. You keep ip route, tcpdump, and your sanity. Wire-to-app latency is not quite DPDK-class, but for line-rate filtering, DDoS shedding, and a gradual migration path, XDP is the pragmatic middle.

What it actually looks like

Bypass is mostly machine preparation: huge pages for the DMA rings, the NIC handed over to a user-space driver, and cores pinned so the poller never shares. A minimal DPDK host setup:

# 1) reserve huge pages for the DMA rings (512 x 2 MiB = 1 GiB)
echo 512 | sudo tee /sys/kernel/mm/hugepages/hugepages-2048kB/nr_hugepages
sudo mkdir -p /mnt/huge
sudo mount -t hugetlbfs none /mnt/huge

# 2) take the feed NIC away from the kernel driver
sudo dpdk-devbind.py --bind=vfio-pci 0000:3b:00.0

# 3) run: core 0 = main, core 4 = RX poller, core 6 = strategy
#    (all three on the same NUMA node as the NIC)
sudo ./build/feedhandler -l 0,4,6 -n 4 -- --rxq 1

The -l 0,4,6 flag is the whole philosophy in one line: core 0 runs control, core 4 does nothing but poll the receive queue, core 6 runs strategy. One queue, one poller, one owner.

On NICs that stay attached to the kernel (the Onload order-entry side, say), the equivalent tuning kills interrupt coalescing and trims the ring buffers:

# no interrupt coalescing, modest rings on the order-entry NIC
sudo ethtool -C eth1 adaptive-rx off rx-usecs 0 rx-frames 1
sudo ethtool -G eth1 rx 512 tx 512

# steer whatever IRQs remain onto housekeeping cores 0-3 (hex mask)
echo 0f | sudo tee /proc/irq/131/smp_affinity

Gotchas that eat your first month

Validate like the business depends on it (it does)

The ruleIf you cannot account for where each microsecond went, you do not have a latency number — you have a rumor.
Further reading
Key takeaways

Keep reading