Kernel Bypass Basics for HFT Networks
Interrupts, copies, and syscalls make the stock kernel path cost 10–30 µs a packet — the whole tick-to-trade budget, spent before your code runs. DPDK, Solarflare Onload, and XDP explained with real numbers, real config, and the gotchas that eat your first month.
Why the kernel path costs too much
The Linux kernel network stack is excellent engineering — for the job it was designed for: serving everything from laptops to file servers through one general-purpose code path. High-frequency trading is not that job. In a colocation cage your competition reacts to the same market-data tick in single-digit microseconds, and the stock kernel path can burn that entire budget before your application sees a single byte. Walk an inbound packet through it and count the tolls:
| Step | What happens | Typical cost |
|---|---|---|
| NIC interrupt | The frame lands and the NIC raises an IRQ. Default interrupt coalescing may deliberately wait, batching interrupts to save CPU. | ~1–3 µs, plus 10s of µs of coalescing delay if enabled |
| Softirq + protocol processing | The driver builds an sk_buff and pushes it up through Ethernet, IP, and UDP/TCP: route lookup, socket lookup, checksums, filters. | ~2–5 µs |
| Queueing + wakeup | The packet waits in the socket receive queue while the kernel wakes your thread and the scheduler finds it a core. | ~1–5 µs — unbounded under load |
| Copy to user space | copy_to_user moves the payload from the kernel buffer into your application's buffer. | ~0.5–1 µs per small packet, plus cache eviction |
| Syscall overhead | recvmsg() entry and exit, including modern Spectre-era mitigations. | ~0.5–1 µs per call |
Add it up: a realistic kernel path costs 10–30 µs one-way on a tuned, idle box — and the tail under load is far worse, because nearly every step involves scheduling somebody else. That is the number you have to beat.
The bypass model: three ideas
Kernel bypass removes the kernel from the data path entirely. Three ideas do all the work:
- User-space drivers. Your application links a poll-mode driver that owns the NIC's queues directly; the kernel driver is unbound from the device. The NIC DMA-writes frames into ring buffers living in your process's own memory.
- Busy-polling. No interrupts, no wakeups, no scheduler. A dedicated core spins in a tight loop checking the ring — packet or no packet, thousands of times per microsecond. It sounds wasteful; it is what determinism costs.
- Zero-copy. The frame never moves. Your code reads it in place in the DMA buffer, decides, and releases the slot. No
sk_buff, nocopy_to_user, no evicting your strategy's data from cache.
The result is wire-to-application latency around 1–2 µs, repeatably. The price: one core pegged at 100% forever, a NIC the kernel can no longer see (no ping, no tcpdump on that port), and a server that has become a single-purpose appliance.
The three approaches worth knowing
DPDK: you own everything above the wire
The Data Plane Development Kit hands you raw frames and gets out of the way. That freedom is total — including the freedom to reimplement ARP, IP, UDP, and TCP yourself, or to license a commercial stack. DPDK fits UDP multicast market data and custom feed handlers, where TCP was never invited anyway. Budget for real software engineering, not just tuning.
Solarflare Onload and TCP Direct: keep your sockets
Onload accelerates the standard POSIX sockets API in user space on Solarflare (now AMD) NICs: most of your existing TCP/UDP application runs unchanged, just faster. TCP Direct is its leaner sibling — a purpose-built API that skips some sockets generality for lower latency. This is the classic fit for order entry: TCP sessions to the exchange where reliability matters and rewriting the app in DPDK would be madness. The trade is vendor lock-in on the NIC.
XDP: the halfway house
XDP runs an eBPF program inside the NIC driver at the earliest possible moment — before the kernel allocates an sk_buff. Programs can drop, rewrite, or redirect packets at line rate, and AF_XDP sockets deliver frames to user space while the kernel stack keeps handling everything else. You keep ip route, tcpdump, and your sanity. Wire-to-app latency is not quite DPDK-class, but for line-rate filtering, DDoS shedding, and a gradual migration path, XDP is the pragmatic middle.
What it actually looks like
Bypass is mostly machine preparation: huge pages for the DMA rings, the NIC handed over to a user-space driver, and cores pinned so the poller never shares. A minimal DPDK host setup:
# 1) reserve huge pages for the DMA rings (512 x 2 MiB = 1 GiB)
echo 512 | sudo tee /sys/kernel/mm/hugepages/hugepages-2048kB/nr_hugepages
sudo mkdir -p /mnt/huge
sudo mount -t hugetlbfs none /mnt/huge
# 2) take the feed NIC away from the kernel driver
sudo dpdk-devbind.py --bind=vfio-pci 0000:3b:00.0
# 3) run: core 0 = main, core 4 = RX poller, core 6 = strategy
# (all three on the same NUMA node as the NIC)
sudo ./build/feedhandler -l 0,4,6 -n 4 -- --rxq 1
The -l 0,4,6 flag is the whole philosophy in one line: core 0 runs control, core 4 does nothing but poll the receive queue, core 6 runs strategy. One queue, one poller, one owner.
On NICs that stay attached to the kernel (the Onload order-entry side, say), the equivalent tuning kills interrupt coalescing and trims the ring buffers:
# no interrupt coalescing, modest rings on the order-entry NIC
sudo ethtool -C eth1 adaptive-rx off rx-usecs 0 rx-frames 1
sudo ethtool -G eth1 rx 512 tx 512
# steer whatever IRQs remain onto housekeeping cores 0-3 (hex mask)
echo 0f | sudo tee /proc/irq/131/smp_affinity
Gotchas that eat your first month
- NUMA amnesia. NIC on node 0, application pinned to node 1: every frame crosses the inter-socket link and picks up microseconds of jitter. Check
/sys/class/net/ethX/device/numa_nodeand pin cores and memory on the NIC's node. - Un-isolated pollers. If the scheduler can park a cron job on your polling core, it will — mid-trade. Boot with
isolcpus,nohz_full, andrcu_nocbson the poller set, steer every remaining IRQ elsewhere, and disable deep C-states: a core waking from C6 can take tens of microseconds to become useful. - Missing huge pages. At millions of packets per second, walking rings mapped in 4 KiB pages thrashes the TLB. 2 MiB or 1 GiB pages make the ring effectively free to walk.
- Receive-side trade-offs are now yours. Polling burns a core at 100% whether the feed is alive or dead. Small rings drop frames under microbursts; big rings turn drops into queueing delay. And market-data UDP has no retransmission — a dropped sequence is a gap, and gap recovery (retransmit channels, snapshot feeds) becomes your code's problem.
- "Just add more cores" fails. Tick-to-trade is a serial chain: propagation, queue, decode, decide, send. Parallel cores do not shorten a serial path, and threads bouncing cache lines between sockets make it slower. One queue with one owner beats eight cores sharing a lock.
- Software timestamps lie politely. A timestamp taken in your driver measures your stack, not the wire. Without NIC hardware timestamping disciplined by PTP, every validation number you produce is a rumor.
Validate like the business depends on it (it does)
- Timestamp at the NIC. Compare hardware ingress timestamps against the moment your strategy finishes deciding. That difference — not a software timer — is your host latency.
- One number rules the review: tick-to-trade. Time from tick on the wire to order on the wire, in a single PTP clock domain. Build the full path budget in the tick-to-trade calculator and see which component owns your microseconds. That clock domain has its own budget, though — every boundary clock and link asymmetry spends a piece of it, so run the chain through the PTP timing budget calculator and confirm the timestamps you're validating against can actually hold the accuracy you claim. A static budget is only the starting point, though: clocks wander, so run the same chain through the PTP/SyncE network-wide timing simulator and watch whether the servo keeps every hop inside its allocation over time.
- Measure tails, not averages. Record p50 / p99 / p99.9 histograms over a full trading day. The average will look great on the day you lose money.
- Change one variable at a time. Same feed, same box: kernel socket vs Onload vs DPDK. Brochure numbers were measured on the vendor's workload, not yours.
- The Network Engineer's Guide to Latency — the four latency components this article builds on.
- TCP Internals Every Network Engineer Should Know — what the kernel stack does with your packets, and why it costs what it costs.
- Low-Latency HFT Network Engineer: Where Nanoseconds Are the Job — the career path this skill belongs to.
- Tick-to-Trade Latency Calculator — model the whole path, component by component.
- Bandwidth-Delay Product Calculator — window sizing and throughput bounds for the long-haul side.
- A stock kernel path costs ~10–30 µs one-way — interrupts, softirq processing, queueing, copies, and syscalls each take their cut.
- Kernel bypass = user-space driver + busy-polling + zero-copy DMA rings. Wire-to-app drops to ~1–2 µs, deterministically.
- DPDK owns the whole stack, Onload keeps the sockets API, XDP is the halfway option that keeps your kernel tooling.
- Pin cores and memory to the NIC's NUMA node, use huge pages, and validate with hardware timestamps — never averages.