The Network Engineer's Guide to Latency
Light in fiber moves at 200,000 km/s and the market doesn't care. Latency broken into its four honest components — and how trading and voice engineers actually shave them.
Why latency became a career
In 2010, a new fiber route between Chicago and New York cut the trading round-trip from roughly 16 ms to 13 ms — and that single millisecond-plus advantage was immediately worth more than the cost of laying the fiber. Low latency is not just a trading curiosity, though. Voice calls start feeling wrong past ~150 ms, gaming and remote surgery live and die on it, and every TCP throughput formula on earth has RTT in the denominator.
The four components
Every delay you will ever chase decomposes into exactly four pieces:
- Propagation delay — the speed-of-light floor. Distance times ~5 µs/km one-way in fiber. Only shorter paths fix this — straight-line rights of way, microwave links (faster than fiber because air beats glass), and eventually free-space optics.
- Serialization delay — how long it takes to push bits onto the wire: packet size ÷ link rate. A 1,500-byte frame takes 12 µs on 1 Gbps and only 120 ns on 100 Gbps. Smaller packets and faster links both shrink it.
- Processing delay — the per-device cost: store-and-forward switching reads a whole frame (latency similar to serialization), cut-through starts forwarding after the header (~1–2 µs per hop), routers add lookup time, firewalls and DPI add much more.
- Queuing delay — the variable one. When a packet arrives to a busy interface, it waits in the buffer. Empty queue: ~0. Full 10 ms buffer: 10 ms. Queuing is the reason latency dances while bandwidth sits still.
Why averages lie
A link can have 2 ms average latency and ruin a trade on the one spike that matters. HFT and voice care about the tail: p99, p99.9. The tail comes from micro-bursts — tiny, millisecond-scale bursts that momentarily fill a buffer and add milliseconds of jitter. Tools that sample once a second never see them.
Two more average-killers: interrupt coalescing on NICs (the host batches interrupts, trading latency for throughput), and power-saving states (a CPU waking from C-state can add tens of microseconds). Production low-latency checklists disable both — along with queued-before-you forwarding behavior you cannot measure on a quiet box.
TCP makes latency multiply
Every round trip taxes TCP twice: once to fill the pipe and once to recover from loss. The rule of thumb is the bandwidth-delay product: BDP = bandwidth × RTT. At 1 Gbps with 100 ms RTT, you need roughly 12.5 MB of TCP window in flight — without window scaling (which caps at 64 KB and must be enabled to grow past it), you get 1/200th of the line rate. Run the numbers on our bandwidth-delay calculator with your own link; most "slow links" over long distances are window problems, not capacity problems.
Measure like the network owes you money
- ping/traceroute are the blunt instruments. ICMP responses from routers can be rate-limited or handled slowly in software; the path can also differ per-protocol. Use them to find hops, not to price latency.
- mtr over minutes beats traceroute once. It surfaces per-hop packet loss and jitter, which is the interesting part.
- Applications are the truth. Capture with Wireshark and read the timestamps between request and response on the wire. That number is what your users pay.
- One-way delay demands synced clocks. Without PTP (or at least solid NTP) on both ends, "one-way 4 ms" is a guess. OWAMP/TWAMP exist precisely to do this honestly. And when the sync itself has a hard number to hit — MiFID II's 100 µs traceability to UTC, say — the chain has to be budgeted, not assumed: add up every boundary clock and link with the PTP timing budget calculator before you trust the measurement. And once the budget adds up, simulate the chain over time — servo wander, per-hop asymmetry, holdover drift — in the PTP/SyncE network-wide timing simulator to see whether it actually stays inside the budget.
- Hardware timestamps. NICs that timestamp at the wire remove the host's jitter from the measurement. On software timestamps, the OS is part of the number.
The practical knobs
Before throwing hardware at latency: keep queues shallow (a giant buffer trades loss for delay — bufferbloat), keep MTUs consistent end-to-end to avoid fragmentation, prefer cut-through switches on the hot path, and remove middleboxes you do not need (each one is serialization plus processing plus its own queue). For trading-grade work, the playbook goes deeper — kernel bypass (DPDK) or FPGA-based feeds — but 90% of real networks never need it. Most need fewer queues, shorter distances, and honest measurements.
- Latency has four components: propagation, serialization, processing, and queuing.
- Fiber gives you ~5 ms of one-way delay per 1,000 km — physics sets the floor.
- Averages lie: jitter and worst-case queues decide real performance.
- Measure application latency with real timestamps, not just ping.