Packet Path/

ECMP Explained: How Data Centers Spread Traffic Across Every Link

ECMP turns a Clos fabric's dozens of equal-cost paths into real bandwidth — until identical hashes polarize traffic onto a few hot links. How the 5-tuple hash works, why polarization happens, and how to break it.

Routing · Intermediate · 9 min · October 2, 2026

Illustration of a spine-leaf fabric with flows spreading across many equal-cost links

One best path is a wasted fabric

Picture a modern spine-leaf data center: every leaf switch connects to every spine. Between any two racks there aren't one or two paths — there are dozens, all the same length, all the same cost. Now ask the router how it forwards. Classic routing picks exactly one best path and programs it into the forwarding table. The other thirty-one links become expensive decoration. You've built a 32-lane highway and you're driving every car down a single lane.

Equal-cost multi-path routing — ECMP — is the fix. When the routing table holds several equal-cost routes to the same destination, the router installs all of them and spreads traffic across every link. In hyperscale fabrics ECMP isn't an optimization; it is the forwarding model. And understanding its failure modes is the difference between a fabric that hums and one that mysteriously congests at 40% utilization.

What ECMP actually does

It starts in the control plane. OSPF, IS-IS, or BGP computes routes; when several routes to a prefix tie — same OSPF cost, or BGP multipath policy — the RIB keeps all of them instead of one winner. The forwarding plane programs an ECMP group (vendors call it a next-hop group) into the FIB: one destination, N next-hops, N egress interfaces.

Each arriving packet is assigned to one member of the group. The assignment happens in hardware, per packet, at line rate — and it must be consistent: every packet of a flow has to take the same path, or TCP sees reordering, assumes loss, and collapses its congestion window. So the switch doesn't round-robin packets. It hashes flows.

The hash: five fields in, one link out

The classic ECMP hash input is the 5-tuple: source IP, destination IP, protocol, source port, destination port. Platforms vary — some mix in the ingress port or VLAN ID. The hash output picks the egress member, conceptually member = hash(fields) mod N.

Two properties carry the whole design. Determinism: the same 5-tuple always selects the same member, so flows never reorder. Distribution: across many flows, each member should receive roughly 1/N of them.

Burn this distinction in: ECMP balances flows, not bytes. Ten thousand small flows spread beautifully. Three elephant flows can land on the same member and saturate it while sibling links idle. "Even hash distribution" and "even link utilization" are not the same thing — and confusing them is the root of half of all ECMP troubleshooting sessions.

Hash polarization: the silent killer

Now the failure mode that burns experienced engineers. The hash is deterministic — and on many platforms it is the same deterministic function, with the same seed, on every switch. Follow one flow across three hops: the leaf hashes its 5-tuple and picks member 2; the spine hashes the identical 5-tuple with the identical function and picks its member 2; the next leaf does it again. Not just this flow — every flow with a similar 5-tuple rides the same relative link at every tier. Traffic polarizes onto a handful of links instead of spreading across the fabric. You get hot spots, drops, and retransmissions while most of the fabric sits idle — and nothing in any routing table looks wrong.

Polarization is vicious because every device is behaving correctly in isolation. The bug is emergent: identical hash behavior everywhere. It presents as persistent, unexplained congestion on a few links in an otherwise healthy fabric, and it survives reboots, because determinism is the point.

The rulePolarization is a property of the fabric, not of any single device. If every switch hashes identically, your 32-way ECMP is effectively 1-way for correlated flows.

Breaking the polarization

Every fix introduces per-device variation into the hash:

See it with your own eyes: our ECMP Hash Polarization Visualizer simulates the hash hop by hop across a fabric. Toggle identical versus seeded hash functions and watch flows collapse onto the same links — or spread the way you intended.

ECMP in the wild: protocols and overlays

In the underlay, OSPF and IS-IS give you ECMP almost for free: equal-cost paths are installed by default up to the platform's max-paths limit — check it, because 8, 16, 32, and 64 are all common ceilings, and silently exceeding one is a classic. BGP needs explicit multipath configuration, and at hyperscale the interesting knobs are multipath relax and Add-Paths for backup diversity.

Overlays add a wrinkle that bites everyone once: VXLAN. The underlay ECMP hash sees only the outer headers — outer source and destination IP, outer UDP ports. If the encapsulation stamps the same outer source port on every inner flow, the underlay sees one giant flow and pins it to a single link. The fix is UDP source-port entropy (RFC 7348): the VTEP hashes the inner 5-tuple into the outer source port, restoring per-flow distribution to the underlay. If your VXLAN fabric shows one hot spine link, check entropy before you check anything else.

Troubleshooting lopsided links

When link utilization is uneven and the routing table is innocent, work this list:

Where this fits

ECMP is the forwarding half of the hyperscaler story; BGP is the control-plane half. Read them together: BGP: How the Internet Actually Finds a Path covers how those equal-cost routes get computed in the first place, and the Hyperscaler Network Engineer track puts both in career context. For capacity planning around what ECMP actually delivers, the Data-Center Capacity Planner is the companion tool.

Further reading
Key takeaways

Keep reading