Packet Path/

Why AI Clusters Run Two Networks: Front-End vs Back-End Fabrics Explained

AI clusters don't run one network — they run two. The front-end fabric handles management, storage, and inference traffic; the back-end fabric exists for one job: GPU-to-GPU collective communication at microsecond latency. Here's why operators split them.

Data Center · Intermediate · 10 min · October 6, 2026

Split illustration: a calm blue enterprise data center network on the left and a blazing amber GPU cluster fabric on the right.

The city with two transportation systems

Imagine a city that needs to move two very different things: commuters, mail, and food deliveries — and, separately, a million tons of steel between factories. You would never run the freight trains down Main Street at rush hour. You would build a dedicated rail line for the steel and leave the streets for everything else.

An AI cluster does exactly the same thing with packets. It runs two networks: a front-end fabric that acts as the city streets, and a back-end fabric that acts as the freight rail. If you have ever wondered why AI data center diagrams always show two sets of switches, this is why.

In plain EnglishAn AI cluster runs two networks. The front-end fabric carries management traffic, storage, and user requests — ordinary Ethernet/IP, the kind of network you have built before. The back-end fabric exists for exactly one job: letting GPUs talk to each other during training, at microsecond latency, without dropping a single packet. They stay separate because the two traffic types have nothing in common — different sizes, different urgency, different failure modes — and mixing them makes both worse.

This month, the split became official

The two-fabric design stopped being insider knowledge recently. Cisco's Secure AI Factory with NVIDIA — expanded to a rack-scale architecture in August 2026 — is built around it: Cisco's own Silicon One switches on the front-end fabric, NVIDIA Spectrum-X-based switches (the Cisco N9100: 64 ports of 800G, 51.2 terabits per second of switching) on the back-end, all managed as one system under Cisco Nexus One. The Supermicro rack-scale systems in that architecture became orderable this month, October 2026, targeting enterprises, neocloud providers, and sovereign cloud operators.

And the compute side explains why the network side has to be special: a Vera Rubin NVL72 rack draws 190–230 kW and requires 100% liquid cooling — Supermicro ships four 110 kW power shelves per rack. When one rack pulls as much power as a city block, its network is not an afterthought. (If you want to see this gear in person, the OCP Global Summit runs October 12–15 in San Jose.)

What the front-end fabric carries

The front end is the city streets. It carries:

It is ordinary Ethernet/IP — the kind of network you have already built. It can run a vendor NOS or an open one like SONiC (see our SONiC explainer). Reliability matters here, but a dropped packet is just a retransmit. Nothing catches fire.

FRONT-END FABRIC management · storage · inference BACK-END FABRIC GPU-to-GPU collectives only users storage mgmt/API switch 1 switch 2 switch 3 WAN /users GPU 0 GPU 1 GPU 2 GPU 3 RoCE · lossless · microsecond latency
The two fabrics side by side: the front end serves people and storage; the back end serves only the GPUs.

What the back-end fabric carries

The back end is the freight rail, and it hauls exactly one cargo: collective communication. During training, every GPU computes gradients (its share of what the model just learned) and must exchange them with every other GPU — operations called all-reduce and all-gather that move gigabytes in tight synchronization, over and over, thousands of times per training run.

Three things make this fabric different from any LAN you have managed:

Ring all-reduce: each GPU sends a chunk, adds its neighbor's, and forwards it. GPU 0 GPU 1 GPU 2 GPU 3 add · forward · repeat
Ring all-reduce, the workhorse collective. The ring advances at the speed of the slowest link — one congested link stalls all four GPUs. That is the straggler effect.

Why not one big network?

Five reasons operators keep the fabrics apart:

  1. The traffic patterns are opposites. Front-end traffic is millions of small, bursty, latency-tolerant flows. Back-end traffic is a small number of enormous, synchronized, latency-intolerant flows. No single tuning serves both.
  2. Blast-radius isolation. A broadcast storm or a misconfigured ACL on the management network must never be able to stall a training run worth thousands of GPU-hours.
  3. Independent scaling. Adding inference capacity? Grow the front end. Adding GPUs? Grow the back end. Neither forces the other to change.
  4. Congestion control conflicts. The back end needs lossless behavior via PFC; running PFC across a mixed network invites head-of-line blocking and PFC storms that cascade across unrelated traffic.
  5. Separate failure domains. You can reboot the front-end fabric's control plane without touching a running training job.
One slow link stretches the whole training step GPU 0 GPU 1 GPU 2 GPU 3 sync point — everyone waits here congested link
Synchronous training waits for the slowest GPU. One congested back-end link stretches the entire step — this is why the back-end fabric is engineered for tail latency, not average latency.

The power forcing function

The two-network design is not only about packets. NVL72-class racks at 190–230 kW each mean the data center itself is being redesigned around liquid cooling and new power distribution — a new 800V DC power architecture is replacing the old 48V distribution in these facilities. The network rides the same rack-scale logic: short, dense, purpose-built fabrics instead of one general-purpose cloud network. The building, the power, the cooling, and the network are all being co-designed now.

What this means for your skills

If you work in data center networking, the back-end fabric is where the demand is moving:

For the career view, see the hyperscaler network engineer career path.

Three mistakes to avoid

  1. Running collectives over the front-end fabric "temporarily." The temporary becomes permanent, and one noisy neighbor stretches every training step.
  2. Treating the back end like a LAN — same SNMP polling, same thresholds. It needs microburst-level telemetry, or you are blind to the exact events that stall training.
  3. One congestion policy everywhere. PFC enabled on mixed traffic invites head-of-line blocking; keep the lossless domain tight and the front end lossy.

Which fabric does it belong on?

Related guides
Key takeaways

Keep reading