Module Capstone — Draw and Defend Your First Mini Data Center
Hyperscaler Network Engineer · Module 1: Data Centers from Zero
Lesson 8 of 8
Prerequisites: What Is a Data Center, Really?, IP Addresses and Subnets for Absolute Beginners, Your First Lab — Make Two Servers Talk
What you'll be able to do: Sketch a working 2-rack mini data center on paper, justify every choice in plain words, and name what still breaks it.
You've toured the kitchen, met the cooks, and learned every recipe — now it's opening night and the floor plan is yours to draw. Two dining rooms, twelve cooks, one promise to keep: dinner service cannot stop, even when something breaks.
Where does every cook stand? Which doors connect the rooms? And if the one shared walk-in fridge dies at 7 p.m., who still gets fed? Designing a tiny data center is exactly this exercise — a handful of racks, a web of cables, and the honest question of what survives the first failure.
Here's the puzzle.
Scenario. A Tulsa startup hands you $15,000 and a small empty room: build them a mini data center — 12 servers across 2 racks — that stays up when parts fail. The founder isn't technical; you'll defend every choice to her in plain words. The room has limits, the budget is fixed, and perfection isn't for sale. Your design must be honest about what still breaks.
Given artifacts. The design brief (your constraints) and the price list:
THE BRIEF
Budget: $15,000 total — every dollar accounted for below
Servers: 12 (each with two power supplies)
Racks: 2 (6 servers each — the room fits exactly two)
Power: ONE wall power feed per rack (no second feed available)
Cooling: one room air conditioner for both racks
Internet: the building gives you ONE internet cable (one handoff)
Stakeholder: the founder — explain it like she's your cousin who runs a bakery
PRICE LIST (must total ≤ $15,000)
Server: $800 each × 12 = $9,600
Switch: $1,500 each × 2 = $3,000
Rack cabinet: $500 each × 2 = $1,000
Cables + PDUs: $1,400 allowance (plenty — cables are cheap)
TOTAL: $15,000 exactly — there is no money left for extra boxes
Your task: (1) State your rack layout — how many servers per rack and where each switch sits. (2) State your cable plan — which servers plug into which switch, and how the racks reach the internet. (3) Name the TWO single points of failure remaining in your design, and for each, either give the $0 fix (reusing only what the budget already bought) or honestly admit what fixing it would cost.
Workspace: Analyze-and-answer. Sketch the layout on paper first (that's the real work), then type your three answers: layout, cable plan, and the two remaining single points of failure with a $0 fix or an honest cost admission for each.
Hint 1 — where to look
List every component in your design that appears exactly once — one internet cable, one switch per rack, one power feed per rack. Anything with no backup is a candidate. The question is which two hurt the most.Hint 2 — what to compare
For each single item, ask: "if this dies at 7 p.m. on a Friday, what stops working?" Compare the blast radius — how many of the 12 servers go dark? The two with the biggest blast radius are your answers.Hint 3 — the mechanism
A $0 fix never buys anything — it rearranges what you own. Cables are already in the budget and switches have spare ports: is there a way to wire things so one dead box only wounds the design instead of killing it? And when no rearrangement helps, the honest answer is "this costs money" — say so.Commitment ritual: ☐ "I've attempted this challenge and thought it through." Check the box (your design above is recorded either way) and the reference worked answer in S7 reveals. Nothing is graded — the struggle is the point.
Checking the box reveals the worked answer in S7 below. Returning learners stay unlocked.
The spare tire principle
Your car carries a spare tire you've hopefully never used. That tire is redundancy (having a backup ready so that one failure doesn't stop everything — the spare tire, the second engine on a plane, the backup goalie). Redundancy feels wasteful right up until the moment it saves you: you pay for a tire that mostly sits in the trunk, and the one night you need it, it's the cheapest money you ever spent. Every data center is designed the same way — nothing important exists exactly once. Hyperscalers take this further than anyone: they assume every single component will fail eventually, and they design as if the failure already happened.
Why this matters for the challenge: your $15,000 buys limited redundancy. The design question is never "can I afford backups for everything?" — it's "which single failures am I willing to survive, and which am I admitting?"
One of everything is one failure away
The evil twin of redundancy has a name: a single point of failure (any component whose death alone stops the whole system — one fridge for the whole restaurant, one bridge into town). Engineers shorten it to SPOF ("spoff") and hunt them the way health inspectors hunt rats: list everything that appears once, then ask what dies with it. One internet cable? The whole room goes dark if it's cut. One switch per rack? A dead switch silences its rack. Finding SPOFs isn't pessimism — it's the job. The founder is paying you to be the person who asks "and then what breaks?" before opening night, not after.
Why this matters for the challenge: part three of your deliverable is a SPOF hunt on your own design. The two you name are the two you'd brief the founder on first.
The $0 fix mindset
Here's the hyperscaler twist: at small scale you can't buy your way out of every SPOF, so you rearrange your way out. A failover (the backup automatically taking over when the primary dies — the spare tire rolling on without a pit stop) doesn't always need new hardware. Sometimes it needs a cable you already own plugged into a different port: two switches connected to each other, servers split across both, and suddenly one dead switch wounds you instead of killing you. The money was already spent — the resilience was hiding in the wiring diagram. When no rearrangement helps, the professional move is the honest admission: "this one costs money, here's the price, here's the risk of skipping it."
Why this matters for the challenge: the deliverable demands a $0 fix or an honest cost admission for each SPOF. Both are correct engineering — bluffing is the only wrong answer.
Drawing it so your cousin gets it
The founder runs a bakery, not a network. She doesn't need to know what a subnet mask is — she needs to know that if the internet cable is cut on Friday night, the shop is closed until Monday, and that a second cable costs real money every month. Explaining a design to a non-technical stakeholder is a core engineering skill, not a soft extra: every budget approval, every outage postmortem, every "why did we spend this?" conversation runs on plain words. Rule of thumb: one everyday comparison per choice ("this switch is the traffic cop for its rack"), then the consequence in her language — money, downtime, customers. If she can repeat your explanation back to you accurately, you did it right; if she can't, the design isn't the problem — the telling is.
Why this matters for the challenge: your written answers are the defense. If your cousin the baker couldn't follow your cable plan, rewrite it until she can.
- Step 1 of 5: Two empty racks, two switches at the top, twelve servers sliding in — six per rack. This is the build-up your design brief describes.
- Step 2 of 5: Cables draw themselves: every server to its rack's switch (gray), the two switches to each other (amber), and one green uplink carrying both racks to the internet.
- Step 3 of 5: Friday, 7 p.m. Switch 1 dies. Watch it turn red — and its entire rack dim with it. Six servers, gone in one failure.
- Step 4 of 5: That is a single point of failure made visible: one box, one death, half the data center dark. Rack 2 doesn't even notice — failures are local, which is exactly the problem.
- Step 5 of 5: The design recovers and the loop restarts. Your challenge: rearrange this picture — with $0 of new spending — so the next red box wounds instead of kills.
🔒 Revealed after the commitment ritual in S2 — attempt the challenge first. (Honor system: the page hides this until you check the box.)
The reference design — two racks, six servers each, one switch atop each, wired so no single death is total:
Layout. Rack 1: servers 1–6; Rack 2: servers 7–12; one switch atop each. (Wrong turn: both switches in one rack "to save cable" — one rack power event then kills both.)
Cable plan. Inter-switch link: Switch 1 ↔ Switch 2. Each rack's servers split 3-and-3 across the switches — 1–3 → SW1, 4–6 → SW2, 7–9 → SW2, 10–12 → SW1 (aisle-crossing cables are already in the budget). One uplink from Switch 1 to the building's internet handoff; addressing per Lesson 1.6 (one /24, two /25s).
SPOF 1 — one switch per rack (FIXED for $0). Switch 1 dies → servers 1–3 and 10–12 go dark, but 4–9 stay up on Switch 2. The $0 fix is the cross-aisle cabling: no new boxes, just budget-already-bought cables — one dead switch now halves each rack instead of killing one.
SPOF 2 — the single internet uplink (ADMITTED). One cable, one building handoff: cut it and all 12 servers go dark — no rearrangement of your own gear changes that. Tell the founder: "A second internet connection costs real money every month, and it's the one thing that takes the whole shop offline. My $0 move: a mobile hotspot as emergency uplink plus a 10-minute switchover runbook — degraded, not dead."
Verify it worked: 7 p.m. test — kill each box, count survivors. Expected: any single switch death leaves 6 of 12 up; only the uplink or a power feed takes everything. Any other single death taking all 12 means a hidden SPOF — redraw.
Check yourself — nothing here is graded. Wrong answers are the useful ones; each explains why.
Question 1. Your mini data center is live. At 2 a.m. the monitoring shows Rack 1's servers unreachable but Rack 2 fine, and the uplink is up. What's your first check?
Question 2. The founder asks: 'We have $2,000 left over — should we buy a third switch or a second internet connection?' What's the better advice?
Question 3. Why does the reference design cable half of each rack's servers to the *other* rack's switch?
Question 4. A well-meaning intern 'tidies' the cabling: all of Rack 1's servers now plug into Switch 1, all of Rack 2's into Switch 2. Nothing is unplugged that was working. What's wrong?
Question 5. The founder (your baker cousin) asks what happens 'if the internet cable gets cut.' What's the honest, plain-words answer?
- Redundancy is a spare tire: it looks wasteful until the night it saves you — design so nothing important exists exactly once.
- A single point of failure is found by listing everything that appears once and asking what dies with it; the 7 p.m. test makes the blast radius concrete.
- The $0 fix rearranges instead of rebuying — cross-wiring you already own can turn a total outage into a half outage.
- When no rearrangement helps, the professional answer is an honest cost admission plus a $0 mitigation — bluffing is the only wrong answer.
- Designs are defended in plain words to non-technical stakeholders: one everyday comparison per choice, then the consequence in money, downtime, and customers.
Next: Why the Old Network Designs Broke — you've just built the classic small design and mapped exactly where it breaks; Module 2 opens by showing why those same breakage patterns forced the entire industry to reinvent data-center networking. An honest note: this capstone closes Module 1, and Module 1 is the free taste — Module 2 is where the subscription begins, because that's where the designs get big enough to need it.