BGP: How the Internet Actually Finds a Path
Every autonomous network on the planet speaks BGP to its neighbors. Here's how path-vector routing really works — and why policy beats distance.
The problem BGP solves
Inside one company, routers can run OSPF or EIGRP and agree on one map of the network. The internet has no single map. It is roughly 75,000 independent networks — autonomous systems (ASes) — run by ISPs, cloud providers, and large enterprises, and none of them trusts the others' routing math.
BGP, the Border Gateway Protocol, is the diplomacy layer. Each AS publishes reachability — "I can carry traffic to these prefixes" — and every other AS decides, by its own policy, which path to believe. It is the only protocol whose answer to "what's the best route?" is legitimately "it depends on who you ask."
Two jobs: eBGP and iBGP
BGP looks like one protocol but plays two different roles.
| eBGP | iBGP | |
|---|---|---|
| Runs between | Different autonomous systems | Routers inside one AS |
| Transport | TCP port 179, single-hop by default | TCP port 179, often on loopbacks |
| Loop safety | AS_PATH check — drops routes already containing its own AS | Split-horizon — iBGP-learned routes are not re-advertised to other iBGP peers |
| Typical use | Peering with ISPs and other networks | Carrying internet routes across a provider backbone |
Path vector, not shortest path
OSPF computes math; BGP tells stories. Each BGP route carries a path vector — the list of autonomous systems it has crossed — plus attributes you use to express policy. The big four:
- LOCAL_PREF — local to your AS. Highest wins. This is how you say "prefer ISP A over ISP B for everything."
- AS_PATH — transit history. Shortest usually wins — and long paths also look like you are prepending (a valid way to make a path less attractive).
- MED (Multi-Exit Discriminator) — used between neighboring ASes with multiple links. Lowest wins, and only compared between paths from the same AS.
- Communities — attached labels (like
no-export, or custom tags) that let peers agree on meaning: "blackhole this," "local-only," "backup path."
Best-path selection, simplified
When BGP learns several ways to reach one prefix, it runs a best-path election. The full list is vendor-long, but the first hits decide most elections:
- Highest weight (Cisco-local tiebreak; vendors without weight start here)
- Highest LOCAL_PREF
- Shortest AS_PATH (and locally originated routes first)
- Lowest origin type: IGP < EGP < incomplete
- Lowest MED, when the candidate paths came from the same neighboring AS
- Prefer eBGP-learned over iBGP-learned
- Lowest IGP metric to the next hop, then oldest route, then lowest router ID
The practical upshot: if traffic takes the "wrong" way, the cause is almost always a LOCAL_PREF or MED knob someone turned on purpose. And when two paths genuinely tie, routers can install both and load-share across them — equal-cost multipath. That last step is where forwarding quietly goes wrong: the per-flow hash spreading traffic over those links can polarize a Clos fabric onto a handful of paths, which the ECMP hash polarization visualizer makes painfully visible.
Loop prevention is built in
Because every route lists the ASes it crossed, a router that sees its own AS number in AS_PATH drops the route — the loop is impossible by construction. Elegant, and it scales: no one would maintain a loop-free map of 75,000 networks by hand.
iBGP has no such check, so it uses a blunt rule instead: split-horizon — routes learned from one iBGP peer are never advertised to another iBGP peer. That means a full mesh (or route reflectors) between iBGP speakers. The classic day-one surprise: routers A, B, C in one AS; A peers with B, B peers with C — and routes from A never reach C. That is not a bug; add a direct A–C session or a route reflector.
BGP in four commands
IOS-style, at the level you need for your first lab. AS 65001 peering with AS 65002 over a direct link:
router bgp 65001
bgp router-id 10.0.0.1
neighbor 203.0.113.2 remote-as 65002
neighbor 203.0.113.2 description ISP-A-primary
exit
network 198.51.100.0 mask 255.255.255.0
Then verify before you believe anything:
show bgp summary ! session state should read Established
show ip bgp neighbor 203.0.113.2 advertised-routes
show ip bgp 198.51.100.0 ! which path won, and why
That verify-it-don't-trust-it habit scales: paste several devices' configs into the network config verifier and it checks the fleet in seconds — BGP neighbor AS mismatches, duplicate IPs, OSPF disagreements, unreachable subnets, and dead ACL rules — the exact class of typo that turns a routine change into a 2 a.m. incident.
Idle → Connect → Active → OpenSent → OpenConfirm → Established. Anything stuck in Active usually means TCP/179 is blocked or the neighbor's AS number is wrong. Defaults: keepalive 60 s, hold time 180 s.Common gotchas
- Different loopbacks, no multihop: eBGP between loopbacks needs
ebgp-multihop, because default TTL is 1. - next-hop unreachable: iBGP passes the eBGP next hop through unchanged — use
neighbor x.x.x.x next-hop-selfor make sure the IGP carries the route. - The network command is exact:
network 10.0.0.0 mask 255.255.0.0only advertises what is already in your table with that exact mask. - Using your ISP's MED against you: MED is compared between links to the same neighbor AS. Two ISPs' MEDs are not comparable — a classic accidental misroute.
Your next steps
Spin up a three-AS topology in GNS3 or EVE-NG: two ISPs, your AS in the middle. Get both eBGP sessions Established, then practice the two moves every network engineer does in production: steer outbound traffic with LOCAL_PREF, and steer inbound traffic with AS_PATH prepending. When what you predict matches show ip bgp, you understand BGP.
- BGP is a path-vector protocol: routers choose routes by policy, not by shortest hop.
- eBGP runs between autonomous systems over TCP/179; iBGP runs inside one.
- Best-path selection (simplified): Local Preference, then AS-path length, then MED.
- The AS_PATH attribute prevents loops automatically; iBGP uses split-horizon instead.