V
vijayananda
Guest
The Default Isn't Always the Right Answer
Walk into almost any modern data center design review, and you'll hear the same recommendation: VXLAN-EVPN. It's become the de facto standard overlay for good reasons — it's vendor-neutral, runs over plain IP fabrics, integrates cleanly with BGP for control-plane scale, and every major switch vendor supports it.
But "default" and "best fit" aren't the same thing. In telco cloud, NFV infrastructure, and certain high-density compute fabrics, MPLS-native designs still outperform IP-in-IP-style overlays (VXLAN included) on three concrete axes: encapsulation overhead, traffic engineering, and fast reroute. This article breaks down why, with the mechanics underneath each claim.
Quick Primer: What We're Actually Comparing
- IP-in-IP / VXLAN-style overlays: encapsulate the original frame inside a UDP/IP header (VXLAN) or a raw IP header (plain IP-in-IP). Forwarding decisions at each hop are made by looking up the outer IP destination in a standard IP routing table.
- MPLS: prepends a short, fixed-length label (or label stack) to the packet. Forwarding decisions are made by a simple label lookup — swap the label, forward out the mapped interface — without re-examining the IP header at every hop.
The distinction matters because it changes what each hop has to do per packet, and what the network can express about how traffic should flow.
1. Encapsulation Overhead
| Encapsulation | Outer header overhead | Notes |
|---|---|---|
| VXLAN (IP-in-UDP) | 50 bytes (14 Eth + 20 IP + 8 UDP + 8 VXLAN) | Plus outer Ethernet FCS considerations; commonly pushes past 1500 MTU |
| Plain IP-in-IP | 20 bytes (IP-in-IP) | Lighter, but lacks entropy for ECMP hashing |
| MPLS (2-label stack) | 8 bytes (2 × 4-byte labels) | Typically riding directly on Ethernet, no extra IP/UDP header |
VXLAN's 50-byte tax isn't just a bandwidth curiosity — it forces MTU planning across the entire fabric (commonly bumping to 9216B jumbo frames to absorb the overhead without fragmenting tenant traffic), and it eats into the effective payload on every single packet at scale. In a network built for maximum flow density (large NFV compute clusters, or line-rate telco data planes), that persistent per-packet tax compounds.
MPLS's label stack is a fraction of the size and rides directly on the existing Ethernet frame — no extra IP or UDP header required. For workloads sensitive to per-packet overhead (small-packet VoIP/RTP streams, telemetry-heavy 5G user-plane traffic), this isn't marginal.
Why this matters specifically for NFV/telco: Virtual Network Functions (vFirewalls, vEPC/UPF, vRouters) frequently process high volumes of small packets. Overhead percentage scales inversely with packet size — a 50-byte tax on a 1500-byte packet is ~3.3%; on a 128-byte VoIP packet, it's closer to 40%. MPLS's 8-byte tax barely registers by comparison.
2. Traffic Engineering: Explicit Paths vs. Hash-and-Hope
This is where the architectural gap is widest.
IP-in-IP/VXLAN overlays inherit whatever the underlay's IP routing gives them. Path selection is left to ECMP hashing over shortest-path IGP routes (OSPF/IS-IS/BGP underlay). You can influence this at the margins — weighted ECMP, some SDN-driven path steering — but you cannot natively say "this specific flow takes this specific physical path, and this other flow takes a completely different one, regardless of IGP cost."
MPLS gives you that natively, via:
- RSVP-TE: signals explicit, resource-reserved LSPs (Label Switched Paths) with bandwidth guarantees, computed via CSPF against real-time link utilization — not just topological shortest path.
- Segment Routing (SR-MPLS): encodes an explicit path as a stack of segment labels, letting a headend steer a flow through specific nodes/links without per-hop signaling state, while still keeping the lightweight MPLS forwarding plane.
bash
Code:
# Example: SR-MPLS explicit path via segment list (conceptual)
# Steer flow through Node-B, then Link-C, bypassing congested Link-A
segment-list PATH_VIA_B_THEN_C
index 10 mpls label 16002 # Node-SID for Node-B
index 20 mpls label 24011 # Adj-SID for Link-C
For an NFV service chain — traffic that must traverse vFirewall → vDPI → vLoad-Balancer in a specific order, on specific paths with guaranteed capacity — this isn't a nice-to-have, it's the entire point of the design. VXLAN-EVPN has no native primitive for "this flow must transit these exact nodes in this exact order with these bandwidth guarantees." You'd bolt on an SDN controller and NSH-style service function chaining to approximate it — which is exactly the layer MPLS gives you for free.
3. Fast Reroute: Sub-50ms Convergence, Natively
MPLS fast reroute (FRR), whether via RSVP-TE's one-to-one/facility backup or SR-TI-LFA (Topology-Independent Loop-Free Alternate), pre-computes backup paths before a failure happens. When a link or node fails, the local node immediately reroutes onto the pre-signaled backup — typically in under 50ms, without waiting for the control plane to reconverge globally.
IP-in-IP/VXLAN overlays are at the mercy of underlay IGP convergence. When a link fails:
- The IGP (OSPF/IS-IS) detects the failure.
- It floods the topology change.
- Every node recomputes SPF.
- New forwarding entries get installed.
Even well-tuned IGPs (BFD-assisted, sub-second timers) typically land in the hundreds of milliseconds to low seconds range for full reconvergence — an order of magnitude slower than MPLS FRR's local, pre-computed failover.
For telco data planes carrying voice, video, or 5G user-plane traffic with strict SLA commitments (some URLLC use cases target single-digit-millisecond disruption budgets), that gap between "under 50ms" and "hundreds of milliseconds" is the difference between an imperceptible blip and a dropped call.
So When Should You Actually Reach for MPLS?
MPLS-native fabrics tend to win when most of the following are true:
- You're building telco cloud / NFV infrastructure with service function chaining requirements.
- Workloads are latency- and jitter-sensitive at a level where sub-50ms failover materially matters (voice, video, URLLC/5G UPF).
- You need explicit, bandwidth-guaranteed traffic engineering — not just best-effort ECMP.
- Packet sizes skew small, making per-packet encapsulation overhead a real cost, not a rounding error.
- Your team already operates MPLS/SR expertise (telco/SP networks commonly do) — the operational learning curve is lower than for a green-field enterprise DC team.
VXLAN-EVPN remains the right default when:
- You're running general-purpose enterprise or cloud compute — East-West microservice traffic without strict per-flow TE requirements.
- Multi-tenant L2/L3 segmentation over a plain IP Clos fabric is the primary requirement.
- You want vendor-neutral, merchant-silicon-friendly hardware support (VXLAN is essentially universal on modern DC switch ASICs).
- Your team is IP/BGP-native and doesn't want to carry MPLS control-plane operational overhead (RSVP-TE state, label distribution protocols).
A Hybrid Middle Ground: SR-MPLS with an EVPN Control Plane
It's worth noting these aren't strictly either/or. EVPN as a control plane can run over an MPLS data plane instead of VXLAN — you get EVPN's BGP-based MAC/IP learning and multi-tenancy model, but forwarding uses MPLS labels instead of VXLAN encapsulation. This is common in SP/telco DC interconnect designs that want EVPN's operational model with MPLS's overhead and TE characteristics.
text
Code:
EVPN (control plane: BGP MAC/IP advertisement, multi-tenancy)
│
├──► Data plane option A: VXLAN encap (typical enterprise DC)
└──► Data plane option B: MPLS/SR encap (typical telco/SP DC)
This decouples the "which overlay control plane" decision from the "which encapsulation" decision — worth evaluating before assuming VXLAN is a package deal with EVPN.
Wrapping Up
VXLAN-EVPN earned its default status honestly — it's simpler to operate, hardware-ubiquitous, and sufficient for the vast majority of enterprise and cloud data center traffic patterns. But "default" shouldn't mean "unconsidered." Where the workload genuinely demands explicit traffic engineering, sub-50ms failover, or minimal per-packet overhead — telco cloud and NFV being the clearest examples — MPLS-native fabrics (particularly SR-MPLS, which keeps the operational simplicity advantage while adding TE and FRR) deliver capabilities that IP-in-IP-style overlays simply don't have a native answer for.
The right question isn't "MPLS or VXLAN?" as a blanket policy — it's "what does this specific fabric's traffic actually require?"