Lesson 05

Load Balancing

L4 vs L7 balancers, consistent-hashing routing, health checks, hot spots.

Plays in the sticky player at the bottom of the page

Transcript

Lesson 5: Load Balancing

A reference-style deep dive into L4 vs L7, consistent hashing routing, health checks, hot spots. Read this as an article, not a transcript. The accompanying audio is the spoken companion; the article below is the canonical written reference.

Audience. Engineers designing, building, or operating distributed systems who want a clear mental model rather than a checklist of tools.

Prerequisites. Working knowledge of HTTP, basic SQL, and the idea of running more than one server behind a load balancer.

Table of contents

  1. Why load balancing matters
  2. The core mental model and where to start 3-12. In-depth sections below (see "Lesson body")

Lesson diagram

Lesson 5 diagram — Load Balancing

Figure 1. The canonical load balancing topology and control flow covered in this lesson.


1. Why load balancing matters

A load balancer is the front door of any horizontally-scaled service. It spreads requests across replicas, hides failures, and provides a single stable endpoint. The choice between L4 and L7 is one of the most consequential architectural decisions in a system. L4 is fast and dumb; L7 is slow and smart.

Three properties make load balancing essential:

  • It hides the backend topology. Clients see one endpoint; the LB routes to N backends.
  • It absorbs failure. When a backend dies, the LB stops sending it traffic.
  • It enables scaling. Add backends; the LB picks them up automatically.

2. L4 vs L7 — the fundamental choice

L4 (Transport layer)

Operates on TCP/UDP packets. Does not inspect the payload.

def route_l4(packet):
    # No HTTP parsing; just look at src IP, dst port, protocol
    backend = pick_backend(src_ip=packet.src_ip, dst_port=packet.dst_port)
    forward(packet, backend)

Properties:

  • Speed. No payload inspection = decisions in microseconds.
  • Simplicity. Stateless; no need to buffer requests.
  • Limitation. Cannot route based on URL path, headers, cookies.

Used by: HAProxy (TCP mode), AWS NLB, IPVS.

L7 (Application layer)

Operates on HTTP/gRPC requests. Inspects headers, paths, cookies, body.

def route_l7(req):
    backend = pick_backend(
        host=req.headers['Host'],
        path=req.path,
        cookies=req.cookies,
        method=req.method,
    )
    return forward(req, backend)

Properties:

  • Smarter routing. Path-based routing, weighted routing, sticky sessions.
  • Higher latency. Must buffer the request headers (and sometimes the body).
  • TLS termination at the LB simplifies backend config.

Used by: HAProxy (HTTP mode), AWS ALB, Envoy, NGINX.

When to use which

Use caseLayerWhy
Pure TCP serviceL4No need for HTTP awareness
HTTP API with path routingL7Need to inspect paths
gRPC streamingL7Multiplexing requires HTTP/2 awareness
High-throughput internal serviceL4Latency matters
External-facing APIL7TLS termination + smart routing

3. Load balancing algorithms

The algorithm determines how the LB picks a backend for each request.

Round robin

class RoundRobin:
    def __init__(self, backends):
        self.backends = backends
        self.idx = 0
    def pick(self):
        backend = self.backends[self.idx % len(self.backends)]
        self.idx += 1
        return backend

Simple, fair, ignores backend capacity. Good default for similar-capacity backends.

Least connections

class LeastConnections:
    def __init__(self, backends):
        self.backends = backends
    def pick(self):
        return min(self.backends, key=lambda b: b.active_connections)

Routes to the backend with the fewest active connections. Better for long-lived connections (WebSocket, gRPC streaming).

Weighted round robin

class WeightedRoundRobin:
    def __init__(self, backends, weights):
        self.backends = backends
        self.weights = weights
    def pick(self):
        # Smooth weighted round-robin (Nginx algorithm)
        ...

Higher-capacity backends get more traffic. Use for heterogeneous fleets (some boxes are bigger).

Consistent hashing

class ConsistentHash:
    def __init__(self, backends, vnodes=200):
        self.ring = {}
        for b in backends:
            for v in range(vnodes):
                self.ring[hash(f"{b}-{v}")] = b
    def pick(self, key):
        h = hash(key)
        for node in sorted(self.ring.keys()):
            if h <= node:
                return self.ring[node]
        return self.ring[min(self.ring.keys())]

Routes the same key to the same backend. Maximizes cache hit ratio. Used for stateful services and session affinity.

Power of two choices

class PowerOfTwo:
    def __init__(self, backends):
        self.backends = backends
    def pick(self):
        candidates = random.sample(self.backends, 2)
        return min(candidates, key=lambda b: b.load)

Pick 2 random backends; route to the less loaded one. Near-optimal without coordination. The recommended default for uniform workloads.

IP hash

def pick(req):
    backend = ring[hash(req.src_ip) % len(backends)]
    return backend

Routes the same client to the same backend. Used when session affinity is required and you cannot modify the backend.

4. Health checks

A load balancer is only as good as its health checks. The LB must actively detect failed backends and remove them from rotation.

Active health checks

The LB periodically sends a synthetic request to each backend.

async def health_check(backend):
    try:
        resp = await http.get(f"http://{backend}/health", timeout=2)
        return resp.status == 200
    except:
        return False

The /health endpoint should:

  • Be cheap (no DB queries, no external calls).
  • Test the real dependency path (DB connection, downstream service).
  • Return 200 only when the backend can serve real traffic.

Passive health checks

The LB observes real traffic for failures.

def on_response(backend, resp):
    if resp.status >= 500:
        backend.failures += 1
        if backend.failures > threshold:
            remove_from_rotation(backend)
    else:
        backend.failures = 0

Cheaper than active checks; reacts faster to failures. Risk: a burst of legitimate 500s causes the LB to remove a healthy backend.

Combined

Production LBs use both: passive for fast reaction, active for periodic verification.

5. TLS termination

The LB can terminate TLS on behalf of the backends. Pros and cons:

  • Pro: Backends do not manage certificates; one cert on the LB.
  • Pro: LB can decrypt and re-encrypt, enabling L7 routing by path/header.
  • Con: Internal traffic between LB and backend is unencrypted (unless you re-encrypt).
  • Con: LB is a SPOF for the TLS handshake (mitigated by active-active LBs).

The standard pattern: terminate TLS at the LB, re-encrypt to the backend (mTLS) for internal traffic. Best of both worlds.

6. Connection draining

During deployments, backends need time to finish in-flight requests before shutting down.

# LB behavior on SIGTERM:
1. Stop accepting new requests from LB
2. Wait up to N seconds for in-flight requests to complete
3. After N seconds, force-close remaining connections

Without draining, users see errors during deployments. With draining, deployments are seamless.

7. Active-active vs active-passive

For the LB itself, you need redundancy:

Active-active

Both LBs serve traffic. DNS or anycast routes clients to both.

  • Pro: full utilization; any single LB failure loses only half capacity.
  • Con: requires state sharing (session affinity can break if a client switches LBs).

Active-passive

One LB serves; the other waits as a hot spare. Failover is automatic (VRRP, keepalived).

  • Pro: no state sharing; primary owns all sessions.
  • Con: spare capacity is wasted.

Most production deployments use active-active with anycast (Cloudflare) or DNS-based load balancing (Route 53).

8. Global load balancing

For multi-region deployments, a global LB routes traffic across regions.

  • Geo-based — route by client country. Predictable; not adaptive.
  • Latency-based — route by measured latency. Adaptive; can be wrong.
  • Weighted — shift traffic between regions gradually. Used for canary or drain.
  • Anycast — same IP advertised globally; BGP routes clients to nearest. Best for HTTP.

Tools: AWS Route 53, GCP LB, Cloudflare Load Balancing, NS1.

9. Service discovery

The LB needs to know which backends exist. In static environments, you configure them by hand. In dynamic environments, backends register themselves with a service registry.

# Backend registers on startup
registry.register("api", f"{ip}:8080", ttl=30)

# LB queries registry
backends = registry.lookup("api")

# Backend heartbeats
while running:
    registry.heartbeat("api", f"{ip}:8080")
    sleep(10)

Service registries: Consul, etcd, ZooKeeper, Kubernetes DNS.

10. Common patterns

Blue-green deployment

Deploy new version (green) alongside old (blue). Shift LB weight from blue to green in one step.

# HAProxy config
backend blue
    server app1 10.0.0.1:8080 weight 100
backend green
    server app2 10.0.0.2:8080 weight 0
# Shift weight from blue to green to deploy

Canary deployment

Route 1-5% of traffic to the new version. If metrics are healthy, ramp up.

if random.random() < 0.05:
    return green_backend()
return blue_backend()

A/B testing

Route by user attribute (cookie, header) to different versions. Used for experiments.

if req.cookies.get('experiment') == 'B':
    return backend_b
return backend_a

11. Anti-patterns

  • No health checks. "The LB will figure it out" — it won't. Dead backends stay in rotation.
  • Synchronous health checks. A health check that takes 100ms is a bug. Keep it under 10ms.
  • Single LB. A single LB is a SPOF. Use at least two in active-active.
  • No connection draining. Deployments break in-flight requests.
  • TLS re-encryption skipped. Internal traffic in plaintext is a security hole.

12. Key takeaways

  • L4 is fast and simple; L7 is smart but adds latency.
  • Algorithm choice matters: power of two is the safe default; consistent hash for cache affinity.
  • Health checks must be cheap and test the real path.
  • Active-active LB pairs are standard; never run a single LB.
  • Connection draining is essential for zero-downtime deploys.
  • Global LB adds geo-routing for multi-region deployments.

Appendix: terms

  • L4 / L7 — Layer 4 (Transport) vs Layer 7 (Application) of the OSI model.
  • Active health check — LB sends synthetic requests to verify backend health.
  • Passive health check — LB observes real traffic for failures.
  • Connection draining — period during which a backend completes in-flight requests before shutdown.
  • TLS termination — LB decrypts TLS on behalf of backends.
  • Active-active — both LBs serve traffic simultaneously.

Appendix: source dialogue excerpt

The audio for this lesson was synthesized from the following Cantonese dialogue (verbatim, not translated):

  • M: 各位同學早晨, 我係子謙。歡迎收聽系統架構課程第五課。今日嘅主題係 Load Balancing。…
  • F: 大家好, 我係曉晴。Load balancer 係將 incoming request 分配去多個 backend server 嘅 component, 提升 throughput 同 availability。今日我哋會拆解 L4 同 L7 load balancer 嘅分別, 同埋常見嘅 routing algo…
  • M: 首先講解基本概念。Load balancer 嘅核心功能係將 request 喺多個 backend 之間分配, 確保每個 backend 嘅 load 唔超過 capacity, 同時 client 唔需要知道 backend 嘅存在。Load balancer 可以係 hardware appliance 例如 F…
  • F: Load balancer 嘅另一個 function 係 TLS termination。即係 load balancer 處理 HTTPS, decrypt 之後 forward 純 HTTP 去 backend。Backend 唔需要處理 TLS, 節省 CPU 同 certificate management。…
  • M: 好, 第一個 important concept 係 L4 load balancer。L4 即係 transport layer, load balancer 只睇 IP 同 port, 唔睇 request content。例如 TCP connection 嘅 destination port 係 80, rou…
  • F: L4 load balancer 嘅 limitation 係冇 content awareness。即係唔可以根據 URL path 或者 header 嘅 content 嚟 routing。例如將 /api path route 去 API server, 將 /static path route 去 stati…

Full dialogue contains 31 segments; see script_raw.json in the source folder.

Lesson quiz · 30 questions

Question 1 of 30Answered 0 / 30
Question 1 of 30

L4 load balancing operates at which OSI layer?

Pick an answer to lock it in. We'll tell you immediately whether you got it right and show an explanation. Then press Enter or click Next to continue.

Shortcuts:ABCDpick answer on current questionEntergo to next unanswered
30 unanswered