Skip to content

Load Balancer

What it does

A load balancer distributes incoming client requests across computing resources (application servers, databases, etc.), then returns each resource's response to the appropriate client.

Load balancers are effective at:

  • Preventing requests from going to unhealthy servers.
  • Preventing overloading of resources.
  • Helping eliminate a single point of failure.

Can be implemented with hardware (expensive) or software (e.g., HAProxy, NGINX).

Additional benefits

  • SSL termination — decrypt incoming requests and encrypt responses so backend servers don't do this expensive work, and you don't need to install X.509 certificates on each server.

  • Session persistence — issue cookies to route a client's requests to the same instance when the web app doesn't track sessions itself.

Routing strategies

A load balancer can route based on:

  • Random
  • Least loaded
  • Session / cookies
  • Round robin or weighted round robin
  • Layer 4 (transport layer)
  • Layer 7 (application layer)

Layer 4 vs Layer 7

Layer 4 (transport layer)

  • Looks at source/destination IP addresses and ports (not packet contents).
  • Forwards packets using Network Address Translation (NAT).
  • Less flexible, but faster and cheaper to compute.

Layer 7 (application layer)

  • Looks at header/message/cookie contents to make routing decisions.
  • Terminates the connection, reads the message, decides, then opens a connection to the selected server.

  • Example: direct video traffic to video servers, and billing traffic to security-hardened servers.

Layer 4 Layer 7
Inspects IP + ports Headers, messages, cookies
Speed Faster Slower
Flexibility Low High

Horizontal scaling

Load balancers enable horizontal scaling (scaling out with more commodity machines) vs vertical scaling (scaling up one bigger, pricier server). Horizontal scaling is more cost-efficient and gives higher availability.

Disadvantages of horizontal scaling

  • Introduces complexity; requires cloning servers.
  • Servers should be stateless — no user data like sessions or profile pics.
  • Store sessions in a central data store (DB or persistent cache like Redis/Memcached).

  • Downstream servers (caches, DBs) must handle more simultaneous connections as upstream scales out.

Disadvantages of load balancers

  • Can become a performance bottleneck if under-resourced or misconfigured.
  • Adds complexity.
  • A single load balancer is itself a single point of failure — so you add multiple (active-passive or active-active), which adds further complexity.

Key takeaways

  • Load balancers solve distribution and availability problems, not correctness.

  • Stateless servers + externalized sessions are prerequisites for clean horizontal scaling.

  • L4 vs L7 is a speed-vs-intelligence trade-off.

Common Interview Questions

Design a system that scales to millions of users on AWS

Why it's asked here: the load balancer is the linchpin of horizontal scaling.

Key points to discuss:

  • Place an LB in front of stateless web/app tiers.
  • Routing algorithms (round-robin, least-connections) and why they matter.
  • L4 vs L7: when you need header/cookie-aware routing (L7) vs raw speed (L4).
  • Multiple LBs (active-active) to avoid the LB being a single point of failure.
flowchart LR
    U[Users] --> LB[ELB]
    LB --> W1[Web 1]
    LB --> W2[Web 2]
    LB --> WN[Web N]
    W1 & W2 & WN --> C[(Redis)]
    W1 & W2 & WN --> DB[(RDS)]
Design the application tier of a web service

Why it's asked here: the L4/L7 choice and session handling are core LB decisions.

Key points to discuss:

  • Session persistence (sticky sessions) vs stateless + a central session store.
  • SSL termination at the LB.
  • Health checks to stop routing to unhealthy nodes.
  • How the LB can become a bottleneck and how to scale it out.
flowchart LR
    C[Client] --> L4[Layer 4 LB]
    L4 -->|SSL termination| L7[Layer 7 proxy]
    L7 --> A[App 1]
    L7 --> B[App 2]
    A --> S[(Session store)]

Further reading