Load Balancer¶
What it does¶
A load balancer distributes incoming client requests across computing resources (application servers, databases, etc.), then returns each resource's response to the appropriate client.
Load balancers are effective at:
- Preventing requests from going to unhealthy servers.
- Preventing overloading of resources.
- Helping eliminate a single point of failure.
Can be implemented with hardware (expensive) or software (e.g., HAProxy, NGINX).
Additional benefits¶
-
SSL termination — decrypt incoming requests and encrypt responses so backend servers don't do this expensive work, and you don't need to install X.509 certificates on each server.
-
Session persistence — issue cookies to route a client's requests to the same instance when the web app doesn't track sessions itself.
Routing strategies¶
A load balancer can route based on:
- Random
- Least loaded
- Session / cookies
- Round robin or weighted round robin
- Layer 4 (transport layer)
- Layer 7 (application layer)
Layer 4 vs Layer 7¶
Layer 4 (transport layer)¶
- Looks at source/destination IP addresses and ports (not packet contents).
- Forwards packets using Network Address Translation (NAT).
- Less flexible, but faster and cheaper to compute.
Layer 7 (application layer)¶
- Looks at header/message/cookie contents to make routing decisions.
-
Terminates the connection, reads the message, decides, then opens a connection to the selected server.
-
Example: direct video traffic to video servers, and billing traffic to security-hardened servers.
| Layer 4 | Layer 7 | |
|---|---|---|
| Inspects | IP + ports | Headers, messages, cookies |
| Speed | Faster | Slower |
| Flexibility | Low | High |
Horizontal scaling¶
Load balancers enable horizontal scaling (scaling out with more commodity machines) vs vertical scaling (scaling up one bigger, pricier server). Horizontal scaling is more cost-efficient and gives higher availability.
Disadvantages of horizontal scaling¶
- Introduces complexity; requires cloning servers.
- Servers should be stateless — no user data like sessions or profile pics.
-
Store sessions in a central data store (DB or persistent cache like Redis/Memcached).
-
Downstream servers (caches, DBs) must handle more simultaneous connections as upstream scales out.
Disadvantages of load balancers¶
- Can become a performance bottleneck if under-resourced or misconfigured.
- Adds complexity.
- A single load balancer is itself a single point of failure — so you add multiple (active-passive or active-active), which adds further complexity.
Key takeaways¶
-
Load balancers solve distribution and availability problems, not correctness.
-
Stateless servers + externalized sessions are prerequisites for clean horizontal scaling.
-
L4 vs L7 is a speed-vs-intelligence trade-off.
Common Interview Questions¶
Design a system that scales to millions of users on AWS
Why it's asked here: the load balancer is the linchpin of horizontal scaling.
Key points to discuss:
- Place an LB in front of stateless web/app tiers.
- Routing algorithms (round-robin, least-connections) and why they matter.
- L4 vs L7: when you need header/cookie-aware routing (L7) vs raw speed (L4).
- Multiple LBs (active-active) to avoid the LB being a single point of failure.
flowchart LR
U[Users] --> LB[ELB]
LB --> W1[Web 1]
LB --> W2[Web 2]
LB --> WN[Web N]
W1 & W2 & WN --> C[(Redis)]
W1 & W2 & WN --> DB[(RDS)]
Design the application tier of a web service
Why it's asked here: the L4/L7 choice and session handling are core LB decisions.
Key points to discuss:
- Session persistence (sticky sessions) vs stateless + a central session store.
- SSL termination at the LB.
- Health checks to stop routing to unhealthy nodes.
- How the LB can become a bottleneck and how to scale it out.
flowchart LR
C[Client] --> L4[Layer 4 LB]
L4 -->|SSL termination| L7[Layer 7 proxy]
L7 --> A[App 1]
L7 --> B[App 2]
A --> S[(Session store)]