Reverse Proxy (Web Server)¶
What it is¶
A reverse proxy is a web server that centralizes internal services and provides a unified interface to the public. Client requests are forwarded to a server that can fulfill them; the reverse proxy returns that server's response to the client.
Benefits¶
-
Increased security — hide backend server info, blacklist IPs, limit connections per client.
-
Increased scalability & flexibility — clients only see the proxy's IP, so you can scale or reconfigure servers freely.
-
SSL termination — decrypt/encrypt at the edge so backends don't need to (removes need for X.509 certs on each server).
-
Compression — compress responses.
- Caching — return responses for cached requests.
- Static content — serve HTML/CSS/JS, photos, videos directly.
Load balancer vs reverse proxy¶
| Load balancer | Reverse proxy | |
|---|---|---|
| Best when | You have multiple servers serving the same function | Even a single web/app server (still gains security, caching, SSL, etc.) |
| Role | Distribute load | Centralize/unify access + add edge features |
- Tools like NGINX and HAProxy can do both layer-7 reverse proxying and load balancing.
Disadvantages¶
- Introduces complexity.
- A single reverse proxy is a single point of failure; configuring multiple (with failover) adds more complexity.
Key takeaways¶
-
A reverse proxy sits in front of servers and acts on their behalf; a load balancer distributes among servers.
-
In practice the two roles frequently overlap in the same software.
Common Interview Questions¶
Design an API gateway
Why it's asked here: an API gateway is a reverse proxy with extra routing/edge concerns.
Key points to discuss:
- Centralize auth, rate limiting, SSL termination, and logging at the edge.
- Route by path/header to internal services (layer-7 reverse proxying).
- Caching and compression as reverse-proxy features.
- Trade-off: a single entry point = single point of failure (add failover).
flowchart LR
C[Client] --> G[API Gateway]
G -->|authenticate| A[Auth service]
G -->|route| S1[Service A]
G -->|route| S2[Service B]
G -->|rate limit| R[(Redis)]
Design a web crawler's fetching layer
Why it's asked here: crawler fetchers commonly use proxies and respect robots/rate limits.
Key points to discuss:
- Reverse proxy vs forward proxy roles in fetching.
- Throttling per domain, retries, and timeouts at the fetch layer.
- How a proxy centralizes connection management and hides origin identity.
- Connection pooling to reduce overhead.
flowchart LR
F[Fetch worker] --> P[Proxy]
P -->|throttle per domain| T[Target site]
F --> Q[(URL frontier queue)]