Performance vs Scalability¶
Definitions¶
-
Scalable — a service is scalable if adding resources results in a proportional increase in performance. "Performance" usually means serving more units of work, but it can also mean handling larger units of work (for example, as datasets grow).
-
Performance — how fast a system is for a single unit of work.
A simple way to tell them apart¶
| Problem | Symptom |
|---|---|
| Performance problem | The system is slow for a single user. |
| Scalability problem | The system is fast for a single user, but slow under heavy load. |
Key takeaways¶
-
Adding resources (servers, memory, cores) should yield roughly linear gains — that's the definition of scalability.
-
You can have a perfectly performant system that falls over when load increases (poor scalability), and vice versa.
-
Scalability is about the relationship between added resources and added throughput, not about raw speed.
Common Interview Questions¶
Design a system that scales to millions of users on AWS
Why it's asked here: the classic "scale from 1 user to 1M+ users" question tests whether you understand the difference between making a system fast (performance) and making it scale.
Key points to discuss:
-
Start single-server, then grow: separate DB, add read replicas, add a cache, add a load balancer, scale the web tier horizontally (stateless clones).
-
Vertical scaling (a bigger box) hits a ceiling; horizontal scaling (more boxes) is the scalable path.
-
Emphasize the proportional relationship: doubling servers should roughly double capacity.
-
Show that each step removes a specific bottleneck (CPU, DB reads, sessions, etc.).
flowchart LR
U[Users] --> DNS[DNS]
DNS --> LB[Load Balancer]
LB --> W1[Web 1]
LB --> W2[Web 2]
LB --> WN[Web N]
W1 & W2 & WN --> C[(Redis cache)]
W1 & W2 & WN --> M[(DB master)]
M --> R[(Read replicas)]
Design a system that serves data from multiple data centers
Why it's asked here: multi-DC design is fundamentally a scalability/performance story — you trade consistency for the ability to serve more users in more places.
Key points to discuss:
- Route users to the nearest DC (geo/latency-based DNS) to cut latency.
- Replicate data across DCs (eventual consistency) vs paying for strong consistency.
- How added DCs increase throughput but introduce replication lag and conflicts.
- Tie back to "performance = fast for one user; scalability = fast under load".
flowchart LR
U[Users] --> DNS[Geo DNS]
DNS --> DC1[DC US]
DNS --> DC2[DC EU]
DNS --> DC3[DC Asia]
DC1 <-->|async replication| DC2
DC2 <-->|async replication| DC3