Skip to content

Performance vs Scalability

Definitions

  • Scalable — a service is scalable if adding resources results in a proportional increase in performance. "Performance" usually means serving more units of work, but it can also mean handling larger units of work (for example, as datasets grow).

  • Performance — how fast a system is for a single unit of work.

A simple way to tell them apart

Problem Symptom
Performance problem The system is slow for a single user.
Scalability problem The system is fast for a single user, but slow under heavy load.

Key takeaways

  • Adding resources (servers, memory, cores) should yield roughly linear gains — that's the definition of scalability.

  • You can have a perfectly performant system that falls over when load increases (poor scalability), and vice versa.

  • Scalability is about the relationship between added resources and added throughput, not about raw speed.

Common Interview Questions

Design a system that scales to millions of users on AWS

Why it's asked here: the classic "scale from 1 user to 1M+ users" question tests whether you understand the difference between making a system fast (performance) and making it scale.

Key points to discuss:

  • Start single-server, then grow: separate DB, add read replicas, add a cache, add a load balancer, scale the web tier horizontally (stateless clones).

  • Vertical scaling (a bigger box) hits a ceiling; horizontal scaling (more boxes) is the scalable path.

  • Emphasize the proportional relationship: doubling servers should roughly double capacity.

  • Show that each step removes a specific bottleneck (CPU, DB reads, sessions, etc.).

flowchart LR
    U[Users] --> DNS[DNS]
    DNS --> LB[Load Balancer]
    LB --> W1[Web 1]
    LB --> W2[Web 2]
    LB --> WN[Web N]
    W1 & W2 & WN --> C[(Redis cache)]
    W1 & W2 & WN --> M[(DB master)]
    M --> R[(Read replicas)]
Design a system that serves data from multiple data centers

Why it's asked here: multi-DC design is fundamentally a scalability/performance story — you trade consistency for the ability to serve more users in more places.

Key points to discuss:

  • Route users to the nearest DC (geo/latency-based DNS) to cut latency.
  • Replicate data across DCs (eventual consistency) vs paying for strong consistency.
  • How added DCs increase throughput but introduce replication lag and conflicts.
  • Tie back to "performance = fast for one user; scalability = fast under load".
flowchart LR
    U[Users] --> DNS[Geo DNS]
    DNS --> DC1[DC US]
    DNS --> DC2[DC EU]
    DNS --> DC3[DC Asia]
    DC1 <-->|async replication| DC2
    DC2 <-->|async replication| DC3

Further reading