Availability vs Consistency (CAP Theorem)¶
The three guarantees¶
In a distributed computer system you can only support two of the following three guarantees at once:
| Guarantee | Meaning |
|---|---|
| Consistency | Every read receives the most recent write, or an error. |
| Availability | Every request receives a response, without guaranteeing it contains the most recent data. |
| Partition Tolerance | The system keeps operating despite arbitrary network partitions/failures. |
Why it's really a two-way choice¶
Networks aren't reliable, so you must support partition tolerance.
Because partitions (network failures between nodes) will happen, you can't sacrifice partition tolerance in practice. That leaves a software tradeoff between consistency and availability.
So the real question is: CP or AP?
CP — Consistency + Partition Tolerance¶
-
When a partition occurs, the system waits for a response from the partitioned node, which may result in a timeout error.
-
Choose CP when your business needs require atomic reads and writes (e.g., banking, inventory).
AP — Availability + Partition Tolerance¶
-
Responses return the most readily available version of the data on any node, which might not be the latest.
-
Writes may take time to propagate once the partition is resolved.
- Choose AP when the business can tolerate eventual consistency, or when the system must keep working despite external errors (e.g., social feeds, DNS, shopping carts).
Key takeaways¶
-
CAP is often taught as "pick two of three," but in real distributed systems partition tolerance is non-negotiable, so the practical choice is between C and A.
-
This maps directly to the consistency patterns in 04-consistency-patterns.md.
Common Interview Questions¶
Design a key-value store like Redis
Why it's asked here: a KV store forces you to state which CAP trade-off you choose and defend it.
Key points to discuss:
-
State CP vs AP explicitly. Redis is typically CP (strong within a node); Dynamo is AP (eventual).
-
Explain what happens during a network partition in each choice.
-
Replication + quorums (read/write quorums) to tune the consistency/availability dial.
-
Writes: single master vs multi-master; the availability you gain vs the consistency you lose.
flowchart LR
C[Client] --> M[Master]
M -->|sync write| R1[Replica 1]
M -.async write.-> R2[Replica 2]
R1 -->|CP strong| C
R2 -->|AP eventual| C
Design a globally distributed database
Why it's asked here: multi-region data forces the CAP decision front and center.
Key points to discuss:
- You can't avoid partitions across regions → pick CP or AP per use case.
- CP: block/error on the minority side (e.g., Spanner-style with consensus).
- AP: keep serving stale data and reconcile later (e.g., Dynamo with vector clocks).
- Quantify the cost of each: stale reads (AP) vs unavailability (CP).
flowchart LR
U[User] --> DC1[DC A]
U --> DC2[DC B]
DC1 <-->|quorum replication| DC2
DC1 -->|CP| E[Error on minority]
DC2 -->|AP| S[Serve stale]