Skip to content

Availability vs Consistency (CAP Theorem)

The three guarantees

In a distributed computer system you can only support two of the following three guarantees at once:

Guarantee Meaning
Consistency Every read receives the most recent write, or an error.
Availability Every request receives a response, without guaranteeing it contains the most recent data.
Partition Tolerance The system keeps operating despite arbitrary network partitions/failures.

Why it's really a two-way choice

Networks aren't reliable, so you must support partition tolerance.

Because partitions (network failures between nodes) will happen, you can't sacrifice partition tolerance in practice. That leaves a software tradeoff between consistency and availability.

So the real question is: CP or AP?

CP — Consistency + Partition Tolerance

  • When a partition occurs, the system waits for a response from the partitioned node, which may result in a timeout error.

  • Choose CP when your business needs require atomic reads and writes (e.g., banking, inventory).

AP — Availability + Partition Tolerance

  • Responses return the most readily available version of the data on any node, which might not be the latest.

  • Writes may take time to propagate once the partition is resolved.

  • Choose AP when the business can tolerate eventual consistency, or when the system must keep working despite external errors (e.g., social feeds, DNS, shopping carts).

Key takeaways

  • CAP is often taught as "pick two of three," but in real distributed systems partition tolerance is non-negotiable, so the practical choice is between C and A.

  • This maps directly to the consistency patterns in 04-consistency-patterns.md.

Common Interview Questions

Design a key-value store like Redis

Why it's asked here: a KV store forces you to state which CAP trade-off you choose and defend it.

Key points to discuss:

  • State CP vs AP explicitly. Redis is typically CP (strong within a node); Dynamo is AP (eventual).

  • Explain what happens during a network partition in each choice.

  • Replication + quorums (read/write quorums) to tune the consistency/availability dial.

  • Writes: single master vs multi-master; the availability you gain vs the consistency you lose.

flowchart LR
    C[Client] --> M[Master]
    M -->|sync write| R1[Replica 1]
    M -.async write.-> R2[Replica 2]
    R1 -->|CP strong| C
    R2 -->|AP eventual| C
Design a globally distributed database

Why it's asked here: multi-region data forces the CAP decision front and center.

Key points to discuss:

  • You can't avoid partitions across regions → pick CP or AP per use case.
  • CP: block/error on the minority side (e.g., Spanner-style with consensus).
  • AP: keep serving stale data and reconcile later (e.g., Dynamo with vector clocks).
  • Quantify the cost of each: stale reads (AP) vs unavailability (CP).
flowchart LR
    U[User] --> DC1[DC A]
    U --> DC2[DC B]
    DC1 <-->|quorum replication| DC2
    DC1 -->|CP| E[Error on minority]
    DC2 -->|AP| S[Serve stale]

Further reading