Availability Patterns¶
Two complementary patterns support high availability:
- Fail-over — a standby takes over when the primary fails.
- Replication — keep redundant copies of data.
Fail-over¶
Active-passive (a.k.a. master-slave)¶
- Heartbeats are sent between the active and the passive (standby) server.
-
If the heartbeat is interrupted, the passive server takes over the active's IP address and resumes service.
-
Only the active server handles traffic.
- Downtime length depends on whether the passive is in hot standby (already running) or cold standby (must start up).
Active-active (a.k.a. master-master)¶
- Both servers handle traffic, spreading the load between them.
- If public-facing, the DNS must know both public IPs.
- If internal-facing, the application logic must know about both servers.
Disadvantages of fail-over¶
- Adds more hardware and complexity.
- Risk of data loss if the active system fails before newly written data is replicated to the passive.
Replication¶
Redundant copies of data. Discussed in depth in the database notes:
Availability in numbers¶
Availability is quantified by uptime/downtime as a percentage of time. Measured in "nines" — 99.99% availability = "four nines".
99.9% availability (three nines)¶
| Duration | Acceptable downtime |
|---|---|
| Per year | 8h 45min 57s |
| Per month | 43m 49.7s |
| Per week | 10m 4.8s |
| Per day | 1m 26.4s |
99.99% availability (four nines)¶
| Duration | Acceptable downtime |
|---|---|
| Per year | 52min 35.7s |
| Per month | 4m 23s |
| Per week | 1m 5s |
| Per day | 8.6s |
In sequence vs in parallel¶
For components with availability < 100%, overall availability depends on whether they are in sequence or in parallel.
In sequence — availability decreases:
Availability(Total) = Availability(Foo) * Availability(Bar)
Two 99.9% components in sequence → 99.9% × 99.9% = 99.8%.
In parallel — availability increases:
Availability(Total) = 1 - (1 - Availability(Foo)) * (1 - Availability(Bar))
Two 99.9% components in parallel → 1 - (0.001 × 0.001) = 99.9999%.
Common Interview Questions¶
Design a system with 99.99% availability
Why it's asked here: availability patterns (failover, replication, redundancy) are the tools that answer this.
Key points to discuss:
- Compute allowed downtime: four nines = ~52 min/year.
-
Redundancy in parallel raises availability; serial dependencies lower it (show the formulas from this page).
-
Active-passive vs active-active failover, and failover detection (heartbeats).
- Eliminate single points of failure; replicate data; use multiple AZs/regions.
flowchart LR
LB[Load Balancer] --> A[Server A active]
LB --> B[Server B standby]
A -.heartbeat.-> B
B -->|failover on failure| A
Design a system that scales to millions of users
Why it's asked here: scaling out requires the availability patterns on this page.
Key points to discuss:
- Stateless servers + externalized sessions enable horizontal scaling.
- Replication (master-slave) for read availability; failover for write availability.
- Availability "in numbers": calculate the combined availability of the chain.
- Trade hardware/complexity (more nines) against cost.
flowchart LR
U[Users] --> DNS[DNS]
DNS --> LB[Load Balancer]
LB --> W[Web tier - stateless]
W --> C[(Cache)]
W --> M[(DB master)]
M --> R[(Replicas)]
Key takeaways¶
- Redundancy (parallel components) improves availability.
-
Serial dependencies reduce availability — a chain is only as available as the product of its parts.
-
Each "nine" roughly translates to order-of-magnitude less downtime.