Skip to content

Availability Patterns

Two complementary patterns support high availability:

  1. Fail-over — a standby takes over when the primary fails.
  2. Replication — keep redundant copies of data.

Fail-over

Active-passive (a.k.a. master-slave)

  • Heartbeats are sent between the active and the passive (standby) server.
  • If the heartbeat is interrupted, the passive server takes over the active's IP address and resumes service.

  • Only the active server handles traffic.

  • Downtime length depends on whether the passive is in hot standby (already running) or cold standby (must start up).

Active-active (a.k.a. master-master)

  • Both servers handle traffic, spreading the load between them.
  • If public-facing, the DNS must know both public IPs.
  • If internal-facing, the application logic must know about both servers.

Disadvantages of fail-over

  • Adds more hardware and complexity.
  • Risk of data loss if the active system fails before newly written data is replicated to the passive.

Replication

Redundant copies of data. Discussed in depth in the database notes:

Availability in numbers

Availability is quantified by uptime/downtime as a percentage of time. Measured in "nines" — 99.99% availability = "four nines".

99.9% availability (three nines)

Duration Acceptable downtime
Per year 8h 45min 57s
Per month 43m 49.7s
Per week 10m 4.8s
Per day 1m 26.4s

99.99% availability (four nines)

Duration Acceptable downtime
Per year 52min 35.7s
Per month 4m 23s
Per week 1m 5s
Per day 8.6s

In sequence vs in parallel

For components with availability < 100%, overall availability depends on whether they are in sequence or in parallel.

In sequence — availability decreases:

Availability(Total) = Availability(Foo) * Availability(Bar)

Two 99.9% components in sequence → 99.9% × 99.9% = 99.8%.

In parallel — availability increases:

Availability(Total) = 1 - (1 - Availability(Foo)) * (1 - Availability(Bar))

Two 99.9% components in parallel → 1 - (0.001 × 0.001) = 99.9999%.

Common Interview Questions

Design a system with 99.99% availability

Why it's asked here: availability patterns (failover, replication, redundancy) are the tools that answer this.

Key points to discuss:

  • Compute allowed downtime: four nines = ~52 min/year.
  • Redundancy in parallel raises availability; serial dependencies lower it (show the formulas from this page).

  • Active-passive vs active-active failover, and failover detection (heartbeats).

  • Eliminate single points of failure; replicate data; use multiple AZs/regions.
flowchart LR
    LB[Load Balancer] --> A[Server A active]
    LB --> B[Server B standby]
    A -.heartbeat.-> B
    B -->|failover on failure| A
Design a system that scales to millions of users

Why it's asked here: scaling out requires the availability patterns on this page.

Key points to discuss:

  • Stateless servers + externalized sessions enable horizontal scaling.
  • Replication (master-slave) for read availability; failover for write availability.
  • Availability "in numbers": calculate the combined availability of the chain.
  • Trade hardware/complexity (more nines) against cost.
flowchart LR
    U[Users] --> DNS[DNS]
    DNS --> LB[Load Balancer]
    LB --> W[Web tier - stateless]
    W --> C[(Cache)]
    W --> M[(DB master)]
    M --> R[(Replicas)]

Key takeaways

  • Redundancy (parallel components) improves availability.
  • Serial dependencies reduce availability — a chain is only as available as the product of its parts.

  • Each "nine" roughly translates to order-of-magnitude less downtime.