Avançado13 min de leituraNível 3

CAP Theorem

When the network splits, you must choose consistency or availability.

Introdução

The CAP theorem states that a distributed data store can guarantee at most two of three properties: Consistency, Availability, and Partition tolerance. Since network partitions are a fact of life you cannot opt out of, the real, practical choice CAP forces is: during a partition, do you sacrifice consistency or availability?

Por que existe

Distributed systems replicate data across nodes for durability and scale. But nodes communicate over networks that fail — links drop, packets are lost, nodes can't reach each other. CAP formalizes the unavoidable trade-off that appears the moment two replicas can't talk: you either refuse to answer (to stay consistent) or answer with possibly-stale data (to stay available). It exists to stop engineers from promising all three.

Analogia

Two shop branches share one stock ledger by phone. If the phone line dies (partition), a customer wants the last item. Branch A can refuse to sell until it confirms with B (consistency, but unavailable to the customer), or sell it and reconcile later — risking both branches selling the same item (available, but inconsistent). You cannot both guarantee correctness and always sell while the line is down.

Como funciona

The three properties: Consistency (every read sees the most recent write, or an error — a single up-to-date view); Availability (every request gets a non-error response, though possibly stale); Partition tolerance (the system keeps operating despite dropped messages between nodes).

Because partitions will happen, P is non-negotiable for any real distributed system. So systems are effectively CP or AP. A CP system (e.g. a system using quorum writes, ZooKeeper, HBase) refuses requests it can't serve consistently during a partition — it sacrifices availability to avoid stale/conflicting data. An AP system (e.g. Dynamo-style stores, Cassandra by default) keeps serving on both sides of the partition and reconciles later, accepting temporary inconsistency.

The refinement is PACELC: else (when there's no partition, E), you still trade latency vs consistency. Even in healthy operation, stronger consistency costs coordination and latency.

During a partition

  1. 1

    A network partition separates replicas into two groups that can't communicate.

  2. 2

    A write arrives on one side and cannot be confirmed with the other side.

  3. 3

    CP choice: reject/stall the request until consistency can be guaranteed (lose availability).

  4. 4

    AP choice: accept the write locally and continue serving (lose consistency temporarily).

  5. 5

    When the partition heals, AP systems reconcile divergent versions (last-write-wins, vector clocks, CRDTs).

Interativo

Arquitetura

A partition forces the choice

Clients

Replicas, network split between them

Node Aone side
Node Bother side

The choice during the split

CP: refuseconsistent, unavailable
AP: serve staleavailable, inconsistent

Partitions aren't optional, so every distributed store is effectively CP or AP during a split.

No código

TypeScriptQuorum reads/writes: tuning toward CP or AP
// With N replicas, a read/write is "consistent" when R + W > N.
const N = 3

// Strong-ish (CP-leaning): W=3, R=1 → writes need all replicas; a partition blocks writes.
// Available (AP-leaning): W=1, R=1 → always writes locally; may read stale, R + W <= N.
function isStronglyConsistent(R: number, W: number) {
  return R + W > N   // overlapping quorums guarantee a read sees the latest write
}

isStronglyConsistent(2, 2) // true  (Cassandra QUORUM/QUORUM)
isStronglyConsistent(1, 1) // false (fast + available, but can read stale)

No mundo real

Onde você já viu isso

A shopping cart (Amazon Dynamo's original use case) is AP: it must always accept 'add to cart' even during a partition, and conflicting versions are merged later — a slightly wrong cart is better than a broken button. A bank's ledger or a config/coordination service (ZooKeeper, etcd) is CP: it would rather return an error than hand out inconsistent balances or two leaders.

Quando usar

Recorra a isso quando

  • Choose CP when correctness beats availability: money, inventory truth, coordination, leader election.
  • Choose AP when availability beats freshness: carts, feeds, likes, presence, catalogs.
  • Use quorum tuning (N/R/W) to slide along the spectrum per workload.

Evite quando

  • Don't force strong consistency on data that tolerates staleness — you pay latency and availability for nothing.
  • Don't choose AP for operations where divergence is unacceptable (double-spend).
  • Don't treat CAP as a system-wide switch; it's per-operation/per-dataset.

Trade-offs

Vantagens

  • Gives a precise vocabulary for an unavoidable trade-off.
  • Guides datastore selection and quorum configuration.
  • Forces explicit decisions about behavior under failure.

Desvantagens

  • Often oversimplified: C/A/P aren't binary and apply per-operation.
  • Ignores the no-partition latency trade-off (addressed by PACELC).
  • 'Two of three' framing misleads — P isn't really optional.

CAP's core trade is consistency vs availability under partition; PACELC adds consistency vs latency without partition. Strong consistency (CP, high quorums) means correct reads but errors during partitions and higher latency always; high availability (AP, low quorums) means the system always answers fast but you must handle staleness and conflict resolution. The mature stance is per-dataset: keep money and coordination CP, keep engagement and content AP, and tune quorums to taste.

CP vs APCP (consistency)AP (availability)
During partitionRejects/stalls unsafe requestsKeeps serving, may be stale
RiskDowntime / errorsStale & conflicting data
ReconciliationNot needed (never diverges)LWW, vector clocks, CRDTs
ExamplesZooKeeper, etcd, HBaseCassandra, DynamoDB (tunable)
FitsMoney, inventory, coordinationCarts, feeds, presence, catalogs

Erros comuns

Fique atento a

  • Claiming a system is 'CA' — impossible for a real distributed system that must tolerate partitions.
  • Treating CAP as one global choice rather than per-dataset/operation.
  • Assuming 'eventually consistent' means 'usually wrong' — it usually converges in milliseconds.
  • Forgetting the latency cost of strong consistency even when the network is healthy.

Pensando em falhas

O que quebra

A network partition splits your database cluster. What does a CP vs AP system do?

A CP system keeps only the side with quorum writable and returns errors on the minority side — consistent but partially unavailable. An AP system keeps both sides serving and reconciles on heal — available but temporarily inconsistent (two carts, conflicting values). The right choice depends entirely on whether stale/conflicting data or downtime is more harmful for that operation.

Two nodes both think they're the leader (split brain). Why, and how is it prevented?

A partition let each side elect its own leader, so both accept writes and diverge. CP systems prevent this with quorum/majority: only the side holding a majority can act as leader, and fencing tokens stop a stale leader's writes from being accepted after it's demoted. This is the availability sacrifice in action — the minority side goes read-only or errors.

Pense como sênior

Visão de Engenheiro(a) Sênior

'CA' does not exist for distributed systems — the moment you have more than one node over a network, you must tolerate partitions. If someone offers all three, they're describing a single node.

Visão de Engenheiro(a) Sênior

Never answer CAP with a global label. The senior answer is 'it depends on the operation': the same product keeps its ledger CP and its activity feed AP, and states the consistency requirement per feature before choosing.

Lembre-se

Partitions are not optional, so the real choice is C vs A during a partition.

Lembre-se

CAP is a per-operation decision, not a whole-system switch.

Questões de entrevista

1

Explain the CAP theorem and what it really forces you to choose.

Recordação ativa

Teste-se

CAP Theorem · Questão 1 / 1Médio

During a network partition, a CP system chooses to…

Desafios práticos

Arquitetura1–3 hours

Same product, two consistency models

For a social commerce app, classify each dataset as CP or AP and justify quorum settings.

Datasets: wallet balance, order state, product catalog, like counts, user presence, and inventory.

Requisitos

  • Label each dataset CP or AP.
  • State the consistency requirement in one sentence each.
  • Pick N/R/W or a store type per dataset.

Critérios de aceite

  • Money and inventory are CP with a clear rationale.
  • Engagement/presence are AP with conflict handling described.

Casos extremos

  • Inventory that must be correct but also highly available under load.

Bônus

  • Describe how each behaves during a partition and after it heals.

Reflexão

  • Which datasets could move from CP to AP if the business accepted a small error rate?

Desafio de arquitetura

You're designing a ticketing platform. The seat-reservation path must never oversell; the 'events near you' browse path must always load fast even during a datacenter link failure. Decide CP vs AP for each path, pick concrete quorum settings, and describe exactly what a user experiences on each path during a partition.

Flashcards

Flashcards1 / 4

Resumo

CAP says that under a network partition a distributed store must choose between consistency and availability, because partition tolerance isn't optional. Real systems are CP or AP per dataset, tuned via quorums, and PACELC adds the consistency-vs-latency trade even when healthy. The expert habit is to decide per operation and to reject the impossible 'CA' and the lazy global label.

Pontos-chave

  • At most two of C, A, P — and P is mandatory for real distributed systems.
  • The practical choice is C vs A during a partition: CP rejects, AP serves stale.
  • Decide per dataset/operation, not once for the whole system.
  • PACELC: even without partitions, strong consistency costs latency.

Suas anotações

Salvo neste dispositivo