--:--
kafka/
KAFKA

Kafka, live

Three brokers, two topics, two consumer groups, all running right here. Take a guided scenario, or break things yourself: stop a broker, cut its power, add a consumer, drag the timeline back to see it again, or type kafka-topics --describe in the terminal.

Simulated time 0:00
Guided scenarios

Short stories played on a live cluster, taken from Kafka design documents and loss tests. At each step, guess what will happen, then watch it happen.

Basics

  • The life of one record

    Follow one record from the producer to three brokers to a consumer.

Durability

  • acks=1 loses acked writes

    The leader confirms a write no follower has, then dies.

  • acks=all is only as strong as the ISR

    With min.insync.replicas=1, "all" can mean one broker.

  • Unclean leader election

    Availability or durability, when the only in-sync replica dies.

KIP-101: leader epochs

  • KIP-101, part 1: a restart loses an acked record (Kafka 0.10)

    Followers used to cut back to their own high watermark. Watch why that was wrong.

  • KIP-101, part 2: the same restart with leader epochs (Kafka 0.11+)

    Followers ask the leader where their epoch ended instead of trusting their own HW.

  • KIP-101, part 3: logs diverge after a power cut (Kafka 0.10)

    Two replicas end up with different records at the same offset, and both look in sync.

  • KIP-101, part 4: the power cut with leader epochs

    Epochs keep the replicas identical, though the unflushed write is still gone.

Consumer groups

  • Eager vs cooperative rebalancing

    Two groups get a new member at the same moment. One stops, one keeps reading.

  • A crash is not a goodbye

    Clean shutdown vs crash, and why at-least-once means duplicates.

Keys and the log

  • Keys, order and partitions

    Same key, same partition; until the partition count changes.

  • Retention and compaction

    Old data is deleted on a schedule, read or not; compaction keeps the latest per key.

Producers

checkout

→ orders · 3/s · acks=all

A few keys

sent 0 · acked 0 · failed 0

clicks

→ orders · 2/s · acks=1

No keys (sticky)

sent 0 · acked 0 · failed 0

profile-svc

→ profiles · 1.5/s · acks=all

A few keys

sent 0 · acked 0 · failed 0

Cluster

Broker 1

Controller

Running

In0 B/s
Out0 B/s
Disk0 B
Cache0 B

Broker 2

Running

In0 B/s
Out0 B/s
Disk0 B
Cache0 B

Broker 3

Running

In0 B/s
Out0 B/s
Disk0 B
Cache0 B

Consumer groups

billing

Rebalancing

range · gen 0 · Total lag 0

  • billing-13/s

    No partitions (more members than partitions, or rebalancing)

  • billing-23/s

    No partitions (more members than partitions, or rebalancing)

search-index

Rebalancing

cooperative-sticky · gen 0 · Total lag 0

  • indexer8/s

    No partitions (more members than partitions, or rebalancing)

How to read this
  • A record; the colour is its key · No key
  • L Leader replica: takes writes and serves reads
  • F Follower replica: copies the leader
  • × Out of the ISR: too far behind
  • High watermark · A follower's own high watermark
  • Only in the page cache (not yet on disk)
  • Differs from the leader at the same offset

orders · partition 0

Leader
3
Leader epoch
0
Replicas
3, 2, 1
In-sync replicas
3, 2, 1
Log start offset
0
High watermark
0
Log end offset
0

Click a record to follow its journey.

Segment files on broker 3 Leader

  • 00000000000000000000.log0 records0 B.index .timeindex

Segment files on broker 2 ISR

own HW 0

  • 00000000000000000000.log0 records0 B.index .timeindex

Segment files on broker 1 ISR

own HW 0

  • 00000000000000000000.log0 records0 B.index .timeindex

HW: consumers read below this line · LEO: next offset to write · only in the page cache · differs from the leader at the same offset

NOTES

Kafka notes

The concepts behind what you see above. They also live under Notes, tagged #kafka.