Kafka stores a sequence of records across several servers so applications can publish events and other applications can read them later. A producer writes a record. The server currently accepting writes for that part of the sequence is the leader, and other servers keep follower copies.
At 10:00:02 a producer writes order 5502 and receives success. At 10:00:08
two servers fail. A stale third server becomes leader, reuses the same position
for order 6011, and the acknowledged order disappears from every copy. This
chapter follows that one partition byte by byte and then shows three other ways
an acknowledged tail can be lost.
Kafka's documentation states the guarantee precisely: "a committed message will not be lost, as long as there is at least one in sync replica alive, at all times." Every loss in this piece comes from one of two places: "committed" meaning less than the producer assumed, or "at least one in sync replica alive" not holding.
For data that cannot be recreated, a common baseline is replication factor 3, min.insync.replicas=2, producers with acks=all, unclean.leader.election.enable=false, and eligible leader replicas on for current KRaft clusters. Spread the replicas across independent failure domains. The trade-off is deliberate: the partition stops accepting writes when it cannot keep two copies in sync.
How to read the diagrams
- The log: offsets, the high watermark, what consumers can read
- Replicas in the in-sync set (ISR)
- Clients: producers and consumers
- Settings and operator actions
- Loss: truncated or overwritten acknowledged data
Before the failures: the pieces in one picture
A topic is a named stream. Kafka divides it into partitions, and each partition is an ordered log whose entries have increasing offsets. A broker is one Kafka server. For each partition, one broker hosts the leader replica and other brokers host follower replicas.
The producer sends each record to the leader. Followers copy the leader's log. The controller chooses a new leader when a broker fails. Modern Kafka stores controller metadata and runs controller elections through KRaft, Kafka's built-in consensus system, instead of ZooKeeper.
Write
Producer1
acks=all
Partition orders-7
Broker 0 · leader2
offsets 0 … 100
Broker 1 · follower2
offsets 0 … 100
Broker 2 · follower2
offsets 0 … 80
Metadata
KRaft controller3
leader and ISR metadata
What "committed" means
Each partition has one leader and some followers. Producers write to the leader; followers fetch from it. The leader tracks which replicas are keeping up. A follower that has not fetched recently enough is removed from the in-sync replica set (ISR).
The log end offset (LEO) is the number the replica will assign next. If the last stored record is offset 100, the LEO is 101. The high watermark (HW) is the first offset that is not committed; ordinary consumers can read records below it. This “next offset” convention avoids an off-by-one trap: HW = 101 means offsets through 100 are committed.
The important word is "currently". The ISR is not the set of all replicas; it shrinks when replicas fall behind and grows when they catch up. Here is a partition with a replication factor of 3 at one moment:
| replica | role | last stored offset | LEO | in ISR |
|---|---|---|---|---|
| broker 0 | leader | 100 | 101 | yes |
| broker 1 | follower | 100 | 101 | yes |
| broker 2 | follower | 80 | 81 | no |
Here ISR = {0, 1} and HW = 101. Offsets 81 through 100 are committed because both current ISR members have them. Broker 2 does not. Every loss path below reuses partition orders-7, brokers 0–2, and this range of offsets.
The three settings
Producer
acks
- 0
- Don't wait for anything
- 1
- The leader has written it to its own log
- all
- Every replica currently in the ISR has it
- Default
- all from Kafka 3.0; idempotence defaulted on in 3.0.1, 3.1.1, and 3.2.0
Broker or topic
min.insync.replicas
- Meaning
- With acks=all, refuse the write if the ISR is smaller than this
- Error
- NotEnoughReplicas or NotEnoughReplicasAfterAppend; a retry can duplicate the record when producer idempotence is off
- Default
- 1: the leader alone counts as enough
Broker or topic
unclean.leader.election.enable
- Meaning
- When no ISR member is alive, allow a replica outside the ISR to become leader
- Default
- false since Kafka 0.11 (KIP-106); true before that
- Managed
- AWS MSK defaults it to true on clusters without tiered storage
With min.insync.replicas left at its default of 1, acks=all can mean “acknowledged by one machine.” The producer waits for every current ISR member; the minimum setting decides whether an ISR of size one is allowed to acknowledge at all.
Managed services choose different defaults, so check yours. Confluent Cloud sets min.insync.replicas to 2 and doesn't allow unclean election to be enabled. AWS MSK sets min.insync.replicas to 2 on three-zone clusters but 1 on two-zone clusters, and leaves unclean leader election enabled on clusters without tiered storage.
The table below is a starting policy, not a substitute for testing the availability and recovery needs of one workload.
- use case
- payments or orders
- producer and topic settings
- RF 3 · acks=all · min ISR 2 · unclean election off · ELR on
- reasoning
- prefer rejected writes and an offline partition over an acknowledged record disappearing
- use case
- application logs used for audit
- producer and topic settings
- same durability baseline; add an independent archive if required
- reasoning
- a replica set is not a backup against operator error or correlated loss
- use case
- re-creatable metrics
- producer and topic settings
- acks=all is still the current producer default; a lower durability policy needs an explicit loss budget
- reasoning
- availability may be worth more than a small, measurable gap when the source can replay
Loss 1: acks=1 and leadership moves before replication
With acks=1, the leader acknowledges as soon as the write is in its own log. The producer documentation describes the consequence directly: "should the leader fail immediately after acknowledging the record but before the followers have replicated it then the record will be lost." Failure is one trigger; a planned leader move can expose the same gap. The clean follower takes leadership without the acknowledged record, and the old leader truncates its divergent tail when it rejoins.
Loss 1 · the leader replies before replication
- 10:00:00.000
All replicas end at offset 80
leader broker 0 · ISR {0, 1, 2} · HW 81
- 10:00:00.012
Broker 0 appends order 5501 at offset 81
The record is in broker 0's log and page cache. Followers have not fetched it yet.
- 10:00:00.014
Producer receives success
acks=1 requires only the leader's local append.
- 10:00:00.019
A planned move elects broker 1
Broker 1 was in the ISR, so this is a clean election. Its next offset is still 81.
- 10:00:00.140
Broker 0 follows the new leader
The old leader truncates its unmatched offset 81. The only copy disappears without a power failure.
| moment | broker 0 | broker 1 | broker 2 |
|---|---|---|---|
| 10:00:00.014 · producer saw success | |||
| offset 81 | order 5501 | not replicated | not replicated |
| 10:00:00.140 · broker 0 has rejoined | |||
| offset 81 | truncated | absent · next write can use 81 | absent |
Disabling unclean election doesn't help here, because nothing unclean happened: the new leader was in the ISR. An open Kafka issue reports the same loss during planned leader changes on Kafka 4.1.2 with acks=1, and none with acks=all under the same test (KAFKA-20554). acks=1 has been the wrong default for anything you can't afford to lose since Kafka 3.0 changed it, but producers configured before then, or copied from old examples, still use it.
Loss 2: the ISR shrinks to one
acks=all waits for every replica in the ISR, and the ISR can contain only the leader. Kyle Kingsbury's Jepsen analysis in 2013 demonstrated this on a three-node cluster using the equivalent of acks=all: the ISR shrank to the leader, which then acknowledged writes stored on that broker alone. When that leader was partitioned away, another node was promoted and the acknowledged writes stayed with the old leader.
Loss 2 · acks=all, but ‘all’ is one broker
- 10:01:00
Brokers 1 and 2 stop fetching
A network or follower stall lasts beyond replica.lag.time.max.ms, 30 seconds by default, so the leader removes both from the ISR.
ISR changes from {0, 1, 2} to {0}
- 10:01:31
Broker 0 appends offsets 81–100
min.insync.replicas is 1, so the one-member ISR is allowed to acknowledge.
- 10:01:32
Producer receives acks=all
Every current ISR member has the records. There is only one current member.
- 10:01:40
Broker 0 fails
No live broker has offsets 81–100.
- 10:01:41
Availability requires a stale leader
If unclean election is allowed, broker 1 or 2 can lead from offset 80. If it is disabled, the partition remains offline.
| moment | broker 0 | broker 1 | broker 2 |
|---|---|---|---|
| 10:01:32 · acknowledged | |||
| offsets 81–100 | present · sole ISR member | absent | absent |
| 10:01:41 · broker 0 is unavailable | |||
| offsets 81–100 | offline | absent | absent |
- test
- Jepsen, Kafka 0.8 pre-release (2013)
- settings
- 3 nodes, all in-sync brokers must acknowledge, unclean election allowed (the only option then)
- acknowledged
- 987
- lost
- 520 (a 52.7% loss rate)
- test
- Jack Vanlightly, chaos test (2018)
- settings
- RF 3, acks=1, unclean election on
- acknowledged
- 495,103
- lost
- 144,375
- test
- Jack Vanlightly, chaos test (2018)
- settings
- RF 3, acks=all, unclean election on
- acknowledged
- 495,776
- lost
- 194,524
In Vanlightly's second test, acks=all lost more than acks=1 did because a one-broker ISR still satisfied “all.” The fix is min.insync.replicas=2 with a replication factor of 3: the leader refuses writes (with a retriable error) when fewer than two replicas are in sync, so every acknowledged write is on at least two machines. The price is availability: with two of three replicas out of sync, that partition stops accepting writes until one catches up.
Loss 3: unclean leader election
Now the case the setting is named for. Start from the partition above, with acks=all and min.insync.replicas=2, so offsets 81 to 100 are on brokers 0 and 1.
orders-7 · unclean leader election
- 10:02:00
Broker 2 falls behind and leaves the ISR
It has not fetched recently enough. The ISR is now {0, 1}.
UnderReplicatedPartitions = 1 · broker 2 ends at 80
- 10:02:31
Offsets 81–100 are written and acknowledged
Two replicas are in sync, which meets min.insync.replicas = 2. The high watermark becomes 101.
acks=all satisfied by brokers 0 and 1
- 10:02:40
Brokers 0 and 1 fail
No ISR member is alive.
the partition has no leader
- 10:02:41
Unclean election: broker 2 becomes leader
With unclean.leader.election.enable = true, the controller elects the only live replica even though it is outside the ISR.
UncleanLeaderElectionsPerSec > 0
- 10:02:42
New writes reuse the offsets
Broker 2 appends new records starting at offset 81. Offsets 81–100 can now name different records.
- 10:02:50
Consumers see a different history
A consumer that had read through 100 finds its position truncated or pointing at different records. Clients from Kafka 2.3 can detect truncation with leader epochs (KIP-320).
offset reset or LogTruncationException, depending on client configuration
- 10:03:00
Brokers 0 and 1 return and truncate
They follow broker 2's new history and cut their divergent tails.
orders-7 diverged after offset 80
| moment | broker 0 | broker 1 | broker 2 |
|---|---|---|---|
| 10:02:31 · after the acknowledged writes | |||
| offset 81 | order 5501 | order 5501 | not replicated |
| offset 82 | order 5502 | order 5502 | not replicated |
| 10:02:42 · broker 2 accepts new writes | |||
| offset 81 | offline | offline | order 6010 |
| offset 82 | offline | offline | order 6011 |
| 10:03:00 · brokers 0 and 1 rejoin and truncate | |||
| offset 81 | order 6010 | order 6010 | order 6010 |
| offset 82 | order 6011 | order 6011 | order 6011 |
The consumer API's documentation for the resulting exception puts it bluntly: "In the event of an unclean leader election, the log will be truncated, previously committed data will be lost, and new data will be written over these offsets." That last clause is what makes this worse than a simple loss. Any system that stored "processed up to offset 100" now has a position that refers to records it has never seen.
With unclean election disabled, broker 2 cannot take leadership while brokers 0 and 1 are unavailable. The partition stays offline and producers receive errors. Kafka's design documentation frames this as a choice with no free option. Jay Kreps, then at LinkedIn, put it more memorably in his response to the Jepsen post: "Is it better to be alive and wrong or right and dead?" Kafka changed the default to "right and dead" in 0.11, in 2017.
Datadog experienced temporary data loss after an unclean election on a cluster running a version older than 0.11; a second cluster held the recoverable copy. Honeycomb's report on a December 2025 incident describes the other side: cleanup after a disaster-recovery exercise destroyed brokers and left partitions with no leader. The resulting dirty elections could roll offsets back. Roughly a third of partitions were affected and one lost all of its data, on a cluster spread across three availability zones.
If you ever do need availability back more than you need the lost tail, make it a deliberate, one-time action rather than a setting: kafka-leader-election.sh --election-type unclean runs one unclean election for the partitions you name (KIP-460 added it because the setting is easy to forget to turn off again). On KRaft clusters, the setting itself, when enabled dynamically, only takes effect when a periodic check runs, every five minutes by default.
Loss 4: the last replica standing loses its page cache
Here is the case that surprises people who have done everything above correctly. Kafka writes to the operating system's page cache and, by default, never forces data to disk: log.flush.interval.messages defaults to the largest possible number. The documentation explains the choice: "we do not want to require the use of fsync on every write for our consistency guarantees as this can reduce performance by two to three orders of magnitude." Durability comes from replication instead: a crashed broker recovers the data it lost from the other replicas.
That works as long as some replica still has the data. The design proposal for Kafka's fix, KIP-966, describes what happens when it doesn't:
Unclean election disabled, and data is lost anyway
- 10:04:00
All three replicas acknowledge offsets 81–100
The tail is committed, but some bytes still live only in operating-system page caches.
- 10:04:05
Brokers 0 and 1 stop fetching long enough to leave the ISR
A network or follower stall exceeds replica.lag.time.max.ms. They retain their existing logs, while broker 2 becomes the last replica standing.
after the 30 s default lag timeout · ISR = {2}
- 10:04:32
Broker 2 loses power
The dirty pages disappear, so its on-disk log ends at 80.
- 10:05:00
Broker 2 restarts and is elected leader
It is still recorded as the ISR member. This is a clean election, so disabling unclean election does not stop it.
- 10:05:10
Brokers 0 and 1 rejoin and truncate
They follow the elected leader's shorter log and remove their longer tails.
local loss on broker 2 becomes loss on all replicas
| moment | broker 0 | broker 1 | broker 2 |
|---|---|---|---|
| 10:04:05 · committed tail, ISR has shrunk | |||
| offsets 81–100 | present but outside ISR | present but outside ISR | present · sole ISR member |
| 10:05:00 · broker 2 restarts as leader | |||
| offsets 81–100 | present on stale follower | present on stale follower | missing from elected leader |
| 10:05:10 · followers match the leader | |||
| offsets 81–100 | truncated | truncated | missing |
In KIP-966's words, the "last replica standing" state "when combined with a data loss unclean shutdown event can turn a local data loss scenario into a global data loss scenario, i.e., committed data can be removed from all replicas." Kafka's issue tracker has reports of exactly this, including a last in-sync replica that came back with an empty disk and took the partition's data with it (KAFKA-15495).
The fix is eligible leader replicas (ELR), part one of KIP-966. It changes two rules: the high watermark only advances while the ISR has at least min.insync.replicas members, and replicas that were in sync when the ISR shrank stay eligible for leadership. After an unclean shutdown, the controller can elect an eligible replica that still has the data instead of the one that lost it. ELR became available in Kafka 4.0 and is enabled by default for new clusters from 4.1. When ELR is enabled, min.insync.replicas must exist at cluster level, cannot be removed, and cannot be altered at broker level; changing the cluster or topic value clears tracked ELR state. With min.insync.replicas=1, ELR has no extra replica to preserve.
The other defence is physical. A process crash doesn't lose the page cache; losing power or crashing the operating system does. Jack Vanlightly's analysis of why Kafka doesn't need fsync makes the dependency explicit: "if all brokers lose power simultaneously then there is a risk of data loss." Spreading replicas across independent failure domains reduces the chance that one power or host event removes every cached copy; it does not protect against region-wide or operator-correlated loss. For the few topics where even that isn't enough, the flush settings can force writes to disk, at a large cost in throughput, and an independent replay source or archive may still be required.
Kafka's fixes over time
This history was checked against Apache Kafka's KIPs and current 4.3 documentation in September 2026. Kafka 4.3.1 is the current stable patch. Some entries changed a default; others changed the replication protocol, so a client setting alone cannot supply the fix.
- version
- 0.8.2 (2015)
- change
- min.insync.replicas; unclean election can be disabled per topic
- what it fixed
- Acknowledged writes on one broker only; no way to prefer consistency
- version
- 0.11 (2017)
- change
- KIP-101: truncation by leader epoch instead of high watermark
- what it fixed
- A restarting follower could truncate committed data, and replicas could diverge after correlated crashes
- version
- 0.11 (2017)
- change
- KIP-106: unclean election off by default
- what it fixed
- Clusters on default settings silently choosing availability over data
- version
- 2.0 (2018)
- change
- KIP-279: fix divergence after fast leader failover
- what it fixed
- Log divergence even with clean elections
- version
- 2.1–2.3 (2018–19)
- change
- KIP-320: broker fencing, then client truncation detection and reset support
- what it fixed
- Consumers silently reading past a truncation
- version
- 3.0–3.2 (2021–22)
- change
- KIP-679: acks=all in 3.0; idempotence enabled by default in 3.0.1, 3.1.1, and 3.2.0
- what it fixed
- New producers defaulting to weaker delivery settings
- version
- 4.0 / 4.1 (2025)
- change
- KIP-966 part 1: ELR available in 4.0 and default for new clusters in 4.1
- what it fixed
- Last replica standing losing its page cache; still relevant in 4.3
Seeing it happen
The names below are complete JMX object names from Apache Kafka's monitoring documentation, checked in September 2026. The count metrics should normally be zero. A rate can briefly become non-zero during a planned broker event, so alert thresholds should match the recovery time your cluster has demonstrated.
- JMX object name
- kafka.server:type=ReplicaManager,name=UnderReplicatedPartitions
- means
- Partitions whose ISR is smaller than their replica set
- why it matters here
- The first condition in losses 2, 3, and 4
- JMX object name
- kafka.server:type=ReplicaManager,name=UnderMinIsrPartitionCount
- means
- Partitions below min.insync.replicas
- why it matters here
- acks=all writes are being refused
- JMX object name
- kafka.server:type=ReplicaManager,name=AtMinIsrPartitionCount
- means
- Partitions exactly at the minimum
- why it matters here
- One more replica loss will stop writes
- JMX object name
- kafka.server:type=ReplicaManager,name=IsrShrinksPerSec
- means
- Rate of replicas leaving ISRs
- why it matters here
- Repeated shrinking means followers cannot keep up or stay connected
- JMX object name
- kafka.controller:type=ControllerStats,name=UncleanLeaderElectionsPerSec
- means
- Rate of unclean elections
- why it matters here
- A stale replica was elected; investigate possible truncation immediately
- JMX object name
- kafka.controller:type=ControllerStats,name=ElectionFromEligibleLeaderReplicasPerSec
- means
- Rate of elections from ELR
- why it matters here
- Confirms the controller used an eligible replica outside the current ISR
Consumers show the same events from the other side. LinkedIn's engineers wrote in 2016 that "one typical cause of offset resets is unclean leader election", and that "it is a good practice to monitor the cluster's unclean leader election rate." A sudden offset reset on a consumer group is worth treating as a possible data-loss event until proven otherwise.
A checklist
Producers
- acks=all and idempotence enabled (verify patch-level defaults on older 3.x clients)
- No producers still configured with acks=1 for data that matters
- delivery.timeout.ms sized to ride out a short loss of in-sync replicas
Topics
- Replication factor 3
- min.insync.replicas = 2, set explicitly rather than inherited
- unclean.leader.election.enable = false, checked on managed services
Cluster
- Replicas spread across availability zones
- KRaft with eligible leader replicas on (4.1+ default for new clusters)
- Unclean elections only as a deliberate one-off command
Detection
- Alerts on under-replicated, under-min-ISR, and offline partitions
- Alert on any unclean leader election
- Consumer offset resets investigated as possible loss
Sources
- Apache Kafka 4.3, Design: replication and unclean leader election, broker configs, producer configs, monitoring, eligible leader replicas, and LogTruncationException
- KIPs: KIP-101, KIP-106, KIP-279, KIP-320, KIP-679, KIP-966
- Kyle Kingsbury, Jepsen: Kafka (2013); Jay Kreps, A few notes on Kafka and Jepsen (2013)
- Jack Vanlightly, How to lose messages on a Kafka cluster, part 2 (2018), KIP-966: fixing the last replica standing issue (2023), and Why Apache Kafka doesn't need fsync to be safe (2023)
- Datadog, Lessons learned from running Kafka at Datadog (2019); Honeycomb, Incident report: exercises, cleanups, and evacuations (2026); LinkedIn, Kafkaesque days at LinkedIn, part 1 (2016)
- AWS, MSK default configuration; Confluent, Manage topics in Confluent Cloud