Why a Kafka broker restart can leave a perfectly healthy-looking consumer silently dead — and how a client is supposed to survive it.
A Kafka cluster is a set of brokers. A topic is split into partitions, and each partition lives on one broker (the leader) with copies on others (replicas). Producers write to the leader; consumers read from it.
Consumers that share a group divide the partitions among themselves — each partition is read by exactly one member. The group is choreographed by a group coordinator: one specific broker that tracks membership, hands out partition assignments, and stores each member's committed offset (how far it has read).
Two clocks keep a member "alive" in the group: a background heartbeat (heartbeat.interval.ms) and a session timeout (session.timeout.ms). Miss heartbeats past the session timeout and the coordinator evicts you and reassigns your partitions to someone else. This eviction-and-reassignment dance is a rebalance.
In Kubernetes, a node drain evicts every pod off a node so it can be patched, upgraded, scaled down, or reclaimed (spot/preemptible). If a Kafka broker pod is on that node, it is terminated and rescheduled elsewhere — new pod, new IP, brief unavailability.
This is not exotic. It happens on a routine cadence:
From the client's side, every one of these looks the same: the TCP connection to a broker drops, and possibly the broker that was its group coordinator or a partition leader just disappeared.
None of the above should cause data loss or a stuck consumer. The protocol has a recovery path for each failure, and a correct client walks it automatically:
The key phrase is "a correct client walks it automatically." The protocol provides the recovery path — but the client library and your code have to actually take it. That is exactly where things break.
A consumer-group client runs a loop: join the group, receive an assignment, read messages, commit offsets, heartbeat. When the coordinator connection breaks, the library surfaces this as the end of a session — the current assignment is revoked and the read loop returns.
A resilient client treats "session ended" as normal and immediately loops back to step 1 to rejoin. The bug is when the application instead treats it as terminal: it logs an error and lets the consume goroutine exit. The process keeps running — its HTTP server, its other goroutines — but the consumer is gone. It never rejoins the group.
Now the group has a member count of zero. Offsets are still stored, the topic keeps accumulating messages, but nothing is assigned to read them. Lag climbs forever.
Messages still arrive at the topic. The read side is severed — but nothing crashed.
The reason this is dangerous rather than merely annoying: every signal a platform normally watches says the service is fine.
The process is alive. The consume goroutine exited, but the main process and HTTP server keep running. Exit code never fires.
Liveness/readiness probes hit an HTTP endpoint. They test "can you serve a request?", not "is your Kafka consumer joined?". Both stay green.
Because k8s sees a healthy pod, it never recycles it. The dead consumer persists for days across no restarts.
Upstream keeps producing successfully. The gap only shows as an absence — work that quietly stops arriving. You notice when someone asks "why are there no alerts?"
A detection pipeline kept firing rules and writing them to Kafka for the better part of a day, while the consumer that turns them into incidents sat dead — and nothing paged anyone.
The fix is a posture, not a patch: assume the connection will break, repeatedly, forever — and structure the consumer so that breaking is a no-op it recovers from on its own.
The single most common Go/Sarama mistake: calling Consume() once. It returns on every rebalance — that is by design. If you don't call it again in a loop, you leave the group the first time anything moves.
Only exit when you are shutting down (context cancelled). Every other return is a transient that should loop straight back into a rejoin, ideally with a short backoff.
Give the readiness probe something real to check: a timestamp of the last successful poll or commit, or live group membership. If the consumer hasn't made progress within a threshold, fail readiness so k8s recycles the pod. A self-healing restart beats a silent stall.
session.timeout.ms / heartbeat.interval.ms — long enough to ride out a broker blip, short enough to evict a truly dead member quicklymax.poll.interval.ms — if your handler is slow, this is what silently kicks you out of the group; size it to your real processing timeMonitor consumer-group lag and member count. "Lag rising with zero active members" is the unambiguous fingerprint of this failure — and the one signal that would have caught DS-5684 in minutes instead of a day.
When work mysteriously stops flowing, describe the group. The state is unmistakable:
Lag > 0 with an empty CONSUMER-ID = the group exists, offsets are stored, but no member is attached. That is a stalled consumer, not a slow one. The stop-gap is to restart the client so it rejoins and drains the backlog — but the durable fix is the supervised loop above, so it never needs a human.
Kafka gives you the recovery path for free. Your client only has to keep walking it — every time, forever, without being asked.
Daniel Brasileiro