Distributed Systems // Kafka

Drains & Reconnection

Why a Kafka broker restart can leave a perfectly healthy-looking consumer silently dead — and how a client is supposed to survive it.

01 // The setup

Brokers, groups, and a coordinator

A Kafka cluster is a set of brokers. A topic is split into partitions, and each partition lives on one broker (the leader) with copies on others (replicas). Producers write to the leader; consumers read from it.

Consumers that share a group divide the partitions among themselves — each partition is read by exactly one member. The group is choreographed by a group coordinator: one specific broker that tracks membership, hands out partition assignments, and stores each member's committed offset (how far it has read).

Two clocks keep a member "alive" in the group: a background heartbeat (heartbeat.interval.ms) and a session timeout (session.timeout.ms). Miss heartbeats past the session timeout and the coordinator evicts you and reassigns your partitions to someone else. This eviction-and-reassignment dance is a rebalance.

02 // What a node drain is

The ground moves under the brokers

In Kubernetes, a node drain evicts every pod off a node so it can be patched, upgraded, scaled down, or reclaimed (spot/preemptible). If a Kafka broker pod is on that node, it is terminated and rescheduled elsewhere — new pod, new IP, brief unavailability.

This is not exotic. It happens on a routine cadence:

From the client's side, every one of these looks the same: the TCP connection to a broker drops, and possibly the broker that was its group coordinator or a partition leader just disappeared.

03 // What is supposed to happen

Kafka is built to survive this

None of the above should cause data loss or a stuck consumer. The protocol has a recovery path for each failure, and a correct client walks it automatically:

1
Connection drops
The client gets a disconnect, or an error like NOT_LEADER_FOR_PARTITION / COORDINATOR_NOT_AVAILABLE.
2
Refresh metadata
The client asks any reachable broker "who is the leader / coordinator now?" Leadership has already failed over to an in-sync replica.
3
Reconnect & rejoin
The client reconnects to the new broker, re-runs FindCoordinator, and rejoins the group — triggering a rebalance.
4
Resume from last offset
Partitions are reassigned and consumption continues from the committed offset. No messages lost; at worst a few re-delivered.

The key phrase is "a correct client walks it automatically." The protocol provides the recovery path — but the client library and your code have to actually take it. That is exactly where things break.

04 // Where it goes wrong

The consumer that never comes back

A consumer-group client runs a loop: join the group, receive an assignment, read messages, commit offsets, heartbeat. When the coordinator connection breaks, the library surfaces this as the end of a session — the current assignment is revoked and the read loop returns.

A resilient client treats "session ended" as normal and immediately loops back to step 1 to rejoin. The bug is when the application instead treats it as terminal: it logs an error and lets the consume goroutine exit. The process keeps running — its HTTP server, its other goroutines — but the consumer is gone. It never rejoins the group.

Now the group has a member count of zero. Offsets are still stored, the topic keeps accumulating messages, but nothing is assigned to read them. Lag climbs forever.

Producer
→
Topic
×
Consumer
(exited)
×
Downstream

Messages still arrive at the topic. The read side is severed — but nothing crashed.

05 // Why it is silent

Healthy on the outside

The reason this is dangerous rather than merely annoying: every signal a platform normally watches says the service is fine.

No crash

The process is alive. The consume goroutine exited, but the main process and HTTP server keep running. Exit code never fires.

Probes pass

Liveness/readiness probes hit an HTTP endpoint. They test "can you serve a request?", not "is your Kafka consumer joined?". Both stay green.

No restart

Because k8s sees a healthy pod, it never recycles it. The dead consumer persists for days across no restarts.

No downstream signal

Upstream keeps producing successfully. The gap only shows as an absence — work that quietly stops arriving. You notice when someone asks "why are there no alerts?"

Field example — DS-5684
Kafka brokers restarted~18–20h before discovery
relayer-http pod1/1 Running, up 6d18h, HTTP healthy
consumer group members0 (CONSUMER-ID empty)
detector still publishingyes — 4 alerts, "sent successfully"
alerts delivered to IRIS0 — silent outage

A detection pipeline kept firing rules and writing them to Kafka for the better part of a day, while the consumer that turns them into incidents sat dead — and nothing paged anyone.

06 // Handling it correctly

Make reconnection the default state

The fix is a posture, not a patch: assume the connection will break, repeatedly, forever — and structure the consumer so that breaking is a no-op it recovers from on its own.

1 — Supervise the consume loop

The single most common Go/Sarama mistake: calling Consume() once. It returns on every rebalance — that is by design. If you don't call it again in a loop, you leave the group the first time anything moves.

Naive — exits on first rebalance
// runs once, then the // goroutine is done forever err := group.Consume( ctx, topics, handler, ) if err != nil { log.Error(err) return // <- dead }
Supervised — rejoins forever
for { // Consume returns on every // rebalance — just loop err := group.Consume( ctx, topics, handler, ) if err != nil { log.Warn(err) } if ctx.Err() != nil { return // only on shutdown } // loop -> rejoin the group }

Only exit when you are shutting down (context cancelled). Every other return is a transient that should loop straight back into a rejoin, ideally with a short backoff.

2 — Tie health to the consumer, not just HTTP

Give the readiness probe something real to check: a timestamp of the last successful poll or commit, or live group membership. If the consumer hasn't made progress within a threshold, fail readiness so k8s recycles the pod. A self-healing restart beats a silent stall.

3 — Tune the timeouts, and process idempotently

4 — Alert on the gap directly

Monitor consumer-group lag and member count. "Lag rising with zero active members" is the unambiguous fingerprint of this failure — and the one signal that would have caught DS-5684 in minutes instead of a day.

07 // Quick diagnosis

One command tells you

When work mysteriously stops flowing, describe the group. The state is unmistakable:

kubectl exec -n kafka kafka-0-<pod> -c kafka -- bash -c \ 'export KAFKA_OPTS=""; /opt/kafka/bin/kafka-consumer-groups.sh \ --bootstrap-server kafka-headless.kafka:29092 \ --describe --group <your-group>' GROUP TOPIC CURRENT END LAG CONSUMER-ID your-group the-topic 37 41 4 - <- empty = dead

Lag > 0 with an empty CONSUMER-ID = the group exists, offsets are stored, but no member is attached. That is a stalled consumer, not a slow one. The stop-gap is to restart the client so it rejoins and drains the backlog — but the durable fix is the supervised loop above, so it never needs a human.

Assume the disconnect. Survive it.

Kafka gives you the recovery path for free. Your client only has to keep walking it — every time, forever, without being asked.

Daniel Brasileiro