linkedin / linkedin/Burrow

Polling the HTTP /lag endpoint causes strange results

Open
#833 2 comments 0 reactions 0 assignees View on GitHub
Dominant language
Go
Stars
4k
Forks
818
Avg merge
1h 14m
Merged PRs (30d)
1

Description

This maybe my limited understanding, here, but it's worth me writing this down to see if I'm misguided or have found an issue. My setup is 1 x Burrow instance consuming a single Kafka cluster. I have a polling script that calls the `/lag` endpoint periodically and forwards the retrieved JSON payload to a timeseries DB for visualisation and alerting.

After a period of ~36 hours I start to see a large increase in the calculated `current_lag` value per Topic Partition. I can find no natural explanation for this lag in terms of an increase of messages produced, or a slowdown in the Consumers within the Group. What I have found is:

- Restarting the poller makes the `current_lag` reset to zero.
- Starting a second poller alongside causes different `current_lag` values to be reported from the first poller.

To be clear: I have no fancy logic in the pollers - they simply perform a HTTP GET on `v3/kafka/$CLUSTER/consumer/$CONSUMER_GROUP/lag` and report the `current_lag` per Partition.

Have I misconfigured something?

This is Burrow v1.8.0

Contributor guide

No contributing guide indexed for this repository

Research direction

Start with Burrow v1.8.0 and the v3/kafka/$CLUSTER/consumer/$CONSUMER_GROUP/lag endpoint, reproducing repeated polling with one and two pollers. Compare the reported current_lag values over time; the investigation is complete when the cause of the reset and divergent values is identified and consistent lag reporting is demonstrated.

Written by the indexing model from the issue text.

Assessment

Tech stack
go, kafka
Domain
observability, stream-processing
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.