slow kafka sink when upstream write data
Nobody has claimed this yet.
- Dominant language
- Go
- Stars
- 56
- Forks
- 63
- Avg merge
- 2d 20h
- Merged PRs (30d)
- 34
Description
What did you do?
- start a changefeed with kafka sink
- write some data (100mb/s)
- stop writing the data
What did you expect to see?
No response
What did you see instead?
After stopping to write data, the sink write bytes increase
Versions of the cluster
Upstream TiDB cluster version (execute SELECT tidb_version(); in a MySQL client):
(paste TiDB cluster version here)
Upstream TiKV version (execute tikv-server --version):
(paste TiKV version here)
TiCDC version (execute cdc version):
release-8.5
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Reproduce the Kafka sink behavior in the issue using a changefeed on TiCDC release-8.5: write about 100 MB/s upstream, stop writing, and observe whether sink write bytes continue increasing. First collect the missing TiDB and TiKV versions, then trace the changefeed and Kafka sink entry points to identify the cause. Done means the post-write increase is explained and the corrected behavior is covered by a regression check.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- go, kafka
- Domain
- data-engineering, distributed-systems
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Quiet
- Clarity
- Needs clarification
- Newbie friendliness
- 45/100