pingcap / pingcap/ticdc

slow kafka sink when upstream write data

Open
#4,895 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

type/bug
Dominant language
Go
Stars
56
Forks
63
Avg merge
2d 20h
Merged PRs (30d)
34

Description

What did you do?
  1. start a changefeed with kafka sink
  2. write some data (100mb/s)
  3. stop writing the data
What did you expect to see?

No response

What did you see instead?

After stopping to write data, the sink write bytes increase
Image

Image Image Image
Versions of the cluster

Upstream TiDB cluster version (execute SELECT tidb_version(); in a MySQL client):

(paste TiDB cluster version here)

Upstream TiKV version (execute tikv-server --version):

(paste TiKV version here)

TiCDC version (execute cdc version):

release-8.5

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Reproduce the Kafka sink behavior in the issue using a changefeed on TiCDC release-8.5: write about 100 MB/s upstream, stop writing, and observe whether sink write bytes continue increasing. First collect the missing TiDB and TiKV versions, then trace the changefeed and Kafka sink entry points to identify the cause. Done means the post-write increase is explained and the corrected behavior is covered by a regression check.

Written by the indexing model from the issue text.

Assessment

Tech stack
go, kafka
Domain
data-engineering, distributed-systems
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.