redpanda-data / redpanda-data/console

Hardcoded context timeouts in GetClusterInfo/GetTopicsOverview cause failures on large clusters

Open
#2,410 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
TypeScript
Stars
4.3k
Forks
432
Avg merge
3d 6h
Merged PRs (30d)
40

Description

Errors observed:

WARN "failed to fetch topic configs to return cleanup.policy" error="context deadline exceeded"
WARN "failed to describe log dirs from some shards" failed_shards=1
WARN "shard error for describing log dirs" broker_id=X error="the internal broker struct chosen to issue this request has died"

Description:

GetClusterInfo https://github.com/redpanda-data/console/blob/534690f8b52944feff76c0c596fc38c7e236ed3d/backend/pkg/console/cluster_info.go#L56 uses a hardcoded 6s timeout for DescribeLogDirs + Metadata.
and GetTopicsOverview https://github.com/redpanda-data/console/blob/534690f8b52944feff76c0c596fc38c7e236ed3d/backend/pkg/console/topic_overview.go#L80 uses a 5s timeout for DescribeConfigs.

These are not configurable.

On large AWS MSK clusters (thousands of partitions) with IAM authentication, these timeouts could be insufficient due to overhead of sts token exchange with IAM?

The same cluster works fine with mTLS auth .

Debug Logs

12:34:30.625 wrote DescribeConfigs v4  broker=2  bytes_written=214330
12:34:35.566 read  DescribeConfigs v4  broker=2  bytes_read=0  time_to_read=4.941s  err="context deadline exceeded"
12:34:35.566 read from broker errored, killing connection  broker=2  successful_reads=0
12:34:35.571 failed to describe log dirs from some shards  failed_shards=1
12:34:35.571 shard error for describing log dirs  broker_id=2  error="broker struct has died"
12:33:10.460 read Metadata v12  bytes_read=505207  time_to_read=5.790s
12:34:30.561 read Metadata v12  bytes_read=505207  time_to_read=3.770s

Should these timeouts configurable via the kafka config section?
I see the code comment that says "shorter timeout because otherwise we'll potentially have very long response times in case of a single broker being down"
But shouldn't that be left to user decision?

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start in backend/pkg/console/cluster_info.go around line 56 and backend/pkg/console/topic_overview.go around line 80, then inspect how the kafka config section is loaded and how the existing timeout comments guide behavior. Done means the timeout behavior is configurable and the large-cluster failures shown in the issue no longer result from fixed deadlines.

Written by the indexing model from the issue text.

Assessment

Tech stack
go
Domain
backend
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
48/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.