apache / apache/pinot

POST /tables blocks HTTP thread for minutes with high-partition-count SASL/SSL Kafka realtime table

Open
#18,743 0 comments 0 reactions 0 assignees View on GitHub
ingestion kafka real-time troubleshooting
Dominant language
Java
Stars
6.1k
Forks
1.5k
Avg merge
2d 55m
Merged PRs (30d)
182

Description

**Environment:** Pinot controller, realtime table, large partition count (e.g. 100+), multiple replica groups, Kafka over SASL/SSL.

**Symptom:** `POST /tables` takes several minutes to respond when the Kafka topic has a large number of partitions. The client times out, but the controller eventually completes the work correctly and writes the ideal state.

**Root cause (traced from source):**

`addTable` is fully synchronous on the HTTP thread. After `InstanceAssignmentDriver` finishes (ZK-only, fast), `PinotLLCRealtimeSegmentManager.setUpNewTable()` calls `getNewPartitionGroupMetadataList()`, which ends up in `StreamMetadataProvider.computePartitionGroupMetadata()`. That method loops over all partitions sequentially - for each partition it constructs a new `KafkaConsumer` (full SASL/SSL handshake) and calls `fetchStreamPartitionOffset`. With SASL/SSL, each handshake takes several seconds, so the total time scales linearly with partition count and blocks the HTTP thread throughout.

Call chain:
```
POST /tables (PinotTableRestletResource.java:262)
→ PinotHelixResourceManager.addTable() (line 1866)
→ PinotLLCRealtimeSegmentManager.setUpNewTable() (line 379)
→ getNewPartitionGroupMetadataList()
→ PinotTableIdealStateBuilder.getPartitionGroupMetadataList()
→ PartitionGroupMetadataFetcher.call()
→ StreamMetadataProvider.computePartitionGroupMetadata()
→ for i in 0..N: ← serial, no parallelism
new KafkaPartitionLevelConnectionHandler(...)
→ new KafkaConsumer<>() ← SASL/SSL handshake per partition
→ fetchStreamPartitionOffset()
```

**Questions:**
1. Is this expected? Is there a known workaround for large partition counts with SASL/SSL?
2. Is there a path to parallelize the per-partition offset fetch in `computePartitionGroupMetadata`?

Contributor guide

Open the contributing guide

Research direction

Start at PinotTableRestletResource.java:262 and follow PinotHelixResourceManager.addTable() into PinotLLCRealtimeSegmentManager.setUpNewTable(). Read StreamMetadataProvider.computePartitionGroupMetadata() and PartitionGroupMetadataFetcher.call(), focusing on the per-partition KafkaPartitionLevelConnectionHandler and offset fetches. Done means the high-partition-count setup no longer blocks the HTTP request for minutes while preserving correct partition-group metadata and ideal-state creation.

Written by the indexing model from the issue text.

Assessment

Tech stack
java, kafka
Domain
backend, distributed-systems, stream-processing
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
48/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.