apache / apache/pinot

Event-driven push mechanism for near real-time granular segment metadata

Open
#18,712 16 comments 0 reactions 0 assignees View on GitHub
PEP-Request
Dominant language
Java
Stars
6.1k
Forks
1.5k
Avg merge
2d 55m
Merged PRs (30d)
182

Description

**Introduction:**
Currently, there is no straightforward way to get in-depth, near real-time metadata for segments as they are committed to the deep store _(I say committed and not ingested, since we do not want to bottleneck ingestion)_ in Pinot.
While ZooKeeper's external view provides basic segment-level metadata (timestamps, total docs, CRCs), more granular physical metadata; like Bloom filter states, dictionary sizes, and specific index configurations only exist on servers.

To access this today, we have to rely on the`/segments/{tableName}/metadata?columns=` API.
This presents a few roadblocks:

- It is not performant for tables with even a moderate number of columns.
- It requires heavy, synchronous polling, which puts unnecessary load on the servers.
- It completely prevents near real-time availability of segment metadata for downstream systems, this can be attributed to introducing multiple bottlenecks with this approach (disk, network)

I would like to propose having a (configurable/optional) event driven mechanism that pushes complete segment metadata to a sink (maybe a kafka topic) once a segment is committed to the deep store.

Capturing metadata at this granular level via an event stream would enable:

- **Better observability into Pinot operations:** Better visibility into index storage footprints, anomaly detection, and pipeline health at a very granular level without hammering the Controller/Server/Zookeeper APIs.
- Managing TTL'd / Cold-Tier Segments: Pushing this data to a separate Meta table would allow a user to maintain a permanent catalog of segments even after their TTL has expired and they are dropped from the active cluster. While this data is present in the deep store, it is not easily queryable as it would be if it resided on Pinot

_This just serves as an issue to gather community interest, if sufficient interest is generated, will come up with a PEP._

Contributor guide

Open the contributing guide

Research direction

Start by reviewing the existing /segments/{tableName}/metadata?columns= API and how segment commits reach the deep store. Compare the proposed configurable event stream and possible Kafka sink with the stated observability and cold-tier catalog goals; a concrete scope and PEP would be needed before implementation can be considered done.

Written by the indexing model from the issue text.

Assessment

Tech stack
java, kafka
Domain
data-engineering, distributed-systems, observability
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.