apache / apache/pinot

Allow updating segment partition metadata in consuming segments

Open
#9,149 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Java
Stars
6.1k
Forks
1.5k
Avg merge
1d 21h
Merged PRs (30d)
189

Description

We were trying to test partitioning data in REALTIME tables and found consuming segments were never getting queried when applying the filter `where = X`. We eventually traced it down to this:
- we have our own stream ingestion plugin that is ingesting data from multiple kafka sources
- to do so, it uses an algorithm we wrote to convert the upstream config into a pinot partition int id
- therefore, `murmur(partition_column) % num_partitions` will never actually equal partition id

But what we saw is, after the segment is sealed, pinot actually knows to update the partition metadata with our made up id plus the actual partition it saw in the data. Having this functionality on the consuming segment would be really useful for us.

For redundancy, we actually want to repartition our upstream data into 4 separate sources not just 1. So we want each `` to exist on 4 pinot partitions. Since pinot already knows to update the segment metadata on seal, can it just do it in realtime for consuming segments? There should only be 1 thread touching a consuming segment at a time anyway.

The other proposal is to offer disabling of partitioning on the consuming segment. But this really defeats the purpose of what we're trying to do as it greatly increases the number of segments we query.

Contributor guide

Open the contributing guide

Research direction

Start by tracing consuming segments, partition metadata updates, and the segment-sealing path described in the issue, including how the Kafka stream ingestion plugin supplies partition IDs. Compare metadata behavior before and after sealing. Done means consuming segments update metadata for observed partition IDs in realtime while supporting the stated repartitioning case.

Written by the indexing model from the issue text.

Assessment

Tech stack
java, kafka
Domain
databases
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.