apache / apache/paimon

[Feature] Introduce statistics in Paimon

Open
#760 0 comments 0 reactions 0 assignees View on GitHub
enhancement
Dominant language
Java
Stars
3.4k
Forks
1.4k
Avg merge
1d 11h
Merged PRs (30d)
396

Description

### Search before asking

- [X] I searched in the [issues](https://github.com/apache/incubator-paimon/issues) and found nothing similar.

### Motivation

1. Introduce partition statistics
2. Introduce column statistics

This is a big feature and should consider how to collect statistics efficiently if there are many files in a snapshot

### Solution

_No response_

### Anything else?

_No response_

### Are you willing to submit a PR?

- [ ] I'm willing to submit a PR!

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reviewing Paimon's snapshot and file-handling concepts, then clarify how partition and column statistics should be collected efficiently across many files in a snapshot. Done should be an agreed design for both kinds of statistics and their efficient collection.

Written by the indexing model from the issue text.

Assessment

Tech stack
java
Domain
data-engineering
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.