apache / apache/paimon

[umbrella] Paimon 2.0 Planning

Open
#6,866 2 comments 12 reactions 0 assignees View on GitHub
enhancement
Dominant language
Java
Stars
3.4k
Forks
1.4k
Avg merge
1d 11h
Merged PRs (30d)
396

Description

Apache Paimon, as a high-performance data lake format, has achieved significant results in the field of structured data processing.
With the deepening of AI application scenarios, the demand for fusion analysis of multimodal data (such as text, images, audio, video, etc.)
is increasing day by day. We have invested a lot of development work in cross modal retrieval, unified storage, and efficient analysis.

Paimon 2.0 development, focusing on breaking through multimodal data support and creating a more powerful data lake solution.

Core objective:
- Unified Storage for structure & multimodal & vector.
- Efficient Search for data and vectors and full text.
- PyPaimon integrate to AI system like Ray, Pytorch.

Contributor guide

No contributing guide indexed for this repository

Research direction

This is an umbrella planning issue rather than a scoped implementation task, and it names no files, tests, or entry points. Start by reviewing the Paimon 2.0 objectives for unified multimodal, vector, and structured storage, search, and PyPaimon integration, then identify a concrete subtask with explicit completion criteria before beginning work.

Written by the indexing model from the issue text.

Assessment

Tech stack
java, pytorch
Domain
data-engineering, machine-learning, search
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.