apache / apache/seatunnel

[Feature] Connector prepare for RAG

Open
#9,713 3 comments 4 reactions 0 assignees View on GitHub
feature good first issue help wanted llm
Dominant language
Java
Stars
9.7k
Forks
2.4k
Avg merge
3d 17h
Merged PRs (30d)
210

Description

### Search before asking

- [x] I had searched in the [feature](https://github.com/apache/seatunnel/issues?q=is%3Aissue+label%3A%22Feature%22) and found no similar feature requirement.

### Description

As a multimodal data integration tool, we hope that SeaTunnel can support parsing complex file types, converting their contents into structured file streams, and ultimately writing them into a vector library through embedding. This issue tracks related tasks.

For chunking please refer
Please refer https://docs.dify.ai/en/guides/knowledge-base/create-knowledge-and-upload-documents/chunking-and-cleaning-text
and
https://docs.llamaindex.ai/en/stable/examples/node_parsers/semantic_chunking/

### Usage Scenario

_No response_

### Related issues

_No response_

### Are you willing to submit a PR?

- [ ] Yes I am willing to submit a PR!

### Code of Conduct

- [x] I agree to follow this project's [Code of Conduct](https://www.apache.org/foundation/policies/conduct)

Contributor guide

No contributing guide indexed for this repository

Research direction

Start with the linked Dify chunking and LlamaIndex semantic-chunking references to define the expected behavior; the issue names no SeaTunnel connector, file, test, or entry point. Done is not specified beyond parsing complex files, producing structured streams, embedding the content, and writing it to a vector library.

Written by the indexing model from the issue text.

Assessment

Tech stack
java
Domain
ai, data-engineering
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.