Add support for partitioned datasets
Open
- Dominant language
- Julia
- Stars
- 312
- Forks
- 78
- PR merge metrics
- No merged PRs in 30d
Description
It would be very useful to support reading (and potentially writing) partitioned datasets. The implementation could be similar to JuliaIO/Parquet.jl#138/JuliaIO/Parquet.jl#142 and would bring Arrow.jl up to par with [Pyarrow's dataset feature](https://arrow.apache.org/docs/python/dataset.html).
Contributor guide
No contributing guide indexed for this repository
Research direction
This issue names no files, tests, or entry points. Start by reviewing the referenced Parquet.jl issues and PyArrow's dataset feature, then clarify whether reading, writing, or both are in scope. Done means Arrow.jl supports the agreed partitioned-dataset behavior with appropriate coverage.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- julia
- Domain
- data-engineering
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100