huggingface / huggingface/datatrove
Spark support
Open
enhancement
- Dominant language
- Python
- Stars
- 3.3k
- Forks
- 302
- Avg merge
- 2h 18m
- Merged PRs (30d)
- 2
Description
I'm wondering if it is possible to add support for other popular large-scale data processing frameworks like spark, since most operations are compatible with the map operation in spark. This would greatly improve the efficiency and scability of the processing pipeline when working with large datasets.
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.