hdfswriter-小文件问题
Open
- Dominant language
- Java
- Stars
- 17.4k
- Forks
- 5.7k
- PR merge metrics
- No merged PRs in 30d
Description
使用hdfswriter,如果源表按照字段切分之后(例如mysql主键),channel很多的情况下,hdfswriter可能会写非常多的小文件,datax是否有考虑这个问题?有没有解决方案呢?
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by tracing how hdfswriter maps field-based source-table splits, such as MySQL primary-key splits, across many channels and produces HDFS files. Reproduce the case with many channels and inspect the resulting small-file count; done requires an agreed solution that addresses excessive small files.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- hadoop, java, mysql
- Domain
- data-engineering, distributed-systems
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100