alibaba / alibaba/DataX

hdfswriter-小文件问题

Open
#215 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Java
Stars
17.4k
Forks
5.7k
PR merge metrics
No merged PRs in 30d

Description

使用hdfswriter,如果源表按照字段切分之后(例如mysql主键),channel很多的情况下,hdfswriter可能会写非常多的小文件,datax是否有考虑这个问题?有没有解决方案呢?

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by tracing how hdfswriter maps field-based source-table splits, such as MySQL primary-key splits, across many channels and produces HDFS files. Reproduce the case with many channels and inspect the resulting small-file count; done requires an agreed solution that addresses excessive small files.

Written by the indexing model from the issue text.

Assessment

Tech stack
hadoop, java, mysql
Domain
data-engineering, distributed-systems
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.