alibaba / alibaba/DataX

hdfswriter orc file,数据丢失

Open
#1,345 2 comments 0 reactions 0 assignees View on GitHub
Dominant language
Java
Stars
17.4k
Forks
5.7k
PR merge metrics
No merged PRs in 30d

Description

hdfs加载组件加载数据时,选择orc格式,抽取的记录数有一千八百万行左右,但是落地orc格式文件后,使用load data将orc格式加载至目标表时,数据变少了,大约有三百万左右

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reproducing the HDFS extraction using ORC output and the subsequent load-data import, comparing record counts at each step. Done means identifying and correcting the loss of roughly 15 million records so the target table contains the extracted total.

Written by the indexing model from the issue text.

Assessment

Tech stack
java
Domain
data-engineering
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.