alibaba / alibaba/DataX

Hive表字段有大量null,识别字段数不同

Open
#306 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Java
Stars
17.4k
Forks
5.7k
PR merge metrics
No merged PRs in 30d

Description

在使用datax过程中`hive同步到RDS`过程,hive有大量null字段,在导出时,有三十多个字段,显示只有9个字段,导致数据导出失败。

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reproducing the Hive-to-RDS synchronization described in the issue, using a Hive table with many null fields and more than thirty columns. Inspect how DataX identifies the source fields and compare that result with the destination schema; done means all fields are recognized consistently and the export no longer fails.

Written by the indexing model from the issue text.

Assessment

Tech stack
java
Domain
data-engineering, databases
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.