AbsaOSS / AbsaOSS/cobrix

NOT DIVISIBLE by the RECORD SIZE

未关闭
#182 7 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
主要语言
Scala
星标
170
派生
96
平均合并
57 分钟
30 天内合并 PR
2

描述

Hi @yruslan ,

I am trying to read a file with the following command:
`spark.read.format("cobol").option("copybook", file://BOOK.txt").load("DATASET.dat").count()`

The error output is:

> ERROR FileUtils$: File hdfs://NHA/user/big/DATASET.dat IS NOT divisible by 200.
> java.lang.IllegalArgumentException: There are some files in DATASET.dat that are NOT DIVISIBLE by the RECORD SIZE calculated from the copybook (200 bytes per record). Check the logs for the names of the files.
>

After trying some options, the following worked:
`spark.read.format("cobol").option("copybook", "file://BOOK.txt").option("file_start_offset", "1").load("DATASET.dat ").count()`

But the line count of the file shows only 20.314.408 lines while the expected value is 38.694.112 lines and the column values are completely wrong.

Thanks for your help !

贡献指南

这个仓库没有索引到贡献指南

评估

这个 Issue 还没有评估数据。

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。