AbsaOSS / AbsaOSS/cobrix

NOT DIVISIBLE by the RECORD SIZE

Offen
#182 7 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen
Vorherrschende Sprache
Scala
Sterne
170
Forks
96
Ø Merge
57 Min.
Gemergte PRs (30 T.)
2

Beschreibung

Hi @yruslan ,

I am trying to read a file with the following command:
`spark.read.format("cobol").option("copybook", file://BOOK.txt").load("DATASET.dat").count()`

The error output is:

> ERROR FileUtils$: File hdfs://NHA/user/big/DATASET.dat IS NOT divisible by 200.
> java.lang.IllegalArgumentException: There are some files in DATASET.dat that are NOT DIVISIBLE by the RECORD SIZE calculated from the copybook (200 bytes per record). Check the logs for the names of the files.
>

After trying some options, the following worked:
`spark.read.format("cobol").option("copybook", "file://BOOK.txt").option("file_start_offset", "1").load("DATASET.dat ").count()`

But the line count of the file shows only 20.314.408 lines while the expected value is 38.694.112 lines and the column values are completely wrong.

Thanks for your help !

Beitragsleitfaden

Für dieses Repository ist kein Beitragsleitfaden indexiert

Bewertung

Dieses Issue wurde noch nicht bewertet.

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.