Make convert-csv work with the input filename which starts with a period or an numeric
- Vorherrschende Sprache
- Java
- Sterne
- 3.1k
- Forks
- 1.6k
- Ø Merge
- 3 T. 12 Std.
- Gemergte PRs (30 T.)
- 33
Beschreibung
I ran parquet-cli's `convert-csv` with an input file which name starts with a numeric character without `--schema` option and got the following error:
```Java
$ java -cp 'target/*:target/dependency/*' org.apache.parquet.cli.Main convert-csv 0sample.csv -o sample.parquet
Unknown error
shaded.parquet.org.apache.avro.SchemaParseException: Illegal initial character: 0sample
at shaded.parquet.org.apache.avro.Schema.validateName(Schema.java:1498)
at shaded.parquet.org.apache.avro.Schema.access$200(Schema.java:86)
at shaded.parquet.org.apache.avro.Schema$Name.(Schema.java:645)
at shaded.parquet.org.apache.avro.Schema.createRecord(Schema.java:182)
at shaded.parquet.org.apache.avro.SchemaBuilder$RecordBuilder.fields(SchemaBuilder.java:1805)
at org.apache.parquet.cli.csv.AvroCSV.inferSchemaInternal(AvroCSV.java:158)
at org.apache.parquet.cli.csv.AvroCSV.inferNullableSchema(AvroCSV.java:78)
at org.apache.parquet.cli.commands.ConvertCSVCommand.run(ConvertCSVCommand.java:160)
at org.apache.parquet.cli.Main.run(Main.java:147)
at org.apache.hadoop.util.ToolRunner.run(ToolRunner.java:70)
at org.apache.parquet.cli.Main.main(Main.java:177)
```
This is because that `convert-csv` uses the input file name as the name for the output schema, while Avro requires its schema name to match the regex pattern `[A-Za-z_][A-Za-z0-9_]*`.
So users have to change the input file name or use the `--schema` option explicitly, but it's not so obvious from the error message.
It'd be nice if the message were improved, or the schema name were automatically replaced with valid characters to avoid this problem.
**Reporter**: [Kengo Seki](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=sekikn) / @sekikn
**Assignee**: [Kengo Seki](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=sekikn) / @sekikn
#### PRs and other links:
- [GitHub Pull Request #652](https://github.com/apache/parquet-mr/pull/652)
**Note**: *This issue was originally created as [PARQUET-1598](https://issues.apache.org/jira/browse/PARQUET-1598). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*
Beitragsleitfaden
Für dieses Repository ist kein Beitragsleitfaden indexiert
Rechercherichtung
Beginne mit org.apache.parquet.cli.csv.AvroCSV.inferSchemaInternal und verwende den Stacktrace sowie ConvertCSVCommand.java als Einstiegspunkte. Prüfe Pull Request #652, um festzustellen, ob die beabsichtigte Lösung eine klarere Fehlermeldung oder die automatische Handhabung von Schemanamen ist, und füge Tests hinzu, die zeigen, dass convert-csv Dateinamen, die wie spezifiziert mit einem Punkt oder einer Zahl beginnen, verarbeitet.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Bewertung
- Tech-Stack
- java
- Bereich
- cli, data
- Issue-Typ
- Bug
- Schwierigkeit
- 3/5
- Geschätzter Aufwand
- 1-2 Tage
- Aktivitätsstatus
- Veraltet
- Klarheit
- Größtenteils klar
- Anfängerfreundlichkeit
- 25/100