apache / apache/beam

Implement a CSV file reader

Open
#17,960 0 comments 0 reactions 0 assignees View on GitHub
ideas io new feature P3
Dominant language
Java
Stars
8.7k
Forks
4.7k
Avg merge
2d 2h
Merged PRs (30d)
205

Description

We should implement a CSV-based source.

One possibility would be to support the same options as BigQuery. https://cloud.google.com/bigquery/preparing-data-for-bigquery#dataformats These options are:

fieldDelimiter: allowing a custom delimiter... csv vs tsv, etc. My guess is this is critical. One common delimiter that people use is 'thorn' (þ).

quote: Custom quote char. By default, this is '"', but this allows users to set it to something else, or, perhaps more commonly, remove it entirely (by setting it to the empty string). For example, tab-separated files generally don't need quotes.

allowQuotedNewlines: whether you can quote newlines. In the official CSV RFC, newlines can be quoted.. that is, you can have "a", "b\n", "c" in a single line. This makes splitting of large csv files impossible, so we should disallow quoted newlines by default unless the user really wants them (in which case, they'll get worse performance).

allowJaggedRows: This allows inferring null if not enough columns are specified. Otherwise we give an error for the row.

ignoreUnknownValues: The opposite of allowJaggedRows, this means that if a user has _too_ many values for the schema, we will ignore the ones we don't recognize, rather than reporting an error for the row.

skipHeaderRows: How many header lines are in the file.

encoding: UTF8-vs latin1, etc.
compression: gzip, bzip, etc.

Imported from Jira [BEAM-51](https://issues.apache.org/jira/browse/BEAM-51). Original Jira may contain additional context.
Reported by: dhalperi.

Contributor guide

Open the contributing guide

Research direction

The issue names no files, tests, or entry points. Start by reading the linked BigQuery data-format options and the original BEAM-51 Jira context; done means implementing a CSV source with an agreed, testable subset of the listed options.

Written by the indexing model from the issue text.

Assessment

Domain
data-engineering
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.