apache / apache/arrow

Allow ConvertOptions.timestamp_parsers for date types

Open
#33,357 5 comments 1 reaction 0 assignees View on GitHub
Component: Python Type: enhancement
Dominant language
C++
Stars
17.1k
Forks
4.3k
Avg merge
3d 13h
Merged PRs (30d)
88

Description

Currently, the timestamp_parsers option of the CSV reader only works for timestamp datatypes.

If one wants to immediately read dates as date32 objects (in my use case csv data is read and stored as parquet files with correct types), one has to cast the table to a schema with date32 types after the fact.
This snipped shows that loading the data fails when specifying the date type:
```java

import pyarrow as pa
from pyarrow import csv

def open_bytes(b, **kwargs):
    return csv.open_csv(pa.py_buffer(b), **kwargs)
def read_bytes(b, **kwargs):
    return open_bytes(b, **kwargs).read_all()

rows = b"a,b\n1970/01/01,1980-01-01 00\n1970/01/02,1980-01-02 00\n"
schema = pa.schema([("a", pa.timestamp("ms")), ("b", pa.string())])
opts = csv.ConvertOptions(column_types=schema, timestamp_parsers=["%Y/%m/%d"])
table = read_bytes(rows, convert_options=opts)
assert table.schema == schema # works

schema = pa.schema([("a", pa.date32()), ("b", pa.string())])
opts = csv.ConvertOptions(column_types=schema, timestamp_parsers=["%Y/%m/%d"])
table = read_bytes(rows, convert_options=opts) # error here
assert table.schema == schema
---------------------------------------------------------------------------
ArrowInvalid                              Traceback (most recent call last)
Input In [134], in ()
     20 schema = pa.schema([("a", pa.date32()), ("b", pa.string())])
     21 opts = csv.ConvertOptions(column_types=schema, timestamp_parsers=["%Y/%m/%d"])
---> 22 table = read_bytes(rows, convert_options=opts)
     23 assert table.schema == schemaInput In [134], in read_bytes(b, **kwargs)
      9 def read_bytes(b, **kwargs):
---> 10     return open_bytes(b, **kwargs).read_all()Input In [134], in open_bytes(b, **kwargs)
      5 def open_bytes(b, **kwargs):
----> 6     return csv.open_csv(pa.py_buffer(b), **kwargs)File ~/.virtualenvs/ogi/lib/python3.9/site-packages/pyarrow/_csv.pyx:1273, in pyarrow._csv.open_csv()File ~/.virtualenvs/ogi/lib/python3.9/site-packages/pyarrow/_csv.pyx:1137, in pyarrow._csv.CSVStreamingReader._open()File ~/.virtualenvs/ogi/lib/python3.9/site-packages/pyarrow/error.pxi:144, in pyarrow.lib.pyarrow_internal_check_status()File ~/.virtualenvs/ogi/lib/python3.9/site-packages/pyarrow/error.pxi:100, in pyarrow.lib.check_status()ArrowInvalid: In CSV column #0: CSV conversion error to date32[day]: invalid value '1970/01/01'
```
It would be useful to allow the timestamp_parsers for date types as well (or add an analogous argument for dates), such that such errors don't occur and the resulting table has the required datatypes without a casting step.

 

A little bit more context is in the comments of https://issues.apache.org/jira/browse/ARROW-10848 (26/Oct/22).

**Reporter**: [Tim Loderhose](https://issues.apache.org/jira/browse/ARROW-18166)
**Watchers**: [Rok Mihevc](https://issues.apache.org/jira/browse/ARROW-18166) / @rok
#### Related issues:
- [[C++] Strptime issues umbrella](https://github.com/apache/arrow/issues/31324) (is a child of)

**Note**: *This issue was originally created as [ARROW-18166](https://issues.apache.org/jira/browse/ARROW-18166). Please see the [migration documentation](https://github.com/apache/arrow/issues/14542) for further details.*

Contributor guide

Open the contributing guide

Research direction

Start at the CSV reader entry point exposed by csv.open_csv and the ConvertOptions.timestamp_parsers handling, then trace conversion for date32 alongside timestamp types. Use the reproducer in the issue as the initial check; done means the custom date format parses successfully and the resulting table retains the requested date32 schema.

Written by the indexing model from the issue text.

Assessment

Tech stack
cpp, python
Domain
data, data-engineering
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.