apache / apache/arrow

[Python] write_to_dataset, coerce_timestamps issue with stored binary schema

Open
#45,062 0 comments 0 reactions 0 assignees View on GitHub
Component: Python Type: bug
Dominant language
C++
Stars
17.1k
Forks
4.3k
Avg merge
3d 13h
Merged PRs (30d)
88

Description

### Describe the bug, including details regarding any error messages, version, and platform.

When using coerce_timestamps=us, it seems that the parquet metadata are correclty being set as datetime[us] however, the stored binary arrow schema seems to still be a datetime[ns] creating mismatch of type depending on the engine you use to read the data.

```python
import pandas as pd
import polars as pl
from pyarrow.parquet import ParquetFile, ParquetWriter

df = pd.DataFrame({"date": [pd.Timestamp.now()]})
df.to_parquet("us.parquet", coerce_timestamps="us", allow_truncated_timestamps=True, index=False)

pqf = ParquetFile("us.parquet")
writer = ParquetWriter("us_pyarrow.parquet", schema=pqf.schema.to_arrow_schema())
writer.write_table(pqf.read())
writer.close()

pl.read_parquet("us.parquet") # Gives datetime[ns]
pl.read_parquet("us_pyarrow.parquet") # Gives datetime[us]

ParquetFile("us.parquet").metadata.schema.to_arrow_schema() # gives datetime[us]
ParquetFile("us_pyarrow.parquet").metadata.schema.to_arrow_schema() # gives datetime[us]
```

Polars is probably leveraging the binary arrow schema stored in the metadata and therefore interpret the column as datetime[ns].

Running the following prevent the mismatch when using polars and we indeed get datetime[us]

```python
df.to_parquet("us.parquet", coerce_timestamps="us", allow_truncated_timestamps=True, index=False, store_schema=False)
```

The issue is that store_schema is not supported in the `write_to_dataset` function as this parameter is not available in `ParquetFileWriteOptions` but only in the `write_table` function and `ParquetWriter` class.

Therefore, running the following doesn't work.

```python
df.to_parquet("us.parquet", coerce_timestamps="us", allow_truncated_timestamps=True, index=False, store_schema=False, partition_cols=[])
```

I guess the stored binary arrow schema should be stored with datetime[us] instead of datetime[ns] when using `coerce_timestamps` parameter but the `store_schema` parameter should also be made available to the `write_to_dataset` function

### Component(s)

Python

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.