apache / apache/arrow

[Python][Parquet] pyarrow.parquet.write_table ignores version and coerce_timestamps arguments for Time64 fields

Open
#50,142 10 comments 0 reactions 0 assignees View on GitHub
Component: Parquet Component: Python Type: bug
Dominant language
C++
Stars
17.1k
Forks
4.3k
Avg merge
3d 13h
Merged PRs (30d)
88

Description

### Describe the bug, including details regarding any error messages, version, and platform.

**OS**: Ubuntu 24.04 WSL
**Python**: 3.12.3
**pyarrow**: 24.0.0

Suppose we have a pyarrow table with a nanosecond resolution time column:

``` python
import pyarrow as pa
t = pa.table({'A': [40984E9]}, schema=pa.schema({'A': pa.time64('ns')}))
```

I want to save this to an older Parquet version, to support older readers such as Azure Synapse, which [doesn't support nanosecond time resolution](https://learn.microsoft.com/en-us/azure/synapse-analytics/sql/develop-openrowset#type-mapping-for-parquet).

According to the [pyarrow.parquet.write_table docs](https://arrow.apache.org/docs/python/generated/pyarrow.parquet.write_table.html), if `coerce_timestamps` is not specified:

> defaults are chosen depending on version. For version='1.0' and version='2.4', nanoseconds are cast to microseconds (‘us’)

So, I should be able to save this to either version '1.0' or version '2.4', and it should automatically cast nanoseconds to microseconds:

## Parquet 2.4

``` python
pa.parquet.write_table(t, 'test-pyarrow-2.4.parquet', version='2.4')
```

However, inspecting the metadata, we can immediately see that the `version` argument has been ignored:

``` python
pa.parquet.ParquetFile('test-pyarrow-2.4.parquet').metadata
#
# created_by: parquet-cpp-arrow version 24.0.0
# num_columns: 1
# num_rows: 1
# num_row_groups: 1
# format_version: 2.6
# serialized_size: 383
```

Additionally, we can see that the time field has _not_ been coerced to microseconds:

``` python
pa.parquet.ParquetFile('test-pyarrow-2.4.parquet').schema
#
# required group field_id=-1 schema {
# optional int64 field_id=-1 A (Time(isAdjustedToUTC=false, timeUnit=nanoseconds));
# }
```

If we add the (optional) argument `coerce_timestamps='us'`, the effect is the same:

``` python
pa.parquet.write_table(t, 'test-pyarrow-2.4-coerce.parquet', version='2.4', coerce_timestamps='us')
pa.parquet.ParquetFile('test-pyarrow-2.4-coerce.parquet').metadata
#
# created_by: parquet-cpp-arrow version 24.0.0
# num_columns: 1
# num_rows: 1
# num_row_groups: 1
# format_version: 2.6
# serialized_size: 383

pa.parquet.ParquetFile('test-pyarrow-2.4-coerce.parquet').schema
#
# required group field_id=-1 schema {
# optional int64 field_id=-1 A (Time(isAdjustedToUTC=false, timeUnit=nanoseconds));
# }
```

## Parquet 1.0

If we try setting the Parquet version to '1.0', we are similarly unable to coerce the timestamps to microseconds:

``` python
pa.parquet.write_table(t, 'test-pyarrow-1.0.parquet', version='1.0')
pa.parquet.ParquetFile('test-pyarrow-1.0.parquet').metadata
#
# created_by: parquet-cpp-arrow version 24.0.0
# num_columns: 1
# num_rows: 1
# num_row_groups: 1
# format_version: 1.0
# serialized_size: 382

pa.parquet.ParquetFile('test-pyarrow-1.0.parquet').schema
#
# required group field_id=-1 schema {
# optional int64 field_id=-1 A (Time(isAdjustedToUTC=false, timeUnit=nanoseconds));
# }
```

This time, we at least get the correct Parquet version. However, the time field still has nanosecond resolution.
Similarly, the `coerce_timestamps='us'` argument has the same result.

### Component(s)

Python, Parquet

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.