apache / apache/arrow

Reading parquet files with timestamp column containing 9999-12-23 23:59:59 yields 1816-03-22 05:56:07.066277376

Open
#44,112 2 comments 0 reactions 0 assignees View on GitHub
Component: Python Type: bug
Dominant language
C++
Stars
17.1k
Forks
4.3k
Avg merge
3d 13h
Merged PRs (30d)
88

Description

### Describe the bug, including details regarding any error messages, version, and platform.

I've stumbled upon a weird issue, where I don't get the underlying isse.

I read a parquet file which contains a timestamp column. This timestamp column contains the value `9999-12-23 23:59:59`. When I read this file using `pyarrow` (or with `pandas` and `pyarrow` engine an dtype_backend), the rows with `9999-12-23 23:59:59` show the value `1816-03-22 05:56:07.066277376`.

I'm pretty certain that `9999-12-23 23:59:59` is the correct value, because this is much more plausible (and that's what `duckdb` and `Impala` say as well).

When I write the respective row to parquet using `duckdb` and read this file using `pyarrow`, I get the correct value of `9999-12-23 23:59:59`.

I've already checked if this is a problem with the parquet version, but both files are version `1.0`. What else might cause this?

Unfortunately, I can't share the parquet file in question because it contains confidential data.

### Component(s)

Python

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.