datalake_fdw: timestamp columns in units other than microseconds
- Dominant language
- C
- Stars
- 1.4k
- Forks
- 247
- Avg merge
- 4d 3h
- Merged PRs (30d)
- 39
Description
### Summary
The Parquet reader in `contrib/datalake_fdw` (#1951) reads only microsecond timestamps (`tsu:`), which is what Iceberg defines and what this module writes. Files written by other systems carry other units, and today every one of them is refused with `a column stored as Arrow type "tsm:..." cannot be read as timestamp`.
### Cases
- **Millisecond columns** (`TIMESTAMP_MILLIS`): common in Parquet written by Spark with `spark.sql.parquet.outputTimestampType=TIMESTAMP_MILLIS`, by Hive, and by many ETL tools. Multiplying by 1000 loses nothing. The question is whether a lake table should read a file whose type is not the table's type; Iceberg's spec says data files carry the table's types, so accepting them is a lenience, not a requirement.
- **Nanosecond columns** (`TIMESTAMP_NANOS`, Iceberg v3 `timestamp_ns`): dividing by 1000 truncates. Refuse, or truncate and say so.
- **INT96** is already handled by coercing to microseconds. One caveat, Arrow's rather than ours (reproduced with pyarrow 21 and the same setting): Arrow's microsecond conversion assumes the nanos-of-day half is non-negative, which Spark/Hive/Impala guarantee; pyarrow's deprecated INT96 writer stores a negative one for instants before 1970 and those read wrong. Coercing to nanoseconds instead would fix that one case and break every date outside 1677..2262, including the 9999-12-31 sentinels warehouses keep. Worth an upstream report.
### Where
`format/arrow_decode.c`: `dl_arrow_decode_check()` decides what a `timestamp` column accepts, `dl_arrow_decode_value()` converts. The `TIMESTAMP`/`TIMESTAMPTZ` case already distinguishes zoned from unzoned by whether the format string names a zone.
Deferred from #1951 on purpose.
Contributor guide
Assessment
This issue has not been assessed yet.