NVIDIA / NVIDIA/cudf

[FEA] ORC writer: allow setting writerTimezone instead of hardcoding "UTC"

Open
#23,422 5 comments 0 reactions 1 assignee Claimed by @vuule View on GitHub
feature request
Dominant language
C++
Stars
9.8k
Forks
1.1k
Avg merge
3d 6m
Merged PRs (30d)
278

Description

## Summary

libcudf's ORC writer never records the actual writer timezone — it hardcodes `writerTimezone="UTC"` at `cpp/src/io/orc/writer_impl.cu:2686` with no writer-side option. The file therefore misdeclares its own timezone, and any spec-conforming ORC reader in a non-UTC zone is guaranteed to read it shifted.

Note this is **not** a missing field. The field is present and set to `"UTC"`, which is precisely why Apache ORC Java *actively* applies an `off(UTC) - off(reader)` correction rather than skipping conversion.

## The asymmetry

The reader options already expose timezone control; the writer options expose none.

| | timezone control |
|---|---|
| `orc_reader_options` | :white_check_mark: `ignore_timezone_in_stripe_footer` (`cpp/include/cudf/io/orc.hpp`) |
| `orc_writer_options` | :x: nothing |

That gap is the entire issue. Everything below follows from it.

## Why it breaks interop

ORC's plain `timestamp` is a wall-clock type; readers apply `off(writerTZ) - off(readerTZ)`. Hive/Spark carry instants through it by tagging the footer with the writing JVM's zone so the correction cancels. libcudf stores the identical values but cannot tag anything but UTC.

Writer zone `A` (offset `a`), reader zone `B` (offset `b`), true instant `I`:

| writer | stored `V` | footer | ORC Java reads | vs Spark CPU |
|---|---|---|---|---|
| Spark CPU | `I` | `A` | `I + a - b` | — |
| **libcudf** | `I` | `"UTC"` | `I - b` | **`-a`** |

The stored values are byte-identical. Only the footer differs, and the entire error is `-a`.

Measured with `A = B = Asia/Shanghai` (`a = b = 8h`), input epoch `0s, 1s, 2s`:

```
CPU-written, read by Spark CPU: 0 1000000 2000000
libcudf-written, read by Spark CPU: -28800000000 -28799000000 -28798000000
```

Exactly `-a`. Full C++ repro and Java reader harness in https://github.com/rapidsai/cudf/issues/23422#issuecomment-5068203734.

Testing with C++ tooling alone will **not** reproduce this — liborc and cudf's own reader share cudf's assumption. The divergence only appears with the ORC **Java** reader, which is what Spark, Hive, and Trino use.

## Requested fix

Add a `writerTimezone` setter to `orc_writer_options` / its builder, defaulting to the current `"UTC"` so existing behavior is unchanged.

That makes libcudf output byte-identical to the Spark CPU writer. The two alternatives are both worse:

- **Emit `TIMESTAMP_INSTANT` instead.** Semantically cleaner, but Spark writes plain `TIMESTAMP` for `TimestampType` (`OrcUtils.scala:379-382`), so GPU files would carry a different schema than CPU files for identical input.
- **Pre-shift values caller-side.** Produces `V = I + a` under a footer still claiming `UTC` — wrong values *and* wrong metadata, correct only for readers in `A`.

## Current cost

NVIDIA/cudf-spark#14544 has to **disable GPU ORC timestamp writes entirely in any non-UTC JVM** and fall back to CPU, purely because of this one unsettable field:

```scala
if (types.exists(GpuOverrides.isOrContainsTimestamp) &&
!GpuOverrides.isUTCTimezone(ZoneId.systemDefault())) {
meta.willNotWorkOnGpu("Writing ORC timestamps is only supported in the UTC timezone ...")
}
```

The read path in that PR was fixed properly — it now reproduces `java.util.TimeZone` semantics for arbitrary writer/reader zone pairs. The write path cannot be fixed the same way from the caller side.

## Environment

cudf 25.12 · Apache ORC Java 1.9.1 (Spark 3.5.0) · JDK 17

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.