datalake_fdw: NUMERIC/DECIMAL in the Parquet format layer
- Dominant language
- C
- Stars
- 1.4k
- Forks
- 247
- Avg merge
- 4d 3h
- Merged PRs (30d)
- 39
Description
### Summary
`contrib/datalake_fdw`'s Parquet format layer (#1951) refuses `numeric` columns at `CREATE TABLE ... USING iceberg`. DECIMAL is the most common column type in real lake tables, so this is the first type to add.
### What has to be decided
- **Storage form.** Parquet stores DECIMAL four ways (INT32, INT64, FIXED_LEN_BYTE_ARRAY, BYTE_ARRAY). Iceberg's spec fixes the writer to `decimal(P,S)` with P <= 38 backed by fixed-length bytes; the reader has to accept all four.
- **Unconstrained `numeric`.** `atttypmod = -1` has no precision or scale. Computing `((typmod - 4) >> 16) & 65535` without checking gives precision 65535 / scale 65531, so bare `numeric` columns silently match nothing (the trap in lithium-tech/tea `validate.cpp:59-61`). Iceberg needs P and S, so the likely answer is to refuse bare `numeric` at `CREATE TABLE` and require `numeric(P,S)`.
- **Precision above 38.** PostgreSQL allows up to 1000; `decimal128` caps at 38. Refuse at `CREATE TABLE`.
- **Conversion.** `NumericVar` <-> two's-complement int128, the way tea's `bridge.cpp:269-277` and `numeric_var.cpp` do it, with `NaN`/`Infinity` refused on write.
### Where
- `format/format_types.h` / `arrow_support.cpp`: `dl_format_type_refusal()` decides what `CREATE TABLE` accepts; the mapping to `arrow::decimal128(P, S)` goes next to the other types.
- `arrow_builder.cpp` / `arrow_decode.c`: append and decode. The decoder reads the Arrow C data interface directly; the format string is `d:P,S`.
- Tests in `test/automation/sqlrepo/smoke/format_parquet/`, plus `CREATE TABLE` refusals in `iceberg_am_reject.sql`.
Deferred from #1951 on purpose: the framework there is settled, the type work is separate.
Contributor guide
Assessment
This issue has not been assessed yet.