apache / apache/cloudberry

datalake_fdw: Iceberg name mapping for data files without field ids

Open
#1,989 0 comments 0 reactions 1 assignee Claimed by @MisterRaindrop View on GitHub
Dominant language
C
Stars
1.4k
Forks
247
Avg merge
4d 3h
Merged PRs (30d)
39

Description

### Summary

Since #1951 the Parquet reader in `contrib/datalake_fdw` matches a table's columns to a file's by Iceberg field id (`PARQUET:field_id` in the Parquet schema). A file whose columns carry no field id can never be matched: every projected column reads as NULL. That is correct per spec for a file the table never wrote, and wrong for the one case the spec covers: tables created over existing Parquet data (`add_files`, migrated Hive tables), whose files predate the ids.

### What the spec says

Iceberg resolves such files through the table property `schema.name-mapping.default`: a JSON mapping from field ids to the column names to look for in the file, including nested names and multiple names per id (for renamed columns). A reader applies it only to columns that have no field id.

### What has to happen

- The metadata engine has to surface the property to the access method.
- `ProjectionSet` (`format/format.h`) needs a way to hand the reader a name mapping alongside the field ids -- a second array of names per id, or a pointer to a parsed mapping.
- `parquet_read.cpp`'s `parquet_project()` matches by id first and falls back to the mapping for id-less columns. Nothing changes in the decoder.
- A file with neither ids nor a mapping match stays NULL, as now.

Deferred from #1951 on purpose: the framework there is settled, and the property needs the metadata engine, which is a later PR.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.