apache / apache/iceberg-python

[Feature Request] Add pyiceberg.catalog.hadoop.HadoopCatalog (filesystem-only catalog)

Đang mở
#3,897 0 bình luận 0 reaction 0 người được giao Xem trên GitHub
Ngôn ngữ chính
Python
Star
1.1k
Fork
581
Merge trung bình
1 ngày 17 giờ
Pull request đã merge (30 ngày)
77

Mô tả

## Is your feature request related to a problem? Please describe.

Java Iceberg ships a filesystem-only `HadoopCatalog` / `HadoopTables`, where table metadata lives under `/.db//metadata/` with no external metastore. PyIceberg currently has no equivalent — the available catalog types are `rest / hive / glue / dynamodb / sql / in-memory / bigquery`.

This gap matters in two ways:

1. **Interop with Java-side HadoopCatalog tables.** Tables created by Java `HadoopCatalog` (common in lightweight deployments without a metastore) cannot be opened through any supported PyIceberg catalog. Users must fall back to `StaticTable.from_metadata` and resolve the latest `metadata.json` themselves, which loses catalog semantics (no namespace listing, no create/commit).

2. **Downstream projects already assume the module exists.** Daft's Gravitino integration imports `from pyiceberg.catalog.hadoop import HadoopCatalog` and calls `HadoopCatalog("gravitino_reader", props).load_table(table_dir)` to open a table from a storage location (`daft/catalog/__gravitino/_catalog.py`). Against PyIceberg 0.11.x this raises `TypeError: HadoopCatalog.__init__() takes 2 positional arguments but 3 were given`, and against versions without the module it fails at import time.

## Describe the solution you'd like

A `pyiceberg.catalog.hadoop.HadoopCatalog` (subclassing `MetastoreCatalog`) implementing Java HadoopCatalog semantics:

- `warehouse` property as the root location
- table dir = `//`
- metadata at `/metadata/v{n}.metadata.json` plus a `version-hint.text` holding the current version
- latest-version resolution: read `version-hint.text`, fall back to scanning `metadata/` for the max `v{n}` (matching Java behavior)
- namespace/table create/list/commit driven purely by the warehouse filesystem (no metastore calls)

## Additional context / pitfalls observed while prototyping

Happy to contribute a PR if this is in scope. A few notes from an internal prototype:

1. **Metadata file naming.** Java HadoopCatalog uses `v{n}.metadata.json` + `version-hint.text`, but tables created by JDBC/REST catalogs use `00000-.metadata.json` with no version-hint. To open those as well, the scan fallback should accept both patterns (`v(\d+)\.metadata\.json` and `\d{5}-.*\.metadata\.json`), or at least document the limitation.

2. **Filesystem abstraction.** `__init__` should derive the filesystem from the catalog's FileIO (`PyArrowFileIO`) instead of hardcoding `pyarrow.fs.HadoopFileSystem.from_uri(warehouse)`. The JVM-backed `HadoopFileSystem` only supports `hdfs://` and fails for object-store schemes (`s3://`, and custom schemes), so routing through FileIO keeps it scheme-agnostic.

3. **Atomicity.** `create_table` / `commit_table` use create-if-absent on `v{n}.metadata.json` for optimistic concurrency — safe on HDFS but not atomic on plain object stores (S3 has no create-if-absent guarantee). Java has the same caveat; worth documenting or using a conditional-write primitive where available.

## Adapting a custom storage scheme (Tencent Cloud TBDSFS as a concrete case)

A related gap surfaced while prototyping against Tencent Cloud TBDS's distributed filesystem scheme `tbdsfs:///` (exposed by a Python client, plus a JVM `fs.tbdsfs.impl`):

- `PyArrowFileIO._initialize_fs(scheme, netloc)` only understands a fixed set of schemes (`hdfs / s3 / gs / file / abfs / ...`), so any `tbdsfs://...` location raises `ValueError: Unrecognized filesystem type in URI: tbdsfs`.
- There is no public, documented way to plug in a custom filesystem. The only workaround today is monkey-patching a private method:

1. Implement a `pyarrow.fs.FileSystemHandler` subclass wrapping the TBDSFS Python client, wrap it in `pyarrow.fs.PyFileSystem`, then patch `PyArrowFileIO._initialize_fs` to return that filesystem when `scheme == "tbdsfs"`.
2. pyarrow 21's `PyFileSystem` callback also has non-obvious contracts any custom handler must satisfy: single paths arrive as one-element lists; `get_file_info` must return a one-element list; and for scheme'd URIs the netloc is prepended into the path (e.g. `internal/usr/...`, without a leading slash).

This works, but it depends on patching a private API (`_initialize_fs`), which is brittle across PyIceberg releases.

**Suggested improvement:** a documented, public extension point for registering an arbitrary `pyarrow.fs.FileSystem` (or a custom `FileIO`) per scheme — e.g. a `register_file_system(scheme, factory)` helper, or a `scheme → FileSystem` mapping read from FileIO/catalog properties — so non-standard object stores and filesystems can be integrated without touching internals. This would also naturally address the `HadoopCatalog.__init__` filesystem-abstraction point above.

## References

- Java: `org.apache.iceberg.hadoop.HadoopCatalog` / `HadoopTables`
- Downstream usage that currently breaks: `daft/catalog/__gravitino/_catalog.py` → `_open_iceberg_table`

Hướng dẫn đóng góp

Chưa lập chỉ mục được hướng dẫn đóng góp cho kho mã nguồn này

Hướng nghiên cứu

Bắt đầu bằng cách đọc các triển khai catalog hiện có và MetastoreCatalog, sau đó kiểm tra PyArrowFileIO._initialize_fs cùng cách sử dụng downstream của daft/catalog/__gravitino/__catalog.py được nêu trong issue. So sánh hành vi được yêu cầu với các tham chiếu HadoopCatalog và HadoopTables của Java Iceberg. Được xem là hoàn tất khi các thao tác namespace và bảng chỉ sử dụng filesystem, việc phân giải metadata tương thích với Java, và có một đường dẫn được ghi chép hoặc công khai cho các scheme filesystem tùy chỉnh.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Đánh giá

Công nghệ
java, python
Lĩnh vực
data-engineering, databases
Loại issue
Tính năng
Độ khó
5/5
Thời gian dự kiến
Hơn một tuần
Mức độ hoạt động
Sôi nổi
Độ rõ ràng
Khá rõ ràng
Mức phù hợp với người mới
42/100

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.