apache / apache/iceberg-python

[Feature Request] Add pyiceberg.catalog.hadoop.HadoopCatalog (filesystem-only catalog)

Aperta
#3,897 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub
Lingua principale
Python
Stelle
1.1k
Fork
581
Merge medio
1g 17h
PR unite (30g)
78

Descrizione

## Is your feature request related to a problem? Please describe.

Java Iceberg ships a filesystem-only `HadoopCatalog` / `HadoopTables`, where table metadata lives under `/.db//metadata/` with no external metastore. PyIceberg currently has no equivalent — the available catalog types are `rest / hive / glue / dynamodb / sql / in-memory / bigquery`.

This gap matters in two ways:

1. **Interop with Java-side HadoopCatalog tables.** Tables created by Java `HadoopCatalog` (common in lightweight deployments without a metastore) cannot be opened through any supported PyIceberg catalog. Users must fall back to `StaticTable.from_metadata` and resolve the latest `metadata.json` themselves, which loses catalog semantics (no namespace listing, no create/commit).

2. **Downstream projects already assume the module exists.** Daft's Gravitino integration imports `from pyiceberg.catalog.hadoop import HadoopCatalog` and calls `HadoopCatalog("gravitino_reader", props).load_table(table_dir)` to open a table from a storage location (`daft/catalog/__gravitino/_catalog.py`). Against PyIceberg 0.11.x this raises `TypeError: HadoopCatalog.__init__() takes 2 positional arguments but 3 were given`, and against versions without the module it fails at import time.

## Describe the solution you'd like

A `pyiceberg.catalog.hadoop.HadoopCatalog` (subclassing `MetastoreCatalog`) implementing Java HadoopCatalog semantics:

- `warehouse` property as the root location
- table dir = `//`
- metadata at `/metadata/v{n}.metadata.json` plus a `version-hint.text` holding the current version
- latest-version resolution: read `version-hint.text`, fall back to scanning `metadata/` for the max `v{n}` (matching Java behavior)
- namespace/table create/list/commit driven purely by the warehouse filesystem (no metastore calls)

## Additional context / pitfalls observed while prototyping

Happy to contribute a PR if this is in scope. A few notes from an internal prototype:

1. **Metadata file naming.** Java HadoopCatalog uses `v{n}.metadata.json` + `version-hint.text`, but tables created by JDBC/REST catalogs use `00000-.metadata.json` with no version-hint. To open those as well, the scan fallback should accept both patterns (`v(\d+)\.metadata\.json` and `\d{5}-.*\.metadata\.json`), or at least document the limitation.

2. **Filesystem abstraction.** `__init__` should derive the filesystem from the catalog's FileIO (`PyArrowFileIO`) instead of hardcoding `pyarrow.fs.HadoopFileSystem.from_uri(warehouse)`. The JVM-backed `HadoopFileSystem` only supports `hdfs://` and fails for object-store schemes (`s3://`, and custom schemes), so routing through FileIO keeps it scheme-agnostic.

3. **Atomicity.** `create_table` / `commit_table` use create-if-absent on `v{n}.metadata.json` for optimistic concurrency — safe on HDFS but not atomic on plain object stores (S3 has no create-if-absent guarantee). Java has the same caveat; worth documenting or using a conditional-write primitive where available.

## Adapting a custom storage scheme (Tencent Cloud TBDSFS as a concrete case)

A related gap surfaced while prototyping against Tencent Cloud TBDS's distributed filesystem scheme `tbdsfs:///` (exposed by a Python client, plus a JVM `fs.tbdsfs.impl`):

- `PyArrowFileIO._initialize_fs(scheme, netloc)` only understands a fixed set of schemes (`hdfs / s3 / gs / file / abfs / ...`), so any `tbdsfs://...` location raises `ValueError: Unrecognized filesystem type in URI: tbdsfs`.
- There is no public, documented way to plug in a custom filesystem. The only workaround today is monkey-patching a private method:

1. Implement a `pyarrow.fs.FileSystemHandler` subclass wrapping the TBDSFS Python client, wrap it in `pyarrow.fs.PyFileSystem`, then patch `PyArrowFileIO._initialize_fs` to return that filesystem when `scheme == "tbdsfs"`.
2. pyarrow 21's `PyFileSystem` callback also has non-obvious contracts any custom handler must satisfy: single paths arrive as one-element lists; `get_file_info` must return a one-element list; and for scheme'd URIs the netloc is prepended into the path (e.g. `internal/usr/...`, without a leading slash).

This works, but it depends on patching a private API (`_initialize_fs`), which is brittle across PyIceberg releases.

**Suggested improvement:** a documented, public extension point for registering an arbitrary `pyarrow.fs.FileSystem` (or a custom `FileIO`) per scheme — e.g. a `register_file_system(scheme, factory)` helper, or a `scheme → FileSystem` mapping read from FileIO/catalog properties — so non-standard object stores and filesystems can be integrated without touching internals. This would also naturally address the `HadoopCatalog.__init__` filesystem-abstraction point above.

## References

- Java: `org.apache.iceberg.hadoop.HadoopCatalog` / `HadoopTables`
- Downstream usage that currently breaks: `daft/catalog/__gravitino/_catalog.py` → `_open_iceberg_table`

Guida per i contributori

Nessuna guida per i contributori indicizzata per questo repository

Direzione di ricerca

Start by reading the existing catalog implementations and MetastoreCatalog, then inspect PyArrowFileIO._initialize_fs and the downstream daft/catalog/__gravitino/__catalog.py usage named in the issue. Compare the requested behavior with Java Iceberg's HadoopCatalog and HadoopTables references. Done means filesystem-only namespace and table operations, Java-compatible metadata resolution, and a documented or public path for custom filesystem schemes.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Stack tecnologico
java, python
Ambito
data-engineering, databases
Tipo di issue
Funzionalità
Difficoltà
5/5
Tempo stimato
Più di una settimana
Stato di attività
Attiva
Chiarezza
Abbastanza chiara
Idoneità per principianti
42/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.