[Python][Parquet] Improve usability of ParquetDataset
- Lingua principale
- C++
- Stelle
- 17.1k
- Fork
- 4.3k
- Merge medio
- 3g 18h
- PR unite (30g)
- 91
Descrizione
In e.g. `pd.read_parquet` implementation, we currently wrap `pq.read_table`. However, I would like to use the dataset implementation to get access to some more features (e.g. reading only the first few rows, defining a filter _after_ opening (discovering the schema) the dataset).
That runs into some issues:
- `pq.read_table()` uses ParquetDataset under the hood, but still has a fallback to plain `ParquetFile().read()` when `pyarrow.dataset` module is not available. While this will definitely be uncommon, I have no clue how important it is to keep supporting this (also on the pandas side, since this is currently a "feature" of `pandas.read_parquet` by means of using `pq.read_table`).
- I would prefer using `pq.ParquetDataset` over `pyarrow.dataset`, because the parquet version is compatible with `pq.read_table` and does all the translation to the `pyarrow.dataset` API for us (constructing the file format object, passing the various keywords in the correct place, mapping some naming differences, etc)
- But `ParquetDataset` is also limited, and currently I am essentially using it as a constructor to then access the underlying `._dataset` (the `pyarrow.dataset` dataset object)
Some ideas that I was having:
- Add a new `pq.open_dataset()`-like function that also has the boilerplate to construct the `pyarrow.dataset` object, but returns that instead of the `ParquetDataset` wrapper
- Would it be technically possible to let `pq.ParquetDataset` inherit from `pyarrow.dataset.Dataset`, so you get those methods that way, while keeping it back-compat
- Add new methods to `pq.ParquetDataset` (`to_table()`, `head()`, etc) to make it look more like a `pyarrow.dataset.Dataset`, but without actually inheriting from it
- Simply "officially" expose the underlying dataset, so that I don't have to use the private `_dataset` attribute
Guida per i contributori
Apri la guida per i contributori
Direzione di ricerca
Inizia esaminando l’integrazione di pandas.read_parquet e pq.read_table, quindi analizza pq.ParquetDataset e il relativo _dataset sottostante insieme all’API pyarrow.dataset. Confronta la funzione open_dataset proposta, l’ereditarietà, i metodi aggiunti e l’esposizione pubblica del dataset; il lavoro è completo quando viene selezionata e documentata una direzione di usabilità compatibile con un’adeguata copertura dell’API.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Valutazione
- Stack tecnologico
- python
- Ambito
- data-engineering
- Tipo di issue
- Funzionalità
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Stato di attività
- Attiva
- Chiarezza
- Da chiarire
- Idoneità per principianti
- 25/100