apache / apache/arrow

[Python] Investigate test_read_multiple_files TODO

Aperta
#39,338 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub
Component: Python
Lingua principale
C++
Stelle
17.1k
Fork
4.3k
Merge medio
3g 23h
PR unite (30g)
101

Descrizione

### Describe the bug, including details regarding any error messages, version, and platform.

When removing `ParquetDatased` legacy code path a test from `python/pyarrow/tests/parquet/test_dataset.py` had to be updated and there was a TODO connected to the "Dataset API skipping bad files" we kept as a comment, see [comment](https://github.com/apache/arrow/pull/39112/files/0cbd03dcbafb76853c0eabb08d74795afd77cc4a#r1431129290). This todo should be looked into and a test for it should be added as a follow up to https://github.com/apache/arrow/pull/39112.

### Component(s)

Python

Guida per i contributori

Apri la guida per i contributori

Direzione di ricerca

Inizia in python/pyarrow/tests/parquet/test_dataset.py, al TODO di test_read_multiple_files, quindi leggi la pull request collegata e il commento della review per comprendere il comportamento previsto dell'API Dataset. Aggiungi un test che verifichi come vengono gestiti i file non validi, con il risultato atteso definito da quel contesto.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Stack tecnologico
python
Ambito
testing
Tipo di issue
Bug
Difficoltà
3/5
Tempo stimato
1-2 giorni
Stato di attività
Ferma
Chiarezza
Abbastanza chiara
Idoneità per principianti
45/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.