apache / apache/datafusion-python

Unable to Read Parquet Files from S3 Bucket

Aperta
#638 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub
bug
Lingua principale
Python
Stelle
604
Fork
174
Merge medio
1g 7h
PR unite (30g)
4

Descrizione

escription:
I'm attempting to read Parquet files from an S3 bucket using DataFusion in Python. Below is the code snippet I'm using:

python
Copy code
import datafusion
from datafusion import SessionContext

s3 = object_store.object_store.AmazonS3("s3://test/", "us-east-2")
ctx = SessionContext()
ctx.register_object_store("s3", s3)
df = ctx.read_parquet("s3://test/00001.parquet")
Error Encountered:
I'm encountering the following error:

css
Copy code
dataFusion error: Internal("No suitable object store found for s3://test/00001.parquet")
Issue Investigation:
I've tried to find relevant documentation or support resources but haven't been successful in locating any.

Resources Reviewed:
While researching, I came across the following Rust documentation which appears relevant but unfortunately doesn't have a corresponding Python counterpart:

[DataFusion Rust Documentation](https://arrow.apache.org/datafusion-python/generated/datafusion.object_store.object_store.html)
[DataFusion ObjectStore S3 Rust Documentation](https://docs.rs/datafusion-objectstore-s3/latest/datafusion_objectstore_s3/)
Request for Assistance:
Could someone please guide me on how to resolve this issue in Python? Any assistance or pointers to relevant documentation would be greatly appreciated. Thank you!

Guida per i contributori

Nessuna guida per i contributori indicizzata per questo repository

Valutazione

Questa issue non è ancora stata valutata.

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.