apache / apache/datafusion-python

Unable to Read Parquet Files from S3 Bucket

Abierto
#638 0 comentarios 0 reacciones 0 asignados Ver en GitHub
bug
Lenguaje dominante
Python
Estrellas
604
Forks
174
Merge medio
1 d 7 h
PR fusionados (30 d)
4

Descripción

escription:
I'm attempting to read Parquet files from an S3 bucket using DataFusion in Python. Below is the code snippet I'm using:

python
Copy code
import datafusion
from datafusion import SessionContext

s3 = object_store.object_store.AmazonS3("s3://test/", "us-east-2")
ctx = SessionContext()
ctx.register_object_store("s3", s3)
df = ctx.read_parquet("s3://test/00001.parquet")
Error Encountered:
I'm encountering the following error:

css
Copy code
dataFusion error: Internal("No suitable object store found for s3://test/00001.parquet")
Issue Investigation:
I've tried to find relevant documentation or support resources but haven't been successful in locating any.

Resources Reviewed:
While researching, I came across the following Rust documentation which appears relevant but unfortunately doesn't have a corresponding Python counterpart:

[DataFusion Rust Documentation](https://arrow.apache.org/datafusion-python/generated/datafusion.object_store.object_store.html)
[DataFusion ObjectStore S3 Rust Documentation](https://docs.rs/datafusion-objectstore-s3/latest/datafusion_objectstore_s3/)
Request for Assistance:
Could someone please guide me on how to resolve this issue in Python? Any assistance or pointers to relevant documentation would be greatly appreciated. Thank you!

Guía de contribución

No hay ninguna guía de contribución indexada para este repositorio

Evaluación

Este issue todavía no se ha evaluado.

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.