googleapis / googleapis/python-genai

Large PDF documents take too long to process

Aperta
#809 5 commenti 0 reazioni 1 assegnatario Rivendicata da @Venkaiahbabuneelam Vedi su GitHub
api: gemini-api priority: p2 type: bug
Lingua principale
Python
Stelle
4k
Fork
1k
Merge medio
2g 11h
PR unite (30g)
40

Descrizione

Although there's a limit of 1000 pages for a PDF document, it becomes very difficult to use it starting after 300 pages. It takes too long for the API to return an answer (>3 minutes). I tried caching it and that also didn't seem to decrease latency. I found that very weird, since including a large file at gemini.google.com and asking something is much faster, usually returning an answer in less than 30s. Why is it that a request containing a large PDF takes too long in the API? Any tips on how to optimize the speed?

#### Environment details

- Programming language: Python
- OS: MacOS
- Language runtime version: 3.11
- Package version: 1.12.1

#### Steps to reproduce

```
chunk_1 = gemini_client.files.get(name="large_doc_part_1.pdf")
chunk_2 = gemini_client.files.get(name="large_doc_part_2.pdf")
response = gemini_client.models.generate_content(
model="gemini-2.5-pro-preview-05-06",
contents=[chunk_1, chunk_2, "Tell me something interesting about this document"],
)
```

Guida per i contributori

Apri la guida per i contributori

Valutazione

Questa issue non è ancora stata valutata.

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.