AlibabaResearch / AlibabaResearch/AdvancedLiterateMachinery

Question about extracting labels of PDF Elements

Aperta
#218 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub
Lingua principale
C++
Stelle
1.8k
Fork
195
Metriche di merge delle PR
Nessuna PR unita negli ultimi 30g

Descrizione

First of all, thank you for making these models available—great work!

I have tried several AI models that extract content from PDFs and identify its type—e.g.,

- text
- title
- list
- table
- figure.

The problem is that I haven’t yet found a model that correctly recognizes the hierarchy of headings, such as H1, H2, and H3. Can any of your models do that? So what I need looking for is a way to detect

- text
- title
- list
- table
- figure.
- H1
- H2
- H3
- H4

Is it possible with one of your model?

Guida per i contributori

Nessuna guida per i contributori indicizzata per questo repository

Valutazione

Questa issue non è ancora stata valutata.

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.