AlibabaResearch / AlibabaResearch/AdvancedLiterateMachinery

Question about extracting labels of PDF Elements

未关闭
#218 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
主要语言
C++
星标
1.8k
派生
195
PR 合并指标
30 天内没有已合并 PR

描述

First of all, thank you for making these models available—great work!

I have tried several AI models that extract content from PDFs and identify its type—e.g.,

- text
- title
- list
- table
- figure.

The problem is that I haven’t yet found a model that correctly recognizes the hierarchy of headings, such as H1, H2, and H3. Can any of your models do that? So what I need looking for is a way to detect

- text
- title
- list
- table
- figure.
- H1
- H2
- H3
- H4

Is it possible with one of your model?

贡献指南

这个仓库没有索引到贡献指南

评估

这个 Issue 还没有评估数据。

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。