acl-org / acl-org/acl-anthology

Screen for mismatched PDFs

未关闭
#7,061 1 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
主要语言
Python
星标
797
派生
408
平均合并
3 天 19 小时
30 天内合并 PR
36

描述

Occasionally the PDFs become mismatched with the paper metadata. Not sure at what point in the pipeline this happens. #7060 is an example.

Could the ingestion pipeline extract text from the PDFs to flag these cases? Suggested by @mbollmann

贡献指南

这个仓库没有索引到贡献指南

调研方向

Start by comparing the mismatched example in issue #7060 and tracing where the ingestion pipeline handles PDFs and paper metadata. Define how extracted text should be compared and how cases should be flagged; done means the pipeline reliably identifies mismatches without disrupting ingestion.

由索引模型根据 Issue 内容生成。

评估

技术栈
python
领域
data-engineering
Issue 类型
功能
难度
5/5
预计耗时
一周以上
活跃度
停滞
描述清晰度
需要澄清
新手友好度
25/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。