acl-org / acl-org/acl-anthology
Screen for mismatched PDFs
未关闭
- 主要语言
- Python
- 星标
- 797
- 派生
- 408
- 平均合并
- 3 天 19 小时
- 30 天内合并 PR
- 36
描述
Occasionally the PDFs become mismatched with the paper metadata. Not sure at what point in the pipeline this happens. #7060 is an example.
Could the ingestion pipeline extract text from the PDFs to flag these cases? Suggested by @mbollmann
贡献指南
这个仓库没有索引到贡献指南
调研方向
Start by comparing the mismatched example in issue #7060 and tracing where the ingestion pipeline handles PDFs and paper metadata. Define how extracted text should be compared and how cases should be flagged; done means the pipeline reliably identifies mismatches without disrupting ingestion.
由索引模型根据 Issue 内容生成。
评估
- 技术栈
- python
- 领域
- data-engineering
- Issue 类型
- 功能
- 难度
- 5/5
- 预计耗时
- 一周以上
- 活跃度
- 停滞
- 描述清晰度
- 需要澄清
- 新手友好度
- 25/100