acl-org / acl-org/acl-anthology

Extract abstracts from PDF

Open
#395 28 comments 2 reactions 1 assignee Claimed by @mbollmann View on GitHub
enhancement help wanted
Dominant language
Python
Stars
797
Forks
408
Avg merge
3d 19h
Merged PRs (30d)
36

Description

The anthology currently only shows the abstracts if there is an authoritative version in the XML. It would be nice if we could scrape the PDF using some off-the-shelf software to extract the abstracts and dump them into a different file (to not tamper with handcrafted information). Having an abstract on the web pages makes quickly searching through literature much faster.

Contributor guide

No contributing guide indexed for this repository

Research direction

The issue names no files, tests, or entry points; begin by tracing the anthology's XML/PDF ingestion and web-page generation paths. Done means abstracts are extracted from PDFs into a separate data file and appear on web pages without changing handcrafted information.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
data
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.