acl-org / acl-org/acl-anthology
Extract abstracts from PDF
- Dominant language
- Python
- Stars
- 797
- Forks
- 408
- Avg merge
- 3d 19h
- Merged PRs (30d)
- 36
Description
The anthology currently only shows the abstracts if there is an authoritative version in the XML. It would be nice if we could scrape the PDF using some off-the-shelf software to extract the abstracts and dump them into a different file (to not tamper with handcrafted information). Having an abstract on the web pages makes quickly searching through literature much faster.
Contributor guide
No contributing guide indexed for this repository
Research direction
The issue names no files, tests, or entry points; begin by tracing the anthology's XML/PDF ingestion and web-page generation paths. Done means abstracts are extracted from PDFs into a separate data file and appear on web pages without changing handcrafted information.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- data
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 25/100