acl-org / acl-org/acl-anthology
Extract abstracts from PDF
Open
enhancement
help wanted
- Dominant language
- Python
- Stars
- 796
- Forks
- 408
- Avg merge
- 3d 13h
- Merged PRs (30d)
- 34
Description
The anthology currently only shows the abstracts if there is an authoritative version in the XML. It would be nice if we could scrape the PDF using some off-the-shelf software to extract the abstracts and dump them into a different file (to not tamper with handcrafted information). Having an abstract on the web pages makes quickly searching through literature much faster.
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.