hypothesis / hypothesis/product-backlog
Capture DOIs and other metadata from text of PDFs
- Dominant language
- No language data
- Stars
- 122
- Forks
- 7
- PR merge metrics
- No merged PRs in 30d
Description
## If creating a new Feature Request use this form, and delete the Bug Report Form above
---
| Fields | Your Response |
| --- | --- |
| **Link to Zendesk ticket:** | https://groups.google.com/a/list.hypothes.is/forum/?utm_medium=email&utm_source=footer#!msg/dev/fXPfcQLAwuA/ZW6S07tgDgAJ |
| **User Name/Company:** | Austrian Academie of Science |
| **Is this from a client (paid user):** | No |
| **What problem is the user trying to solve?** | "I'm currently searching for a solution to annotate a Digital Object Identifier (DOI) for the Austrian Academie of Science.". The need, AIUI, is that the user could annotate one version of a paper and have the annotations appear when viewing a different version of the paper, provided it had the same DOI. It also sounds as if the user may want to search for annotations based on the DOI of the paper. |
| **What is the feature the user is requesting?** | They want a way to associate annotations made on PDFs with the DOI associated with the PDF |
When annotating a web page, the Hypothesis client captures metadata from the page and includes that with the annotation saved to the service. This information is used by the service to establish whether two different URLs refer to the same content. We can also use this information to enable the user to search for annotations based on the metadata of the document that was annotated (eg. title, author, identifier).
The request here is to capture similar information from PDFs.
Contributor guide
No contributing guide indexed for this repository
Research direction
No files, tests, or entry points are named. Start by clarifying which PDF metadata must be captured, including DOI, title, author, and identifier, and how annotations should match or be searched across document versions; done means the required metadata-based association and search behavior are defined.
Written by the indexing model from the issue text.
Assessment
- Domain
- backend
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 35/100