hypothesis / hypothesis/product-backlog

Capture DOIs and other metadata from text of PDFs

Open
#13 3 comments 0 reactions 0 assignees View on GitHub
feature request
Dominant language
No language data
Stars
122
Forks
7
PR merge metrics
No merged PRs in 30d

Description

## If creating a new Feature Request use this form, and delete the Bug Report Form above

---

| Fields | Your Response |
| --- | --- |
| **Link to Zendesk ticket:** | https://groups.google.com/a/list.hypothes.is/forum/?utm_medium=email&utm_source=footer#!msg/dev/fXPfcQLAwuA/ZW6S07tgDgAJ |
| **User Name/Company:** | Austrian Academie of Science |
| **Is this from a client (paid user):** | No |
| **What problem is the user trying to solve?** | "I'm currently searching for a solution to annotate a Digital Object Identifier (DOI) for the Austrian Academie of Science.". The need, AIUI, is that the user could annotate one version of a paper and have the annotations appear when viewing a different version of the paper, provided it had the same DOI. It also sounds as if the user may want to search for annotations based on the DOI of the paper. |
| **What is the feature the user is requesting?** | They want a way to associate annotations made on PDFs with the DOI associated with the PDF |

When annotating a web page, the Hypothesis client captures metadata from the page and includes that with the annotation saved to the service. This information is used by the service to establish whether two different URLs refer to the same content. We can also use this information to enable the user to search for annotations based on the metadata of the document that was annotated (eg. title, author, identifier).

The request here is to capture similar information from PDFs.

Contributor guide

No contributing guide indexed for this repository

Research direction

No files, tests, or entry points are named. Start by clarifying which PDF metadata must be captured, including DOI, title, author, and identifier, and how annotations should match or be searched across document versions; done means the required metadata-based association and search behavior are defined.

Written by the indexing model from the issue text.

Assessment

Domain
backend
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.