internetarchive / internetarchive/openlibrary

Extract ToC from Deutsche Nationalbibliothek

Open
#10,119 0 comments 0 reactions 0 assignees View on GitHub
Lead: @cdrini Module: Table of Contents Needs: Breakdown Priority: 3 Type: Proposal
Dominant language
Python
Stars
6.7k
Forks
2k
Avg merge
2d 19h
Merged PRs (30d)
138

Description

### Proposal

Deutsche Nationalbibliothek has high-quality scans of the table of contents for a large part of their holdings. These can be freely accessed as a PDF. For each OL edition that has a DNB identifier attached, OL could attempt to download the corresponding PDF and extract the ToC text. Note that some of the PDFs have a wildly inaccurate text layer, so it makes sense to run our own OCR.

Example edition page at DNB:
https://d-nb.info/973546166

Example TOC:
https://d-nb.info/973546166/04

See also #8756

### Justification

Problem: OL currently only has a table of contents for a small fraction of editions. This impacts patrons’ ability to learn what a book is about.

Impact: Increase the number of ToCs, especially for German-language books.

Research: I’ve been manually OCRing and/or transcribing a number of TOCs from DNB for use on OL and can attest that the scans are of consistently high quality.

### Breakdown

#### Requirements Checklist

* [ ]

#### Related files

*

#### Stakeholders

*


#### Instructions for Contributors

Please [run these commands](https://github.com/internetarchive/openlibrary/wiki/Git-Cheat-Sheet#working-on-your-branch) to ensure your repository is up to date **before** [creating a new branch](https://github.com/internetarchive/openlibrary/wiki/Git-Cheat-Sheet#making-changes-and-creating-a-pull-request) to work on this issue and **each time after** pushing code to Github, because the pre-commit bot may add commits to your PRs upstream.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.