bigscience-workshop / bigscience-workshop/lam

Add dataset: early_printed_books_font_detection

Open
#45 2 comments 0 reactions 1 assignee Claimed by @davanstrien View on GitHub
dataset good first issue
Dominant language
No language data
Stars
91
Forks
8
PR merge metrics
No merged PRs in 30d

Description

### A URL for this dataset

https://zenodo.org/record/3366686

### Dataset description

>This dataset is composed of photos of various resolution of 35'623 pages of printed books dating from the 15th to the 18th century. Each page has been attributed by experts from one to five labels corresponding to the font groups used in the text, with two extra-classes for non-textual content and fonts not present in the following list: Antiqua, Bastarda, Fraktur, Gotico Antiqua, Greek, Hebrew, Italic, Rotunda, Schwabacher, and Textura.

This dataset offers an image classification dataset that has potential implications for other downstream tasks such as OCR recognition.

A related paper [Dataset of Pages from Early Printed Books with Multiple Font Groups](https://dl.acm.org/doi/10.1145/3352631.3352640)

### Dataset modality

Image

### Dataset licence

Creative Commons Attribution Non Commercial Share Alike 4.0 International

### Other licence

_No response_

### How can you access this data

As a download from a repository/website

### Confirm the dataset has an open licence

- [X] To the best of my knowledge, this dataset is accessible via an open licence

### Contact details for data custodian

_No response_

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.