bigscience-workshop / bigscience-workshop/lam

Add dataset: beyond_words_coco

Open
#37 0 comments 0 reactions 0 assignees View on GitHub
dataset
Dominant language
No language data
Stars
91
Forks
8
PR merge metrics
No merged PRs in 30d

Description

### A URL for this dataset

https://github.com/LibraryOfCongress/newspaper-navigator/tree/master/beyond_words_data

### Dataset description

This is a dataset containing crowdsourced annotastions of the bounding boxes of different types of 'visual content in historic US newspapers. It can be used to train object detection models to extract visual content from historic newspaper collections. Breakdown of annotations:

Category | # in Full Dataset
-- | --
Photograph | 4,254
Illustration | 1,048
Map | 215
Comics/Cartoon | 1,150
Editorial Cartoon | 293
Headline | 27,868
Advertisement | 13,581
Total | 48,409

### Dataset modality

Image

### Dataset licence

Other license

### Other licence

https://chroniclingamerica.loc.gov/about/

### How can you access this data

As a download from a repository/website

### Confirm the dataset has an open licence

- [X] To the best of my knowledge, this dataset is accessible via an open licence

### Contact details for data custodian

_No response_

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.