bigscience-workshop / bigscience-workshop/lam

Add dataset: swiss_federal_council_handwritten_text_recognition

Open
#64 3 comments 1 reaction 1 assignee Claimed by @loleg View on GitHub
dataset
Dominant language
No language data
Stars
91
Forks
8
PR merge metrics
No merged PRs in 30d

Description

### A URL for this dataset

https://doi.org/10.5281/zenodo.4746342

### Dataset description

> This data set is a test set generated to test the capabilities of engines for Optical Character Recognition and Handwritten Text Recognition.
> The data set consists of extracts of the minutes of the Swiss Federal Council. The single lines have been randomly chosen from about 150'000 pages of handwritten minutes.
> For each line, an image file is being provided by the Swiss Federal Archives/Schweizerisches Bundesarchiv [images.tar.gz]. Please cite the images as follows: Excerpts of BAR E1004.1#1000/9#1-215. The images are in the public domain.
> A PageXML file [page.zip] accompanies every image file and indicates the transcription and coordinates of the line.

### Dataset modality

Mixed

### Dataset licence

Creative Commons Attribution 4.0 International

### Other licence

_No response_

### How can you access this data

As a download from a repository/website

### Confirm the dataset has an open licence

- [X] To the best of my knowledge, this dataset is accessible via an open licence

### Contact details for data custodian

_No response_

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.