bigscience-workshop / bigscience-workshop/lam

Add dataset: TexBiG

Open
#84 0 comments 0 reactions 0 assignees View on GitHub
dataset
Dominant language
No language data
Stars
91
Forks
8
PR merge metrics
No merged PRs in 30d

Description

### A URL for this dataset

https://zenodo.org/record/6885144

### Dataset description

> TexBiG (from the German Text-Bild-Gefüge, meaning Text-Image-Structure) is a document layout analysis dataset for historical documents in the late 19th and early 20th century. The dataset provides instance segmentation (bounding boxes and polygons/masks) annotations for 19 different classes with more then 52.000 instances. Annotations are manually annotated by experts and evaluated with Krippendorff's Alpha, for each document image are least two different annotators have labeled the document. Further details can be found in the Paper.

### Dataset modality

Mixed

### Dataset licence

Creative Commons Attribution 4.0 International

### Other licence

_No response_

### How can you access this data

As a download from a repository/website

### size of dataset

>10GB

### Confirm the dataset has an open licence

- [X] To the best of my knowledge, this dataset is accessible via an open licence

### Contact details for data custodian

_No response_

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.