bigscience-workshop / bigscience-workshop/lam
Add dataset: TexBiG
- Dominant language
- No language data
- Stars
- 91
- Forks
- 8
- PR merge metrics
- No merged PRs in 30d
Description
### A URL for this dataset
https://zenodo.org/record/6885144
### Dataset description
> TexBiG (from the German Text-Bild-Gefüge, meaning Text-Image-Structure) is a document layout analysis dataset for historical documents in the late 19th and early 20th century. The dataset provides instance segmentation (bounding boxes and polygons/masks) annotations for 19 different classes with more then 52.000 instances. Annotations are manually annotated by experts and evaluated with Krippendorff's Alpha, for each document image are least two different annotators have labeled the document. Further details can be found in the Paper.
### Dataset modality
Mixed
### Dataset licence
Creative Commons Attribution 4.0 International
### Other licence
_No response_
### How can you access this data
As a download from a repository/website
### size of dataset
>10GB
### Confirm the dataset has an open licence
- [X] To the best of my knowledge, this dataset is accessible via an open licence
### Contact details for data custodian
_No response_
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.