bigscience-workshop / bigscience-workshop/lam

Add dataset: odeuropa_benchmarks_and_corpora

Open
#54 1 comment 0 reactions 0 assignees View on GitHub
candidate-dataset
Dominant language
No language data
Stars
91
Forks
8
PR merge metrics
No merged PRs in 30d

Description

### A URL for this dataset

https://github.com/Odeuropa/benchmarks_and_corpora

### Dataset description

This dataset
> contains the annotations related to olfactory information from the benchmark created for the ODEUROPA project.
> For 7 languages we selected a pool of documents covering different time periods (from 1620 to 1925) and topics (e.g. medicine, law, literature).

This offers an exciting dataset of annotations related to olfactory (smell) information in historical documents. The dataset is interesting because it covers a range of periods but also offers the possibility of utilising ml for a different task than standard entity recognition tasks.

### Dataset modality

Text

### Dataset licence

Other license

### Other licence

_No response_

### How can you access this data

As a download from a repository/website

### Confirm the dataset has an open licence

- [X] To the best of my knowledge, this dataset is accessible via an open licence

### Contact details for data custodian

_No response_

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.