bigscience-workshop / bigscience-workshop/lam

Add dataset: nb-tale

Open
#48 1 comment 0 reactions 0 assignees View on GitHub
dataset
Dominant language
No language data
Stars
91
Forks
8
PR merge metrics
No merged PRs in 30d

Description

### A URL for this dataset

https://www.nb.no/sprakbanken/ressurskatalog/oai-nb-no-sbr-31/

### Dataset description

NB Tale is a basic acoustic-phonetic speech database for Norwegian. The database contains recordings of 380 speakers from 24 different dialect areas. The database is produced for the National Library of Norway by Lingit AS.

The data are divided into three modules:

1) Manuscript-read speech from 240 native speakers of Norwegian, contains audio recordings, phonotypical annotation, informant data and documentation.

2) Extension of module 1, 140 new speakers (20 native speakers, 120 with Norwegian as foreign language), contains audio recordings, phonotypical annotation, informant data and documentation.

3) Recordings of spontaneous speech from speakers in module 1 and 2, on average around 2 minutes each, contains audio recordings, orthographic annotations (sentence-level) and documentation.

### Dataset modality

Audio

### Dataset licence

Creative Commons Zero v1.0 Universal

### Other licence

_No response_

### How can you access this data

As a download from a repository/website

### Confirm the dataset has an open licence

- [X] To the best of my knowledge, this dataset is accessible via an open licence

### Contact details for data custodian

sprakbanken@nb.no

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.