bigscience-workshop / bigscience-workshop/data_tooling

Create dataset iarpa_babel_swahili_language_pack

Open
#128 1 comment 0 reactions 0 assignees View on GitHub
data catalog need custodian permission need data sourcing feedback
Dominant language
HTML
Stars
91
Forks
47
PR merge metrics
No merged PRs in 30d

Description

- uid: iarpa_babel_swahili_language_pack
- type: processed
- description:
- name: IARPA Babel Swahili Language Pack
- description: Swahili ASR Dataset. Official description says "IARPA Babel Swahili Language Pack IARPA-babel202b-v1.0d was developed by Appen for the IARPA (Intelligence Advanced Research Projects Activity) Babel program. It contains approximately 350 hours of Swahili conversational and scripted telephone speech collected from 2012-2014 along with corresponding transcripts."
- homepage: https://doi.org/10.35111/afrp-a637
- validated: True
- languages:
- language_names:
- Niger-Congo
- Swahili
- language_comments:
- language_locations:
- Eastern Africa
- Kenya
- validated: False
- custodian:
- name:
- in_catalogue: linguistic_data_consortium_ldc
- type:
- location:
- contact_name:
- contact_email:
- contact_submitter: False
- additional:
- validated: False
- availability:
- procurement:
- for_download: Yes - after signing a user agreement
- download_url: https://doi.org/10.35111/afrp-a637
- download_email:
- licensing:
- has_licenses: Yes
- license_text: Multiple licenses:
* For-profit use: https://catalog.ldc.upenn.edu/license/iarpa-babel-swahili-agreement-for-profit.pdf
* Non-member: https://catalog.ldc.upenn.edu/license/iarpa-babel-swahili-agreement-non-member.pdf
* Not-for-profit: https://catalog.ldc.upenn.edu/license/iarpa-babel-swahili-agreement-not-for-profit.pdf
- license_properties:
- multiple licenses
- research use
- non-commercial use
- copyright - all rights reserved
- license_list:
- other: Other license
- pii:
- has_pii: Unclear
- generic_pii_likely:
- generic_pii_list:
- numeric_pii_likely:
- numeric_pii_list:
- sensitive_pii_likely:
- sensitive_pii_list:
- no_pii_justification_class: general knowledge not written by or referring to private persons
- no_pii_justification_text:
- validated: False
- processed_from_primary:
- from_primary: Taken from primary source
- primary_availability: No - the dataset curators kept the source data secret
- primary_license:
- primary_types:
- validated: False
- media:
- category:
- audiovisual
- text_format:
- audiovisual_format:
- .WAV
- image_format:
- database_format:
- text_is_transcribed:
- instance_type: 350 hours of telephone conversations, no clue how long each file is. Estimating 15 minutes each, and 350 words per minute?
- instance_count: 10K

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.