bigscience-workshop / bigscience-workshop/data_tooling
Create dataset iarpa_babel_swahili_language_pack
- Dominant language
- HTML
- Stars
- 91
- Forks
- 47
- PR merge metrics
- No merged PRs in 30d
Description
- uid: iarpa_babel_swahili_language_pack
- type: processed
- description:
- name: IARPA Babel Swahili Language Pack
- description: Swahili ASR Dataset. Official description says "IARPA Babel Swahili Language Pack IARPA-babel202b-v1.0d was developed by Appen for the IARPA (Intelligence Advanced Research Projects Activity) Babel program. It contains approximately 350 hours of Swahili conversational and scripted telephone speech collected from 2012-2014 along with corresponding transcripts."
- homepage: https://doi.org/10.35111/afrp-a637
- validated: True
- languages:
- language_names:
- Niger-Congo
- Swahili
- language_comments:
- language_locations:
- Eastern Africa
- Kenya
- validated: False
- custodian:
- name:
- in_catalogue: linguistic_data_consortium_ldc
- type:
- location:
- contact_name:
- contact_email:
- contact_submitter: False
- additional:
- validated: False
- availability:
- procurement:
- for_download: Yes - after signing a user agreement
- download_url: https://doi.org/10.35111/afrp-a637
- download_email:
- licensing:
- has_licenses: Yes
- license_text: Multiple licenses:
* For-profit use: https://catalog.ldc.upenn.edu/license/iarpa-babel-swahili-agreement-for-profit.pdf
* Non-member: https://catalog.ldc.upenn.edu/license/iarpa-babel-swahili-agreement-non-member.pdf
* Not-for-profit: https://catalog.ldc.upenn.edu/license/iarpa-babel-swahili-agreement-not-for-profit.pdf
- license_properties:
- multiple licenses
- research use
- non-commercial use
- copyright - all rights reserved
- license_list:
- other: Other license
- pii:
- has_pii: Unclear
- generic_pii_likely:
- generic_pii_list:
- numeric_pii_likely:
- numeric_pii_list:
- sensitive_pii_likely:
- sensitive_pii_list:
- no_pii_justification_class: general knowledge not written by or referring to private persons
- no_pii_justification_text:
- validated: False
- processed_from_primary:
- from_primary: Taken from primary source
- primary_availability: No - the dataset curators kept the source data secret
- primary_license:
- primary_types:
- validated: False
- media:
- category:
- audiovisual
- text_format:
- audiovisual_format:
- .WAV
- image_format:
- database_format:
- text_is_transcribed:
- instance_type: 350 hours of telephone conversations, no clue how long each file is. Estimating 15 minutes each, and 350 words per minute?
- instance_count: 10K
Contributor guide
No contributing guide indexed for this repository
Research direction
Start with the requested dataset metadata in this issue and the referenced IARPA Babel Swahili Language Pack source. Add the entry as iarpa_babel_swahili_language_pack.json using the repository’s existing dataset-entry conventions; done means the dataset is represented with its source, availability, licensing, language, and media details.
Written by the indexing model from the issue text.
Assessment
- Domain
- data-engineering
- Issue type
- Feature
- Difficulty
- 2/5
- Estimated time
- 1-3 hours
- Activity status
- Stale
- Clarity
- Clearly specified
- Newbie friendliness
- 62/100