bigscience-workshop / bigscience-workshop/data_tooling

Create dataset iarpa_babel_swahili_language_pack

Open Beginner friendly
#128 1 comment 0 reactions 0 assignees View on GitHub
data catalog need custodian permission need data sourcing feedback
Dominant language
HTML
Stars
91
Forks
47
PR merge metrics
No merged PRs in 30d

Description

- uid: iarpa_babel_swahili_language_pack
- type: processed
- description:
- name: IARPA Babel Swahili Language Pack
- description: Swahili ASR Dataset. Official description says "IARPA Babel Swahili Language Pack IARPA-babel202b-v1.0d was developed by Appen for the IARPA (Intelligence Advanced Research Projects Activity) Babel program. It contains approximately 350 hours of Swahili conversational and scripted telephone speech collected from 2012-2014 along with corresponding transcripts."
- homepage: https://doi.org/10.35111/afrp-a637
- validated: True
- languages:
- language_names:
- Niger-Congo
- Swahili
- language_comments:
- language_locations:
- Eastern Africa
- Kenya
- validated: False
- custodian:
- name:
- in_catalogue: linguistic_data_consortium_ldc
- type:
- location:
- contact_name:
- contact_email:
- contact_submitter: False
- additional:
- validated: False
- availability:
- procurement:
- for_download: Yes - after signing a user agreement
- download_url: https://doi.org/10.35111/afrp-a637
- download_email:
- licensing:
- has_licenses: Yes
- license_text: Multiple licenses:
* For-profit use: https://catalog.ldc.upenn.edu/license/iarpa-babel-swahili-agreement-for-profit.pdf
* Non-member: https://catalog.ldc.upenn.edu/license/iarpa-babel-swahili-agreement-non-member.pdf
* Not-for-profit: https://catalog.ldc.upenn.edu/license/iarpa-babel-swahili-agreement-not-for-profit.pdf
- license_properties:
- multiple licenses
- research use
- non-commercial use
- copyright - all rights reserved
- license_list:
- other: Other license
- pii:
- has_pii: Unclear
- generic_pii_likely:
- generic_pii_list:
- numeric_pii_likely:
- numeric_pii_list:
- sensitive_pii_likely:
- sensitive_pii_list:
- no_pii_justification_class: general knowledge not written by or referring to private persons
- no_pii_justification_text:
- validated: False
- processed_from_primary:
- from_primary: Taken from primary source
- primary_availability: No - the dataset curators kept the source data secret
- primary_license:
- primary_types:
- validated: False
- media:
- category:
- audiovisual
- text_format:
- audiovisual_format:
- .WAV
- image_format:
- database_format:
- text_is_transcribed:
- instance_type: 350 hours of telephone conversations, no clue how long each file is. Estimating 15 minutes each, and 350 words per minute?
- instance_count: 10K

Contributor guide

No contributing guide indexed for this repository

Research direction

Start with the requested dataset metadata in this issue and the referenced IARPA Babel Swahili Language Pack source. Add the entry as iarpa_babel_swahili_language_pack.json using the repository’s existing dataset-entry conventions; done means the dataset is represented with its source, availability, licensing, language, and media details.

Written by the indexing model from the issue text.

Assessment

Domain
data-engineering
Issue type
Feature
Difficulty
2/5
Estimated time
1-3 hours
Activity status
Stale
Clarity
Clearly specified
Newbie friendliness
62/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.