bigscience-workshop / bigscience-workshop/data_tooling

Create dataset iarpa_babel_swahili_language_pack

未关闭 适合新手
#128 1 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
data catalog need custodian permission need data sourcing feedback
主要语言
HTML
星标
91
派生
47
PR 合并指标
30 天内没有已合并 PR

描述

- uid: iarpa_babel_swahili_language_pack
- type: processed
- description:
- name: IARPA Babel Swahili Language Pack
- description: Swahili ASR Dataset. Official description says "IARPA Babel Swahili Language Pack IARPA-babel202b-v1.0d was developed by Appen for the IARPA (Intelligence Advanced Research Projects Activity) Babel program. It contains approximately 350 hours of Swahili conversational and scripted telephone speech collected from 2012-2014 along with corresponding transcripts."
- homepage: https://doi.org/10.35111/afrp-a637
- validated: True
- languages:
- language_names:
- Niger-Congo
- Swahili
- language_comments:
- language_locations:
- Eastern Africa
- Kenya
- validated: False
- custodian:
- name:
- in_catalogue: linguistic_data_consortium_ldc
- type:
- location:
- contact_name:
- contact_email:
- contact_submitter: False
- additional:
- validated: False
- availability:
- procurement:
- for_download: Yes - after signing a user agreement
- download_url: https://doi.org/10.35111/afrp-a637
- download_email:
- licensing:
- has_licenses: Yes
- license_text: Multiple licenses:
* For-profit use: https://catalog.ldc.upenn.edu/license/iarpa-babel-swahili-agreement-for-profit.pdf
* Non-member: https://catalog.ldc.upenn.edu/license/iarpa-babel-swahili-agreement-non-member.pdf
* Not-for-profit: https://catalog.ldc.upenn.edu/license/iarpa-babel-swahili-agreement-not-for-profit.pdf
- license_properties:
- multiple licenses
- research use
- non-commercial use
- copyright - all rights reserved
- license_list:
- other: Other license
- pii:
- has_pii: Unclear
- generic_pii_likely:
- generic_pii_list:
- numeric_pii_likely:
- numeric_pii_list:
- sensitive_pii_likely:
- sensitive_pii_list:
- no_pii_justification_class: general knowledge not written by or referring to private persons
- no_pii_justification_text:
- validated: False
- processed_from_primary:
- from_primary: Taken from primary source
- primary_availability: No - the dataset curators kept the source data secret
- primary_license:
- primary_types:
- validated: False
- media:
- category:
- audiovisual
- text_format:
- audiovisual_format:
- .WAV
- image_format:
- database_format:
- text_is_transcribed:
- instance_type: 350 hours of telephone conversations, no clue how long each file is. Estimating 15 minutes each, and 350 words per minute?
- instance_count: 10K

贡献指南

这个仓库没有索引到贡献指南

调研方向

Start with the requested dataset metadata in this issue and the referenced IARPA Babel Swahili Language Pack source. Add the entry as iarpa_babel_swahili_language_pack.json using the repository’s existing dataset-entry conventions; done means the dataset is represented with its source, availability, licensing, language, and media details.

由索引模型根据 Issue 内容生成。

评估

领域
data-engineering
Issue 类型
功能
难度
2/5
预计耗时
1-3 小时
活跃度
停滞
描述清晰度
描述清楚
新手友好度
62/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。