bigscience-workshop / bigscience-workshop/data_tooling

Create dataset uit_viquad

未关闭
#160 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
data catalog need custodian permission
主要语言
HTML
星标
91
派生
47
PR 合并指标
30 天内没有已合并 PR

描述

- uid: uit_viquad
- type: processed
- description:
- name: UIT-ViQuAD – A Vietnamese Dataset for Evaluating Machine Reading Comprehension.
- description: Vietnamese Question Answering Dataset (UIT-ViQuAD), a new
dataset for the low-resource language as Vietnamese to evaluate MRC models. This dataset comprises over 23,000 human-generated question-answer pairs based on 5,109 passages of 174 Vietnamese articles from Wikipedia.
- homepage: https://sites.google.com/uit.edu.vn/uit-nlp/datasets-projects
- validated: True
- languages:
- language_names:
- Vietnamese
- language_comments: unclear
- language_locations:
- South-eastern Asia
- Vietnam
- validated: False
- custodian:
- name: The UIT NLP Group
- in_catalogue:
- type: A university or research institution
- location: Vietnam
- contact_name: Mr. Kiet Nguyen
- contact_email: kietnv@uit.edu.vn
- contact_submitter: True
- additional: The UIT Natural Language Processing Group is a scientific research group on Natural Language Processing and Computational Linguistics. Members in our group are lecturers, undergraduate and postgraduate students from Vietnam National University- Ho Chi Minh City (VNU-HCM). https://sites.google.com/uit.edu.vn/uit-nlp/
- validated: False
- availability:
- procurement:
- for_download: Yes - after signing a user agreement
- download_url: https://sites.google.com/uit.edu.vn/uit-nlp/datasets-projects
- download_email:
- licensing:
- has_licenses: Yes
- license_text: Creative Commons Attribution 4.0 International License
- license_properties:
- open license
- license_list:
- cc-by-nc-sa-4.0: Creative Commons Attribution Non Commercial Share Alike 4.0 International
- pii:
- has_pii: Yes
- generic_pii_likely: very likely
- generic_pii_list:
- names
- website account name or handle
- dates (birth, death, etc.)
- URLs
- physical addresses
- email addresses
- numeric_pii_likely: somewhat likely
- numeric_pii_list:
- sensitive_pii_likely:
- sensitive_pii_list:
- no_pii_justification_class:
- no_pii_justification_text:
- validated: False
- processed_from_primary:
- from_primary: Taken from primary source
- primary_availability: Yes - they are fully available
- primary_license: Yes - the dataset has the same license as the source material
- primary_types:
- web | wiki
- validated: False
- from_primary_entries:
- media:
- category:
- text
- text_format:
- audiovisual_format:
- image_format:
- database_format:
- text_is_transcribed: No
- instance_type:
- instance_count: 10K

贡献指南

这个仓库没有索引到贡献指南

调研方向

使用 issue #160 作为新数据集条目的规范,并使用指定的输出文件 uit_viquad.json 进行工作。保留数据集、语言、负责人、可用性、许可、隐私、来源和媒体部分所提供的元数据。完成标准是:该条目已使用所要求的 UID 和文件名存在,并且在可用的情况下通过仓库的数据集验证流程。

由索引模型根据 Issue 内容生成。

评估

技术栈
json
领域
data
Issue 类型
功能
难度
2/5
预计耗时
1-3 小时
活跃度
停滞
描述清晰度
描述清楚
新手友好度
45/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。