bigscience-workshop / bigscience-workshop/data_tooling

Create dataset uit_viquad

Đang mở
#160 0 bình luận 0 reaction 0 người được giao Xem trên GitHub
data catalog need custodian permission
Ngôn ngữ chính
HTML
Star
91
Fork
47
Chỉ số merge pull request
Không có pull request nào được merge trong 30 ngày

Mô tả

- uid: uit_viquad
- type: processed
- description:
- name: UIT-ViQuAD – A Vietnamese Dataset for Evaluating Machine Reading Comprehension.
- description: Vietnamese Question Answering Dataset (UIT-ViQuAD), a new
dataset for the low-resource language as Vietnamese to evaluate MRC models. This dataset comprises over 23,000 human-generated question-answer pairs based on 5,109 passages of 174 Vietnamese articles from Wikipedia.
- homepage: https://sites.google.com/uit.edu.vn/uit-nlp/datasets-projects
- validated: True
- languages:
- language_names:
- Vietnamese
- language_comments: unclear
- language_locations:
- South-eastern Asia
- Vietnam
- validated: False
- custodian:
- name: The UIT NLP Group
- in_catalogue:
- type: A university or research institution
- location: Vietnam
- contact_name: Mr. Kiet Nguyen
- contact_email: kietnv@uit.edu.vn
- contact_submitter: True
- additional: The UIT Natural Language Processing Group is a scientific research group on Natural Language Processing and Computational Linguistics. Members in our group are lecturers, undergraduate and postgraduate students from Vietnam National University- Ho Chi Minh City (VNU-HCM). https://sites.google.com/uit.edu.vn/uit-nlp/
- validated: False
- availability:
- procurement:
- for_download: Yes - after signing a user agreement
- download_url: https://sites.google.com/uit.edu.vn/uit-nlp/datasets-projects
- download_email:
- licensing:
- has_licenses: Yes
- license_text: Creative Commons Attribution 4.0 International License
- license_properties:
- open license
- license_list:
- cc-by-nc-sa-4.0: Creative Commons Attribution Non Commercial Share Alike 4.0 International
- pii:
- has_pii: Yes
- generic_pii_likely: very likely
- generic_pii_list:
- names
- website account name or handle
- dates (birth, death, etc.)
- URLs
- physical addresses
- email addresses
- numeric_pii_likely: somewhat likely
- numeric_pii_list:
- sensitive_pii_likely:
- sensitive_pii_list:
- no_pii_justification_class:
- no_pii_justification_text:
- validated: False
- processed_from_primary:
- from_primary: Taken from primary source
- primary_availability: Yes - they are fully available
- primary_license: Yes - the dataset has the same license as the source material
- primary_types:
- web | wiki
- validated: False
- from_primary_entries:
- media:
- category:
- text
- text_format:
- audiovisual_format:
- image_format:
- database_format:
- text_is_transcribed: No
- instance_type:
- instance_count: 10K

Hướng dẫn đóng góp

Chưa lập chỉ mục được hướng dẫn đóng góp cho kho mã nguồn này

Hướng nghiên cứu

Use issue #160 as the specification for the new dataset entry and work with the named output file, uit_viquad.json. Preserve the provided metadata for the dataset, language, custodian, availability, licensing, privacy, source, and media sections. Done means the entry is present with the requested UID and filename and passes the repository's dataset validation process, if available.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Đánh giá

Công nghệ
json
Lĩnh vực
data
Loại issue
Tính năng
Độ khó
2/5
Thời gian dự kiến
1-3 giờ
Mức độ hoạt động
Đình trệ
Độ rõ ràng
Đặc tả rõ ràng
Mức phù hợp với người mới
45/100

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.