bigscience-workshop / bigscience-workshop/data_tooling

Create dataset uit_viquad

Ouverte
#160 0 commentaires 0 réactions 0 personnes assignées Voir sur GitHub
data catalog need custodian permission
Langage dominant
HTML
Étoiles
91
Forks
47
Métriques de merge des PR
Aucune PR mergée en 30 j

Description

- uid: uit_viquad
- type: processed
- description:
- name: UIT-ViQuAD – A Vietnamese Dataset for Evaluating Machine Reading Comprehension.
- description: Vietnamese Question Answering Dataset (UIT-ViQuAD), a new
dataset for the low-resource language as Vietnamese to evaluate MRC models. This dataset comprises over 23,000 human-generated question-answer pairs based on 5,109 passages of 174 Vietnamese articles from Wikipedia.
- homepage: https://sites.google.com/uit.edu.vn/uit-nlp/datasets-projects
- validated: True
- languages:
- language_names:
- Vietnamese
- language_comments: unclear
- language_locations:
- South-eastern Asia
- Vietnam
- validated: False
- custodian:
- name: The UIT NLP Group
- in_catalogue:
- type: A university or research institution
- location: Vietnam
- contact_name: Mr. Kiet Nguyen
- contact_email: kietnv@uit.edu.vn
- contact_submitter: True
- additional: The UIT Natural Language Processing Group is a scientific research group on Natural Language Processing and Computational Linguistics. Members in our group are lecturers, undergraduate and postgraduate students from Vietnam National University- Ho Chi Minh City (VNU-HCM). https://sites.google.com/uit.edu.vn/uit-nlp/
- validated: False
- availability:
- procurement:
- for_download: Yes - after signing a user agreement
- download_url: https://sites.google.com/uit.edu.vn/uit-nlp/datasets-projects
- download_email:
- licensing:
- has_licenses: Yes
- license_text: Creative Commons Attribution 4.0 International License
- license_properties:
- open license
- license_list:
- cc-by-nc-sa-4.0: Creative Commons Attribution Non Commercial Share Alike 4.0 International
- pii:
- has_pii: Yes
- generic_pii_likely: very likely
- generic_pii_list:
- names
- website account name or handle
- dates (birth, death, etc.)
- URLs
- physical addresses
- email addresses
- numeric_pii_likely: somewhat likely
- numeric_pii_list:
- sensitive_pii_likely:
- sensitive_pii_list:
- no_pii_justification_class:
- no_pii_justification_text:
- validated: False
- processed_from_primary:
- from_primary: Taken from primary source
- primary_availability: Yes - they are fully available
- primary_license: Yes - the dataset has the same license as the source material
- primary_types:
- web | wiki
- validated: False
- from_primary_entries:
- media:
- category:
- text
- text_format:
- audiovisual_format:
- image_format:
- database_format:
- text_is_transcribed: No
- instance_type:
- instance_count: 10K

Guide de contribution

Aucun guide de contribution indexé pour ce dépôt

Piste de recherche

Use issue #160 as the specification for the new dataset entry and work with the named output file, uit_viquad.json. Preserve the provided metadata for the dataset, language, custodian, availability, licensing, privacy, source, and media sections. Done means the entry is present with the requested UID and filename and passes the repository's dataset validation process, if available.

Rédigé par le modèle d'indexation à partir du texte de l'issue.

Évaluation

Stack technique
json
Domaine
data
Type d'issue
Fonctionnalité
Difficulté
2/5
Temps estimé
1-3 heures
Activité
À l'abandon
Clarté
Clairement spécifiée
Accessibilité débutants
45/100

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.