bigscience-workshop / bigscience-workshop/data_tooling

Create dataset uit_viquad

Offen
#160 0 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen
data catalog need custodian permission
Vorherrschende Sprache
HTML
Sterne
91
Forks
47
PR-Merge-Kennzahlen
Keine gemergten PRs in 30 T.

Beschreibung

- uid: uit_viquad
- type: processed
- description:
- name: UIT-ViQuAD – A Vietnamese Dataset for Evaluating Machine Reading Comprehension.
- description: Vietnamese Question Answering Dataset (UIT-ViQuAD), a new
dataset for the low-resource language as Vietnamese to evaluate MRC models. This dataset comprises over 23,000 human-generated question-answer pairs based on 5,109 passages of 174 Vietnamese articles from Wikipedia.
- homepage: https://sites.google.com/uit.edu.vn/uit-nlp/datasets-projects
- validated: True
- languages:
- language_names:
- Vietnamese
- language_comments: unclear
- language_locations:
- South-eastern Asia
- Vietnam
- validated: False
- custodian:
- name: The UIT NLP Group
- in_catalogue:
- type: A university or research institution
- location: Vietnam
- contact_name: Mr. Kiet Nguyen
- contact_email: kietnv@uit.edu.vn
- contact_submitter: True
- additional: The UIT Natural Language Processing Group is a scientific research group on Natural Language Processing and Computational Linguistics. Members in our group are lecturers, undergraduate and postgraduate students from Vietnam National University- Ho Chi Minh City (VNU-HCM). https://sites.google.com/uit.edu.vn/uit-nlp/
- validated: False
- availability:
- procurement:
- for_download: Yes - after signing a user agreement
- download_url: https://sites.google.com/uit.edu.vn/uit-nlp/datasets-projects
- download_email:
- licensing:
- has_licenses: Yes
- license_text: Creative Commons Attribution 4.0 International License
- license_properties:
- open license
- license_list:
- cc-by-nc-sa-4.0: Creative Commons Attribution Non Commercial Share Alike 4.0 International
- pii:
- has_pii: Yes
- generic_pii_likely: very likely
- generic_pii_list:
- names
- website account name or handle
- dates (birth, death, etc.)
- URLs
- physical addresses
- email addresses
- numeric_pii_likely: somewhat likely
- numeric_pii_list:
- sensitive_pii_likely:
- sensitive_pii_list:
- no_pii_justification_class:
- no_pii_justification_text:
- validated: False
- processed_from_primary:
- from_primary: Taken from primary source
- primary_availability: Yes - they are fully available
- primary_license: Yes - the dataset has the same license as the source material
- primary_types:
- web | wiki
- validated: False
- from_primary_entries:
- media:
- category:
- text
- text_format:
- audiovisual_format:
- image_format:
- database_format:
- text_is_transcribed: No
- instance_type:
- instance_count: 10K

Beitragsleitfaden

Für dieses Repository ist kein Beitragsleitfaden indexiert

Rechercherichtung

Use issue #160 as the specification for the new dataset entry and work with the named output file, uit_viquad.json. Preserve the provided metadata for the dataset, language, custodian, availability, licensing, privacy, source, and media sections. Done means the entry is present with the requested UID and filename and passes the repository's dataset validation process, if available.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
json
Bereich
data
Issue-Typ
Feature
Schwierigkeit
2/5
Geschätzter Aufwand
1-3 Stunden
Aktivitätsstatus
Veraltet
Klarheit
Klar beschrieben
Anfängerfreundlichkeit
45/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.