bigscience-workshop / bigscience-workshop/data_tooling

Create dataset xnli

Abierto
#350 0 comentarios 0 reacciones 0 asignados Ver en GitHub
data catalog
Lenguaje dominante
HTML
Estrellas
91
Forks
47
Métricas de merge de PR
Sin PR fusionados en 30 d

Descripción

- uid: xnli
- type: primary
- description:
- name: XNLI
- description: it is a Cross-lingual Natural Language Inference corpus that
- homepage: https://github.com/facebookresearch/XNLI
- validated: True
- languages:
- language_names:
- English
- French
- Spanish
- Arabic
- Vietnamese
- Chinese
- ar-MSA
- German
- Greek languages
- Bulgarian
- Russian
- Arabic
- Turkish
- Thai
- Hindi
- Swahili (macrolanguage)
- Urdu
- language_comments:
- language_locations:
- validated: False
- custodian:
- name:
- in_catalogue:
- type:
- location:
- contact_name: XNLI
- contact_email:
- contact_submitter: False
- additional:
- validated: False
- availability:
- procurement:
- for_download: Yes - it has a direct download link or links
- download_url: https://dl.fbaipublicfiles.com/XNLI/XNLI-MT-1.0.zip
- download_email:
- licensing:
- has_licenses: Yes
- license_text:
- license_properties:
- open license
- public domain
- license_list:
- pii:
- has_pii: Unclear
- generic_pii_likely:
- generic_pii_list:
- numeric_pii_likely:
- numeric_pii_list:
- sensitive_pii_likely:
- sensitive_pii_list:
- no_pii_justification_class: general knowledge not written by or referring to private persons
- no_pii_justification_text:
- validated: False
- source_category:
- category_type: collection
- category_web:
- category_media:
- validated: False
- media:
- category:
- text
- text_format:
- .CSV
- audiovisual_format:
- image_format:
- database_format:
- .ZIP
- text_is_transcribed: Yes - audiovisual
- instance_type:
- instance_count: 100K

Guía de contribución

No hay ninguna guía de contribución indexada para este repositorio

Línea de trabajo

Start by locating the existing dataset metadata entries and compare their structure with the requested xnli.json record. Add the XNLI metadata using the supplied download URL, language list, licensing details, and media fields; done means the record follows the repository's schema and is included with the other datasets.

Escrito por el modelo de indexación a partir del texto del issue.

Evaluación

Stack tecnológico
json
Área
data
Tipo de issue
Nueva funcionalidad
Dificultad
2/5
Tiempo estimado
1-3 horas
Estado de actividad
Estancado
Claridad
Bastante claro
Aptitud para principiantes
48/100

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.