bigscience-workshop / bigscience-workshop/data_tooling

Create dataset xnli

Ouverte
#350 0 commentaires 0 réactions 0 personnes assignées Voir sur GitHub
data catalog
Langage dominant
HTML
Étoiles
91
Forks
47
Métriques de merge des PR
Aucune PR mergée en 30 j

Description

- uid: xnli
- type: primary
- description:
- name: XNLI
- description: it is a Cross-lingual Natural Language Inference corpus that
- homepage: https://github.com/facebookresearch/XNLI
- validated: True
- languages:
- language_names:
- English
- French
- Spanish
- Arabic
- Vietnamese
- Chinese
- ar-MSA
- German
- Greek languages
- Bulgarian
- Russian
- Arabic
- Turkish
- Thai
- Hindi
- Swahili (macrolanguage)
- Urdu
- language_comments:
- language_locations:
- validated: False
- custodian:
- name:
- in_catalogue:
- type:
- location:
- contact_name: XNLI
- contact_email:
- contact_submitter: False
- additional:
- validated: False
- availability:
- procurement:
- for_download: Yes - it has a direct download link or links
- download_url: https://dl.fbaipublicfiles.com/XNLI/XNLI-MT-1.0.zip
- download_email:
- licensing:
- has_licenses: Yes
- license_text:
- license_properties:
- open license
- public domain
- license_list:
- pii:
- has_pii: Unclear
- generic_pii_likely:
- generic_pii_list:
- numeric_pii_likely:
- numeric_pii_list:
- sensitive_pii_likely:
- sensitive_pii_list:
- no_pii_justification_class: general knowledge not written by or referring to private persons
- no_pii_justification_text:
- validated: False
- source_category:
- category_type: collection
- category_web:
- category_media:
- validated: False
- media:
- category:
- text
- text_format:
- .CSV
- audiovisual_format:
- image_format:
- database_format:
- .ZIP
- text_is_transcribed: Yes - audiovisual
- instance_type:
- instance_count: 100K

Guide de contribution

Aucun guide de contribution indexé pour ce dépôt

Piste de recherche

Start by locating the existing dataset metadata entries and compare their structure with the requested xnli.json record. Add the XNLI metadata using the supplied download URL, language list, licensing details, and media fields; done means the record follows the repository's schema and is included with the other datasets.

Rédigé par le modèle d'indexation à partir du texte de l'issue.

Évaluation

Stack technique
json
Domaine
data
Type d'issue
Fonctionnalité
Difficulté
2/5
Temps estimé
1-3 heures
Activité
À l'abandon
Clarté
Plutôt claire
Accessibilité débutants
48/100

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.