bigscience-workshop / bigscience-workshop/data_tooling

Create dataset AOC

Ouverte
#287 4 commentaires 0 réactions 1 personne assignée Réclamée par @apergo-ai Voir sur GitHub
data catalog need data sourcing feedback
Langage dominant
HTML
Étoiles
91
Forks
47
Métriques de merge des PR
Aucune PR mergée en 30 j

Description

Source: [Masader Project](https://arbml.github.io/masader/)
- uid: arabic_online_commentary
- entry: https://arbml.github.io/masader/card.html?39
- Link: https://github.com/sjeblee/AOC
- License : unknown
- Year: 2011
- Language: ar
- Dialect: other
- Domain: news articles
- Form: text
- Collection Style: crawling and annotation(other)
- Description: a 52M-word monolingual dataset rich in dialectal content
- Volume: 108,000
- Unit: sentences
- Ethical Risks: Low
- Provider: Johns Hopkins University
- Derived From:
- Paper Title: The Arabic Online Commentary Dataset: an Annotated Dataset of Informal Arabic with High Dialectal Content
- Paper Link: https://aclanthology.org/P11-2007.pdf
- Script: Arab
- Tokenized: No
- Host: GitHub
- Access: Free
- Cost:
- Test Split: No
- Tasks: dialect identification
- Evaluation Set?:
- Venue Title: ACL
- Citations: 147
- Venue Type: conference
- Venue Name: Assofications of computation linguisitcs
- authors: Omar Zaidan,Chris Callison-Burch
- affiliations: ,
- abstract: The written form of Arabic, Modern Standard Arabic (MSA), differs quite a bit from the spoken dialects of Arabic, which are the true "native" languages of Arabic speakers used in daily life. However, due to MSA's prevalence in written form, almost all Arabic datasets have predominantly MSA content. We present the Arabic Online Commentary Dataset, a 52M-word monolingual dataset rich in dialectal content, and we describe our long-term annotation effort to identify the dialect level (and dialect itself) in each sentence of the dataset. So far, we have labeled 108K sentences, 41% of which as having dialectal content. We also present experimental results on the task of automatic dialect identification, using the collected labels for training and evaluation.
- Added by :
- Notes:

Guide de contribution

Aucun guide de contribution indexé pour ce dépôt

Évaluation

Cette issue n'a pas encore été évaluée.

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.