huggingface / huggingface/fineweb-2
ARABIC: Wrong language or dialect or script
- Dominant language
- Python
- Stars
- 264
- Forks
- 17
- PR merge metrics
- No merged PRs in 30d
Description
Hey team,
As part of the Argilla FineWeb-C sprint, we are annotating arabic and its dialects. MSA, ARY and ARZ.
The problem for all of them is being miscallafied most of the time.
For example, annotators in arabic report most data is not arabic but rather dialects with a lot of Arabizi (usage of latin script). In dialects, people report that most of the samples are in fact in arabic MSA !
This mismatch leads to labeling most of the data as problematic.
cc: @nataliaElv
Contributor guide
No contributing guide indexed for this repository
Research direction
No file, test, or entry point is named. First clarify how Arabic MSA, ARY, ARZ, and Arabizi should be labeled, then locate the annotation or classification workflow and define corrected labels and validation examples as the completion criteria.
Written by the indexing model from the issue text.
Assessment
- Domain
- data, internationalization
- Issue type
- Bug
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100