huggingface / huggingface/fineweb-2

ARABIC: Wrong language or dialect or script

Open
#1 5 comments 1 reaction 0 assignees View on GitHub
Dominant language
Python
Stars
264
Forks
17
PR merge metrics
No merged PRs in 30d

Description

Hey team,
As part of the Argilla FineWeb-C sprint, we are annotating arabic and its dialects. MSA, ARY and ARZ.
The problem for all of them is being miscallafied most of the time.
For example, annotators in arabic report most data is not arabic but rather dialects with a lot of Arabizi (usage of latin script). In dialects, people report that most of the samples are in fact in arabic MSA !
This mismatch leads to labeling most of the data as problematic.
cc: @nataliaElv

Contributor guide

No contributing guide indexed for this repository

Research direction

No file, test, or entry point is named. First clarify how Arabic MSA, ARY, ARZ, and Arabizi should be labeled, then locate the annotation or classification workflow and define corrected labels and validation examples as the completion criteria.

Written by the indexing model from the issue text.

Assessment

Domain
data, internationalization
Issue type
Bug
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.