bigscience-workshop / bigscience-workshop/data_tooling
Create dataset MultiUN v2
- Dominant language
- HTML
- Stars
- 91
- Forks
- 47
- PR merge metrics
- No merged PRs in 30d
Description
Source: [Masader Project](https://arbml.github.io/masader/)
- uid: multi_un_2
- entry: https://arbml.github.io/masader/card.html?158
- Link: http://www.euromatrixplus.net/multi-un/
- License : unknown
- Year: 2010
- Language: multilingual
- Dialect: ar-MSA: (Arabic (Modern Standard Arabic))
- Domain: other
- Form: text
- Collection Style: human translation
- Description: 6 official languages of the UN, consisting of around 300 million words per language
- Volume: 65,156
- Unit: documents
- Ethical Risks: Low
- Provider: DFKI
- Derived From:
- Paper Title: MultiUN: A Multilingual Corpus from United Nation Documents
- Paper Link: https://www.dfki.de/fileadmin/user_upload/import/4790_686_Paper.pdf
- Script: Arab
- Tokenized: No
- Host: other
- Access: Free
- Cost:
- Test Split: Yes
- Tasks: machine translation
- Evaluation Set?:
- Venue Title: LREC
- Citations: 223
- Venue Type: conference
- Venue Name: International Conference on Language Resources and Evaluation
- authors: A. Eisele
- affiliations:
- abstract: This paper describes the acquisition, preparation and properties of a corpus extracted from the official documents of the United Nations (UN). This corpus is available in all 6 official languages of the UN, consisting of around 300 million words per language. We describe the methods we used for crawling, document formatting, and sentence alignment. This corpus also includes a common test set for machine translation. We present the results of a French-Chinese machine translation experiment performed on this corpus.
- Added by : Zaid
- Notes:
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.