bigscience-workshop / bigscience-workshop/data_tooling

Create dataset MultiUN v2

Open
#288 7 comments 0 reactions 2 assignees Claimed by @mariosasko View on GitHub
data catalog language modeling script
Dominant language
HTML
Stars
91
Forks
47
PR merge metrics
No merged PRs in 30d

Description

Source: [Masader Project](https://arbml.github.io/masader/)
- uid: multi_un_2
- entry: https://arbml.github.io/masader/card.html?158
- Link: http://www.euromatrixplus.net/multi-un/
- License : unknown
- Year: 2010
- Language: multilingual
- Dialect: ar-MSA: (Arabic (Modern Standard Arabic))
- Domain: other
- Form: text
- Collection Style: human translation
- Description: 6 official languages of the UN, consisting of around 300 million words per language
- Volume: 65,156
- Unit: documents
- Ethical Risks: Low
- Provider: DFKI
- Derived From:
- Paper Title: MultiUN: A Multilingual Corpus from United Nation Documents
- Paper Link: https://www.dfki.de/fileadmin/user_upload/import/4790_686_Paper.pdf
- Script: Arab
- Tokenized: No
- Host: other
- Access: Free
- Cost:
- Test Split: Yes
- Tasks: machine translation
- Evaluation Set?:
- Venue Title: LREC
- Citations: 223
- Venue Type: conference
- Venue Name: International Conference on Language Resources and Evaluation
- authors: A. Eisele
- affiliations:
- abstract: This paper describes the acquisition, preparation and properties of a corpus extracted from the official documents of the United Nations (UN). This corpus is available in all 6 official languages of the UN, consisting of around 300 million words per language. We describe the methods we used for crawling, document formatting, and sentence alignment. This corpus also includes a common test set for machine translation. We present the results of a French-Chinese machine translation experiment performed on this corpus.
- Added by : Zaid
- Notes:

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.