bigscience-workshop / bigscience-workshop/data_tooling

Create dataset MultiUN v2

オープン
#288 コメント 7 件 リアクション 0 件 担当者 2 名 @mariosasko が担当を希望しています GitHub で見る
data catalog language modeling script
主要言語
HTML
スター
91
フォーク
47
PR マージ指標
30日以内にマージされた PR はありません

説明

Source: [Masader Project](https://arbml.github.io/masader/)
- uid: multi_un_2
- entry: https://arbml.github.io/masader/card.html?158
- Link: http://www.euromatrixplus.net/multi-un/
- License : unknown
- Year: 2010
- Language: multilingual
- Dialect: ar-MSA: (Arabic (Modern Standard Arabic))
- Domain: other
- Form: text
- Collection Style: human translation
- Description: 6 official languages of the UN, consisting of around 300 million words per language
- Volume: 65,156
- Unit: documents
- Ethical Risks: Low
- Provider: DFKI
- Derived From:
- Paper Title: MultiUN: A Multilingual Corpus from United Nation Documents
- Paper Link: https://www.dfki.de/fileadmin/user_upload/import/4790_686_Paper.pdf
- Script: Arab
- Tokenized: No
- Host: other
- Access: Free
- Cost:
- Test Split: Yes
- Tasks: machine translation
- Evaluation Set?:
- Venue Title: LREC
- Citations: 223
- Venue Type: conference
- Venue Name: International Conference on Language Resources and Evaluation
- authors: A. Eisele
- affiliations:
- abstract: This paper describes the acquisition, preparation and properties of a corpus extracted from the official documents of the United Nations (UN). This corpus is available in all 6 official languages of the UN, consisting of around 300 million words per language. We describe the methods we used for crawling, document formatting, and sentence alignment. This corpus also includes a common test set for machine translation. We present the results of a French-Chinese machine translation experiment performed on this corpus.
- Added by : Zaid
- Notes:

コントリビューションガイド

このリポジトリのコントリビューションガイドは索引されていません

評価

この issue はまだ評価されていません。

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。