tensorflow / tensorflow/datasets

newscommentary_v14 of wmt_translate datasets create opposite zh-en label sentence

Open
#2,309 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

bug
Dominant language
Python
Stars
4.6k
Forks
1.6k
Avg merge
3h 54m
Merged PRs (30d)
1

Description

Short description
newscommentary_v14 create opposite zh-en label sentence

Environment information

  • Operating System: windows 7
  • Python version: 3.7
  • tensorflow-datasets/tfds-nightly version: 3.2.1
  • tensorflow/tensorflow-gpu/tf-nightly/tf-nightly-gpu version: 2.3.0

Reproduction instructions

config = tfds.translate.wmt.WmtConfig(
      version=tfds.core.Version('0.0.3'),
      language_pair=("zh", "en"),
      subsets={
        tfds.Split.TRAIN: ['newscommentary_v14']
      }
    )
builder_config={"config": config}
train_examples, val_examples = tfds.load(name='wmt_translate', split=['train[:1%]', 'train[:1%]'], as_supervised=True, builder_kwargs=builder_config)
for en, zh in train_examples.take(3):
  print(en)
  print(zh)
  print('-' * 10)

Link to logs

WARNING:absl:Using custom data configuration zh-en
tf.Tensor(b'The fear is real and visceral, and politicians ignore it at their peril.', shape=(), dtype=string)
tf.Tensor(b'\xe8\xbf\x99\xe7\xa7\x8d\xe6\x81\x90\xe6\x83\xa7\xe6\x98\xaf\xe7\x9c\x9f\xe5\xae\x9e\xe8\x80\x8c\xe5\x86\x85\xe5\x9c\xa8\xe7\x9a\x84\xe3\x80\x82 \xe5\xbf\xbd\xe8\xa7\x86\xe5\xae\x83\xe7\x9a\x84\xe6\x94\xbf\xe6\xb2\xbb\xe5\xae\xb6\xe4\xbb\xac\xe5\x89\x8d\xe9\x80\x94\xe5\xa0\xaa\xe5\xbf\xa7\xe3\x80\x82', shape=(), dtype=string)
----------
tf.Tensor(b'In fact, the German political landscape needs nothing more than a truly liberal party, in the US sense of the word \xe2\x80\x9cliberal\xe2\x80\x9d \xe2\x80\x93 a champion of the cause of individual freedom.', shape=(), dtype=string)
tf.Tensor(b'\xe4\xba\x8b\xe5\xae\x9e\xe4\xb8\x8a\xef\xbc\x8c\xe5\xbe\xb7\xe5\x9b\xbd\xe6\x94\xbf\xe6\xb2\xbb\xe5\xb1\x80\xe5\x8a\xbf\xe9\x9c\x80\xe8\xa6\x81\xe7\x9a\x84\xe4\xb8\x8d\xe8\xbf\x87\xe6\x98\xaf\xe4\xb8\x80\xe4\xb8\xaa\xe7\xac\xa6\xe5\x90\x88\xe7\xbe\x8e\xe5\x9b\xbd\xe6\x89\x80\xe8\xb0\x93\xe2\x80\x9c\xe8\x87\xaa\xe7\x94\xb1\xe2\x80\x9d\xe5\xae\x9a\xe4\xb9\x89\xe7\x9a\x84\xe7\x9c\x9f\xe6\xad\xa3\xe7\x9a\x84\xe8\x87\xaa\xe7\x94\xb1\xe5\x85\x9a\xe6\xb4\xbe\xef\xbc\x8c\xe4\xb9\x9f\xe5\xb0\xb1\xe6\x98\xaf\xe4\xb8\xaa\xe4\xba\xba\xe8\x87\xaa\xe7\x94\xb1\xe4\xba\x8b\xe4\xb8\x9a\xe7\x9a\x84\xe5\x80\xa1\xe5\xaf\xbc\xe8\x80\x85\xe3\x80\x82', shape=(), dtype=string)
----------
tf.Tensor(b'Shifting to renewable-energy sources will require enormous effort and major infrastructure investment.', shape=(), dtype=string)
tf.Tensor(b'\xe5\xbf\x85\xe9\xa1\xbb\xe4\xbb\x98\xe5\x87\xba\xe5\xb7\xa8\xe5\xa4\xa7\xe7\x9a\x84\xe5\x8a\xaa\xe5\x8a\x9b\xe5\x92\x8c\xe5\x9f\xba\xe7\xa1\x80\xe8\xae\xbe\xe6\x96\xbd\xe6\x8a\x95\xe8\xb5\x84\xe6\x89\x8d\xe8\x83\xbd\xe5\xae\x8c\xe6\x88\x90\xe5\x90\x91\xe5\x8f\xaf\xe5\x86\x8d\xe7\x94\x9f\xe8\x83\xbd\xe6\xba\x90\xe7\x9a\x84\xe8\xbf\x87\xe6\xb8\xa1\xe3\x80\x82', shape=(), dtype=string)
----------

Expected behavior
While load the other subdatasets, the first line is in chinese, the second line is in english, but newscommentary_v14 is opposite.

Then I watch the record file in data dir

?
R
zhL
J
HThe fear is real and visceral, and politicians ignore it at their peril.
V
enP
N
L杩欑鎭愭儳鏄湡瀹炶€屽唴鍦ㄧ殑銆?蹇借瀹冪殑鏀挎不瀹朵滑鍓嶉€斿牚蹇с€倉^V賣      ??
?

the 'zh' label match english sentence and the 'en' label match chinese sentence.
I think something wrong here.

Additional context
in _generate_examples of tensorflow_datasets\translate\wmt.py

        if ".tsv" in fname:
          sub_generator = _parse_tsv
        elif ss_name.startswith("newscommentary_v14"):
          sub_generator = functools.partial(
              _parse_tsv, language_pair=self.builder_config.language_pair)

this will pass the language_pair('zh', 'en') to _parse_tsv
but the tsv is like this

1929 or 1989?	1929年还是1989年?
PARIS – As the economic crisis deepens and widens, the world has been searching for historical analogies to help us understand what has been happening.	巴黎-随着经济危机不断加深和蔓延,整个世界一直在寻找历史上的类似事件希望有助于我们了解目前正在发生的情况。

the first sentence is in english, will match the opposite label 'zh'

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start in tensorflow_datasets/translate/wmt.py, especially _generate_examples and _parse_tsv, and reproduce the zh-en newscommentary_v14 load described in the issue. Compare the parser's language-pair handling with the TSV column order; done means the generated zh-en examples consistently place Chinese and English sentences under their matching labels.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, tensorflow
Domain
data
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Clearly specified
Newbie friendliness
48/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.