OpenNMT / OpenNMT/CTranslate2

Difference translation result after convert to ctranslate

Open
#1,781 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
C++
Stars
4.7k
Forks
536
Avg merge
12h 12m
Merged PRs (30d)
4

Description

Hi. I have finetune the Helsinki-NLP/opus-mt-zh-vi model for translating Chinese to Vietnamese. When I convert the model to ctranslate2, the performance is decrease (from 32 sacrebleu with transformer inference to just 28 sacrebleu with ctranslate2 inference). Can anyone explain for me why ? Thank you. Here is my code :

Converted code : ct2-transformers-converter --model /home/hieunq/Documents/VTP/chinsese_translation/train_chinese_vietnamsese_translation/finetune_helsinky_zh_vi/model/checkpoint-36625 --output_dir zh-vi-ct2 --force --copy_files generation_config.json tokenizer_config.json vocab.json source.spm target.spm

Inference code :
import ctranslate2
import transformers
import time
import torch
import evaluate

metric = evaluate.load("sacrebleu")

device = "cuda" if torch.cuda.is_available() else "cpu"
translator = ctranslate2.Translator("/home/hieunq/Documents/VTP/chinsese_translation/train_chinese_vietnamsese_translation/finetune_helsinky_zh_vi/zh-vi-ct2", device=device, compute_type="auto")
tokenizer = transformers.AutoTokenizer.from_pretrained("/home/hieunq/Documents/VTP/chinsese_translation/train_chinese_vietnamsese_translation/finetune_helsinky_zh_vi/zh-vi-ct2")

f_zh = open("/home/hieunq/Documents/VTP/chinsese_translation/data_version_1_and_2/zh/test_zh/test_zh_version_2_data_Thời_trang_nữ.txt","r")
f_vi = open("/home/hieunq/Documents/VTP/chinsese_translation/data_version_1_and_2/vi/test_vi/test_vi_version_2_data_Thời_trang_nữ.txt","r",encoding="utf-8")
texts = f_zh.readlines()

translated_texts = []
start = time.time()
batch_source_tokens = [tokenizer.convert_ids_to_tokens(tokenizer.encode(sentence)) for sentence in texts]

batch_size = 10
results = translator.translate_batch(batch_source_tokens, max_batch_size = batch_size, beam_size = 4)

for i, result in enumerate(results):
target = result.hypotheses[0] # Giả sử chúng ta lấy hypothesis tốt nhất
translated_sentence = tokenizer.decode(tokenizer.convert_tokens_to_ids(target))
translated_texts.append(translated_sentence)

references = f_vi.readlines()

predictions_texts = [pred.strip() for pred in translated_texts]
references_text = [pred.strip() for pred in references]

result = metric.compute(predictions=predictions_texts, references=references_text)
print(result["score"])
print("Time :", time.time() - start)

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reproducing the inline conversion command and inference script with the fine-tuned Helsinki-NLP/opus-mt-zh-vi checkpoint, comparing the Transformer and CTranslate2 SacreBLEU scores. Check the tokenizer conversion, generated hypotheses, decoding, beam settings, and evaluation inputs; done means identifying a reproducible discrepancy and explaining its cause or documenting the missing reproduction details.

Written by the indexing model from the issue text.

Assessment

Tech stack
cpp, python, pytorch
Domain
machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.