Rethinking the Evaluation of Neural Machine Translation

Yan, Jianhao; Wu, Chenming; Meng, Fandong; Zhou, Jie

Computer Science > Computation and Language

arXiv:2106.15217v1 (cs)

[Submitted on 29 Jun 2021 (this version), latest version 8 Oct 2022 (v2)]

Title:Rethinking the Evaluation of Neural Machine Translation

Authors:Jianhao Yan, Chenming Wu, Fandong Meng, Jie Zhou

View PDF

Abstract:The evaluation of neural machine translation systems is usually built upon generated translation of a certain decoding method (e.g., beam search) with evaluation metrics over the generated translation (e.g., BLEU). However, this evaluation framework suffers from high search errors brought by heuristic search algorithms and is limited by its nature of evaluation over one best candidate. In this paper, we propose a novel evaluation protocol, which not only avoids the effect of search errors but provides a system-level evaluation in the perspective of model ranking. In particular, our method is based on our newly proposed exact top-$k$ decoding instead of beam search. Our approach evaluates model errors by the distance between the candidate spaces scored by the references and the model respectively. Extensive experiments on WMT'14 English-German demonstrate that bad ranking ability is connected to the well-known beam search curse, and state-of-the-art Transformer models are facing serious ranking errors. By evaluating various model architectures and techniques, we provide several interesting findings. Finally, to effectively approximate the exact search algorithm with same time cost as original beam search, we present a minimum heap augmented beam search algorithm.

Comments:	Submitted to NeurIPS 2021
Subjects:	Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2106.15217 [cs.CL]
	(or arXiv:2106.15217v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2106.15217

Submission history

From: Jianhao Yan [view email]
[v1] Tue, 29 Jun 2021 09:59:50 UTC (82 KB)
[v2] Sat, 8 Oct 2022 12:08:40 UTC (200 KB)

Computer Science > Computation and Language

Title:Rethinking the Evaluation of Neural Machine Translation

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Rethinking the Evaluation of Neural Machine Translation

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators