Building a Japanese Document-Level Relation Extraction Dataset Assisted by Cross-Lingual Transfer

Ma, Youmi; Wang, An; Okazaki, Naoaki

Computer Science > Computation and Language

arXiv:2404.16506 (cs)

[Submitted on 25 Apr 2024]

Title:Building a Japanese Document-Level Relation Extraction Dataset Assisted by Cross-Lingual Transfer

Authors:Youmi Ma, An Wang, Naoaki Okazaki

View PDF HTML (experimental)

Abstract:Document-level Relation Extraction (DocRE) is the task of extracting all semantic relationships from a document. While studies have been conducted on English DocRE, limited attention has been given to DocRE in non-English languages. This work delves into effectively utilizing existing English resources to promote DocRE studies in non-English languages, with Japanese as the representative case. As an initial attempt, we construct a dataset by transferring an English dataset to Japanese. However, models trained on such a dataset suffer from low recalls. We investigate the error cases and attribute the failure to different surface structures and semantics of documents translated from English and those written by native speakers. We thus switch to explore if the transferred dataset can assist human annotation on Japanese documents. In our proposal, annotators edit relation predictions from a model trained on the transferred dataset. Quantitative analysis shows that relation recommendations suggested by the model help reduce approximately 50% of the human edit steps compared with the previous approach. Experiments quantify the performance of existing DocRE models on our collected dataset, portraying the challenges of Japanese and cross-lingual DocRE.

Comments:	Accepted LREC-COLING 2024
Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2404.16506 [cs.CL]
	(or arXiv:2404.16506v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2404.16506

Submission history

From: Youmi Ma [view email]
[v1] Thu, 25 Apr 2024 10:59:02 UTC (1,628 KB)

Computer Science > Computation and Language

Title:Building a Japanese Document-Level Relation Extraction Dataset Assisted by Cross-Lingual Transfer

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Building a Japanese Document-Level Relation Extraction Dataset Assisted by Cross-Lingual Transfer

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators