Character-Level Neural Translation for Multilingual Media Monitoring in the SUMMA Project

Barzdins, Guntis; Renals, Steve; Gosko, Didzis

Computer Science > Computation and Language

arXiv:1604.01221 (cs)

[Submitted on 5 Apr 2016]

Title:Character-Level Neural Translation for Multilingual Media Monitoring in the SUMMA Project

Authors:Guntis Barzdins, Steve Renals, Didzis Gosko

View PDF

Abstract:The paper steps outside the comfort-zone of the traditional NLP tasks like automatic speech recognition (ASR) and machine translation (MT) to addresses two novel problems arising in the automated multilingual news monitoring: segmentation of the TV and radio program ASR transcripts into individual stories, and clustering of the individual stories coming from various sources and languages into storylines. Storyline clustering of stories covering the same events is an essential task for inquisitorial media monitoring. We address these two problems jointly by engaging the low-dimensional semantic representation capabilities of the sequence to sequence neural translation models. To enable joint multi-task learning for multilingual neural translation of morphologically rich languages we replace the attention mechanism with the sliding-window mechanism and operate the sequence to sequence neural translation model on the character-level rather than on the word-level. The story segmentation and storyline clustering problem is tackled by examining the low-dimensional vectors produced as a side-product of the neural translation process. The results of this paper describe a novel approach to the automatic story segmentation and storyline clustering problem.

Comments:	LREC-2016 submission
Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:1604.01221 [cs.CL]
	(or arXiv:1604.01221v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.1604.01221

Submission history

From: Guntis Barzdins [view email]
[v1] Tue, 5 Apr 2016 11:34:11 UTC (444 KB)

Computer Science > Computation and Language

Title:Character-Level Neural Translation for Multilingual Media Monitoring in the SUMMA Project

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Character-Level Neural Translation for Multilingual Media Monitoring in the SUMMA Project

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators