Mapping Unseen Words to Task-Trained Embedding Spaces

Madhyastha, Pranava Swaroop; Bansal, Mohit; Gimpel, Kevin; Livescu, Karen

Computer Science > Computation and Language

arXiv:1510.02387v1 (cs)

[Submitted on 8 Oct 2015 (this version), latest version 23 Jun 2016 (v2)]

Title:Mapping Unseen Words to Task-Trained Embedding Spaces

Authors:Pranava Swaroop Madhyastha, Mohit Bansal, Kevin Gimpel, Karen Livescu

View PDF

Abstract:We consider the setting in which we train a supervised model that learns task-specific word representations. We assume that we have access to some initial word representations (e.g., unsupervised embeddings), and that the supervised learning procedure updates them to task-specific representations for words contained in the training data. But what about words not contained in the supervised training data? When such unseen words are encountered at test time, they are typically represented by either their initial vectors or a single unknown vector, which often leads to errors. In this paper, we address this issue by learning to map from initial representations to task-specific ones. We present a general technique that uses a neural network mapper with a weighted multiple-loss criterion. This allows us to use the same learned model parameters at test time but now with appropriate task-specific representations for unseen words. We consider the task of dependency parsing and report improvements in performance (and reductions in out-of-vocabulary rates) across multiple domains such as news, Web, and speech. We also achieve downstream improvements on the task of parsing-based sentiment analysis.

Comments:	10 + 3 pages, 3 figures
Subjects:	Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as:	arXiv:1510.02387 [cs.CL]
	(or arXiv:1510.02387v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.1510.02387

Submission history

From: Pranava Swaroop Madhyastha [view email]
[v1] Thu, 8 Oct 2015 16:17:47 UTC (214 KB)
[v2] Thu, 23 Jun 2016 06:24:18 UTC (223 KB)

Computer Science > Computation and Language

Title:Mapping Unseen Words to Task-Trained Embedding Spaces

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Mapping Unseen Words to Task-Trained Embedding Spaces

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators