research-article

AppTechMiner: Mining Applications and Techniques from Scientific Articles

Authors:

Sanyam Agarwal,

Animesh MukherjeeAuthors Info & Claims

WOSP 2017: Proceedings of the 6th International Workshop on Mining Scientific Publications

Pages 1 - 8

https://doi.org/10.1145/3127526.3127527

Published: 15 December 2017 Publication History

Abstract

This paper presents AppTechMiner, a rule-based information extraction framework that automatically constructs a knowledge base of all application areas and problem solving techniques. Techniques include tools, methods, datasets or evaluation metrics. We also categorize individual research articles based on their application areas and the techniques proposed/improved in the article. Our system achieves high average precision (~82%) and recall (~84%) in knowledge base creation. It also performs well in application and technique assignment to an individual article (average accuracy ~66%). In the end, we further present two use cases presenting a trivial information retrieval system and an extensive temporal analysis of the usage of techniques and application areas. At present, we demonstrate the framework for the domain of computational linguistics but this can be easily generalized to any other field of research. We plan to make the codes publicly available.

References

[1]

Sharon A Caraballo. 1999. Automatic construction of a hypernym-labeled noun hierarchy from text. In Proceedings of the 37th annual meeting of the Association for Computational Linguistics on Computational Linguistics. Association for Computational Linguistics, 120--126.

Digital Library

[2]

Laura Chiticariu, Yunyao Li, and Frederick R. Reiss. 2013. Rule-Based Information Extraction is Dead! Long Live Rule-Based Information Extraction Systems!. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, EMNLP 2013, 18-21 October 2013, Grand Hyatt Seattle, Seattle, Washington, USA, A meeting of SIGDAT, a Special Interest Group of the ACL. 827--832. http://aclweb.org/anthology/D/D13/D13-1079.pdf

[3]

Aaron M Cohen and William R Hersh. 2005. A survey of current work in biomedical text mining. Briefings in bioinformatics 6, 1 (2005), 57--71.

[4]

Carol Friedman, Pauline Kra, Hong Yu, Michael Krauthammer, and Andrey Rzhetsky. 2001. GENIES: a natural-language processing system for the extraction of molecular pathways from journal articles. Bioinformatics 17, suppl 1 (2001), S74--S82.

[5]

Ken-ichiro Fukuda, Tatsuhiko Tsunoda, Ayuchi Tamura, Toshihisa Takagi, and others. 1998. Toward information extraction: identifying protein names from biological papers. In Pac symp biocomput, Vol. 707. Citeseer, 707--718.

[6]

Robert Gaizauskas, George Demetriou, Peter J. Artymiuk, and Peter Willett. 2003. Protein structures and information extraction from biological texts: the PASTA system. Bioinformatics 19, 1 (2003), 135--143.

[7]

Sonal Gupta and Christopher D Manning. 2011. Analyzing the Dynamics of Research by Extracting Key Aspects of Scientific Papers. In IJCNLP. 1--9.

[8]

Sonal Gupta and Christopher D Manning. 2014. Spied: Stanford pattern-based information extraction and diagnostics. Sponsor: Idibon 38 (2014).

[9]

Xiaodong He. 2007. Using word dependent transition models in HMM based word alignment for statistical machine translation. In Proceedings of the Second Workshop on Statistical Machine Translation. Association for Computational Linguistics, 80--87.

Digital Library

[10]

Marti A. Hearst. 1992. Automatic Acquisition of Hyponyms from Large Text Corpora. In Proceedings of the 14th Conference on Computational Linguistics - Volume 2 (COLING '92). Association for Computational Linguistics, Stroudsburg, PA, USA, 539--545.

Digital Library

[11]

Yiping Jin, Min-Yen Kan, Jun-Ping Ng, and Xiangnan He. 2013. Mining Scientific Terms and their Definitions: A Study of the ACL Anthology. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, Seattle, Washington, USA, 780--790. http://www.aclweb.org/anthology/D13-1073

[12]

Rosie Jones. 2005. Learning to extract entities from labeled and unlabeled text. Ph.D. Dissertation. Citeseer.

[13]

Su Nam Kim, Olena Medelyan, Min-Yen Kan, and Timothy Baldwin. 2010. Semeval-2010 task 5: Automatic keyphrase extraction from scientific articles. In Proceedings of the 5th International Workshop on Semantic Evaluation. Association for Computational Linguistics, 21--26.

Digital Library

[14]

Philipp Koehn, Hieu Hoang, Alexandra Birch, Chris Callison-Burch, Marcello Federico, Nicola Bertoldi, Brooke Cowan, Wade Shen, Christine Moran, Richard Zens, and others. 2007. Moses: Open source toolkit for statistical machine translation. In Proceedings of the 45th annual meeting of the ACL on interactive poster and demonstration sessions. Association for Computational Linguistics, 177--180.

Digital Library

[15]

Martin Krallinger, Alfonso Valencia, and Lynette Hirschman. 2008. Linking genes to literature: text mining, information extraction, and retrieval applications for biology. Genome biology 9, Suppl 2 (2008), 1--14.

[16]

Patrice Lopez and Laurent Romary. 2010. HUMB: Automatic key term extraction from scientific articles in GROBID. In Proceedings of the 5th international workshop on semantic evaluation. Association for Computational Linguistics, 248--251.

Digital Library

[17]

Yashar Mehdad, Matteo Negri, and José Guilherme C de Souza. 2012. FBK: cross-lingual textual entailment without translation. In Proceedings of the First Joint Conference on Lexical and Computational Semantics-Volume 1: Proceedings of the main conference and the shared task, and Volume 2: Proceedings of the Sixth International Workshop on Semantic Evaluation. Association for Computational Linguistics, 701--705.

Digital Library

[18]

Wolfgang Menzel and Ingo Schröder. 1998. Decision procedures for dependency parsing using graded constraints. In in proceedings of ACL'90. Citeseer.

[19]

Hans-Michael Müller, Arun Rangarajan, Tracy K Teal, and Paul W Sternberg. 2008. Textpresso for neuroscience: searching the full text of thousands of neuroscience research papers. Neuroinformatics 6, 3 (2008), 195--204.

[20]

Vahed Qazvinian and Dragomir R Radev. 2008. Scientific paper summarization using citation summary networks. In Proceedings of the 22nd International Conference on Computational Linguistics-Volume 1. Association for Computational Linguistics, 689--696.

Digital Library

[21]

Vahed Qazvinian, Dragomir R Radev, and Arzucan Özgür. 2010. Citation summarization through keyphrase extraction. In Proceedings of the 23rd International Conference on Computational Linguistics. Association for Computational Linguistics, 895--903.

Digital Library

[22]

Dragomir R. Radev, Pradeep Muthukrishnan, and Vahed Qazvinian. 2009. The ACL Anthology Network Corpus. In Proceedings, ACL Workshop on Natural Language Processing and Information Retrieval for Digital Libraries. Singapore.

Digital Library

[23]

Thomas Schoenemann. 2013. Training Nondeficient Variants of IBM-3 and IBM-4 for Word Alignment. In ACL (1). 22--31.

[24]

Parantu K. Shah, Carolina Perez-Iratxeta, Peer Bork, and Miguel A. Andrade. 2003. Information extraction from full text scientific articles: Where are the keywords? BMC Bioinformatics 4, 1 (2003), 1--9.

[25]

Simone Teufel and others. 2000. Argumentative zoning: Information extraction from scientific text. Ph.D. Dissertation. Citeseer.

[26]

Daya C Wimalasuriya and Dejing Dou. 2010. Ontology-based information extraction: An introduction and a survey of current approaches. Journal of Information Science (2010).

Digital Library

[27]

Chengxiang Zhai and John Lafferty. 2004. A study of smoothing methods for language models applied to information retrieval. ACM Transactions on Information Systems (TOIS) 22, 2 (2004), 179--214.

Digital Library

Cited By

LIAO WHUANG MMA PWANG Y(2021)Extracting Knowledge Entities from Sci-Tech Intelligence Resources Based on BiLSTM and Conditional Random FieldIEICE Transactions on Information and Systems10.1587/transinf.2020BDP0007E104.D:8(1214-1221)Online publication date: 1-Aug-2021
https://doi.org/10.1587/transinf.2020BDP0007
Hou LZhang JWu OYu TWang ZLi ZGao JYe YYao R(2021)Method and dataset entity mining in scientific literature: A CNN + BiLSTM model with self-attentionKnowledge-Based Systems10.1016/j.knosys.2021.107621(107621)Online publication date: Oct-2021
https://doi.org/10.1016/j.knosys.2021.107621
D’Souza JAuer S(2021)Pattern-Based Acquisition of Scientific Entities from Scholarly Article TitlesTowards Open and Trustworthy Digital Societies10.1007/978-3-030-91669-5_31(401-410)Online publication date: 30-Nov-2021
https://doi.org/10.1007/978-3-030-91669-5_31

Index Terms

AppTechMiner: Mining Applications and Techniques from Scientific Articles
1. Information systems
  1. Information systems applications
    1. Data mining

Recommendations

A case study validation of a knowledge-based approach for the selection of requirements engineering techniques

Requirements engineering (RE) is a critical phase in the software engineering process and plays a vital role in ensuring the overall quality of a software product. Recent research has shown that industry increasingly recognizes the importance of good RE ...
The BETTER Cross-Language Datasets
SIGIR '23: Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval

The IARPA BETTER (Better Extraction from Text Through Enhanced Retrieval) program held three evaluations of information retrieval (IR) and information extraction (IE). For both tasks, the only training data available was in English, but systems had to ...
Structured decomposition of adaptive applications

We describe an approach to automate certain high-level implementation decisions in a pervasive application, allowing them to be postponed until runtime. Our system enables a model in which an application programmer can specify the behavior of an ...

Comments

Information & Contributors

Information

Published In

cover image ACM Other conferences

WOSP 2017: Proceedings of the 6th International Workshop on Mining Scientific Publications

December 2017

72 pages

ISBN:9781450353885

DOI:10.1145/3127526

Copyright © 2017 ACM.

Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than ACM must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from [email protected]

In-Cooperation

Oak Ridge National Laboratory
OU: The Open University

Publisher

Association for Computing Machinery

New York, NY, United States

Publication History

Published: 15 December 2017

Permissions

Request permissions for this article.

Request Permissions

Check for updates

Author Tags

Qualifiers

Research-article
Research
Refereed limited

Conference

WOSP 2017

WOSP 2017: 6th International Workshop on Mining Scientific Publications

December 15, 2017

ON, Toronto, Canada

Acceptance Rates

WOSP 2017 Paper Acceptance Rate 11 of 17 submissions, 65%;

Overall Acceptance Rate 149 of 241 submissions, 62%

Contributors

Other Metrics

View Article Metrics

Bibliometrics & Citations

Bibliometrics

Article Metrics

3
Total Citations
View Citations
108
Total Downloads

Downloads (Last 12 months)6
Downloads (Last 6 weeks)0

Reflects downloads up to 13 Jan 2025

Other Metrics

View Author Metrics

Citations

Cited By

LIAO WHUANG MMA PWANG Y(2021)Extracting Knowledge Entities from Sci-Tech Intelligence Resources Based on BiLSTM and Conditional Random FieldIEICE Transactions on Information and Systems10.1587/transinf.2020BDP0007E104.D:8(1214-1221)Online publication date: 1-Aug-2021
https://doi.org/10.1587/transinf.2020BDP0007
Hou LZhang JWu OYu TWang ZLi ZGao JYe YYao R(2021)Method and dataset entity mining in scientific literature: A CNN + BiLSTM model with self-attentionKnowledge-Based Systems10.1016/j.knosys.2021.107621(107621)Online publication date: Oct-2021
https://doi.org/10.1016/j.knosys.2021.107621
D’Souza JAuer S(2021)Pattern-Based Acquisition of Scientific Entities from Scholarly Article TitlesTowards Open and Trustworthy Digital Societies10.1007/978-3-030-91669-5_31(401-410)Online publication date: 30-Nov-2021
https://doi.org/10.1007/978-3-030-91669-5_31

View Options

Login options

Check if you have access through your login credentials or your institution to get full access on this article.

Full Access

Get this Publication

View options

PDF

View or Download as a PDF file.

eReader

View online with eReader.

Media

Figures

Other

Tables

View Table of Contents