Jump to content

Corpus linguistics: Difference between revisions

From Wikipedia, the free encyclopedia
Content deleted Content added
m Gramática corregida
Tags: Mobile edit Mobile app edit iOS app edit
→‎See also: spell out
(28 intermediate revisions by 21 users not shown)
Line 1: Line 1:
{{short description|Branch of linguistics that studies language through examples contained in real texts}}
{{short description|Branch of linguistics that studies language through examples contained in real texts}}
{{Use dmy dates|date=October 2021}}
{{Use dmy dates|date=October 2021}}
'''Corpus linguistics''' is the [[study of language|study of a language]] as that language is expressed in its [[text corpus]] (plural ''corpora''), its body of "real world" text. Corpus linguistics proposes that a reliable analysis of a language is more feasible with corpora collected in the field—the natural context ("realia") of that language—with minimal experimental interference.
'''Corpus linguistics''' is an empirical method for the [[study of language]] by way of a [[text corpus]] (plural ''corpora'').<ref name=":0">{{Cite book |last=Meyer |first=Charles F. |title=English Corpus Linguistics |publisher=Cambridge University Press |year=2023 |edition=2nd |location=Cambridge |pages=4}}</ref> Corpora are balanced, often stratified collections of authentic, "real world", text of speech or writing that aim to represent a given [[linguistic variety]].<ref name=":0" /> Today, corpora are generally machine-readable data collections.


Corpus linguistics proposes that a reliable analysis of a language is more feasible with corpora collected in the field—the natural context ("realia") of that language—with minimal experimental interference. Large collections of text, though corpora may also be small in terms of running words, allow linguists to run quantitative analyses on linguistic concepts that may be difficult to test in a qualitative manner.<ref>{{Citation |last=Hunston |first=S. |title=Corpus Linguistics |date=2006-01-01 |url=https://www.sciencedirect.com/science/article/pii/B0080448542009445 |encyclopedia=Encyclopedia of Language & Linguistics (Second Edition) |pages=234–248 |editor-last=Brown |editor-first=Keith |access-date=2023-10-31 |place=Oxford |publisher=Elsevier |doi=10.1016/b0-08-044854-2/00944-5 |isbn=978-0-08-044854-1}}</ref>
The text-corpus method uses the body of texts written in any natural language to derive the set of abstract rules which govern that language. Those results can be used to explore the relationships between that subject language and other languages which have undergone a similar analysis. The first such corpora were manually derived from source texts, but now that work is automated.


The text-corpus method uses the body of texts in any natural language to derive the set of abstract rules which govern that language. Those results can be used to explore the relationships between that subject language and other languages which have undergone a similar analysis. The first such corpora were manually derived from source texts, but now that work is automated.
Corpora have not only been used for linguistics research, they have also been used to compile [[dictionaries]] (starting with ''[[The American Heritage Dictionary of the English Language]]'' in 1969) and grammar guides, such as ''[[A Comprehensive Grammar of the English Language]]'', published in 1985.

Corpora have not only been used for linguistics research, they have since the 1969 been increasingly used to compile [[dictionaries]] (starting with ''[[The American Heritage Dictionary of the English Language]]'' in 1969) and reference grammars, with ''[[A Comprehensive Grammar of the English Language]]'', published in 1985, as a first.


Experts in the field have differing views about the annotation of a corpus. These views range from [[John McHardy Sinclair]], who advocates minimal annotation so texts speak for themselves,<ref>Sinclair, J. 'The automatic analysis of corpora', in Svartvik, J. (ed.) ''Directions in Corpus Linguistics (Proceedings of Nobel Symposium 82)''. Berlin: Mouton de Gruyter. 1992.</ref> to the [[Survey of English Usage]] team ([[University College, London|University College]], London), who advocate annotation as allowing greater linguistic understanding through rigorous recording.<ref>Wallis, S. 'Annotation, Retrieval and Experimentation', in Meurman-Solin, A. & Nurmi, A.A. (ed.) Annotating Variation and Change. Helsinki: Varieng, [University of Helsinki]. 2007. [http://www.helsinki.fi/varieng/journal/volumes/01/wallis e-Published]</ref>
Experts in the field have differing views about the annotation of a corpus. These views range from [[John McHardy Sinclair]], who advocates minimal annotation so texts speak for themselves,<ref>Sinclair, J. 'The automatic analysis of corpora', in Svartvik, J. (ed.) ''Directions in Corpus Linguistics (Proceedings of Nobel Symposium 82)''. Berlin: Mouton de Gruyter. 1992.</ref> to the [[Survey of English Usage]] team ([[University College, London|University College]], London), who advocate annotation as allowing greater linguistic understanding through rigorous recording.<ref>Wallis, S. 'Annotation, Retrieval and Experimentation', in Meurman-Solin, A. & Nurmi, A.A. (ed.) Annotating Variation and Change. Helsinki: Varieng, [University of Helsinki]. 2007. [http://www.helsinki.fi/varieng/journal/volumes/01/wallis e-Published]</ref>
Line 16: Line 18:
=== English corpora ===
=== English corpora ===


A landmark in modern corpus linguistics was the publication of ''Computational Analysis of Present-Day American English'' in 1967. Written by [[Henry Kučera]] and [[W. Nelson Francis]], the work was based on an analysis of the [[Brown Corpus]], which was a contemporary compilation of about a million American English words, carefully selected from a wide variety of sources.<ref>{{cite book | last1=Francis | first1=W. Nelson | last2=Kučera | first2=Henry | title=Computational Analysis of Present-Day American English | publisher=Brown University Press | date=1 June 1967 | location=Providence | isbn= 978-0870571053}}</ref> Kučera and Francis subjected the Brown Corpus to a variety of computational analyses and then combined elements of linguistics, language teaching, [[psychology]], statistics, and sociology to create a rich and variegated opus. A further key publication was [[Randolph Quirk]]'s "Towards a description of English Usage" in 1960<ref>{{cite journal | last1=Quirk | first1= Randolph | title=Towards a description of English Usage | journal=Transactions of the Philological Society | date=November 1960 | pages=40–61 | volume=59 | issue=1| doi= 10.1111/j.1467-968X.1960.tb00308.x }}</ref> in which he introduced [[Survey of English Usage|the Survey of English Usage]].
A landmark in modern corpus linguistics was the publication of ''Computational Analysis of Present-Day American English'' in 1967. Written by [[Henry Kučera]] and [[W. Nelson Francis]], the work was based on an analysis of the [[Brown Corpus]], which is a structured and balanced corpus of one million words of American English from the year 1961. The corpus comprises 2000 text samples, from a variety of genres.<ref>{{cite book | last1=Francis | first1=W. Nelson | last2=Kučera | first2=Henry | title=Computational Analysis of Present-Day American English | publisher=Brown University Press | date=1 June 1967 | location=Providence | isbn= 978-0870571053}}</ref> The Brown Corpus was the first computerized corpus designed for linguistic research.<ref>{{Citation |last=Kennedy |first=G. |title=Corpus Linguistics |date=2001-01-01 |url=https://www.sciencedirect.com/science/article/pii/B0080430767030564 |encyclopedia=International Encyclopedia of the Social & Behavioral Sciences |pages=2816–2820 |editor-last=Smelser |editor-first=Neil J. |access-date=2023-10-31 |place=Oxford |publisher=Pergamon |isbn=978-0-08-043076-8 |editor2-last=Baltes |editor2-first=Paul B.}}</ref> Kučera and Francis subjected the Brown Corpus to a variety of computational analyses and then combined elements of linguistics, language teaching, [[psychology]], statistics, and sociology to create a rich and variegated opus. A further key publication was [[Randolph Quirk]]'s "Towards a description of English Usage" in 1960<ref>{{cite journal | last1=Quirk | first1= Randolph | title=Towards a description of English Usage | journal=Transactions of the Philological Society | date=November 1960 | pages=40–61 | volume=59 | issue=1| doi= 10.1111/j.1467-968X.1960.tb00308.x }}</ref> in which he introduced [[Survey of English Usage|the Survey of English Usage]]. Quirk's corpus was the first modern corpus to be built with the purpose of representing the whole language.<ref>{{Citation |last=Kennedy |first=G. |title=Corpus Linguistics |date=2001-01-01 |url=https://www.sciencedirect.com/science/article/pii/B0080430767030564 |encyclopedia=International Encyclopedia of the Social & Behavioral Sciences |pages=2816–2820 |editor-last=Smelser |editor-first=Neil J. |access-date=2023-10-31 |place=Oxford |publisher=Pergamon |doi=10.1016/b0-08-043076-7/03056-4 |isbn=978-0-08-043076-8 |editor2-last=Baltes |editor2-first=Paul B.}}</ref>


Shortly thereafter, Boston publisher [[Houghton-Mifflin]] approached Kučera to supply a million-word, three-line citation base for its new ''[[The American Heritage Dictionary of the English Language|American Heritage Dictionary]]'', the first [[dictionary]] compiled using corpus linguistics. The ''AHD'' took the innovative step of combining prescriptive elements (how language ''should'' be used) with descriptive information (how it actually ''is'' used).
Shortly thereafter, Boston publisher [[Houghton-Mifflin]] approached Kučera to supply a million-word, three-line citation base for its new ''[[The American Heritage Dictionary of the English Language|American Heritage Dictionary]]'', the first [[dictionary]] compiled using corpus linguistics. The ''AHD'' took the innovative step of combining prescriptive elements (how language ''should'' be used) with descriptive information (how it actually ''is'' used).
Line 26: Line 28:
The first computerized corpus of transcribed spoken language was constructed in 1971 by the Montreal French Project,<ref>{{cite journal | last1=Sankoff | first1=David | last2=Sankoff | first2=Gillian | title=Sample survey methods and computer-assisted analysis in the study of grammatical variation | journal=Canadian Languages in Their Social Context | location=Edmonton | publisher=Linguistic Research Incorporated | date=1973 | pages=7–63 | editor-last=Darnell | editor-first=R.}}</ref> containing one million words, which inspired [[Shana Poplack]]'s much larger corpus of spoken French in the Ottawa-Hull area.<ref>{{cite journal | last1=Poplack | first1=Shana | title=The care and handling of a mega-corpus | editor-first1=R. | editor-last1=Fasold | editor-first2=D. | editor-last2=Schiffrin | journal=Language Change and Variation | series=Current Issues in Linguistic Theory | location=Amsterdam | publisher=Benjamins | date=1989 | volume=52 |pages=411–451| doi=10.1075/cilt.52.25pop | isbn=978-90-272-3546-6 }}</ref>
The first computerized corpus of transcribed spoken language was constructed in 1971 by the Montreal French Project,<ref>{{cite journal | last1=Sankoff | first1=David | last2=Sankoff | first2=Gillian | title=Sample survey methods and computer-assisted analysis in the study of grammatical variation | journal=Canadian Languages in Their Social Context | location=Edmonton | publisher=Linguistic Research Incorporated | date=1973 | pages=7–63 | editor-last=Darnell | editor-first=R.}}</ref> containing one million words, which inspired [[Shana Poplack]]'s much larger corpus of spoken French in the Ottawa-Hull area.<ref>{{cite journal | last1=Poplack | first1=Shana | title=The care and handling of a mega-corpus | editor-first1=R. | editor-last1=Fasold | editor-first2=D. | editor-last2=Schiffrin | journal=Language Change and Variation | series=Current Issues in Linguistic Theory | location=Amsterdam | publisher=Benjamins | date=1989 | volume=52 |pages=411–451| doi=10.1075/cilt.52.25pop | isbn=978-90-272-3546-6 }}</ref>


=== Multilingual Corpora ===
=== Multilingual corpora ===


In the 1990s, many of the notable early successes on statistical methods in natural-language programming (NLP) occurred in the field of [[machine translation]], due especially to work at IBM Research. These systems were able to take advantage of existing multilingual [[text corpus|textual corpora]] that had been produced by the [[Parliament of Canada]] and the [[European Union]] as a result of laws calling for the translation of all governmental proceedings into all official languages of the corresponding systems of government.
In the 1990s, many of the notable early successes on statistical methods in natural-language programming (NLP) occurred in the field of [[machine translation]], due especially to work at IBM Research. These systems were able to take advantage of existing multilingual [[text corpus|textual corpora]] that had been produced by the [[Parliament of Canada]] and the [[European Union]] as a result of laws calling for the translation of all governmental proceedings into all official languages of the corresponding systems of government.


There are corpora in non-European languages as well. For example, the National Institute for Japanese Language and Linguistics in Japan has built a number of corpora of spoken and written Japanese.
There are corpora in non-European languages as well. For example, the National Institute for Japanese Language and Linguistics in Japan has built a number of corpora of spoken and written Japanese. [[Sign language]] corpora have also been created using video data.<ref>{{Cite web |title=National Center for Sign Language and Gesture Resources at B.U. |url=https://www.bu.edu/asllrp/cslgr/ |access-date=2023-10-31 |website=www.bu.edu}}</ref>


=== Ancient languages corpora ===
=== Ancient languages corpora ===
Line 45: Line 47:
| volume =40
| volume =40
| pages =43–61 [45]
| pages =43–61 [45]
}}</ref><ref>{{Citation | last =Eyland| first =E. Ann| year =1987 | contribution =Revelations from Word Counts | editor-last =Newing | editor-first =Edward G. | editor2-last =Conrad | editor2-first =Edgar W. | title =Perspectives on Language and Text: Essays and Poems in Honor of Francis I. Andersen's Sixtieth Birthday, July 28, 1985 | location =Winona Lake, IN | publisher =[[Eisenbrauns]] | page =51 | isbn =0-931464-26-9 }}</ref> The [[Quranic Arabic Corpus]] is an annotated corpus for the Classical Arabic language of the [[Quran]]. This is a recent project with multiple layers of annotation including morphological segmentation, [[part-of-speech tagging]], and syntactic analysis using dependency grammar.<ref>Dukes, K., Atwell, E. and Habash, N. 'Supervised Collaboration for Syntactic Annotation of Quranic Arabic'. ''Language Resources and Evaluation Journal''. 2011.</ref>
}}</ref><ref>{{Citation | last =Eyland| first =E. Ann| year =1987 | contribution =Revelations from Word Counts | editor-last =Newing | editor-first =Edward G. | editor2-last =Conrad | editor2-first =Edgar W. | title =Perspectives on Language and Text: Essays and Poems in Honor of Francis I. Andersen's Sixtieth Birthday, July 28, 1985 | location =Winona Lake, IN | publisher =[[Eisenbrauns]] | page =51 | isbn =0-931464-26-9 }}</ref> The [[Quranic Arabic Corpus]] is an annotated corpus for the Classical Arabic language of the [[Quran]]. This is a recent project with multiple layers of annotation including morphological segmentation, [[part-of-speech tagging]], and syntactic analysis using dependency grammar.<ref>Dukes, K., Atwell, E. and Habash, N. 'Supervised Collaboration for Syntactic Annotation of Quranic Arabic'. ''Language Resources and Evaluation Journal''. 2011.</ref> The Digital Corpus of Sanskrit (DCS) is a "Sandhi-split corpus of Sanskrit texts with full morphological and lexical analysis... designed for text-historical research in Sanskrit linguistics and philology."<ref>{{cite web |url=http://www.sanskrit-linguistics.org/dcs/#:~:text=The%20Digital%20Corpus%20of%20Sanskrit,in%20Sanskrit%20linguistics%20and%20philology. |title=Digital Corpus of Sanskrit (DCS) |access-date=2022-06-28}}</ref>


=== Corpora from specific fields ===
=== Corpora from specific fields ===


Besides pure linguistic inquiry, researchers had begun to apply corpus linguistics to other academic and professional fields, such as the emerging sub-discipline of [[Law and Corpus Linguistics]], which seeks to understand legal texts using corpus data and tools. The [[DBLP]] Discovery Dataset concentrates on [[computer science]], containing relevant computer science publications with sentient metadata such as author affiliations, citations, or study fields.<ref>{{Cite journal |last1=Wahle |first1=Jan Philip |last2=Ruas |first2=Terry |last3=Mohammad |first3=Saif |last4=Gipp |first4=Bela |date=2022 |title=D3: A Massive Dataset of Scholarly Metadata for Analyzing the State of Computer Science Research |url=https://aclanthology.org/2022.lrec-1.283 |journal=Proceedings of the Thirteenth Language Resources and Evaluation Conference |location=Marseille, France |publisher=European Language Resources Association |pages=2642–2651|arxiv=2204.13384 }}</ref> A more focused dataset was introduced by NLP Scholar, a combination of papers of the [[ACL Anthology]] and [[Google Scholar]] metadata.<ref>{{Cite journal |last=Mohammad |first=Saif M. |date=2020 |title=NLP Scholar: A Dataset for Examining the State of NLP Research |url=https://aclanthology.org/2020.lrec-1.109 |journal=Proceedings of the Twelfth Language Resources and Evaluation Conference |language=English |location=Marseille, France |publisher=European Language Resources Association |pages=868–877 |isbn=979-10-95546-34-4}}</ref> Corpora can also aid in translation efforts<ref>{{Citation |last=Bernardini |first=S. |title=Machine Readable Corpora |date=2006-01-01 |url=https://www.sciencedirect.com/science/article/pii/B0080448542004764 |encyclopedia=Encyclopedia of Language & Linguistics (Second Edition) |pages=358–375 |editor-last=Brown |editor-first=Keith |access-date=2023-10-31 |place=Oxford |publisher=Elsevier |doi=10.1016/b0-08-044854-2/00476-4 |isbn=978-0-08-044854-1}}</ref> or in teaching foreign languages.<ref>{{Cite web |last=Mainz |first=Johannes Gutenberg-Universität |title=Corpus Linguistics {{!}} ENGLISH LINGUISTICS |url=https://www.english-linguistics.uni-mainz.de/corpus-linguistics/ |access-date=2023-10-31 |website=Johannes Gutenberg-Universität Mainz |language=de-DE}}</ref>
Besides pure linguistic inquiry, researchers had begun to apply corpus linguistics to other academic and professional fields, such as the emerging sub-discipline of [[Law and Corpus Linguistics]], which seeks to understand legal texts using corpus data and tools.


== Methods ==
== Methods ==
Line 67: Line 69:
* [[Collocation]]
* [[Collocation]]
* [[Collostructional analysis]]
* [[Collostructional analysis]]
* [[Concordance (publishing)|Concordance]] ([[Key Word in Context|KWIC]])
* [[Concordance (publishing)|Concordance]] ([[Key Word in Context]])
* [[European Language Resources Association|European Language Resource Association]]
* [[Keyword (linguistics)]]
* [[Keyword (linguistics)]]
* [[Linguistic Data Consortium]]
* [[Linguistic Data Consortium]]
Line 81: Line 82:
* [[Translation memory]]
* [[Translation memory]]
* [[Treebank]]
* [[Treebank]]
* [[Word list]]


== Notes and references ==
== Notes and references ==
Line 93: Line 95:
* Facchinetti, R. and Rissanen M. (eds.) ''Corpus-based Studies of Diachronic English''. Bern: Peter Lang, 2006 {{ISBN|3-03910-851-4}}
* Facchinetti, R. and Rissanen M. (eds.) ''Corpus-based Studies of Diachronic English''. Bern: Peter Lang, 2006 {{ISBN|3-03910-851-4}}
* Lenders, W. ''Computational lexicography and corpus linguistics until ca. 1970/1980'', in: Gouws, R. H., Heid, U., Schweickard, W., Wiegand, H. E. (eds.) ''Dictionaries – An International Encyclopedia of Lexicography. Supplementary Volume: Recent Developments with Focus on Electronic and Computational Lexicography''. Berlin: De Gruyter Mouton, 2013 {{ISBN|978-3112146651}}
* Lenders, W. ''Computational lexicography and corpus linguistics until ca. 1970/1980'', in: Gouws, R. H., Heid, U., Schweickard, W., Wiegand, H. E. (eds.) ''Dictionaries – An International Encyclopedia of Lexicography. Supplementary Volume: Recent Developments with Focus on Electronic and Computational Lexicography''. Berlin: De Gruyter Mouton, 2013 {{ISBN|978-3112146651}}
* Fuß, Eric et al. (Eds.): ''Grammar and Corpora 2016'', Heidelberg: Heidelberg University Publishing, 2018. {{DOI| 10.17885/heiup.361.509}} ([https://heiup.uni-heidelberg.de/catalog/book/361?lang=en digital open access]).
* Fuß, Eric et al. (Eds.): ''Grammar and Corpora 2016'', Heidelberg: Heidelberg University Publishing, 2018. {{doi| 10.17885/heiup.361.509}} ([https://heiup.uni-heidelberg.de/catalog/book/361?lang=en digital open access]).
* Stefanowitsch A. 2020. ''Corpus linguistics: A guide to the methodology''. Berlin: Language Science Press. {{ISBN|978-3-96110-225-9}}, {{doi|10.5281/zenodo.3735822}} Open Access https://langsci-press.org/catalog/book/148.
* Stefanowitsch A. 2020. ''Corpus linguistics: A guide to the methodology''. Berlin: Language Science Press. {{ISBN|978-3-96110-225-9}}, {{doi|10.5281/zenodo.3735822}} Open Access https://langsci-press.org/catalog/book/148.


Line 114: Line 116:
==External links==
==External links==
{{commons category}}
{{commons category}}

* [http://martinweisser.org/corpora_site/CBLLinks.html Bookmarks for Corpus-based Linguists – very comprehensive site with categorized and annotated links to language corpora, software, references, etc.]
* [https://web.archive.org/web/20060113235630/http://torvald.aksis.uib.no/corpora/ Corpora discussion list]
* [http://corpus.byu.edu/ Freely-available, web-based corpora (100 million – 400 million words each): American (COCA, COHA), British (BNC), ''Time'', Spanish, Portuguese]
* [http://www.bmanuel.org/index.html Manuel Barbera's overview site]
* [https://web.archive.org/web/20110725203641/http://ifa.amu.edu.pl/~kprzemek/biblios/corpling.zip Przemek Kaszubski's list of references]
* [http://www.askoxford.com/oec/mainpage/oec01/?view=uk AskOxford.com] ''the composition and use of the Oxford Corpus''
* [https://archive.is/20121208123647/http://www.dmcbc.com.cn/ DMCBC.com]
* [https://translate.google.com/translate?hl=en&sl=zh-CN&tl=en&u=https%3A%2F%2Frp.liu233w.com%3A443%2Fhttp%2Fwww.dmcbc.com.cn%2F Datum Multilanguage Corpora Based on chinese free sample download]
* [http://www.corpus4u.org/ Corpus4u Community] a Chinese online forum for corpus linguistics
* [http://www.lancs.ac.uk/fss/courses/ling/corpus McEnery and Wilson's Corpus Linguistics Page]
* [https://groups.google.com/group/corpling-with-r Corpus Linguistics with R mailing list]
* [http://rdues.bcu.ac.uk/ Research and Development Unit for English Studies]
* [http://www.ucl.ac.uk/english-usage/ Survey of English Usage]
* [http://www.corpus.bham.ac.uk/ The Centre for Corpus Linguistics at Birmingham University]
* [http://corpus-analysis.com/ Tools for Corpus Linguistics (annotated list)]
* [http://www.corpus-linguistics.com Gateway to Corpus Linguistics on the Internet]: an annotated guide to corpus resources on the web
* [https://web.archive.org/web/20060920015213/http://compbio.uchsc.edu/corpora/ Biomedical corpora]
* [https://web.archive.org/web/20060830044341/http://www.ldc.upenn.edu/ Linguistic Data Consortium], a major distributor of corpora
* [http://www.ling.upenn.edu/hist-corpora Penn Parsed Corpora of Historical English]
* [http://www.ling.upenn.edu/hist-corpora Penn Parsed Corpora of Historical English]
* [http://corsis.sourceforge.net Corsis]: (formerly Tenka Text) an [[open source software|open-source]] ([[GPL]]ed) corpus analysis tool written in C#
* [http://www.ucl.ac.uk/english-usage/resources/icecup ICECUP] and [http://www.ucl.ac.uk/english-usage/resources/ftfs Fuzzy Tree Fragments]
* [https://web.archive.org/web/20070928002315/http://www.arts-humanities.net/text_mining Discussion group] [[text mining]]
* [https://plus.google.com/u/0/communities/101266284417587206243 Google+ discussion community on corpus linguistics for language learning and teaching]
* A corpus linguistics related conference MAG 2017: You can find some information and events related to [http://www.metadiscourseacrossgenres.com/ Metadiscourse Across Genres by visiting MAG 2017 website].
*[https://digital.lib.hkbu.edu.hk/corpus/index.php Corpus of Political Speeches], Free access to political speeches by American and Chinese politicians, developed by Hong Kong Baptist University Library
* [https://lighttag.io LightTag -Text Annotation Tool], A text annotation tool for machine learning corpus focused on team management
*[[LIVAC Synchronous Corpus]]


{{Portal bar|Languages}}
{{Portal bar|Languages}}

Revision as of 18:22, 10 June 2024

Corpus linguistics is an empirical method for the study of language by way of a text corpus (plural corpora).[1] Corpora are balanced, often stratified collections of authentic, "real world", text of speech or writing that aim to represent a given linguistic variety.[1] Today, corpora are generally machine-readable data collections.

Corpus linguistics proposes that a reliable analysis of a language is more feasible with corpora collected in the field—the natural context ("realia") of that language—with minimal experimental interference. Large collections of text, though corpora may also be small in terms of running words, allow linguists to run quantitative analyses on linguistic concepts that may be difficult to test in a qualitative manner.[2]

The text-corpus method uses the body of texts in any natural language to derive the set of abstract rules which govern that language. Those results can be used to explore the relationships between that subject language and other languages which have undergone a similar analysis. The first such corpora were manually derived from source texts, but now that work is automated.

Corpora have not only been used for linguistics research, they have since the 1969 been increasingly used to compile dictionaries (starting with The American Heritage Dictionary of the English Language in 1969) and reference grammars, with A Comprehensive Grammar of the English Language, published in 1985, as a first.

Experts in the field have differing views about the annotation of a corpus. These views range from John McHardy Sinclair, who advocates minimal annotation so texts speak for themselves,[3] to the Survey of English Usage team (University College, London), who advocate annotation as allowing greater linguistic understanding through rigorous recording.[4]

History

Some of the earliest efforts at grammatical description were based at least in part on corpora of particular religious or cultural significance. For example, Prātiśākhya literature described the sound patterns of Sanskrit as found in the Vedas, and Pāṇini's grammar of classical Sanskrit was based at least in part on analysis of that same corpus. Similarly, the early Arabic grammarians paid particular attention to the language of the Quran. In the Western European tradition, scholars prepared concordances to allow detailed study of the language of the Bible and other canonical texts.

English corpora

A landmark in modern corpus linguistics was the publication of Computational Analysis of Present-Day American English in 1967. Written by Henry Kučera and W. Nelson Francis, the work was based on an analysis of the Brown Corpus, which is a structured and balanced corpus of one million words of American English from the year 1961. The corpus comprises 2000 text samples, from a variety of genres.[5] The Brown Corpus was the first computerized corpus designed for linguistic research.[6] Kučera and Francis subjected the Brown Corpus to a variety of computational analyses and then combined elements of linguistics, language teaching, psychology, statistics, and sociology to create a rich and variegated opus. A further key publication was Randolph Quirk's "Towards a description of English Usage" in 1960[7] in which he introduced the Survey of English Usage. Quirk's corpus was the first modern corpus to be built with the purpose of representing the whole language.[8]

Shortly thereafter, Boston publisher Houghton-Mifflin approached Kučera to supply a million-word, three-line citation base for its new American Heritage Dictionary, the first dictionary compiled using corpus linguistics. The AHD took the innovative step of combining prescriptive elements (how language should be used) with descriptive information (how it actually is used).

Other publishers followed suit. The British publisher Collins' COBUILD monolingual learner's dictionary, designed for users learning English as a foreign language, was compiled using the Bank of English. The Survey of English Usage Corpus was used in the development of one of the most important Corpus-based Grammars, which was written by Quirk et al. and published in 1985 as A Comprehensive Grammar of the English Language.[9]

The Brown Corpus has also spawned a number of similarly structured corpora: the LOB Corpus (1960s British English), Kolhapur (Indian English), Wellington (New Zealand English), Australian Corpus of English (Australian English), the Frown Corpus (early 1990s American English), and the FLOB Corpus (1990s British English). Other corpora represent many languages, varieties and modes, and include the International Corpus of English, and the British National Corpus, a 100 million word collection of a range of spoken and written texts, created in the 1990s by a consortium of publishers, universities (Oxford and Lancaster) and the British Library. For contemporary American English, work has stalled on the American National Corpus, but the 400+ million word Corpus of Contemporary American English (1990–present) is now available through a web interface.

The first computerized corpus of transcribed spoken language was constructed in 1971 by the Montreal French Project,[10] containing one million words, which inspired Shana Poplack's much larger corpus of spoken French in the Ottawa-Hull area.[11]

Multilingual corpora

In the 1990s, many of the notable early successes on statistical methods in natural-language programming (NLP) occurred in the field of machine translation, due especially to work at IBM Research. These systems were able to take advantage of existing multilingual textual corpora that had been produced by the Parliament of Canada and the European Union as a result of laws calling for the translation of all governmental proceedings into all official languages of the corresponding systems of government.

There are corpora in non-European languages as well. For example, the National Institute for Japanese Language and Linguistics in Japan has built a number of corpora of spoken and written Japanese. Sign language corpora have also been created using video data.[12]

Ancient languages corpora

Besides these corpora of living languages, computerized corpora have also been made of collections of texts in ancient languages. An example is the Andersen-Forbes database of the Hebrew Bible, developed since the 1970s, in which every clause is parsed using graphs representing up to seven levels of syntax, and every segment tagged with seven fields of information.[13][14] The Quranic Arabic Corpus is an annotated corpus for the Classical Arabic language of the Quran. This is a recent project with multiple layers of annotation including morphological segmentation, part-of-speech tagging, and syntactic analysis using dependency grammar.[15] The Digital Corpus of Sanskrit (DCS) is a "Sandhi-split corpus of Sanskrit texts with full morphological and lexical analysis... designed for text-historical research in Sanskrit linguistics and philology."[16]

Corpora from specific fields

Besides pure linguistic inquiry, researchers had begun to apply corpus linguistics to other academic and professional fields, such as the emerging sub-discipline of Law and Corpus Linguistics, which seeks to understand legal texts using corpus data and tools. The DBLP Discovery Dataset concentrates on computer science, containing relevant computer science publications with sentient metadata such as author affiliations, citations, or study fields.[17] A more focused dataset was introduced by NLP Scholar, a combination of papers of the ACL Anthology and Google Scholar metadata.[18] Corpora can also aid in translation efforts[19] or in teaching foreign languages.[20]

Methods

Corpus linguistics has generated a number of research methods, which attempt to trace a path from data to theory. Wallis and Nelson (2001)[21] first introduced what they called the 3A perspective: Annotation, Abstraction and Analysis.

  • Annotation consists of the application of a scheme to texts. Annotations may include structural markup, part-of-speech tagging, parsing, and numerous other representations.
  • Abstraction consists of the translation (mapping) of terms in the scheme to terms in a theoretically motivated model or dataset. Abstraction typically includes linguist-directed search but may include e.g., rule-learning for parsers.
  • Analysis consists of statistically probing, manipulating and generalising from the dataset. Analysis might include statistical evaluations, optimisation of rule-bases or knowledge discovery methods.

Most lexical corpora today are part-of-speech-tagged (POS-tagged). However even corpus linguists who work with 'unannotated plain text' inevitably apply some method to isolate salient terms. In such situations annotation and abstraction are combined in a lexical search.

The advantage of publishing an annotated corpus is that other users can then perform experiments on the corpus (through corpus managers). Linguists with other interests and differing perspectives than the originators' can exploit this work. By sharing data, corpus linguists are able to treat the corpus as a locus of linguistic debate and further study.[22]

See also

Notes and references

  1. ^ a b Meyer, Charles F. (2023). English Corpus Linguistics (2nd ed.). Cambridge: Cambridge University Press. p. 4.
  2. ^ Hunston, S. (1 January 2006), "Corpus Linguistics", in Brown, Keith (ed.), Encyclopedia of Language & Linguistics (Second Edition), Oxford: Elsevier, pp. 234–248, doi:10.1016/b0-08-044854-2/00944-5, ISBN 978-0-08-044854-1, retrieved 31 October 2023
  3. ^ Sinclair, J. 'The automatic analysis of corpora', in Svartvik, J. (ed.) Directions in Corpus Linguistics (Proceedings of Nobel Symposium 82). Berlin: Mouton de Gruyter. 1992.
  4. ^ Wallis, S. 'Annotation, Retrieval and Experimentation', in Meurman-Solin, A. & Nurmi, A.A. (ed.) Annotating Variation and Change. Helsinki: Varieng, [University of Helsinki]. 2007. e-Published
  5. ^ Francis, W. Nelson; Kučera, Henry (1 June 1967). Computational Analysis of Present-Day American English. Providence: Brown University Press. ISBN 978-0870571053.
  6. ^ Kennedy, G. (1 January 2001), "Corpus Linguistics", in Smelser, Neil J.; Baltes, Paul B. (eds.), International Encyclopedia of the Social & Behavioral Sciences, Oxford: Pergamon, pp. 2816–2820, ISBN 978-0-08-043076-8, retrieved 31 October 2023
  7. ^ Quirk, Randolph (November 1960). "Towards a description of English Usage". Transactions of the Philological Society. 59 (1): 40–61. doi:10.1111/j.1467-968X.1960.tb00308.x.
  8. ^ Kennedy, G. (1 January 2001), "Corpus Linguistics", in Smelser, Neil J.; Baltes, Paul B. (eds.), International Encyclopedia of the Social & Behavioral Sciences, Oxford: Pergamon, pp. 2816–2820, doi:10.1016/b0-08-043076-7/03056-4, ISBN 978-0-08-043076-8, retrieved 31 October 2023
  9. ^ Quirk, Randolph; Greenbaum, Sidney; Leech, Geoffrey; Svartvik, Jan (1985). A Comprehensive Grammar of the English Language. London: Longman. ISBN 978-0582517349.
  10. ^ Sankoff, David; Sankoff, Gillian (1973). Darnell, R. (ed.). "Sample survey methods and computer-assisted analysis in the study of grammatical variation". Canadian Languages in Their Social Context. Edmonton: Linguistic Research Incorporated: 7–63.
  11. ^ Poplack, Shana (1989). Fasold, R.; Schiffrin, D. (eds.). "The care and handling of a mega-corpus". Language Change and Variation. Current Issues in Linguistic Theory. 52. Amsterdam: Benjamins: 411–451. doi:10.1075/cilt.52.25pop. ISBN 978-90-272-3546-6.
  12. ^ "National Center for Sign Language and Gesture Resources at B.U." www.bu.edu. Retrieved 31 October 2023.
  13. ^ Andersen, Francis I.; Forbes, A. Dean (2003), "Hebrew Grammar Visualized: I. Syntax", Ancient Near Eastern Studies, vol. 40, pp. 43–61 [45]
  14. ^ Eyland, E. Ann (1987), "Revelations from Word Counts", in Newing, Edward G.; Conrad, Edgar W. (eds.), Perspectives on Language and Text: Essays and Poems in Honor of Francis I. Andersen's Sixtieth Birthday, July 28, 1985, Winona Lake, IN: Eisenbrauns, p. 51, ISBN 0-931464-26-9
  15. ^ Dukes, K., Atwell, E. and Habash, N. 'Supervised Collaboration for Syntactic Annotation of Quranic Arabic'. Language Resources and Evaluation Journal. 2011.
  16. ^ "Digital Corpus of Sanskrit (DCS)". Retrieved 28 June 2022.
  17. ^ Wahle, Jan Philip; Ruas, Terry; Mohammad, Saif; Gipp, Bela (2022). "D3: A Massive Dataset of Scholarly Metadata for Analyzing the State of Computer Science Research". Proceedings of the Thirteenth Language Resources and Evaluation Conference. Marseille, France: European Language Resources Association: 2642–2651. arXiv:2204.13384.
  18. ^ Mohammad, Saif M. (2020). "NLP Scholar: A Dataset for Examining the State of NLP Research". Proceedings of the Twelfth Language Resources and Evaluation Conference. Marseille, France: European Language Resources Association: 868–877. ISBN 979-10-95546-34-4.
  19. ^ Bernardini, S. (1 January 2006), "Machine Readable Corpora", in Brown, Keith (ed.), Encyclopedia of Language & Linguistics (Second Edition), Oxford: Elsevier, pp. 358–375, doi:10.1016/b0-08-044854-2/00476-4, ISBN 978-0-08-044854-1, retrieved 31 October 2023
  20. ^ Mainz, Johannes Gutenberg-Universität. "Corpus Linguistics | ENGLISH LINGUISTICS". Johannes Gutenberg-Universität Mainz (in German). Retrieved 31 October 2023.
  21. ^ Wallis, S. and Nelson G. Knowledge discovery in grammatically analysed corpora. Data Mining and Knowledge Discovery, 5: 307–340. 2001.
  22. ^ Baker, Paul; Egbert, Jesse, eds. (2016). Triangulating Methodological Approaches in Corpus-Linguistic Research. New York: Routledge.

Further reading

Books

  • Biber, D., Conrad, S., Reppen R. Corpus Linguistics, Investigating Language Structure and Use, Cambridge: Cambridge UP, 1998. ISBN 0-521-49957-7
  • McCarthy, D., and Sampson G. Corpus Linguistics: Readings in a Widening Discipline, Continuum, 2005. ISBN 0-8264-8803-X
  • Facchinetti, R. Theoretical Description and Practical Applications of Linguistic Corpora. Verona: QuiEdit, 2007 ISBN 978-88-89480-37-3
  • Facchinetti, R. (ed.) Corpus Linguistics 25 Years on. New York/Amsterdam: Rodopi, 2007 ISBN 978-90-420-2195-2
  • Facchinetti, R. and Rissanen M. (eds.) Corpus-based Studies of Diachronic English. Bern: Peter Lang, 2006 ISBN 3-03910-851-4
  • Lenders, W. Computational lexicography and corpus linguistics until ca. 1970/1980, in: Gouws, R. H., Heid, U., Schweickard, W., Wiegand, H. E. (eds.) Dictionaries – An International Encyclopedia of Lexicography. Supplementary Volume: Recent Developments with Focus on Electronic and Computational Lexicography. Berlin: De Gruyter Mouton, 2013 ISBN 978-3112146651
  • Fuß, Eric et al. (Eds.): Grammar and Corpora 2016, Heidelberg: Heidelberg University Publishing, 2018. doi:10.17885/heiup.361.509 (digital open access).
  • Stefanowitsch A. 2020. Corpus linguistics: A guide to the methodology. Berlin: Language Science Press. ISBN 978-3-96110-225-9, doi:10.5281/zenodo.3735822 Open Access https://langsci-press.org/catalog/book/148.

Book series

Book series in this field include:

Journals

There are several international peer-reviewed journals dedicated to corpus linguistics, for example: