The 385 million word corpus of contemporary american english 19902008 design architecture and linguistic insights2009The wacky wide web a collection of very large linguistically processed web crawled corpora2009Googleology as Smart Lexicography: Big Messy Data for Better Regional Labels · Dictionaries2016
The Stanford CoreNLP Natural Language Processing Toolkit2014
Biber Redux: Reconsidering Dimensions of Variation in American English2014
Semi-supervised Graph-based Genre Classification for Web Pages2014
{bs,hr,sr}WaC - Web Corpora of Bosnian, Croatian and Serbian2014
Web Corpus Construction · Synthesis lectures on human language technologies2013
StirWaC: compiling a diverse corpus based on texts from the web for South Tyrolean German · View2013
On the autonomy and homogeneity of Canadian English · World Englishes2012
Tony McEnery and Andrew Hardie. Corpus Linguistics: Method, Theory and Practice. · International Journal of Lexicography2012
Building a 70 billion word corpus of English from ClueWeb2012
Langid.py for better language modelling2012
Using Web Corpora for the Recognition of Regional Variation in Standard German Collocations · edoc (University of Basel)2012
Cross-domain Feature Selection for Language Identification2011
Removing Boilerplate and Duplicate Content from Web Corpora2011
PaddyWaC: A Minimally-Supervised Web-Corpus of Hiberno-English · Research Portal (Queen's University Belfast)2011
Building Webcorpora of Academic Prose with BootCaT2010
The automatic identification of lexical variation between language varieties · Natural Language Engineering2010
A way with words: recent advances in lexical theory and analysis: a festschrift for Patrick Hanks · Ghent University Academic Bibliography (Ghent University)2010
The DANTE Database: Its Contribution to English Lexical Research, and in Particular to Complementing the FrameNet Data.2010
A Corpus Factory for Many Languages2010
Natural Language Processing with Python · CERN Document Server (European Organization for Nuclear Research)2009
Simple Maths for Keywords2009
Overview of the TREC 2009 Web Track · Text REtrieval Conference2009
The Architecture of a MultipurposeAustralian National Corpus2009
Selected proceedings of the 2008 HCSNet workshop on designing the Australian National Corpus: Mustering Languages · Griffith Research Online (Griffith University, Queensland, Australia)2009
Cleaneval: a Competition for Cleaning Web Pages2008
Introducing and evaluating ukWaC, a very large Web-derived corpus of English2008
Googleology is Bad Science · Computational Linguistics2007
WebBootCaT: a Web Tool for Instant Corpora · Institutional Research Information System (Università degli Studi di Trento)2006
The Sketch Engine2004
BootCaT: Bootstrapping Corpora and Terms from the Web2004
The Web as a Parallel Corpus · Computational Linguistics2003
Scaling to very very large corpora for natural language disambiguation2001
Comparing Corpora · International Journal of Corpus Linguistics2001
How dynamic is the Web? · Computer Networks2000
<b>Dictionary of American Regional English:</b> Volumes I-III. Chief Ed., Frederic G. Cassidy, Assoc. Ed., Joan Houston Hall. Cambridge, MA: Belknap Press of Harvard University Press. - Vol. I (A-C), 1985. Pp. clvi, 903. - Vol. II (D-H), 1991. Pp. 1175. - Vol. III (1-0), 1996. Pp. 927. · Language1998
Canadian Oxford Dictionary1998
N-gram-based text categorization1994
The Australian national dictionary: a dictionary of Australianisms on historical principles · Choice Reviews Online1989