Developing a corpus of plagiarised short answers2011Journal of Universal Computer Science · Zenodo (CERN European Organization for Nuclear Research)2020
Language and Law=Linguagem e Direito · Círculo de lingüística aplicada a la comunicación2017
Paraphrase Identification by Using Clause-Based Similarity Features and Machine Translation Metrics · The Computer Journal2015
Cross-Language Urdu-English (CLUE) Text Alignment Corpus: Notebook for PAN at CLEF 2015. · CLEF (Working Notes)2015
Cross-Language Urdu-English (CLUE) Text Alignment Corpus2015
Is This a Paraphrase? What Kind? Paraphrase Boundaries and Typology · Open Journal of Modern Linguistics2014
Detecting Translingual Plagiarism And The Backlash Against Translation Plagiarists · Portuguese National Funding Agency for Science, Research and Technology (RCAAP Project by FCT)2014
A Winning Approach to Text Alignment for Text Reuse Detection at PAN 2014. · CLEF (Working Notes)2014
Plagiarism Meets Paraphrasing: Insights for the Next Generation in Automatic Plagiarism Detection · Computational Linguistics2013
Survey of Text Plagiarism Detection · Computer Engineering and Applications Journal2012
Determining and characterizing the reused text for plagiarism detection · Expert Systems with Applications2012
Text Reuse Detection Using a Composition of Text Similarity Measures · TUbilio (Technical University of Darmstadt)2012
Data Mining: Practical Machine Learning Tools and Techniques · Elsevier eBooks2011
Overview of the 2nd International Competition on Plagiarism Detection2011
The Dissemination of News and the Emergence of Contemporaneity in Early Modern Europe · Media History2011
Plagiarism detection using stopword n-grams · Journal of the American Society for Information Science and Technology2011
An Evaluation Framework for Plagiarism Detection2010
Automatic Detection of Local Reuse · Lecture notes in computer science2010
Evaluating text reuse discovery on the web2010
Text Relatedness Based on a Word Thesaurus · Journal of Artificial Intelligence Research2010
Reflections on measuring text reuse from a copyright law perspective · Cambridge University Press eBooks2010
A corpus-based approach to text reuse in the newsbooks of the Commonwealth · Lancaster EPrints (Lancaster University)2010
Reducing the Plagiarism Detection Search Space on the Basis of the Kullback-Leibler Distance · Lecture notes in computer science2009
Finding text reuse on the web2009
The toolbox for local and global plagiarism detection · Computers & Education2009
The WEKA data mining software · ACM SIGKDD Explorations Newsletter2009
Local text reuse detection2008
Detection of Duplicate Defect Reports Using Natural Language Processing · Proceedings/Proceedings - International Conference on Software Engineering2007
Corpus-based and knowledge-based measures of text semantic similarity · University of North Texas Digital Library (University of North Texas)2006
A Survey of Automatic Urdu Language Processing2006
Plagiarism - A Survey · Zenodo (CERN European Organization for Nuclear Research)2006
RCV1: A New Benchmark Collection for Text Categorization Research · Goldsmiths (University of London)2004
Extracting multiword expressions with a semantic tagger2003
Methods for identifying versioned and plagiarized documents · Journal of the American Society for Information Science and Technology2003
A study in Urdu corpus construction2002
Collection statistics for fast duplicate document detection · ACM Transactions on Information Systems2002
On the resemblance and containment of documents2002
Detecting Short Passages of Similar Text in Large Document Collections · University of Hertfordshire Research Archive (University of Hertfordshire)2001
METER2001
The METER corpus : a corpus for analysing journalistic text reuse · Lancaster EPrints (Lancaster University)2001
Corpus Resources and Minority Language Engineering2000
Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition2000
The decomposition of human-written summary sentences1999
Sim1999
Sim · ACM SIGCSE Bulletin1999
SCAM: A Copy Detection Mechanism for Digital Documents1995
Copy detection mechanisms for digital documents1995
<b>The language of news media</b> . By Allan Bell. (Language in society, 16.) Oxford, UK, & Cambridge, MA: Blackwell, 1991. Pp. xv, 277. Cloth $52.95, paper $18.95. · Language1994
Detection of similarities in student programs1992
Detection of similarities in student programs · ACM SIGCSE Bulletin1992
A vector space model for automatic indexing · Communications of the ACM1975
Selected Tables in Mathematical Statistics · Mathematics of Computation1971
Weighted kappa: Nominal scale agreement provision for scaled disagreement or partial credit. · Psychological Bulletin1968
A Coefficient of Agreement for Nominal Scales · Educational and Psychological Measurement1960
On Sentence-Length as a Statistical Characteristic of Style in Prose: With Application to Two Cases of Disputed Authorship · Biometrika1939