Compilation transcription and usage of a reference speech corpus the case of the slovene corpus gos2013Multext east morphosyntactic resources for central and eastern european languages2012The imp historical slovene language resources2015Korpusi slovenskega jezika Gigafida, KRES, ccGigafida in ccKRES: gradnja, vsebina, uporaba · Znanstvena založba Filozofske fakultete Univerze v Ljubljani (Ljubljana University Press, Faculty of Arts) eBooks2020
Part-of-Speech Tagging for Twitter: Annotation, Features, and Experiments · Figshare2018
Language Identification and Morphosyntactic Tagging: The Second VarDial Evaluation Campaign · Työväentutkimus Vuosikirja2018
Comparing CRF and LSTM performance on the task of morphosyntactic tagging of non-standard varieties of South Slavic languages · International Conference on Computational Linguistics2018
Viri, orodja in metode za analizo spletne slovenščine · Znanstvena založba Filozofske fakultete2018
(Ne)normativnost računalniško posredovane komunikacije v slovenščini: merilo vejice · Znanstvena založba Filozofske fakultete2018
Six Challenges for Neural Machine Translation2017
Legal Framework, Dataset and Annotation Schema for Socially Unacceptable Online Discourse Practices in Slovene2017
Adapting a State-of-the-Art Tagger for South Slavic Languages to Non-Standard Text2017
Learning attention for historical text normalization by learning to pronounce2017
Gender Profiling for Slovene Twitter communication: the Influence of Gender Marking, Content and Style2017
The CLIN27 Shared Task: Translating Historical Text to Contemporary Language for Improving Automatic Linguistic Annotation2017
Multilingual Twitter Sentiment Classification: The Role of Human Annotators · PLoS ONE2016
Phytoplankton across Tropical and Subtropical Regions of the Atlantic, Indian and Pacific Oceans · PLoS ONE2016
Slovene Twitter Analytics2016
Syntactic Annotation of Slovene CMC: First Steps2016
Corpus-Based Diacritic Restoration for South Slavic Languages2016
The DiDi Corpus of South Tyrolean CMC Data: A multilingual corpus of Facebook texts · Accademia University Press eBooks2016
A Multi-domain Corpus of Swedish Word Sense Annotation2016
Corpus vs. Lexicon Supervision in Morphosyntactic Tagging: the Case of Slovene2016
Building a Social Media Adapted PoS Tagger Using FlexTag – A Case Study on Italian Tweets · Accademia University Press eBooks2016
Broad Twitter Corpus: A Diverse Named Entity Recognition Resource · International Conference on Computational Linguistics2016
Overview of the Evalita 2016 SENTIment POLarity Classification Task · Accademia University Press eBooks2016
Private or Corporate? Predicting User Types on Twitter · International Conference on Computational Linguistics2016
Gold-Standard Datasets for Annotation of Slovene Computer-Mediated Communication. · RASLAN2016
Normalising Slovene data: historical texts vs. user-generated content. · KONVENS2016
Bidirectional LSTM-CRF Models for Sequence Tagging · arXiv (Cornell University)2015
Predicting the Level of Text Standardness in User-generated Content · Recent Advances in Natural Language Processing2015
Adding Value to CMC Corpora: CLARINification and Part-of-speech Annotation of the Dortmund Chat Corpus · MADOC (University of Mannheim)2015
Sentiment Analysis: Mining Opinions, Sentiments, and Emotions2015
Sentiment Analysis · Cambridge University Press eBooks2015
Standardizing Tweets with Character-Level Machine Translation · Lecture notes in computer science2014
Stream-based active learning for sentiment analysis in the financial domain · Information Sciences2014
The slWaC 2.0 Corpus of the Slovene Web2014
Analysis of named entity recognition and linking for tweets · Information Processing & Management2014
TweetCaT: a tool for building Twitter corpora of smaller languages2014
Building Linguistic Corpora from Wikipedia Articles and Discussions · LDV-Forum/Journal for language technology and computational linguistics2014
The CoMeRe corpus for French : structuring and annotating heterogeneous CMC genres · LDV-Forum/Journal for language technology and computational linguistics2014
WebAnno: A Flexible, Web-based and Visually Supported System for Distributed Annotations · TUbilio (Technical University of Darmstadt)2013
Optimierung des Stuttgart-Tübingen-Tagset für die linguistische Annotation von Korpora zur internetbasierten Kommunikation: Phänomene, Herausforderungen, Erweiterungsvorschläge · LDV-Forum/Journal for language technology and computational linguistics2013
Natural Language Annotation for Machine Learning · CERN Document Server (European Organization for Nuclear Research)2012
Communication Processes in Participatory Websites · Journal of Computer-Mediated Communication2012
A TEI Schema for the Representation of Computer-mediated Communication · Journal of the Text Encoding Initiative2012
Computational Linguistics · Studies in computational intelligence2012
Sms4science: An International Corpus-Based Texting Project and the Specific Challenges for Multilingual Switzerland · Zurich Open Repository and Archive (University of Zurich)2011
Sms4science: An International Corpus-Based Texting Project and the Specific Challenges for Multilingual Switzerland · Oxford University Press eBooks2011
The JOS Linguistically Tagged Corpus of Slovene2010
Corpus Linguistics An International Handbook2009
VARD2 : a tool for dealing with spelling variation in historical corpora · Lancaster EPrints (Lancaster University)2008
Corpora of computer-mediated communication · MADOC (University of Mannheim)2008
Participative Web And User-Created Content: Web 2.0 Wikis and Social Networking2007
Similarity Measures for Short Segments of Text · Lecture notes in computer science2007
Manatee/Bonito - A Modular Corpus Manager · RASLAN2007
Massive multi lingual corpus compilation: Acquis Communautaire and totale2005
Guidelines for electronic text encoding and interchange · Medical Entomology and Zoology1994
Class-based n -gram models of natural language1992
Content Analysis: An Introduction to its Methodology. · Journal of the American Statistical Association1984