Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
One of the principal findings of research on talk-in-interaction is that speakers design their utterances to perform specific social actions (Searle 1969; Atkinson & Heritage 1984; Levinson 2013). These actions, and the forms with which they are accomplished, do not arise haphazardly, but are instead sequentially organised in response to the speech of others and in keeping with the norms of particular speech events and activities (e.g., Labov 1971; Hymes 1974; Jefferson 1978; Schegloff 1982; Sacks 1987; Finegan & Biber 1994). In this chapter, I apply this basic precept to the analysis of a corpus of responses in online question-and-answer (Q+A) forums across three different topics (Family & Relationships, Politics & Government, and Society & Culture) and four world varieties of English (India, the Philippines, the UK, and the US). Like the other chapters in this volume, my central research question is how language use in these responses differs across topics and varieties. I approach this question through a close qualitative analysis of a small subset of the Q+A response corpus. The subset consists of 12 ‘texts’—one per topic for each of the four varieties-made up of the initial question posed and all of the responses provided. The texts themselves were automatically extracted from a larger corpus using ProtAnt, a tool that identifies the most (lexically) ‘typical’ text in a corpus by identifying texts that contain the most keywords when compared to a reference corpus1 (Anthony & Baker 2015). Table 10.1 lists the files that were identified as most typical in this way. Thus while the subcorpus I analyse in the following section is admittedly rather small (a total of 15,140 words with a mean value of 1,261.7 words per text), it is in a sense representative of the larger corpus from which it is drawn. The benefit of focusing on a smaller subcorpus is that doing so enables a detailed examination of certain aspects of linguistic form and content that would be difficult to explore in a larger sample. That said, it is nevertheless important to highlight that the findings to be discussed are based on a restricted empirical set, and it is therefore necessary to exercise caution when attempting to generalise from any patterns identified.
Cette étude propose d’analyser le point de vue des automobilistes sur la circulation inter-files (CIF) des deux-roues motorisés (2RM). Jamais questionnés jusqu’alors sur ce comportement typique 2RM, c’est pourtant une pratique qui les implique du point de vue opératoire, bien qu’ils n’en soient pas à l’initiative. Pour cela, soixante entretiens semi-directifs auprès d’automobilistes choisis en fonction de 3 critères (ville de mobilité, ancienneté du permis de conduire B et pratique ou non du 2RM) ont été conduits et ont permis de recueillir un corpus lexical riche d’informations. Ce corpus a fait l’objet d’une analyse informatisée grâce au logiciel ALCESTE. Les résultats de cette analyse fine soulignent, entre autres, l’importance de l’expertise des individus dans le domaine du 2RM et l’importance du contexte de circulation et des normes sociales s’y référant sur la pratique et les attitudes vis-à-vis de la CIF.
This paper is an extended description of SemEval-2014 Task 1, the task on the evaluation of Compositional Distributional Semantics Models on full sentences. Systems participating in the task were presented with pairs of sentences and were evaluated on their ability to predict human judgments on (1) semantic relatedness and (2) entailment. Training and testing data were subsets of the SICK (Sentences Involving Compositional Knowledge) data set. SICK was developed with the aim of providing a proper benchmark to evaluate compositional semantic systems, though task participation was open to systems based on any approach. Taking advantage of the SemEval experience, in this paper we analyze the SICK data set, in order to evaluate the extent to which it meets its design goal and to shed light on the linguistic phenomena that are still challenging for state-of-the-art computational semantic systems. Qualitative and quantitative error analyses show that many systems are quite sensitive to changes in the proportion of sentence pair types, and degrade in the presence of additional lexico-syntactic complexities which do not affect human judgements. More compositional systems seem to perform better when the task proportions are changed, but the effect needs further confirmation.
In this paper, we investigate four important issues together for explicit discourse relation labelling in Chinese texts: (1) discourse connective extraction, (2) linking ambiguity resolution, (3) relation type disambiguation, and (4) argument boundary identification. In a pipelined Chinese discourse parser, we identify potential connective candidates by string matching, eliminate non-discourse usages from them with a binary classifier, resolve linking ambiguities among connective components by ranking, disambiguate relation types by a multiway classifier, and determine the argument boundaries by conditional random fields. The experiments on Chinese Discourse Treebank show that the F1 scores of 0.7506, 0.7693, 0.7458, and 0.3134 are achieved for discourse usage disambiguation, linking disambiguation, relation type disambiguation, and argument boundary identification, respectively, in a pipelined Chinese discourse parser.
This paper describes FinnPos, an open-source morphological tagging and lemmatization toolkit for Finnish. The morphological tagging model is based on the averaged structured perceptron classifier. Given training data, new taggers are estimated in a computationally efficient manner using a combination of beam search and model cascade. The lemmatization is performed employing a combination of a rule-based morphological analyzer, OMorFi, and a data-driven lemmatization model. The toolkit is readily applicable for tagging and lemmatization of running text with models learned from the recently published Finnish Turku Dependency Treebank and FinnTreeBank. Empirical evaluation on these corpora shows that FinnPos performs favorably compared to reference systems in terms of tagging and lemmatization accuracy. In addition, we demonstrate that our system is highly competitive with regard to computational efficiency of learning new models and assigning analyses to novel sentences.
We present the development and evaluation of a semantic analysis task that lies at the intersection of two very trendy lines of research in contemporary computational linguistics: (1) sentiment analysis, and (2) natural language processing of social media text. The task was part of SemEval, the International Workshop on Semantic Evaluation, a semantic evaluation forum previously known as SensEval. The task ran in 2013 and 2014, attracting the highest number of participating teams at SemEval in both years, and there is an ongoing edition in 2015. The task included the creation of a large contextual and message-level polarity corpus consisting of tweets, SMS messages, LiveJournal messages, and a special test set of sarcastic tweets. The evaluation attracted 44 teams in 2013 and 46 in 2014, who used a variety of approaches. The best teams were able to outperform several baselines by sizable margins with improvement across the 2 years the task has been run. We hope that the long-lasting role of this task and the accompanying datasets will be to serve as a test bed for comparing different approaches, thus facilitating research.
We propose a method for improving the dependency parsing of complex sentences. This method assumes segmentation of input sentences into clauses and does not require to re-train a parser of one's choice. We represent a sentence clause structure using clause charts that provide a layer of embedding for each clause in the sentence. Then we formulate a parsing strategy as a two-stage process where (i) coordinated and subordinated clauses of the sentence are parsed separately with respect to the sentence clause chart and (ii) their dependency trees become subtrees of the final tree of the sentence. The object language is Czech and the parser used is a maximum spanning tree parser trained on the Prague Dependency Treebank. We have achieved an average 0.97% improvement in the unlabeled attachment score. Although the method has been designed for the dependency parsing of Czech, it is useful for other parsing techniques and languages.
Images play an important role in the representation and acquisition of specialized knowledge. Not surprisingly, terminological knowledge bases (TKBs) often include images as a way to enhance the information in concept entries. However, the selection of these images should not be random, but rather based on specific guidelines that take into account the type and nature of the concept being described. This paper presents a proposal on how to combine the features of images with the conceptual propositions in EcoLexicon, a multilingual TKB on the environment. This proposal is based on the following: (1) the combinatory possibilities of concept types; (2) image types, such as photographs, drawings and flow charts; (3) morphological features or visual knowledge patterns (VKPs), such as labels, colours, arrows, and their effect on the functional nature of each image type. Currently, images are stored in association with concept entries according to the semantic content of their definitions, but they are not described or annotated according to the parameters that guided their selection, which would undoubtedly contribute to the systematization and automatization of the process. First, the images included in EcoLexicon were analyzed in terms of their adequateness, the semantic relations expressed, the concept types and their VKPs. Then, with these data, guidelines for image selection and annotation were created. The final aim is twofold: (1) to systematize the selection of images and (2) to start annotating old and new images so that the system can automatically allocate them in different concept entries based on shared conceptual propositions.
Stemming is a process of reducing a derivational or inflectional word to its root or stem by stripping all its affixes. It is been used in applications such as information retrieval, machine translation, and text summarization, as their pre-processing step to increase efficiency. Currently, there are a few stemming algorithms which have been developed for languages such as English, Arabic, Turkish, Malay and Amharic. Unfortunately, no algorithm has been used to stem text in Hausa, a Chadic language spoken in West Africa. To address this need, we propose stemming Hausa text using affix-stripping rules and reference lookup. We stemmed Hausa text, using 78 affix stripping rules applied in 4 steps and a reference look-up consisting of 1500 Hausa root words. The over-stemming index, under-stemming index, stemmer weight, word stemmed factor, correctly stemmed words factor and average words conflation factor were calculated to determine the effect of reference look-up on the strength and accuracy of the stemmer. It was observed that reference look-up aided in reducing both over-stemming and under-stemming errors, increased accuracy and has a tendency to reduce the strength of an affix stripping stemmer. The rationality behind the approach used is discussed and directions for future research are identified.
This article proposes an ontology design pattern for leading knowledge providers to represent knowledge in more normalized, precise and interrelated ways, hence in ways that help the matching and exploitation of knowledge from different sources. This pattern is a knowledge sharing best practice that is domain and language independent. It can be used as a criteria for measuring the quality of an ontology. This pattern is: using binary relation types directly derived from concept types, especially role types or types of process. The article explains and illustrates this pattern, and relates it to other patterns and general ontology quality criteria. It also provides an ontology for automatically deriving relation types from concept types (e.g., those from lexical ontologies such as those derived from the WordNet lexical database). This derivation helps normalizing knowledge, reduces having to introduce new relation types and helps keeping all the types organized.
Continuous word representations appeared to be a useful feature in many natural language processing tasks.Using fixed-dimension pre-trained word embeddings allows avoiding sparse bag-of-words representation and to train models with fewer parameters.In this paper, we use fixed pre-trained word embeddings as additional features for a neural scoring function in the MST parser.With the multi-layer architecture of the scoring function we can avoid handcrafting feature conjunctions.The continuous word representations on the input also allow us to reduce the number of lexical features, make the parser more robust to out-of-vocabulary words, and reduce the total number of parameters of the model.Although its accuracy stays below the state of the art, the model size is substantially smaller than with the standard features set.Moreover, it performs well for languages where only a smaller treebank is available and the results promise to be useful in cross-lingual parsing.
Hoarding is a complex and impairing psychiatric disorder and a public health problem. Traditionally it is assessed through observation and interview, but recently a new method has been proposed where living quarters of an individual are visually compared with a set of template images ranked according to the “Clutter Image Rating” (CIR) scale from 1 to 9. However, such an assessment is time-consuming, subjective, and weak in repeatability. We propose an automatic method for classifying hoarding images according to the CIR scale. Since clutter in living quarters (e.g., piles of boxes, newspapers, clothing) corresponds to “busy” areas with lots of edges in captured images, we use the histogram-of-gradients (HOG) descriptor to characterize images and estimate the CIR value using two methods: regression and classification. In 4-fold cross-validation on 620 images that we harvested from the internet, both methods result in mean-absolute CIR error of about 1.2. Given the simplicity of our method, this is an encouraging result as it approximates ratings by trained professionals who admit assigning CIR values within ± 1 CIR point.
Recently, these has been a surge on studying how to obtain partially annotated data for model supervision. However, there still lacks a systematic study on how to train statistical models with partial annotation (PA). Taking dependency parsing as our case study, this paper describes and compares two straightforward approaches for three mainstream dependency parsers. The first approach is previously proposed to directly train a log-linear graph-based parser (LLGPar) with PA based on a forest-based objective. This work for the first time proposes the second approach to directly training a linear graph-based parse (LGPar) and a linear transition-based parser (LTPar) with PA based on the idea of constrained decoding. We conduct extensive experiments on Penn Treebank under three different settings for simulating PA, i.e., random dependencies, most uncertain dependencies, and dependencies with divergent outputs from the three parsers. The results show that LLGPar is most effective in learning from PA and LTPar lags behind the graph-based counterparts by large margin. Moreover, LGPar and LTPar can achieve best performance by using LLGPar to complete PA into full annotation (FA).
A known way to improve the accuracy of dependency parsers is to combine several different parsing algorithms, in such a way that the weaknesses of each of the models can be compensated by the strengths of others. For example, voting-based combination schemes are based on variants of the idea of analyzing each sentence with various parsers, and constructing a combined output where the head of each node is determined by “majority vote” among the different parsers. Typically, such approaches combine very different parsing models to take advan- tage of the variability in the parsing errors they make. In this paper, we show that consistent improvements in accuracy can be obtained in a much simpler way by combining a single parser with itself. In particular, we start with a greedy implementation of the Nivre pseudo-projective arc-eager algorithm, a well-known left-to-right transition-based parser, and we combine it with a “mirrored” version of the algorithm that analyzes sentences from right to left. To determine which of the two obtained outputs we trust for the head of each node, we use simple criteria based on the length and position of dependency arcs. Experiments on several datasets from the CoNLL-X shared task and the WSJ section of the English Penn Treebank show that the novel combination system obtains better performance than the baseline arc-eager parser in all cases. To test the generality of the approach, we also perform experiments with a different transition system (arc-standard) and a different search strategy (beam search), obtaining similar improvements in all these settings.
Perirhinal cortex (PrC) has been implicated as a brain region in the medial temporal lobes (MTL) that critically contributes to familiarity-based recognition memory, a process that allows for recognition to occur independently of contextual recollection. Informed by neurophysiological research in non-human primates, fMRI, as well as behavioural work in humans, the current thesis research tests the novel hypothesis that PrC cortex functioning also underlies the ability to assess cumulative lifetime familiarity with object concepts that are characterized by a lifetime of experiences. In Chapter 2, a patient (NB) with a left anterior temporal lobe (ATL) lesion that included PrC as well as an amnesic patient (HC) with a bilateral lesion to the hippocampus were tested on their ability to make lifetime familiarity judgements for object concepts (i.e., concrete nouns). Patient NB made abnormal familiarity ratings for objects concepts relative to matched controls, while patient HC produced ratings that did not differ from control participants. In Chapter 3, I tested healthy young adults on a frequency judgement task and lifetime familiarity task while they underwent fMRI. A region in the left PrC tracked both the perceived frequency of recent laboratory exposure as well as perceived lifetime familiarity. Finally, in Chapter 4, I tested whether indeed lifetime familiarity judgements are based on conceptual processing by making use of an associative priming paradigm. Associatively-related primes increased the perceived familiarity of object concepts while also reducing the latency of these judgements. Overall, the results from all three empirical chapters provides evidence that warrants an extension of PrC functioning to include the cumulative assessment of lifetime familiarity with object concepts.
During the first decade of this century, a new subculture emerged on the Runet (the Russian Internet). It promoted a version of the deliberately distorted Russian language, first of all by systematically altering orthographic and punctuation norms. This “padonkafskii” language (given the misspelled usage of the original word, it seems more accurate to translate it as “basterds,” as in in Tarantino’s movie) has been analyzed by linguists. Its political role and implicit message have not been scrutinized until now. In the article, Ilya Kukulin argues that this <i>basterd</i> language was used and promoted by pro-Kremlin political entrepreneurs to develop new strategies of communication: performances of cynical transgression and symbolic violence through the humiliation of counterparts. Cyberbullying became a distinctive style of seemingly nonconformist subculture, and already in this capacity was used by the regime. Central to this communication strategy was the cult of power, xenophobia, and discarding of any idealistic motivations of human behavior as hypocritical public relations technologies. At the same time, the <i>basterd</i> subculture itself claimed the status of nonconformist sincerity for its members. This image of a novel subculture helped to promote its aggressive xenophobia and support for the authorities among the most socially dynamic groups of Russia’s population. Political mobilization of the <i>basterds</i>’ language relied on the earlier aesthetic and political strategies of Russian mass culture of the 1990s, which were partially reconfigured or further developed. The radical and socially escapist irony underlying the cultural scene of the 1990s (the article uses Russian punk rock as an example) has been recast into the view of society as a total war, perceived from the position of total cynicism and nihilism. Cynicism has become the main mode of social thinking in modern Russia. Once popular with many bloggers, the <i>basterds</i>’ language is now out of vogue. Moreover, experiments with language transgressions seem to have lost their popularity. Instead, the role of transmitters of an ultra-cynical worldview and proponents of symbolic violence has been assumed by officialdom as represented by press secretaries of the president, key ministries, or MPs. They display the same conscious transgression of linguistic norm and moral standards as their <i>Basterd</i> predecessors, who pretended to be antiestablishment and nonconformist. The preponderance of <i>basterd</i> language during the previous decade cleared the ground for hate-speech and symbolic violence in the public sphere as an acceptable, attractive, and even necessary format of public communication.
We address in this paper some theoretical and practical issues relating to generation, processing, and management of Parallel Translation Corpus (PTC) in Indian languages, which is under development in a consortium-mode project (ILCI-II) 1 under the aegis of DeitY, Govt. of India. These issues are discussed here for the first time keeping in mind the ready application of PTC in various domains of linguistics including computational linguistics, Natural Language Processing, applied linguistics, lexicography, translation, language description, etc. In a normative manner, we define what is a PTC; describe the process of its construction; identify its features; exemplify the processes of text alignment in PTC; discuss the methods of text analysis; propose for restructuring of translational units; define the process of extraction of translational equivalents; propose for generation of bilingual lexical database and Term Bank from a structured PTC; and finally identify the areas where a PTC and information extracted from it may be utilized. Since construction of PTC in Indian languages is full of hurdles, we try to construct a roadmap with a focus on techniques and methodologies that may be applied for achieving the task. The issues are brought under focus to justify the present work that is trying to construct PTC for some Indian languages for future reference and application.
Representation of syntactic structure is a core area of research in Computational Linguistics, disambiguating distinctions in meaning that are crucial for correct interpretation of language. Development of algorithms and statistical models over the past three decades has led to systems that are accurate enough to be deployed in industry, playing a key role in products such as Google Search and Apple Siri. However, syntactic parsers today are usually constrained to tree representations of language, and performance is interpreted through a single metric that conveys no linguistic information regarding remaining errors.In this dissertation, we present new algorithms for error analysis and parsing. The heart of our approach to error analysis is the use of structural transformations to identify more meaningful classes of errors, and to enable comparisons across formalisms. For parsing, we combine a novel dynamic program with careful choices in syntactic representation to create an efficient parser that produces graph structured output. Together, these developments allowed us to evaluate the outstanding challenges in parsing and to address a key weakness in current work.First, we present a search algorithm that, given two structures, finds a sequence of modifications leading from one structure to the other. We applied this algorithm to syntactic error analysis, where one structure is the output of a parser, the other is the correct parse, and each modification corresponds to fixing one error. We constructed a tool based on the algorithm and analyzed variations in behavior between parsers, types of text, and languages. Our observations shine light on several assumptions about syntactic errors, showing some to be true and others to be false. For example, prepositional phrase attachment errors are indeed a major issue, while coordination scope errors do not hurt performance as much as expected.Next, we describe an algorithm that builds a parse in one syntactic representation to match a parse in another representation. Specifically, we build phrase structure parses from Combinatory Categorial Grammar derivations. Our approach follows the philosophy of CCG, defining specific phrase structures for each lexical category and generic rules for combinatory steps. The new parse is built by following the CCG derivation bottom-up, gradually building the corresponding phrase structure parse. This produced significantly more accurate parses than past work, and enabled us to compare performance of several parsers across formalisms.Finally, we address a weakness we observed in phrase structure parsers: the exclusion of syntactic trace structures for computational convenience. We present an efficient dynamic programming algorithm that constructs the graph structure that has the highest score under an edge-factored scoring function. We define a parse representation compatible with the algorithm, and show how certain linguistic distinctions dramatically impact coverage. We also show various ways to modify the algorithm to improve performance by exploiting properties of observed linguistic structure. This approach to syntactic parsing is the first to cover virtually all structure encoded in the Penn Treebank.
The rapid accumulation of data in social media (in million and billion scales) has imposed great challenges in information extraction, knowledge discovery, and data mining, and texts bearing sentiment and opinions are one of the major categories of user generated data in social media. Sentiment analysis is the main technology to quickly capture what people think from these text data, and is a research direction with immediate practical value in ‘big data’ era. Learning such techniques will allow data miners to perform advanced mining tasks considering real sentiment and opinions expressed by users in additional to the statistics calculated from the physical actions (such as viewing or purchasing records) user perform, which facilitates the development of real-world applications. However, the situation that most tools are limited to the English language might stop academic or industrial people from doing research or products which cover a wider scope of data, retrieving information from people who speak different languages, or developing applications for worldwide users. More specifically, sentiment analysis determines the polarities and strength of the sentiment-bearing expressions, and it has been an important and attractive research area. In the past decade, resources and tools have been developed for sentiment analysis in order to provide subsequent vital applications, such as product reviews, reputation management, call center robots, automatic public survey, etc. However, most of these resources are for the English language. Being the key to the understanding of business and government issues, sentiment analysis resources and tools are required for other major languages, e.g., Chinese. In this tutorial, audience can learn the skills for retrieving sentiment from texts in another major language, Chinese, to overcome this obstacle. The goal of this tutorial is to introduce the proposed sentiment analysis technologies and datasets in the literature, and give the audience the opportunities to use resources and tools to process Chinese texts from the very basic preprocessing, i.e., word segmentation and part of speech tagging, to sentiment analysis, i.e., applying sentiment dictionaries and obtaining sentiment scores, through step-by-step instructions and a hand-on practice. The basic processing tools are from CKIP Participants can download these resources, use them and solve the problems they encounter in this tutorial. This tutorial will begin from some background knowledge of sentiment analysis, such as how sentiment are categorized, where to find available corpora and which models are commonly applied, especially for the Chinese language. Then a set of basic Chinese text processing tools for word segmentation, tagging and parsing will be introduced for the preparation of mining sentiment and opinions. After bringing the idea of how to pre-process the Chinese language to the audience, I will describe our work on compositional Chinese sentiment analysis from words to sentences, and an application on social media text (Facebook) as an example. All our involved and recently developed related resources, including Chinese Morphological Dataset, Augmented NTU Sentiment Dictionary (aug-NTUSD), E-hownet with sentiment information, Chinese Opinion Treebank, and the CopeOpi Sentiment Scorer, will also be introduced and distributed in this tutorial. The tutorial will end by a hands-on session of how to use these materials and tools to process Chinese sentiment. Content Details, Materials, and Program please refer to the tutorial URL: http://www.lunweiku.com/
My research focuses on the study of grammatical change in the recent history of the English language; in particular, I am currently writing my PhD dissertation on from Early Modern English to Present-Day English. In this PhD project, I analyse and compare the different factors that appear to influence in Present-Day English with earlier stages of the language (Late Modern English).The concept of ellipsis refers to a syntactic strategy in which expected elements have been left unpronounced in certain constructions. This omission triggers a mismatch between meaning (the intended message) and sound (what is in fact uttered). In particular, my research focuses on those examples of (Miller 2011, Miller and Pullum 2013), i.e. ellipsis types that occur after the following licensors (that is, those elements that license ellipsis): modal verbs, auxiliaries be, have and do, infinitival marker to and negator not. The main aim is to carry out an empirical analysis of from Late Modern English to Present-Day English (1700-1914), both quantitatively and qualitatively, by means of data retrieved from the Penn Corpora of Historical English. This project pays attention to syntactic variation, genre distribution and discourse variables (type of anaphora, mismatches in polarity, aspect, voice, modality, tense; comparison of clause types; distance, linking, type of focus).Esta tesis doctoral versa sobre la variacion diacronica de la elipsis sintactica desde ingles moderno temprano hasta la actualidad. La elipsis representa un desajuste entre el significado de lo que se dice (la intencion de un mensaje) y lo que se pronuncia en realidad. En el ambito de la linguistica moderna, la elipsis se estudia en el campo la semantica, la sintaxis, la pragmatica, la psicolinguistica, la linguistica de corpus, etc. y constituye la novedad de esta investigacion el tratamiento empirico de este fenomeno linguistico tratando de juntar variables procedentes de las distintas teorias consultadas. Cabe destacar que no existe un estudio multidisciplinar sobre este concepto sintactico en ingles moderno. Existen solo unos pocos estudios relativamente recientes pero se centran unicamente en el estudio del ingles actual. Por esta razon, se decidio llevar a cabo la investigacion tanto de los aspectos formales como de los funcionales de la elipsis y su evolucion diacronica teniendo en cuenta distintas variables discursivas. Para eso, se creo una base de datos en la que apareceria el ejemplo de Post-Auxiliary Ellipsis (elipsis despues de un auxiliar) con su numero identificador; el genero al que pertenece (de los dieciocho diferentes que existen en el corpus utilizado, el Penn Treebank); el licensor (aquel elemento que posibilita la elipsis); el tipo de union entre el antecedente de la elipse y la clausula eliptica (coordinacion, subordinacion, parataxis, etc.); el tipo de conector entre el antecedente y la parte la distancia existente entre el antecedente y la clausula eliptica (numero de clausulas); contexto sintactico en el que aparece la elipsis (clausulas principales, subordinadas, question-tags, etc); tipo de anafora (anaforico, cataforico, exoforico); categoria del antecedente (sintagma verbal, nominal, adjetival o no constituyente); categoria del material elidido (sintagma verbal, nominal, adjetival o no constituyente); presencia o ausencia de cambio de referente; comparacion del aspecto, la voz, la modalidad y el tiempo del antecedente con respecto a la parte presencia o ausencia de question-tag; comparacion entre el tipo de clausula del antecedente (declarativa, interrogativa o imperativa) y el de la clausula elidida; y tipo de foco de la parte elidida (auxiliary-choice (eleccion de auxiliar), subject-choice (eleccion de sujeto) o ambos).Esta tese de doutoramento versa sobre a variacion diacronica da elipse sintactica dende ingles moderno temperan ata a actualidade. A elipse representa un desaxuste entre o significado do que se di (a intencion dunha mensaxe) e o que se pronuncia en realidade. No ambito da linguistica moderna, a elipse estudase no campo a semantica, a sintaxe, a pragmatica, a psicolinguistica, a linguistica de corpus, etc. e constitue a novidade desta investigacion o tratamento empirico deste fenomeno linguistico tratando de xuntar variables procedentes das distintas teorias consultadas. Cabe destacar que non existe un estudo multidisciplinar sobre este concepto sintactico en ingles moderno. Existen so uns poucos estudos relativamente recentes pero centranse unicamente no estudo do ingles actual. Por esta razon, decidiuse levar a cabo a investigacion tanto dos aspectos formais coma dos funcionais da elipse e a sua evolucion diacronica tendo en conta distintas variables discursivas. Para iso, creouse unha base de datos na que apareceria o exemplo de Post-Auxiliary Ellipsis (elipse despois dun auxiliar) co seu numero identificador; o xenero ao que pertence (dos dezaoito diferentes que existen no corpus utilizado, o Penn Treebank); o licensor (aquel elemento que posibilita a elipse); o tipo de union entre o antecedente da elipse e a clausula eliptica (coordinacion, subordinacion, parataxe, etc.); o tipo de conector entre o antecedente e a parte a distancia existente entre o antecedente e a clausula eliptica (numero de clausulas); contexto sintactico no que aparece a elipse (clausulas principais, subordinadas, question-tags, etc); tipo de anafora (anaforico, cataforico, exoforico); categoria do antecedente (sintagma verbal, nominal, adxectival ou non constituinte); categoria do material elidido (sintagma verbal, nominal, adxectival ou non constituinte); presenza ou ausencia de cambio de referente; comparacion do aspecto, a voz, a modalidade e o tempo do antecedente con respecto a parte presenza ou ausencia de question-tag; comparacion entre o tipo de clausula do antecedente (declarativa, interrogativa ou imperativa) e o da clausula elidida; e tipo de foco da parte elidida (auxiliary-choice (eleccion de auxiliar), subject-choice (eleccion de suxeito) ou ambos).
Despite the proliferation of corpus-based studies focusing on the complex category of discourse markers in the recent years, consensus is yet to be found regarding the most reliable yet informative model to describe their behavior in authentic data. The major and most widely spread frameworks include the Penn Discourse TreeBank (Prasad et al. 2008), Rhetorical Structure Theory (Mann & Thompson 1988), Segmented Discourse Representation Theory (Asher & Lascarides 2003) and the Cognitive approach to Coherence Relations (Sanders et al. 1992). These models disagree both on the top levels (number and type of generic annotation levels, if any) and the specific relations included in them. It is precisely the relation between top levels and corresponding sublevels that will be discussed in this paper, starting from a recent proposal of functional taxonomy applied to the French-English spoken corpus DisFrEn (Crible, in press) and its revision in the framework of the LOCAS-F corpus (Degand, Martin & Simon 2014). In DisFrEn, four top-level functions or "domains" are distinguished, from the revision of existing proposals for both speech and writing (Cuenca 2013, González 2005, Halliday & Hasan 1976, Zufferey & Degand in press): ideational (objective relations), rhetorical (subjective, metadiscursive functions), sequential (structuring functions) and interpersonal (speaker-hearer relationship). In the original model, these four domains include a total of thirty functions, each of them belonging to one – and only one – domain. For instance, the function labeled "CAUSE" is always ideational, while "MOTIVATION" is always rhetorical, etc. This interdependent system is intended to maximize the informativity of the annotation labels (one label for one function, vs. combinations of labels) while giving the opportunity to filter the distribution from thirty to four values, hence more efficient for quantitative purposes. The annotation procedure doesn't specify which decision should be made first, the domain or the function, and it is assumed that both orders are possible although not equally relevant depending on the specific function at stake and/or the research question, annotator's expertise, etc. (see Crible & Degand 2015). An on-going project (Degand & Simon 2015) is currently working on a revision of the DisFrEn taxonomy aiming at reducing the number of options and enhancing the reliability and cognitive validity of the model. The main difference between this revision and the original is that domains and functions are no longer interdependent. On the contrary, it is assumed that most – if not all – functions can be assigned more than one domain, roughly following Sweetser’s (1990) discourse domains. For instance, a [CONTRAST] can be ideational (1), rhetorical (2), and even sequential (3) or interpersonal (4), a proposal which still needs careful investigation and discussion. (1) I wasn’t looking forward to doing it but I am now (DisFrEn EN-phon-01) (2) a rebate is when they send the money back // yes but how do you define it in economic terms (DisFrEn EN-clas-02) (3) (after a digression on the industries in Bristol area) but Bristol itself is a large metropolis (DisFrEn EN-intf-05) (4) I think the Marks is better // actually I’m not sure it is (DisFrEn EN-conv-01) In this presentation, we will address several methodological implications of the interdependence vs. independence of domains and functions, paying particular attention to issues of inter-rater reliability and overall validity and consistency of the taxonomy. In addition, quantitative results will be presented to illustrate the comparison of the two versions on samples of spoken French from the LOCAS-F corpus annotated with both systems.
This thesis describes unexpected constructions based on THIS and THAT by French and Spanishlearners of English. Chapter 1 raises the issue of the study of THIS and THAT as markers in the twomicrosystems of proforms and deictics. Chapter 2 covers different types of analyses of referencewith THIS and THAT in native English and refers to different theoretical frameworks (Cornish,Cotte, Halliday & Hasan, Kleiber, Fraser & Joly, Lapaire & Rotgé). It crossreferencesrepresentations (anaphora/deixis; endophoricity/exophoricity) with an analysis of functionalrealisations. Chapter 3 broaches the issue of interlanguage analysis, and it shows that a dynamicsystemic approach grounded in the functional distinction of the forms is necessary. Chapter 4 givesdetails about existing annotation tagsets for English corpora (Penn Treebank, Claws7, ICEGB). Itshows the need for a finergrained annotation relying on functional tags and for semanticinformation on the positions (subject v. oblique). Chapter 5 describes the multilayer annotationstructure which is implemented for the analysis of different corpora. It also covers the methods usedto automatically annotate functional categories (as well as their evaluation), and it justifies thechoices made to support corpus interoperability. Chapter 6 offers a regression analysis whichprovides evidence on the tendencies of the operationalised variables (the L1, the written or spokenmode of the corpora and the type of reference). Chapter 7 examines the role of the previously codedlinguistic properties of the analysis. With the use of classifiers, it describes a system for automaticerror analysis. Chapter 8 concludes on the methodologies used in the thesis and their implicationsin linguistic analysis.
In many natural language processing based intelligent systems, parsing is the first task to perform. However, in the next stages, many systems often have the capacity of processing a limited number of parsed structures. The problem is to determine what parsed sentences can be recognized by a system. The decision of syntactic structures which can be processed by a system is consider as the task of "classification" of a parsed sentence into one of given classes of recognizable parses. In this paper we deal with this issue by proposing a method for mapping Vietnamese chunked sentences to a set of pre-defined shallow structures. Also, we tag lexicons and chunk phrases of the original sentences using our Functional Part-of-Speech (FPOS) tagset with Apache OpenNLP tools (Tokenizer, POS Tagger, Chunker). Based on the foundation of Functional Grammar, we define new lexical tags and combine with Penn-Treebank tagset to build our FPOS tagset. Due to our set of shallow structures is finite, instead of using a parser, we propose a rule-based algorithm for the mapping process. We establish conversion rules according to the reality experiences when using Vietnamese in common communication. The experiment shows that we converse successfully for the major of testing sentences and the algorithm can be applied for different languages.
Two fundamental components of causality are the Cause and the Result. In linguistic work the distinction between these aspects is commonly blurred, presumably because the primary research focus has been on describing how language encodes causality. The semantic nature of the component events and the constraints on their relationship are seldom discussed; however, the current work aims to shed light on a broader spectrum of features that underlie the concept. This is an essential foundation for understanding how language communicates Result. The present discussion explores and illuminates the nature of this concept focusing on a relatively open-ended set of linguistic elements that can play a role in shaping a discourse relation in addition to discourse connectives. This is in contrast to the majority of the previous research, which has been quite intensely concerned with investigating a limited collection of well-established causality markers. Also, despite the fact that English has been used in studies on causality both as a control language and a metalanguage, there is surprisingly little work on the semantics of the relations that occur specifically in English, let alone Result relations. By borrowing from several cognitively-oriented approaches and combining empirical data from two written corpora (British National Corpus and the Penn Discourse Treebank) with experimental work, the current study systematically investigates the conceptual and linguistic properties of several closely related Result relation types (including Purpose), along with the joint role of discourse connectives and other discourse elements in conveying the intended sense. The findings indicate that linguistic signals of the conceptual structure of the relation seem to play a more significant role in the interpretation than explicit marking. Two factors emerged as more vital cues than the presence of the ambiguous connective so. In Purpose relations, a modal auxiliary conveying an intended effect, and in Result relations the presence/absence of an intentionally acting actor are crucial for disambiguation. The multifunctional connective therefore seems to merely satisfy the mandatory marking requirement related to the intrinsically unrealized (‘nonveridical’) nature of Purpose. In Result the presence of an ambiguous marker is to a great extent optional in English. However, discourse markers can also reflect how language users categorize causal event types. This claim has been confirmed in several cross-linguistic analyses, but the lexicon of English connectives has not been systematically investigated from this vantage point. The few existing studies found that the uses of English connectives are quite unconstrained across causal categories. The present work contributes to this line of research and suggests that two unambiguous markers, as a result and for this reason, indeed cover a wide range of causal event types; however, they also exhibit significant tendencies to occur prototypically in certain relation types. The presence and role of an intentionally acting discourse participant behind both real-world and linguistic causally-related events contributes to these tendencies. The contexts that include such a participant are regarded as intrinsically subjective and have been found to manifest surface expressions of subjectivity in previous work on other languages. The current study confirms similar tendencies in the linguistic construal and marking of Result relations in English, which proves that certain language elements partake in establishing the intended interpretation on a par with discourse connectives. What emerges as a result of this discussion, is therefore an account on how English utilizes the broad category of Result and what linguistic elements are used to convey the array of resultative events.
El acto de destrucción de la lengua que lleva cabo uno de los sujetos de la escritura de Altazor abre el espacio y el tiempo para el surgimiento de nuevos significantes. Para este sujeto, el uso sistemático de la lengua -cumpliendo la norma lingüística- impide la representación (aparición de imágenes) y referencia de correlatos radicalmente nuevos. Este sujeto es discontinuo y coexiste con otros, en especial, con un sujeto voluntarioso, que sigue aferrándose a la tradicional concepción del mundo, fundada en la trascendencia divina. Las operaciones sobre la lengua de este poeta altazoriano se realizan, sobre todo, como utilización paródica de ella, como demolición intencional de la lengua, como transformación de las ruinas idiomáticas en significantes, como uso alegórico de la lengua (en el sentido sugerido por W. Benjamin). Este sujeto altazoriano propone orientar y sostener sentimentalmente la constitución de nuevas imágenes (en el sentido de los universales fantásticos de Vico). La poesía se hace, así, acontecimiento (Ereignis, según lo nombra Heidegger), actividad cuya plenitud se produce en su consumación y consunción, en la mostración de su temporalidad fundante. Pero la energía del poeta no alcanza para la continuidad del acontecimiento poético -en el caso de que sea posible-, para la urgencia y necesidad de su aparición. The act oflanguage destruction carried out by one ofthe subjects in the writing of "Altazor", opens both space and time to the birth ofnew signifiers. For such subject, the systematic use oflanguage sticking to the linguistic norm, hinders the representation (the creation o.fimages and the reference to radically new co-relatives. Such subject is discontinuous and coexists with others, in particular, with a willful subject that keeps its allegiance to the traditional world view, based on divine transcendence. The language operations ofthis altazorean poet are effected, above all, as parody; as transformation ofidiomatic debris into signifiers; and as an allegorical use of language (in the sense suggested by Walter Benjamín). The altarzorean subject propases to orientate and sustain sentimentally the creation of new images (in the sense of the "fantastic universals" of Vico). In this way, poetry becomes 'event' ("Ereignis", as named by Heidegger) an activity whosefullnes is achieved through its realization and consumption, in the exhibition of its founding tenporality.
This dataset introduces a companion reproducibility Java console program, called HESML_vs_SML_test.jar, of the work introduced by Lastra-Díaz and García-Serrano [1]. This latter work introduces the Half-Edge Semantic Measures Library (HESML), and carries-out an experimental survey between HESML V1R2, the Semantic Measures Library (SML) 0.9 [2] and the WNetSS [4] semantic measures libraries. The HESML_vs_SML_test.jar program runs the set of performance and scalability benchmarks detailed in [1] and generates the figures and tables of results reported in the aforementioned work, which are also enclosed as complementary files of this dataset (see files below). Licensing note: The 'HESML_vs_SML_test.jar' program is based on the HESML V1R2 [3], SML 0.9 [2] and WNetSS [4] semantic measures libraries, and it includes these libraries in its distribution, as well as WordNet 3.0 [6] and the SimLex665 [5] dataset. Thus, if you use this dataset, you should also cite the works related to these resources. References: [1] Lastra-Díaz, J. J., and García-Serrano, A. (2016). HESML: a scalable ontology-based semantic similarity measures library with a set of reproducible experiments and a replication dataset. To appear in Information Systems Journal. [2] Harispe, S., Ranwez, S., Janaqi, S., and Montmain, J. (2014). The Semantic Measures Library: Assessing Semantic Similarity from Knowledge Representation Analysis. In E. Métais, M. Roche, & M. Teisseire (Eds.), Proc. of the 19th International Conference on Applications of Natural Language to Information Systems (NLDB 2014) (Vol. 8455, pp. 254–257). Montpelier, France: Springer. http://dx.doi.org/10.1007/978-3-319-07983-7_37 [3] Lastra-Díaz, J. J., & García-Serrano, A. (2016). HESML V1R2 Java software library of ontology-based semantic similarity measures and information content models. Mendeley Data, v2. https://doi.org/10.17632/t87s78dg78.2 [4] Ben Aouicha, M., Taieb, M. A. H., and Ben Hamadou, A. (2016). SISR: System for integrating semantic relatedness and similarity measures. Soft Computing, 1–25. http://dx.doi.org/10.1007/s00500-016-2438-x [5] Hill, F., Reichart, R., & Korhonen, A. (2015). SimLex-999: Evaluating Semantic Models with (Genuine) Similarity Estimation. Computational Linguistics, 41(4), 665–695. http://dx.doi.org/10.1162/COLI_a_00237 [6] Miller, G. A. (1995). WordNet: A Lexical Database for English. Communications of the ACM, 38(11), 39–41. http://dx.doi.org/10.1145/219717.219748
This paper deals with the declension of foreign words and loanwords in Croatian newspapers in the period from 1950 to 1999. The usage of foreign nouns, of a-type declension, with vowel final sounds in nominative in newspaper texts is compared with Croatian literary-linguistic norm. This comparison uncovers which forms are either inconsistent in terms of their usage or not in accordance with the linguistic norm. It can be seen that the grammatical norms of the second half of the 20th century were insufficient for news writing style incessantly inundated with new expressions - particularly foreign names which should be adjusted to Croatian declension system. That also points out the need for more detailed description in the present morphological norm.
108 Objectives To diagnose the dementia subtypes has significant information for determining the treatment strategies and predicting the clinical courses. Recently, many dementia patients undergo dopamine transporter (DAT) imaging, in addition to brain perfusion imaging to investigate the subtypes of dementia. However, patients must wait for tracer decay when using the conventional method, and this is burdensome for dementia patients and often delays the decision making for treatment. We developed a prototype CdTe SPECT system with 4-Pixel Matched Collimator for brain study. This system provides high energy resolution (6.6%), high sensitivity (220 cps/MBq/head) and provide high spatial resolution images of I-123 and Tc-99m simultaneously. The aim of this study was to evaluate findings and quantification on dual isotope study of cerebral blood flow (CBF) and Dopamine transporter (DAT) images with the new SPECT system. Methods We prospectively enrolled 21 patients with cognitive disorder. Every patient underwent imaging examinations on the same schedule. We used 3-head scanner (GCA-9300R, TOSHIBA) as a conventional scanner. Both 99mTc-ECD (CBF) and 123I-IFP (DAT) scans were independently performed with the conventional scanner on separate days, and simultaneously performed with the new system on another day. CBF and DAT images were both visually and quantitatively analyzed.For visual analyses, two nuclear medicine physicians visually interpreted both DAT and CBF images rating the severity into 4 grades (from 0 to 3). Totally 12 (6/hemisphere) supratentorial regions on the 99mTc-ECD images and 2 regions (left and right striatum) on the 123I-IFP were defined for each patient. The correlation of visual analysis results between GCA and SPICA was evaluated by weighted Kappa statistics. For quantitative analyses, we calculated the cerebrum-thalamus count ratio of 99mTc-ECD and the specific-to-background ratios of 123I-IFP. We assessed those regions of Pearson9s correlation coefficient and intraclass correlation coefficient (ICC) between the conventional scanner and the new system. Results The weighted Kappa statistics between GCA and SPICA were 0.67 and 0.68 for 99mTc-ECD and 123I-IFP, respectively, indicating high interrater reliability. The semi-quantitative analyses demonstrated that the 99mTc-ECD cerebrum-thalamus count ratio of SPICA was well correlated to GCA (R=0.81, p Conclusions The findings and quantitative results of both CBF and DAT images of dual isotope study with the new system had excellent correlation with those of single isotope studies with the conventional scanner with two separate days. In conclusion, our new SPECT scanner with semiconductor detectors enables quantitative dual tracer diagnostic imaging of CBF and Dopamine transporter imaging in patients with cognitive disorder. This technique may enable the one-stop imaging diagnosis for dementia and reduce the burden on patients.
The purpose of this paper is to discuss the social legitimacy of the non-dominant variety of French that is used in Belgium (henceforth ‘Belgian French’). As will be detailed, Francophone Belgians’ attitudes have shifted from early 19th c. – late 20th c. purism and subsequent linguistic subjection to France to more recent acceptation of endogenous traits and increasing distance from the Hexagonal model. Nevertheless, these attitudes remain characterized by a “double distance” from both Hexagonal and Belgian French. The idea that French is viewed by Francophone Belgians as a polycentric/polynomic language will thus be questioned: do they really consider that there is a legitimate Belgian variety of French? What is the relevance of the national criterion in the way they define linguistic norms? What other criteria lie behind the definition and legitimization of their linguistic norms?
Semantic analysis of sentences can only be carried out using Dependency Parsing. Dependency parser accepts words in a sentence and builds dependency relation among the words resulting in a unique tree for each sentence. An Indian Panini is the first to develop semantic analysis for Sanskrit using a dependency framework. Western researchers in the near past have also deliberated on dependency parsing so that automated dependency parser can be generated. Dependency Parser is useful in information extraction, question-answering, text summarization etc. For many indian languages namely Bengali, Kannada, Malayalam and Marathi a dependency based Treebank is in the development stage. Also for Hindi, Telugu and Tamil dependency Treebanks are already developed. The Treebank data can be used by a dependency parser generator like Maltparser to develop a Dependency parser. In this paper few dependency parsing algorithms are discussed
INTRODUCTIONThe Strategic Integrated Management Seminar (SIMS) course is mandatory for every senior student in the school of business at a mid-size private university in the northeastern United States. The course allows students to integrate their accumulated knowledge and apply this knowledge to issues from a strategic perspective. It examines a firm from the position of top level management, focusing on the role of the general manager in formulating and implementing corporate and business level strategy. Strategic issues of an entire athletic (hereon, footwear company or company) and industry are analyzed. Students are expected to draw their accumulated knowledge of the functional areas of their majors into a homogenous team effort. Each individual student works on developing his/her ability to analyze information, draw logical conclusions, and offer sound supporting evidence for their arguments in written form and classroom discussions. The course is highly interactive with students taking the lead and the professors sharing knowledge and offering supplementary support.The SIMS course uses the Business Strategy Game (BSG) simulation to enable students to experience a top management team perspective in running a and experiencing competitive conditions in the athletic industry. In the BSG, students compete in teams (each team constitutes a company, hereon team/company will be synonymous) within a global arena that encompasses four regions - Europe-Africa, North America, Asia-Pacific, and Latin America (The Business Strategy Game, 2016). They compete against teams in their individual classes and compare/contrast data with teams/companies worldwide. Each competes head-to-head against companies run by other teams in the course, hence competition plays an important role in the experience. Each sells its brand of to retailers worldwide and to individuals buying online at the company's website.Competing in the BSG requires a series of complex decisions by the students, taking into account the team's strategy for their and the competitive conditions in the industry and the strategies of their competitors. The simulation allows for numerous decisions for each round, requiring students to choose which decisions are most important to implement their strategy and which areas of the business must receive attention in order for their firm to be its most competitive. Beyond overall strategy (corporate, competitive) are several key functional areas for decision making. Decision areas in operations include capacity planning (either adding to existing plants or building new plants in new geographic locations), production quality decisions for the athletic footwear, plant operations efficiency, and labor decisions. Footwear must be shipped to distribution centers around the world and students must choose where it is best to manufacture the and where to ship taking into consideration demand, shipping costs, tariffs, and exchange rates. Marketing decisions include pricing the product in a wholesale and a retail environment, advertising and use or non-use of celebrity endorsements, rebates, and incentives to retailers. Financial decisions include funding the capital structure of the firm using debt, equity, and/or cash. Dividend payouts and stock repurchases may be used by the companies.The simulation has students take control of an athletic that has been in operation for ten years. Teams make in total eight years of decisions (years 11 - 18), approximately one per week. Each decision rollover represents one year and includes many decisions within the decision. The simulation evaluates team performance based on five investor expectation performance targets: Earnings Per Share (EPS), Return on Equity (ROE), credit rating, image rating (a combination of market share and shoe quality), and stock price. Each measure of performance is equally weighted at 20% of the total score (The Business Strategy Game, 2016). …
Nonostante una secolare tradizione lessicografica, la lingua latina manca ancora di risorse lessicali di tipo computazionale aggiornate allo stato dell’arte. Cio e strettamente connesso alla limitata disponibilita di corpora testuali latini annotati linguisticamente, sulla cui base empirica possano essere costruite nuove risorse lessicali. Tuttavia, una serie di progetti mirati allo sviluppo di avanzate risorse linguistiche per il latino (tra cui alcune treebank) e stata avviata nel corso dell’ultimo decennio. In questo articolo, presentiamo Latin Vallex, un lessico di valenza per il latino realizzato in stretta connessione con l’annotazione semantico-pragmatica di due treebank latine comprensive di testi di epoche e generi diversi. Cio consente di connettere biunivocamente le strutture valenziali registrate nel lessico e le loro occorrenze nei dati testuali delle treebank.
Tokenizer, POS Tagger, Lemmatizer and Parser models for all 50 languages of Universal Depenencies 2.0 Treebanks, created solely using UD 2.0 data (http://hdl.handle.net/11234/1-1983). The model documentation including performance can be found at http://ufal.mff.cuni.cz/udpipe/users-manual#universal_dependencies_20_models. To use these models, you need UDPipe binary version at least 1.2, which you can download from http://ufal.mff.cuni.cz/udpipe. In addition to models itself, all additional data and value of hyperparameters used for training are available in the second archive, allowing reproducible training.
We describe results of investigation of a specific type of discontinuous constructions, namely non-projective constructions concerning verbs and their arguments. This topic is especially important for languages with a relatively free word order, such as Czech, which is the language we have primarily worked with. For comparison, we have included some results for English. The corpora used for both languages are the Prague Czech-English Dependency Treebank and the Prague Dependency Treebank, which are both annotated at a dependency syntax level as well as a deep (semantic) level, including verbs and their valency (arguments). We are using traditionally defined non-projectivity on trees with full linear ordering, but the two levels of annotation are innovatively combined to determine if a particular (deep) verb -argument structure is non-projective. As a result, we have identified several types of discontinuities, which we classify either by the verb class or structurally in terms of the verb, its arguments and their dependents. In addition, we have quantitatively compared selected phenomena found in Czech translated texts (in the PCEDT) to the native Czech as found in the original Prague Dependency Treebank.
SOME PROBLEMS OF THE FORMATION PROCESS OF THE LATVIAN LITERARY LANGUAGE Summary 0. Literary language is the most complete variety of language, manifesting its functions in the best way possible and unifying the nation, as well as representing national mentality among other nations and their languages. 0.1. When speaking about the Latvian literary language and its formation, it seems useful to acknowledge that the literary language is (1) the language of the entire nation, (2) is being consciously cultivated and (3) has written form. 0.2. When dealing with the Latvian literary language formation processes, one should take into consideration (a) the specific external (sociopolitical) conditions, (b) the sources of the literary language, (c) aspects of language development and its research. 1. Development of the Latvian language has been affected by external sociopolitical factors and factors of migration of representatives of various cultural layers; these factors have both stimulated and hampered the overall formation of the language both in space and time. 2. Research of the Latvian literary language is complicated, because the most reliable proofs of this process are written texts, which in Latvian appeared only in the 16th century. Therefore both folk-lore and the spoken language can be used as sources. 2.1. From the 16th to the 19th century written texts were mainly produced by German clergymen who in the beginning (in the 16th century) had a poor knowledge of Latvian. Therefore these texts must be properly handled by differentiating the sociopolitical and philological activities of the Germans. Beginning with the 17th century a normative approach has been consciously applied to the language and thus a common variety of the language is being created by maximally keeping aloof of various patois forms. 2.2. The source of analysis of the literary language and the process of its formation is the abundant Latvian folk-lore and especially the folk-songs (dainas), but the folk-lore language changes gradually from generation to generation. 2.3. In studying the formation of the literary language, the spoken language found in written monuments and recorded in special questionnaires can be used. 3. In inquiring about the Latvian literary language and its formation, definite layers of cognition can be traced: 1) statement and description of phonetic, lexical and grammatical phenomena found in the Latvian language and in the Latvian texts written by the German clergymen, 2) historical survey of language phenomena and their comparison with related languages, 3) analysis of the old-Latvian (and later also new-Latvian) written language, 4) analysis of colloquial Latvian and folk-lore, 5) fundamental synchronic and diachronic research in phonetics, word stock and grammar within the framework of the language norm, style and language culture aspects. 3.1. None of these layers give an all-embracing picture of the stages of the development of the Latvian literary language, because in those days no such task was formulated. 3.2. In Latvian linguistics the Latvian literary language formation problem became topical in the 50ies of the 20th century with the emergence of historical research of literary Latvian as an independent branch of linguistics.
Nonostante una secolare tradizione lessicografica, la lingua latina manca ancora di risorse lessicali di tipo computazionale aggiornate allo stato dell’arte. Ciò è strettamente connesso alla limitata disponibilità di corpora testuali latini annotati linguisticamente, sulla cui base empirica possano essere costruite nuove risorse lessicali. Tuttavia, una serie di progetti mirati allo sviluppo di avanzate risorse linguistiche per il latino (tra cui alcune treebank) è stata avviata nel corso dell’ultimo decennio. In questo articolo, presentiamo Latin Vallex, un lessico di valenza per il latino realizzato in stretta connessione con l’annotazione semantico-pragmatica di due treebank latine comprensive di testi di epoche e generi diversi. Ciò consente di connettere biunivocamente le strutture valenziali registrate nel lessico e le loro occorrenze nei dati testuali delle treebank.
Universal Dependencies (UD) are gaining much attention of late for systematic evaluation of cross-lingual techniques for crosslingual dependency parsing. In this paper we present our work in line with UD. Our contribution to this is manifold. We extend UD to Indian languages through conversion of Pān inian Dependencies to UD for the Hindi Dependency Treebank (HDTB). We discuss the differences in annotation in both the schemes, present parsing experiments for both the formalisms and empirically evaluate their weaknesses and strengths for Hindi. We produce an automatically converted Hindi Treebank conforming to the international standard UD scheme, making it useful as a resource for multilingual language technology.
We describe the Corpus of Spoken Icelandic (ÍS-TAL) which is made up of 15 hours of spontaneous naturally occurring conversa-tions, 31 conversations in all. The corpus comprises 184,080 tokens, 14,297 types and 9,221 lemmas. It has been transcribed using standard orthography. We present a list of the 30 most common lemmas in the corpus and compare it to a list of the most frequent lemmas in the written language, concluding that the differences between the two lists are smaller than expected. We have tagged the corpus morphologically with a statistical tagger that had been trained on written texts. The results are much better than we expected, and the tagging accuracy is as least as high as for the written texts. The final part of the paper is a report on a work in progress. We have been experimenting with converting the morphological tagging into a shallow syntactic markup by applying a few simple hand-written rules. Even though the analysis we get by using this procedure is bound to be incomplete and contain several errors, we conclude that the results are promising and we can use this method to build a simple yet useful treebank with minimal effort. 1.
Estimates of the prevalence of sensitive attributes obtained through direct questions are prone to being distorted by untruthful responding. Indirect questioning procedures such as the Randomized Response Technique (RRT) aim to control for the influence of social desirability bias. However, even on RRT surveys, some participants may disobey the instructions in an attempt to conceal their true status. In the present study, we experimentally compared the validity of two competing indirect questioning techniques that presumably offer a solution to the problem of nonadherent respondents: the Stochastic Lie Detector and the Crosswise Model. For two sensitive attributes, both techniques met the “more is better” criterion. Their application resulted in higher, and thus presumably more valid, prevalence estimates than a direct question. Only the Crosswise Model, however, adequately estimated the known prevalence of a nonsensitive control attribute.
We present the results of the joint student response analysis (SRA) and 8th recognizing textual entailment challenge. The goal of this challenge was to bring together researchers from the educational natural language processing and computational semantics communities. The goal of the SRA task is to assess student responses to questions in the science domain, focusing on correctness and completeness of the response content. Nine teams took part in the challenge, submitting a total of 18 runs using methods and features adapted from previous research on automated short answer grading, recognizing textual entailment and semantic textual similarity. We provide an extended analysis of the results focusing on the impact of evaluation metrics, application scenarios and the methods and features used by the participants. We conclude that additional research is required to be able to leverage syntactic dependency features and external semantic resources for this task, possibly due to limited coverage of scientific domains in existing resources. However, each of three approaches to using features and models adjusted to application scenarios achieved better system performance, meriting further investigation by the research community.
Dataset for TACL submission "The Galactic Dependencies Treebanks: Getting More Data by Synthesizing New Languages".<br> The scripts and model parameters for replicating this dataset are available at https://github.com/gdtreebank/gdtreebank.
The article deals with the phenomenon of diglossia in the context of development of Greek language in the Hellenistic period, observes in diachrony the correlation between the linguistic norm of Atticism and Koine Greek – from the beginning of theory of Atticist mimesis to the times of second sophistry; determines the specificity of Koine Greek of the New Testament and its relation to, on the one hand, Atticism canon and literary Koine Greek, and on the other hand, – to the language of the early Patristic literature.
espanolCon el desarrollo de la informatica, en la investigacion del lenguaje se introdujo la teoria y metodologia de redes complejas, que transforma el sistema de la lengua en las redes complejas compuestas de nodos y enlaces para hacer un analisis cuantitativo de la estructura de la lengua. El desarrollo de la gramatica de dependencias proporciona un apoyo teorico a la construccion del corpus anotado (treebank), por lo que el analisis estadistico con las redes complejas se hace posible. Este articulo presenta la teoria y metodologia de las redes complejas y construye las redes sintacticas de dependencia a base del corpus anotado (treebank) de las expresiones orales del examen EEE-4 (Examen del Espanol como Especialidad - Nivel 4). Mediante el analisis de las caracteristicas generales de las redes, incluyendo el numero de nodos, los enlaces, el grado medio, la longitud media de los caminos, la distribucion de grados y la centralizacion, tiene como objetivo descubrir la diferencia y similitud potencial entre las expresiones orales de distintos niveles. Ademas, con el analisis de conglomerados, esta investigacion pretende demostrar la capacidad discriminatoria de las variables de las redes complejas y proporcionar una referencia potencial para el trabajo de calificacion. EnglishWith the development of information technology, the theory and methodology of complex network has been introduced to the language research, which transforms the system of language in a complex networks composed of nodes and edges for the quantitative analysis about the language structure. The development of dependency grammar provides theoretical support for the construction of a treebank corpus, making possible a statistic analysis of complex networks. This paper introduces the theory and methodology of the complex network and builds dependency syntactic networks based on the treebank of speeches from the EEE-4 oral test. According to the analysis of the overall characteristics of the networks, including the number of edges, the number of the nodes, the average degree, the average path length, the network centrality and the degree distribution, it aims to find in the networks potential difference and similarity between various grades of speaking performance. Through clustering analysis, this research intends to prove the network parameters’ discriminating feature and provide potential reference for scoring speaking performance.
This paper presents a novel high-order dependency parsing framework that targets non-projective treebanks. It imitates how a human parses sentences in an intuitive way. At every step of the parse, it determines which word is the easiest to process among all the remaining words, identifies its head word and then folds it under the head word. Further, this work is flexible enough to be augmented with other parsing techniques.
This paper analyzes the content of the proceedings of the Language Resources and Evaluation Conference (LREC) over the past 17 years (1998–2014), with the goal of gaining a picture of the LREC community and the topics that are most relevant to the field. We follow the methodology used in similar studies, including the survey of the IEEE ICASSP conference proceedings from 1976 to 1990, the survey of the Association of Computational Linguistics conference proceedings over 50 years, and the survey of the proceedings of the conferences contained in the ISCA Archive over 25 years (1987–2012). We expand on results originally presented at LREC 2014, but include the proceedings of LREC 2014 itself in the study together with an analysis of various citation graphs. We show the evolution over time of the number of papers and authors, including their distribution by gender and affiliation, as well as collaborations and citation patterns among authors and papers, funding sources for reported research, and plagiarism and reuse in LREC papers; results for LREC are compared with similar results for major conferences in related fields. We also consider the evolution of research topics over time and identify the authors who introduced key terms. Finally, we propose and apply a measure of a researcher’s notability and provide the results for LREC authors. The study uses NLP methods that have been published in the corpus considered in the study. In addition to providing a revealing characterization of the LRE community, the study also demonstrates the need for establishing a system for unique identification of authors, papers and other sources to facilitate this type of analysis.