Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
The recent success of statistical parsing methods has made treebanks become important resources for building good parsers. However, constructing highquality annotated treebanks is a challenging task. We utilized two publicly available parsers, Berkeley and MST parsers, for feedback on improving the quality of part-of-speech tagging for the Vietnamese Treebank. Analysis of the treebank and parsing errors revealed how problems with the Vietnamese Treebank influenced the parsing results and real difficulties of Vietnamese parsing that required further improvements to existing parsing technologies. 1
The article discusses some inaccurate uses of pronominal forms, certain terms, syntactical and lexical mistakes in papers on welding. The article also provides the examples of mistakes found in welding papers. The use of nonpronominal forms in the denoted terms is one of the typical morphological mistakes in treating the process of welding. Another problem is the inaccurate use and wrong formation of terminology. The main lexical mistakes are semantic mistakes and calque. The article also focuses on the syntactical mistakes – the mistaken expression of purpose, the wrong use of preposition prie, redundant use of abstract nouns in treatises, also inaccurate use of infinitive with conjuction kad in dependant clauses. Causes for language mistakes and correction of errors are given.
The French words relating to everyday, colloquial, popular, vulgar and argo: notional and \nterminological correlations \nThe language is one of the elements of the culture. Cités, which are mainly located in the suburbs of the major \ncities of France are often theater production of specific cultural norms. This language is called familiar, \ndomestic, popular, vulgar, slang. It is good to mark that language and vocabulary, in particular, arise from the \nsame natural sources, that is to say French and Latin, provincial and foreign. A choice of vocabulary is realized, \nof course, depending on the situation of communication. But there exists some ambiguity in identifying the words \nthat make up spoken French. These lexemes are nominated as current (courants), conversational (familiers), folk \n(populaires), vulgar (vulgaires), argo (argotiques). The paper interprets of the terms: vocabulary, \nconversational vocabulary, everyday language, slang, taking into consideration the development of linguistic \nresearch in the field of French lexicology. Clarifying this problematic issue to students of higher educational \nestablishments is up to date as vocabulary competence is always relevant for would-be linguists. However, it is \nnecessary to know how to identify these words for linguistic research students who are keenly interested in the \nstudy of language of young French. This is based on the language of the suburb city and is a mixture of slang, \nwords borrowed from French, slang old or different cultures that coexist in the city. Thus, clarify the status of \ncolloquial language in contemporary French, define the place of familiar lexicon in the lexical system of \ncontemporary French vocabulary, to differentiate the language of cities in France, different from standard \nvocabulary, will be the subject of study in this article. The structural and semantic peculiarities of the French \nlexicon constitute the scope of the search. Its purpose is, therefore, to define the linguistic properties of words \nbelonging to the familiar register. \nKey words: conversational French, vocabulary, slang, vulgar vocabulary, argo.
Human communication relies on words—spoken or written labels for the concepts we intend to convey. These linguistic units map meanings onto forms that can be recognized and produced by others within a shared communication system. To make this possible, words are stored in long-term memory within what is often called the mental lexicon. This repository includes orthographic, phonological, morphological, and semantic information, and enables retrieval whenever comprehension or production demands it. The act of retrieving such information is what researchers describe as lexical access. In reading, the orthographic stimulus must be matched with its stored representation, just as the phonological form of the acoustic signal must be matched during speech comprehension. In production, by contrast, the intended meaning serves as the entry point, giving access to the phonological or orthographic form required for speech or writing. The term lexical access was first popularized in studies of visual word recognition, but its use has since expanded. In current literature, especially on word recognition, alternative terms such as lexical retrieval or lexical processing are often preferred, since access implies a discrete lexical entry that can be “looked up.” This assumption is at odds with many contemporary models, which favor distributed, parallel-activation accounts where sublexical units such as letters, phonemes, or morphemes contribute dynamically to recognition. In contrast, in word production the notion of access is less contentious, because selecting the correct lexical item from meaning necessarily requires pinpointing a specific representation. Research into lexical processing has focused on two main questions: the nature of the stored representations in the mental lexicon and the cognitive procedures through which they are retrieved. Much of this work has been conducted by cognitive psychologists, leading to an emphasis on mechanisms of retrieval rather than linguistic content. Moreover, explanations have focused on the cognitive rather than the neural level, though psycholinguistic theories are increasingly informed by neuroscience. Indeed, although the present article emphasizes cognitive perspectives—as suggested by its title—key findings on neural and electrophysiological correlates of lexical processing are also acknowledged, since they have substantially contributed to refining and constraining cognitive theories. The bulk of empirical research has concentrated on visual word recognition, not only because reading experiments offer precise control and measurement, but also because of their pedagogical and societal importance. Nevertheless, the same theoretical questions extend to spoken word recognition, speech production, and writing, each of which poses its own challenges for models of lexical access. In recent years, important advances have reshaped the field: large-scale megastudies and open-access lexical databases now allow researchers to examine the joint influence of multiple lexical and semantic variables, moving beyond traditional factorial designs. Computational modeling has also become more diverse, integrating Bayesian frameworks, hybrid connectionist approaches, and deep learning architectures, while empirical work has expanded to a wider range of languages and writing systems. The references selected throughout this article represent either foundational studies that shaped the field or recent contributions that capture the current state of debate, providing a framework for understanding how humans connect word forms to meanings in real time. Updated in September 2025 by Maria Fernández-López.
A growing body of behavioral studies has demonstrated that women’s hemispheric specialization varies as a function of their menstrual cycle, with hemispheric specialization enhanced during their menstruation period. Our recent high-density electroencephalogram (EEG) study with lateralized emotional versus neutral words extended these behavioral results by showing that hemispheric specialization in men, but not in women under birth-control, depends upon specific EEG resting brain states at stimulus arrival, suggesting that hemispheric specialization may be pre-determined at the moment of the stimulus onset. To investigate whether EEG brain resting state for hemispheric specialization could vary as a function of the menstrual phase, we tested 12 right-handed healthy women over different phases of their menstrual cycle combining high-density EEG recordings and the same lateralized lexical decision paradigm with emotional versus neutral words. Results showed the presence of specific EEG r)
Abstract This chapter looks at more sophisticated versions of the ambiguity theory. We might say that “ought” is context sensitive rather than lexically ambiguous. And we can try wide-scoping. There are many ways of making (T) and (J) both come out true. But what we really want is to accept the norms. We don’t really just want the truth of the two sentences on some interpretation or another. But even in its more sophisticated versions, all the ambiguity theory can provide is the truth of the two sentences. The norms themselves remain inconsistent.
Traditionally, language processing has been attributed to a separate system in the brain, which supposedly works in an abstract propositional manner. However, there is increasing evidence suggesting that language processing is strongly interrelated with sensorimotor processing. Evidence for such an interrelation is typically drawn from interactions between language and perception or action. In the current study, the effect of words that refer to entities in the world with a typical location (e.g., sun, worm) on the planning of saccadic eye movements was investigated. Participants had to perform a lexical decision task on visually presented words and non-words. They responded by moving their eyes to a target in an upper (lower) screen position for a word (non-word) or vice versa. Eye movements were faster to locations compatible with the word's referent in the real world. These results provide evidence for the importance of linguistic stimuli in directing eye movements, even if the wor)
This article presents a comparative study on how textbooks of Portuguese and Spanish language handle mode and tense variation.. We took as theoretical background both studies on variation and teaching by Labov (1972, 1978 and 2003), and what official documents states about the teaching of foreign languages in Brazil. The research followed explicit instructions we developed for the analysis of norm/use, change, linguistic and extra-linguistic constraints, the use of authentic texts and variation between verbal tenses. The results obtained underlined that the themes in question were handled only partially by the textbooks: the text is still used as a pretext; when there is a theoretical explanation on variation, it appears only in the chapter at issue, not been applied throughout the book; the effects of meanings of the diverse verbal forms in the communicative context are also not substantiated. However, in the textbook Hacia el español, we were able to verify an effort in the sense of making students aware of the linguistic variation in phonetic and phonologic, lexical and morphosyntactic levels.
Jurislinguistic interpretation of the law contributes to clarifying the meaning of the norm, a better understanding of the content of the text, competent and understanding the contents of logical statements to study the documents. The content of the legal norm only in the process of understanding and clarifi cation of comments is becoming clear and accessible. Success of comments depends on the lexical-semantic analysis of a particular text with specifi c content. Yurislinguistics interpretation can be divided on extented and restrictive, which combine the interpretation and the result of legal education. During comments interpreter must take into account that the highest idea of justice is refl ected in the law. Therefore, the law should be interpreted and expressed in the relevant statements, which is precisely the result of legal interpretation. Currently, Tajik law-badly needs linguistic processing and refi nement, which is the main objective In order to improve the quality of legislation and paperwork the author makes a number of series of concrete proposals.
This article presents two case studies to explore whether and how web corpora can be used to automatically acquire lexical-semantic knowledge from distributional information. For this purpose, we compare three German web corpora and a traditional newspaper corpus on modelling two types of semantic relatedness: (1) Assuming that free word associations are semantically related to their stimuli, we explore to which extent stimulus– associate pairs from various associations norms are available in the corpus data. (2) Assuming that the distributional similarity between a noun–noun compound and its nominal constituents corresponds to the compound’s degree of compositionality, we rely on simple corpus co-occurrence features to predict compositionality. The case studies demonstrate that the corpora can indeed be used to model semantic relatedness, (1) covering up to 73/77% of verb/noun–association types within a 5-word window of the corpora, and (2) predicting compositionality with a correlation of ρ = 0.65 against human ratings. Furthermore, our studies illustrate that the corpus parameters domain, size and cleanness all have an effect on the semantic tasks.
Using the event-related optical signal (EROS) technique, this study investigated the dynamics of semantic brain activation during sentence comprehension. Participants read sentences constituent-by-constituent and made a semantic judgment at the end of each sentence. The EROSs were recorded simultaneously with ERPs and time-locked to expected or unexpected sentence-final target words. The unexpected words evoked a larger N400 and a late positivity than the expected ones. Critically, the EROS results revealed activations first in the left posterior middle temporal gyrus (LpMTG) between 128 and 192 ms, then in the left anterior inferior frontal gyrus (LaIFG), the left middle frontal gyrus (LMFG), and the LpMTG in the N400 time window, and finally in the left posterior inferior frontal gyrus (LpIFG) between 832 and 864 ms. Also, expected words elicited greater activation than unexpected words in the left anterior temporal lobe (LATL) between 192 and 256 ms. These results suggest that the )
The article focuses on the moral and ethical code of the Abkhaz people, Apsuara, and on the meaning that the Abkhazians invest in such concepts as alamys, anamys and apatu as core components of the Apsuara, and how valuable these concepts are for them. The Apsuara has no lexical equivalent in Russian language. The term is difficult to describe, to define and to analyze, and its literal meaning — «abhazstvo» — is very conventional and fi gurative. The article analyzes the main components of the Apsuara structure: worldview of Abkhazians, norms of social behaviour, rules of life, a system of upbringing and etiquette, and prohibiting categories. It compares the Apsuara with the Circassians` ethical system, the Adygage (“dejstvo”). Since 1960 Abkhaz scholars such as Sh D. and Inal-Ipa, EK Adzhindzhal and others have been trying to interpret the concept of Apsuara and its system by providing their own interpretations. This article analyzes several interpretations of the concepts Apsuara.
The paper considers the content and the ratio of the concepts of competence, competence, competence. The peculiarities of the linguistic of students of technical universities in Ukraine are analyzed. The analysis of the basic concept competence shows that it is treated in two ways: as a given rule, a requirement for specialist training and considered as the prevailing quality, the result of the learning activity of the student. Linguistic is assimilation, comprehension of language norms that have developed historically in phonetics, vocabulary, grammar, pronunciation, semantics, stylistics, and adequate use of them in a specific language. Speech and linguistic competences are inseparable because the ability to speak (speech competence) is based on grammatical, lexical and phonetic knowledge and skills (linguistic competence). The structure of linguistic includes phonological, lexical, grammatical, and spelling competences. The levels of linguistic (low, medium and high) are defined. The low level is characterized by mastery of basic language skills, grammar and general professional vocabulary. The middle level is characterized by the ability to produce professionally oriented language material, the high level corresponds to ability of using specific lexical units in dialogue and monologue speech on professional topics. It is revealed that the effectiveness of the mechanism of the formation of autonomous Ukrainian speech depends on how well students distinguish language means in Russian and Ukrainian languages, how they differentiate the two language systems. In addition, the efficiency of developing skills to communicate in the Ukrainian language depends on the language environment in which the young people are placed (at home, at school and out of it).
Reviewed by: Arguments as relations by John Bowers Diane Massam Arguments as relations. By John Bowers. (Linguistic inquiry monograph 58.) Cambridge, MA: MIT Press, 2010. Pp. xii, 239. ISBN 9780262514330. $25. We can get so comfortable with certain ideas that we forget why we hold them, until someone makes a proposal that turns things upside down: there is no D-structure, or control is raising, to mention two such proposals. John Bowers's proposal in this book fits into this category, literally turning some of our long-held views upside down. In sum, he argues that agents are merged very low, near the verb, with themes merged above them, and he explores the consequences of this idea. B's proposal appears to run counter to the Aristotelian view of sentences as consisting of a subject and a predicate [VO] (cf. Baker's 2001 verb object constraint), yet predication remains at the core of his work (Bowers 1993, 2001). The locus of predication has moved around over time. Since the VP-internal subject hypothesis (VPISH), there have been two potential sites for predication, one involving merge positions, with agent as the subject of a transitive verb, and the other involving grammatical positions, with an EPP-determined subject for the sentence. Some have suggested that the subject-predicate relation might exist only within vP in some languages (e.g. Massam 2001a), whereas B here is suggesting that what remains of D-structure is a verbal root with an upwardly extending ordered string of uniformly introduced arguments, so predication takes place only at the higher level through Agree and/or EPP. His view of argument structure evokes nonconfigurationality, in which there is also no VP, yet unlike such analyses (e.g. Jelinek 1984), for B, phrasal arguments constitute the true arguments of the clause and they are strictly ordered according to grammatical principles. B's work rests on two key points: first, that all argument structure is built through the ordered merging of functional heads, each taking a thematically specific argument in its specifier; and second, that the order of argument merge is universally fixed, with agents merging below themes. B's analysis of a basic transitive clause depends on his claim that the higher merged argument (theme) is local for the lower case relation (accusative from Voi (= Voice)), leaving the lower merged argument (agent) free to raise via EPP to PrP (Pr = Pred), and then undergo Agree with T (T = Tense), thus surfacing as the subject of the clause. His book is a set of arguments for this point of view, examining a range of constructions such as the passive (Ch. 2), affectee constructions (Ch. 3), applicative constructions (Ch. 4), and derived nominals (Ch. 5). The book also contains a brief appendix (Appendix A) that provides a compositional semantics for his analysis and another (Appendix B) that discusses the formal aspects of labeling and selection. In the rest of this review I outline each chapter of the book in turn, ending with some potential problems for B's view of argument structure. In his introductory chapter, B presents an overview of his 'radically different idea' (1) in which all arguments and modifiers are introduced uniformly by functional heads in accordance with a UNIVERSAL ORDER OF MERGE (UOM). Primary arguments are Agent, Theme, and Affectee, which are merged in this order, opposite to the norm. In addition to these arguments, there are secondary arguments (e.g. Instruments), and modifiers (e.g. Manner), also merged in accordance with the UOM. B outlines and counters the reasons why agents are traditionally merged high. His view is post-government and binding, in that syntactic structures provide the lexical semantics of the sentence, rather than being projected from it (as in Borer 2005). His approach here brings to mind construction grammar, where a given meaning is rigidly associated with a particular syntactic configuration. In the final section of this chapter B argues, against Marantz (1984), that subject idioms do exist (e.g. the lovebug bit NP, cf. Postal 2004). This chapter ends with a brief overview of the UOM and works through sample derivations of transitive, intransitive, passive, locative, [End Page 354] and expletive sentences, using...
In Japanese, vowel duration can distinguish the meaning of words. In order for infants to learn this phonemic contrast using simple distributional analyses, there should be reliable differences in the duration of short and long vowels, and the frequency distribution of vowels must make these differences salient enough in the input. In this study, we evaluate these requirements of phonemic learning by analyzing the duration of vowels from over 11 hours of Japanese infant-directed speech. We found that long vowels are substantially longer than short vowels in the input directed to infants, for each of the five oral vowels. However, we also found that learning phonemic length from the overall distribution of vowel duration is not going to be easy for a simple distributional learner, because of the large base-rate effect (i.e., 94% of vowels are short), and because of the many factors that influence vowel duration (e.g., intonational phrase boundaries, word boundaries, and vowel height). )
Psychophysiological evidence suggests that music and language are intimately coupled such that experience/training in one domain can influence processing required in the other domain. While the influence of music on language processing is now well-documented, evidence of language-to-music effects have yet to be firmly established. Here, using a cross-sectional design, we compared the performance of musicians to that of tone-language (Cantonese) speakers on tasks of auditory pitch acuity, music perception, and general cognitive ability (e.g., fluid intelligence, working memory). While musicians demonstrated superior performance on all auditory measures, comparable perceptual enhancements were observed for Cantonese participants, relative to English-speaking nonmusicians. These results provide evidence that tone-language background is associated with higher auditory perceptual performance for music listening. Musicians and Cantonese speakers also showed superior working memory capacity )
The article presents the outcome of research on 30 books of Quranic interpretations for sura al- Ghasyiah, verses 17-26, which are strongly assumed to contain pedagogic meanings, concepts, and values that can be formulated into an instructional model. The research was conducted by analyzing the keywords of the verses lexically, contextually, and hermeneutically. Then, the meanings gained were categorized, compared, contrasted, and abstracted, so that they were eventually synthesized into a main idea as a hypothetical model, termed M-3 Model. The model consists of three main instructional activities, represented in the terms munazharah, mudzakarah, and muhasabah. The three activities are a mutually completing and supporting cycle for the achievement of various instructional objectives, ultimately to improve the ability and skills of students to think systematically, logically, creatively, and innovatively through the development of potentials and fi trah (human norm). Specifi cally, munazharah activity is expected to result in cognitivistic knowledge (ainal yaqin), mudzarakah to develop knowledge, experience, and values into faith-based knowledge (‘ilm al-yaqin), and muhasabah to encourage the achievement of knowledge and values whose truths have been proven (haqqul yaqin), so that they will be the driving force for various activities based on law, moral, and ethics. Because the model taught by God to human beings is still hypothetical and theoretical in nature, it is suggested that the model be empirically tested to be more valid. Keywords: Instructional Model, Munazharah, Mudzakarah, Muhasabah
espanolPresentamos un sistema de normalizacion de tweets en espanol, que usa reglas de preproceso, un modelo de distancias de edicion adecuado al dominio y modelos de lengua para seleccionar candidatos de correccion segun el contexto. El sistema obtuvo resultados superiores a la media en la tarea Tweet-Norm de SEPLN 2013. EnglishWe present a system to normalize Spanish tweets, which uses preprocessing rules, a domain-appropriate edit-distance model, and language models to select correction candidates based on context. The system’s results at SEPLN 2013 Tweet-Norm task were above-average.
In article questions of a theoretical and practical lexicology on the example of the linguistic analysis of publicist texts with psychological semantics are considered. Features of author's generation and reader's potential perception of this type of value in heading components of texts of modern means of communication are defined. Psychological semantics as part of an information field of words and the text can be initial, the main in relation to event, and also increment, received in the course of communication. It is capable to become a leading sign of the general contents or positionally to be staticized depending on a speech situation, an intellectual and emotional condition of communicators, their life experience. Studying of functional and semantic opportunities of the text gets anthropocentric approach to studying of its language form. The comparative analysis of headings, their options, and also observance of norms of the literary language is an example of it. Studying of lexical components of the text with psychological semantics can serve understanding of features of communicative process in modern living conditions of the language personality that allows to express more fully positive and negative emotions, to create the identity to society.
Purpose: this article discusses which implicit meanings can be uncovered by means of the etymological analysis of the lexical units containing a numerical component as seen in the example of the lexemes which describe the degree of drunkenness in the English language. Methodology: method of continuous sampling; descriptive method, etymological analysis of the lexical units. Results: The analysis of the numerical component in the DRINKING concept helps one to better study the inner form of the latter and understand its cultural meanings. The ideas taken as a basis for the nomination of some phrases are impossible to clarify, however all the expressions describe a different degree of violation of the norm. Practical implications: lectures on stylistics and lexicology; teaching English as a foreign language; dictionary compiling; interpreter preparation. DOI: http://dx.doi.org/10.12731/2218-7405-2013-7-6
Web 2.0 provides user-friendly tools that allow persons to create and publish content online. User generated content often takes the form of short texts (e.g., blog posts, news feeds, snippets, etc). This has motivated an increasing interest on the analysis of short texts and, specifically, on their categorisation. Text categorisation is the task of classifying documents into a certain number of predefined categories. Traditional text classification techniques are mainly based on word frequency statistical analysis and have been proved inadequate for the classification of short texts where word occurrence is too small. On the other hand, the classic approach to text categorization is based on a learning process that requires a large number of labeled training texts to achieve an accurate performance. However labeled documents might not be available, when unlabeled documents can be easily collected. This paper presents an approach to text categorisation which does not need a pre-classified set of training documents. The proposed method only requires the category names as user input. Each one of these categories is defined by means of an ontology of terms modelled by a set of what we call proximity equations. Hence, our method is not category occurrence frequency based, but highly depends on the definition of that category and how the text fits that definition. Therefore, the proposed approach is an appropriate method for short text classification where the frequency of occurrence of a category is very small or even zero. Another feature of our method is that the classification process is based on the ability of an extension of the standard Prolog language, named Bousi~Prolog , for flexible matching and knowledge representation. This declarative approach provides a text classifier which is quick and easy to build, and a classification process which is easy for the user to understand. The results of experiments showed that the proposed method achieved a reasonably useful performance.
The paper considers consequences of the unprecedented phenomenon in the world's system of languages, i.e. the transformation of the English language into the language of global communication as the result of information revolution and globalization of all aspects of human activities, which have changed generally accepted ideas about foreign languages and the notion of literacy. The emergence of a new educational paradigm according to which the English language is no longer considered as ''foreign'', but as a sine qua non of ensuring the participation of all Europeans in the new knowledge society has been recognized by the European Commission in its document issued in 2005, ''A New Framework Strategy for Multilingualism''. The paper presents an analysis of reasons which have brought about the acquiring by the English language of the status of the language of world communication, and takes on the question of what the language of world communication is and in which way it differs from the national variants of the English language. One of the characteristic features of the English language as a lingua franca is its high variability, which is not confined to differences in grammatical and lexical structures of the two major variants of the English language that have developed historically in the course of the emergence of the North American standard of the English language. The paper next considers the problem of the world standard of English as the global language and argues the unfoundedness of the thesis about the need of recognizing the right of developing language norms by the users of English as the second language, even though their number at present significantly exceeds the number of native speakers. Summing up, a conclusion can be made that the transformation of English into the language of global communication calls forth a revision of the traditional approach to teaching foreign languages. English as the global lingua franca has lost its status as a foreign language. This demands a reorganization of language education based on a transition to multilingual teaching and learning in which English language teaching is envisioned not as traditional teaching of one of the national variants of English, which leads to acquiring the corresponding dominant cultures, but as teaching English as the language of world communication used for overcoming interlingual and intercultural barriers in the globalizing world. The transition to multilingual teaching is based on regarding English as a necessary condition for entry into the world economic, political and cultural areas and includes, alongside with the native (or state) language and English as a language of global communication, teaching at least one of the foreign languages offered by the educational system.
Unlike some varieties of English in Southeast Asia, the notion that there is a ‘Thai English’ is debatable. This paper examines distinctive non-native features of a lexicon found in contemporary Thai writing in English to ascertain if English in this Expanding Circle country is developing its own linguistic norms. An analysis of features of lexical creativity in five short stories and novels is carried out to determine whether the characteristics found indicate that a Thai English vocabulary exists. An ‘integrated framework’ which combines concepts in World Englishes by Braj B, Kachru, Peter Strevens, and Edgar W. Schneider is adopted in this study. It appears that certain categories of lexical creativity in the fiction examined represent five indicators of Thai English - contextualization, innovation, nativization, transcultural creativity, and localization - and reveals a developing non-native variety of English.
In a large scale study on 843 transcripts of Technology, Entertainment and Design (TED) talks, the authors address the relation between word usage and categorical affective ratings of lectures by a large group of internet users. Users rated the lectures by assigning one or more predefined tags which relate to the affective state evoked in the audience (e. g., ‘fascinating’, ‘funny’, ‘courageous’, ‘unconvincing’ or ‘long-winded’). By automatic classification experiments, they demonstrate the usefulness of linguistic features for predicting these subjective ratings. Extensive test runs are conducted to assess the influence of the classifier and feature selection, and individual linguistic features are evaluated with respect to their discriminative power. In the result, classification whether the frequency of a given tag is higher than on average can be performed most robustly for tags associated with positive valence, reaching up to 80.7% accuracy on unseen test data.
This paper presents a reranking approach to combining constituent and dependency parsing, aimed at improving parsing performance on both sides. Most previous combination methods rely on complicated joint decoding to integrate graph- and transition-based dependency models. Instead, our approach makes use of a high-performance probabilistic context free grammar (PCFG) model to output k-best candidate constituent trees, and then a dependency parsing model to rerank the trees by their scores from both models, so as to get the most probable parse. Experimental results show that this reranking approach achieves the highest accuracy of constituent and dependency parsing on Chinese treebank (CTB5.1) and a comparable performance to the state of the art on English treebank (WSJ).
The problem of Vietnamese syntactic parsing, especially constituency parsing, has recently been tackled by several research groups. A common effort of the Vietnamese language processing community has allowed the creation of VietTreebank, a reference parsed corpus containing about 10,000 sentences for the constituency parsing task. In this paper, we present our work to build a reference treebank, based on VietTreebank, for the dependency parsing task, which has not yet been very well studied for Vietnamese. First we define a dependency label set by adapting the dependency schema developed by the NLP group at Stanford university and taking into account the particularities of Vietnamese grammar. Then we propose an algorithm to convert a constituency treebank to a dependency one. The algorithm is tested on a set of 100 sentences of VietTreebank corpus and gives very good results. Finally, we carry out an experiment on Vietnamese dependency parsing using MaltParser tool and the dependency treebank converted from VietTreebank.
This paper sets out to study the letters of Gaston B., a French prisoner of war held in captivity in the camp of Münster (Germany) from the beginning of the First World War until its end. These letters make possible a relativisation of linguistic macrohistory through microhistory, by focussing on the grassroots level and by using as sources the traces of people with no significant name or identity. They shed important light on how a member of a lower class acquired the prescriptive linguistic norm through his schooling at the end of the nineteenth century and how this affected his subsequent linguistic behaviour. An individual is exposed to the political and social dimension of language planning, and his language reflects its level of success, but also reveals what grammatical tools and rules have been focused on during his schooling.
In recent years much progress has been made in developing systematic protocols for finding linguistic metaphors in authentic language data. The description of conceptual structures, however, has not been placed on equally firm footing. One existing proposal, known as the five-step method, introduces systematicity to the process of determining conceptual structures of metaphors in discourse. However, it does not take sufficient steps to minimize intuition and to maximize transparency. This paper seeks to reduce these weaknesses by introducing the systematic use of dictionaries and a lexical database. The result is a more transparent and constrained method.
Sylvain Kahane est Professeur en Sciences du Langage à l'Université Paris Ouest - Nanterre et membre du laboratoire Modyco (CNRS UMR 7114). Support de présentation de Sylvain Kahane: PDF Podcast: Résumé de l'intervention: Nous présenterons les différentes couches d'annotation du treebank Rhapsodie, un corpus de français parlé richement annoté. Le corpus contient plusieurs niveaux de segmentation indépendants: en unités illocutoires pour la macrosyntaxe, en unités rectionnelles pour la mi...
This paper reports an effort to annotate modality in the Penn Chinese Treebank. We introduce the modals and features that were annotated, and describe the phases of our working process. Along with this, we address the issues in the preparation of annotation guidelines, and present the preliminary results of the first pass. Finally, we analyze the types of disagreement, and propose directions to improve consistency. 1
In this work, we present a data-driven method to enhance syntax trees with additional dependencies as defined in the wellknown Stanford Dependencies scheme, so as to give more information about the structure of the sentence. This hybrid method utilizes both machine learning and a rule-based approach, and achieves a performance of 93.1 % in F1-score, as evaluated using an existing treebank of Finnish. The resulting tool will be integrated into an existing Finnish parser and made publicly available at the address
We describe a novel approach to detecting empty categories (EC) as represented in de-pendency trees as well as a new metric for measuring EC detection accuracy. The new metric takes into account not only the position and type of an EC, but also the head it is a dependent of in a dependency tree. We also introduce a variety of new features that are more suited for this approach. Tested on a sub-set of the Chinese Treebank, our system im-proved significantly over the best previously reported results even when evaluated with this more stringent metric. 1
BACKGROUND: Previous studies investigating speech recognition in adverse listening conditions have found extensive variability among individual listeners. However, little is currently known about the core underlying factors that influence speech recognition abilities. PURPOSE: To investigate sensory, perceptual, and neurocognitive differences between good and poor listeners on the Perceptually Robust English Sentence Test Open-set (PRESTO), a new high-variability sentence recognition test under adverse listening conditions. RESEARCH DESIGN: Participants who fell in the upper quartile (HiPRESTO listeners) or lower quartile (LoPRESTO listeners) on key word recognition on sentences from PRESTO in multitalker babble completed a battery of behavioral tasks and self-report questionnaires designed to investigate real-world hearing difficulties, indexical processing skills, and neurocognitive abilities. STUDY SAMPLE: Young, normal-hearing adults (N = 40) from the Indiana University community participated in the current study. DATA COLLECTION AND ANALYSIS: Participants' assessment of their own real-world hearing difficulties was measured with a self-report questionnaire on situational hearing and hearing health history. Indexical processing skills were assessed using a talker discrimination task, a gender discrimination task, and a forced-choice regional dialect categorization task. Neurocognitive abilities were measured with the Auditory Digit Span Forward (verbal short-term memory) and Digit Span Backward (verbal working memory) tests, the Stroop Color and Word Test (attention/inhibition), the WordFam word familiarity test (vocabulary size), the Behavioral Rating Inventory of Executive Function-Adult Version (BRIEF-A) self-report questionnaire on executive function, and two performance subtests of the Wechsler Abbreviated Scale of Intelligence (WASI) Performance Intelligence Quotient (IQ; nonverbal intelligence). Scores on self-report questionnaires and behavioral tasks were tallied and analyzed by listener group (HiPRESTO and LoPRESTO). RESULTS: The extreme groups did not differ overall on self-reported hearing difficulties in real-world listening environments. However, an item-by-item analysis of questions revealed that LoPRESTO listeners reported significantly greater difficulty understanding speakers in a public place. HiPRESTO listeners were significantly more accurate than LoPRESTO listeners at gender discrimination and regional dialect categorization, but they did not differ on talker discrimination accuracy or response time, or gender discrimination response time. HiPRESTO listeners also had longer forward and backward digit spans, higher word familiarity ratings on the WordFam test, and lower (better) scores for three individual items on the BRIEF-A questionnaire related to cognitive load. The two groups did not differ on the Stroop Color and Word Test or either of the WASI performance IQ subtests. CONCLUSIONS: HiPRESTO listeners and LoPRESTO listeners differed in indexical processing abilities, short-term and working memory capacity, vocabulary size, and some domains of executive functioning. These findings suggest that individual differences in the ability to encode and maintain highly detailed episodic information in speech may underlie the variability observed in speech recognition performance in adverse listening conditions using high-variability PRESTO sentences in multitalker babble.
This paper has two main objectives. The first is to provide an overview of the CDT annotation design with special emphasis on the modeling of the interface between syntactic and morphological structure. Against this background, the second objective is to explain the basic fundamentals of how CDT is marked-up with semantic relations in accordance with the dependency principles governing the annotation on the other levels of CDT. Specifically, focus will be on how Generative Lexicon theory has been incorporated into the unitary theoretical dependency framework of CDT by developing an annotation scheme for lexical semantics which is able to account for the lexico-semantic structure of complex NPs.
ABSTRACT The aim of this article is to carry out a structural-functional analysis of the formation of Old English adjectives by means of affixation. By analysing the rules and operations that produce the 3,356 adjectives which the lexical database of Old English Nerthus (www.nerthusproject.com) turns out as affixal derivatives, a total of fourteen derivational functions have been identified. Additionally, the analysis yields conclusions concerning the relationship between affixes and derivational functions, the patterns of recategorization present in adjective formation and recursive word-formation.
According dependency syntactic theory this paper gave Tibetan typed dependencies and its hierarchy,and then we analyzed some problems in building Tibetan dependency Treebank.We proposed a mode to construct dependency tree semi-automatically,it includes word-pairs dependency classification model and dependency edges annotation model with rich features template based on Tibetan language grammar.And we implemented visualized tool which used to build and proofreading 11thousand sentences Treebank.On the baseline system the experimental results show that,the dependency recognition accuracy obtains an improvement of 3%.
This paper tries to give answers for successful receptive multilingualism (RM) but also for its failure. It is mainly based on the results of two projects, one on inter-dialectal communication in the Baltic area during the era of the Hanseatic League and the other analyses inter-Scandinavian communication today. The main purpose of this survey is to outline the essential preconditions for successful RM, from a linguistic, social and environmental perspective. The historical project about communication in the Baltic focuses on long-term language contact based on common mutual trading interests whilst the contemporary project highlights the cultural factors (among others Pan-Scandinavism) as a common basis for using one's own mother tongue in transnational communication. Moreover, other relevant issues belonging to successful RM are touched upon, such as diglossia (i.e. the functional distribution of different languages/varieties in various settings), oral face-to-face communication, the absence of written norms and the non-existence of standardised forms, which result not only in a greater flexibility in communication but also support openness for divergent varieties. Disfavouring factors for RM are, however, taken into consideration as well, such as nationalism and the suppression of minorities (and thus indirectly multilingualism), the enforcement of strict linguistic norms by the society and finally the use of a lingua franca such as Latin in the Middle Ages or English today.
Given the enormity of the app market and the velocity with which new apps arrive, it is extremely challenging for apps to reach the intended audience, or any audience at all, in fact. An entire "app promotion" industry exists to help publishers achieve post-launch app success. This is done by understanding relationships between app success and a variety of app attributes (like category, price, etc.). In this paper, we study a dimension not addressed thus far - the timing of app launch. Specifically, we study a large data set to uncover relationships between app launch times and its subsequent commercial success, or lack thereof. A number of interesting findings are revealed in this study. Users are generally less price-sensitive around holiday seasons, especially around Christmas and New Year's, in stark contrast, they are extra price sensitive during weekends. Specifically, more expensive apps released on weekends tend to get a higher negative word-of-mouth (review valence) rating. In addition, our results indicate that apps released in the latter part of the week tend to fare better than do apps released earlier. Furthermore, Thursday is the optimal day to release an app when considering review sentiments. Finally, it does tend to get a higher number of reviews on weekends.
This study examines verb modes and tenses as well as the predicate structure in «Libro decimosexto» of Bartolomé Jiménez Patón’s Comentarios de erudición in an effort to demonstrate how the text straddles the line between Medieval and Golden Age norms. Jiménez Patón, thus, combines traits already considered archaic at the time, due perhaps to his solid training in grammar and his linguistic awareness, with those modern solutions which, towards the end of the 1500’s, were forging the new Spanish language (that of the «plain style», which he championed), one that was gradually being refined in order to take its place as the language of culture.