Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
We have developed EyeMap, a freely available software system for visualizing and analyzing eye movement data specifically in the area of reading research. As compared with similar systems, including commercial ones, EyeMap has more advanced features for text stimulus presentation, interest area extraction, eye movement data visualization, and experimental variable calculation. It is unique in supporting binocular data analysis for unicode, proportional, and nonproportional fonts and spaced and unspaced scripts. Consequently, it is well suited for research on a wide range of writing systems. To date, it has been used with English, German, Thai, Korean, and Chinese. EyeMap is platform independent and can also work on mobile devices. An important contribution of the EyeMap project is a device-independent XML data format for describing data from a wide range of reading experiments. An online version of EyeMap allows researchers to analyze and visualize reading data through a standard Web browser. This facility could, for example, serve as a front-end for online eye movement data corpora.
This stylebook is an updated version of Telljohann et al. (2006). It describes the design principles and the annotation scheme for the German treebank TüBa-D/Z developed by the Division of Computational Linguistics (Lehrstuhl Prof. Hinrichs) at the Department of Linguistics (Seminar für Sprachwis-
The social and cultural ‘turn’ in language education of recent years has helped move language teaching and curriculum design away from many of the more rigid dogmas of earlier generations, but the issue of the roles of the learners’ first language (L1) in language pedagogy and classroom interaction is far from settled. Some follow a strict ‘exclusive target language’ pedagogy, while others ‘resort to’ the use of the L1 for a variety of purposes (see ACTFL 2008). Underlying these competing views is the perspective of the L1 as an impediment to second language learning. Following sociocultural theory and ecological perspectives of language and learning and based on the findings of research on classroom code-switching and code choice, this paper lays out an approach to the language classroom as a multilingual social space in which learners and teacher study, negotiate, and co-construct code choice norms toward the dynamic, creative, and pedagogically effective use of both the target language and the learners’ L1(s). Learner use of the L1 for the purpose of grammatical or lexical learning is also considered, and some examples for instruction are offered.
This work describes how derivation tree fragments based on a variant of Tree Adjoining Grammar (TAG) can be used to check treebank consistency. Annotation of word sequences are compared both for their internal structural consistency, and their external relation to the rest of the tree. We expand on earlier work in this area in three ways. First, we provide a more complete description of the system, showing how a naive use of TAG structures will not work, leading to a necessary refinement. We also provide a more complete account of the processing pipeline, including the grouping together of structurally similar errors and their elimination of duplicates. Second, we include the new experimental external relation check to find an additional class of errors. Third, we broaden the evaluation to include both the internal and external relation checks, and evaluate the system on both an Arabic and English treebank. The evaluation has been successful enough that the internal check has been integrated into the standard pipeline for current English treebank construction at the
We present the ongoing development of MCG, a linguistically deep and precise grammar for Mandarin Chinese together with its accompanying treebank, both based on the linguistic framework of HPSG, and using MRS as the semantic representation. We highlight some key features of our grammar design, and review a number of challenging phenomena, with comparisons to alternative linguistic treatments and implementations. One of the distinguishing characteristics of our approach is the tight integration of grammar and treebank development. The two-step treebank annotation procedure benefits from the efficiency of the discriminant-based annotation approach, while giving the annotators full freedom of producing extra-grammatical structures. This not only allows the creation of a precise and full-coverage treebank with an imperfect grammar, but also provides prompt feedback for grammarians to identify the errors in the grammar design and implementation. Preliminary evaluation and error analysis shows that the grammar already covers most of the core phenomena for Mandarin Chinese, and the treebank annotation procedure reaches a stable speed of 35 sentences per hour with satisfying quality. Keywords:Grammar Engineering, Treebank Annotation, Syntax 1.
The article covers analysis of lexical units, which represent the concept of the child in folk culture. At present dialectal cultural linguistics, called to model dialectal linguistic world-image, is quickly developing. The urgency of studying folk dialects and dialectal word's cultural meanings is caused by social aspiration for self-cognition, which among other things is achieved by means of traditional culture exploration. The present research has been based on the material of Middle Priobie dialects. The object of the research is a dialectal word, which contains a cultural component in its semantic structure. The approach to the lexical dialect system is realized through characterizing the concept of the child from the positions of cultural linguistics. The choice of this concept as the research object is caused by the fact that from the scientific point of view childhood is a particular phenomenon, which, when studied, shows the world of ''adult'' culture and makes it possible to remodel its world-view principles. In the peasants' world family is the basic community unit, which explains the village society's great attention to inter-family relationship. Families, in which parents have children together, are considered the standard. The deviation from the norm is registered by means of the language. The article covers lexical representation of three situations closely connected with the child's family status: belonging to one of the spouses only, orphanhood and illegitimacy. On traditional mind's mental level there is an opposition of related and unrelated children, which is realized in syntagmatic expansion. The ''related unrelated'' opposition is a special case of ''one's own somebody else's'' opposition, according to which everything that is not ''one's own'' or ''together'' is estranged. The idea of ''extraneity'' models stereotyped images of the unjustly oppressed stepson and stepdaughter. The denomination ''orphan'' marks a child who lost parents out of other children's mass. In folk mind the image of the orphan is closely connected with the notion of fate. Orphan's emotional deprivation is explained by his loneliness. Compassionate treatment of orphans is represented on the language level through the ability of this word to have a diminutive form and its semantic support. The natural child in traditional culture is a marginal creature by birth. Since he was born out of wedlock, outside the law, he is unprotected before the society. Negative attitude towards the woman who gave birth out of wedlock and her child is realized in abusive, insulting designations. There are also denominations formed from reputed loci of the natural child's conception or birth. Their semantics is opposed to the idea of the house as ''one's own'' space. It is remarkable that there are no particular appellations for legitimate children since legitimacy is considered as a norm and is not marked by linguistic means. Thus, the consideration of linguistic realisation of the notion of the child's family status allows us model a fragment of the native dialect speaker's value worldimage. It is possible to make a conclusion that family occupies one of the fundamental places in the world-view constants of traditional culture.
The subject of the article is the actual state and functioning of modern Brasilian Portuguese language norm. The European Portuguese language norm, considered as a standard in Brasil, in the recent decades has been losing its social base, while the national (Brasilian) forms of speech are being actively recognized in grammatical and lexical descriptions.
This study was designed to examine the relationship between deficits in empathy, emotional responsivity, and social behavior in adults with severe traumatic brain injury (TBI). A total of 21 patients with severe TBI and 25 control participants viewed six film clips containing pleasant, unpleasant, and neutral content whilst facial muscle responses, skin conductance, and valence and arousal ratings were measured. Emotional empathy (the Balanced Emotional Empathy Scale, BEES: self-report) and changes in drive and control in social situations (The Current Behaviour Scale, CBS: relative report) were also assessed. In comparison to control participants, those in the TBI group reported less ability to empathize emotionally and had reduced facial responding to both pleasant and unpleasant films. They also exhibited lowered autonomic arousal, as well as abnormal ratings of valence and arousal, particularly to unpleasant films. Relative reported loss of emotional control was significantly associated with heightened empathy, while there was a trend to suggest that impaired drive (or motivation) may be related to lower levels of emotional empathy. The results represent the first to suggest that level of emotional empathy post traumatic brain injury may be associated with behavioral manifestations of disorders of drive and control.
Anhedonia, a reduced capacity for pleasure, is viewed as a trait-like vulnerability marker for schizophrenia and depressive disorders. To date there are scarce data from the Arab world on anhedonia as a symptom, and even less on the psychometric properties of instruments designed to assess it. This study examines the internal consistency of an Arabic version of the Snaith Hamilton Pleasure Scale (SHAPS), and its convergence with real-time hedonic responses to emotional stimuli. A correlational study design is used; undergraduate students ( N = 113) in the United Arab Emirates (UAE) completed the SHAPS, and also undertook an expanded version of the picture rating task (PRT). The PRT required participants to rate a series of pleasant and unpleasant images in terms of emotional valence. Levels of anhedonia as assessed by SHAPs were similar to those observed in nonclinical populations in other countries. Internal consistency for the Arabic version of SHAPs was very good; α =.86. Furthermore, SHAPS scores were correlated with lower valence ratings for pleasant images ( r =.36), and uncorrelated with unpleasant images. The SHAPS appears to be a useful instrument for assessing anhedonia in the present UAE student population.
En psychologie tout comme en traitement automatique des langues, les normes qui portent sur des proprietes semantiques des mots, comme le degre d’abstraction, l’imagerie ou la polarite, sont importantes. Ces normes ont systematiquement ete obtenues en demandant a des juges d’evaluer les mots sur des echelles, allant par exemple de tres concret a tres abstrait. Ce mode de recolte etant lent et couteux, des methodes de construction automatique ont vu le jour. Elles peuvent etre divisees en deux types: celles qui se basent sur des ressources linguistiques et celles qui se basent sur des corpus. Notre objectif est de comparer, pour une meme methode d’accroissement de normes lexicales basee sur les similarites entre les mots, l’utilisation d’un corpus et d’une ressource lexicale (WordNet) pour estimer ces similarites. Nous montrons que les similarites calculees a partir d’informations sur les cooccurrences des mots dans les textes sont plus efficaces, et ce pour 4 des 5 normes etendues. Nous montrons egalement que le choix du corpus influence peu les resultats, du moins pour des corpus generaux.
Cette thèse porte sur une étude de la variation et du changement lexicaux des mots référant aux notions de « véhicule automobile » et de « travail rémunéré » dans le français de l’Outaouais, une variété de français laurentien caractérisée par le bilinguisme équilibré et stable et le contact intense avec l’anglais. La thèse est réalisée dans le cadre de la sociolinguistique variationniste labovienne combinée avec des méthodes quantitatives et des techniques analytiques multivariationnelles des règles variables. Cette étude se base sur les données empiriques recueillies dans les communautés francophones de la région de la capitale canadienne parmi les locuteurs nés entre 1846 et 1994 (RFQ, Ottawa-Hull, FdO).\nLe chapitre 2 suit l’évolution sémantique des termes lexicaux, étudie un système d’interaction des facteurs historiques et ceux socialement motivés, et examine l’hypothèse du développement interne du vocabulaire du français canadien. Les chapitres 3 et 4 examinent la corrélation des facteurs liés au bilinguisme et au contact avec l’anglais avec la fréquence d’emploi des variables lexicales; et la marque sociale des variantes lexicales.\nCette thèse: i) met en valeur la méthodologie variationniste quantitative dans l’étude de la variation lexicale; ii) approfondit plusieurs réflexions théoriques et des patrons classiques sur la théorie variationniste; iii) caractérise le lien dynamique entre le parler des locuteurs et les normes de la communauté à laquelle ils se rattachent; iv) contribue à la meilleure compréhension de la dynamique lexicale en fonction du statut du français en situation de contact de langues.
Written texts, particularly published ones, are widely perceived to have legitimacy beyond that of the spoken word in literate societies. One reason perhaps is that written text generally has more staying power as a concrete and tangible documentation of thought, intention, information, agreements and so on than does the spoken word. Furthermore, in many cases written text assumes a larger, more unifi ed, identifi able audience that is refl ective of those who share some subset of cultural norms and values. With legitimacy and cultural norms as a backdrop, the reason for code-switching in written discourse becomes an interesting subject of inquiry. Why switch between languages in a medium where one has ample time and resources to produce a monolingual text per the expected norm? Researchers have identifi ed this phenomenon in texts ranging from blogs to historical documents and have proposed various accounts, some of which are presented in this volume. On the surface, the switches found in written text may look and read like typical oral code-switches, where two or more languages are used, at times inter-sententially, at times intra-sententially and occasionally intra-lexically, with bound and free morphemes of two (or more) languages collaborating to create a discourse.
The paper studies Hausa film language through the analysis of three communication strategies, namely proverbs, imperatives and forms of address. It shows that Hausa film creates a new discourse by reflecting modern and traditional Hausa society. The films preserve some accepted cultural norms of behavior and norms of communication in order to please the more conservative public. On the other hand, combination of traditional and modern Hausa lifestyle evokes changes in the discourse. The paper shows that proverbs are commonly used as communication strategy for indirectness, rather than a specialized language. It also discovers that imperatives are used as communication strategy in close relations between interlocutors (no matter what their social status is) to express the direct message. As for forms of address, traditional and borrowed terms reflect the changing style of life. The examples extracted from the Hausa films are to show how the regular grammatical and lexical means change their discourse function in new social context.
Translating legal texts is a highly complex and multilayered process, especially for non-professionals. The first obstacle is the discipline itself. Legal systems differ from one another and each has its own specific norms, which is especially reflected at the lexical level, or in the terminology. Translating legal texts is thus primarily a type of comparative law because we continually ponder to what extent the terms in the target text correspond to the terms in the source text; they only rarely match completely in terms of content and they most often differ from one another to various degrees. Lexical gaps are also frequent; for example, a term exists in the source legal system, but there is no equivalent for it in the target system. An important criterion is also the text type; different text types (e.g., normative, descriptive, etc.) demand different translation approaches and strategies. The text type also determines whether a translation equivalent is easy or difficult to obtain. This article presents some of the main problems that translation students encounter in the elective translation module Translating Legal Texts. Each group of problems first includes the theoretical premises for the issue in translation studies, followed by specific illustrative examples and recommended translation strategies.
Abstract Scholars have struggled with the meaning of the word אֲנָךְ, which appears twice in Amos 7:7 and twice in Amos 7:8. Traditionally, many scholars have defined אֲנָךְ as “lead” and understood it with reference to a plumb line. In recent years, however, the majority of readers have instead defined אֲנָךְ as “tin.” This has spawned a variety of interpretations that try to make sense of how exactly tin might fit into the context of Amos 7:7-9. This paper critiques the notion that אֲנָךְ means “tin” rather than “lead.” It shows that, contrary to the present interpretative norm, there are no lexical grounds for defining אֲנָךְ exclusively as “tin.” From a lexicographical perspective, the definition “lead” remains the most viable option.
This is a pilot study investigating the role of phrasal fixedness in the development of a standardised text type. The linguistic material comes from the Edinburgh Corpus of Older Scots (ECOS), consisting of samples of administrative records from 15th-century Scotland. The corpus has been searched for re-occurring lemmatic bundles, which are the indicators of emerging patterns and standardising usage in the records, developing in the context of linguistic standardisation of Scots. The findings are interpreted with regard to their semantics and function in the records, and indicate that the text type as such was not yet fully standardised in its repertoire of fixed phrases serving a specific purpose. In individual locations, however, one finds a greater degree of consistency and a tendency to develop a local norm. Similarly, in some specific textual functions the lexical fixedness may be present to a larger extent than in others.
In this paper, we introduce the syntactic annotation of the CReST corpus, a corpus of natural language dialogues obtained from humans performing a cooperative, remote search task. The corpus contains the speech signals as well as transcriptions of the dialogues, which are additionally annotated for dialogue structure, disfluencies, and for syntax. The syntactic annotation comprises POS annotation, Penn Treebank style constituent annotations, dependency annotations, and combinatory categorial grammar annotations. The corpus is the first of its kind, providing parallel syntactic annotation based on three different grammar formalisms for a dialogue corpus. All three annotations are manually corrected, thus providing a high quality resource for linguistic comparisons, but also for parser evaluation across frameworks.
The article analyzes the problem of setting the norm for a number of new language elements in modern Russian. The given Russian material provides reviewing the functioning of the triad “the norm a variation of the norm speech error”; different codification degree is being revealed in the examples granted. Criteria to define a language fact as normative or wrong have been discussed until the present day. The objective of this study is the analysis of new lexical units of modern Russian as to their correspondence to the norm. The test object makes up the analysis of abstracts from printed and electronic Russian media sources through continuous sampling. The article treats non-typical Russian lexical units such as contaminated complexes making loan transitions of English words meanings. As is known, the Russian language borrows not only lexical units, but word-building patterns of other languages. The question of codification of such language facts is still overt, but the increasing trend of their usage, the polysemy growth makes it possible to speak of their fixation in dictionaries. The results of the analysis could be used in composing new lexical units dictionaries.
The last known printed work of Šimun Kožičić Benja’s printing house in Rijeka is ''Od bitija redovničkoga knjižice''. At the end of the booklet there is a colophon with the date of its completion: ''dan 27. maja miseca: leto od Krstova rojstva 1531.''. A facsimile reprint of the only preserved specimen (which is the last of the six known Glagolitic publications of Kožičić's printing house) was published in Rijeka in 2009. This publication has finally returned this valuable booklet in situ and made it available to the wider scientific and cultural community. This is also due to the fact that this facsimile edition contains the transcript of the Glagolitic text with an introduction and a glossary, written by the academician Anica Nazor, an eminent researcher and promoter of Kožičić's Glagolitic work. In the introduction, which precedes the Latin transcription of the Glagolitic text, Anica Nazor writes about the language of this booklet saying it is "strongly imbued by Church-Slavonic elements". The linguistic features of ''Od bitija redovničkoga knjižice'' are the central part of this paper, and the emphasis is on the selected phonological, morphological and lexical features. These linguistic features are observed in the context of other Kožičić's publications and the former knowledge of his literary and linguistic concept, especially in regard to his treatment of Church-Slavonic and Croatian linguistic norm. The results of the conducted analysis confirm that ''Od bitija redovničkoga knjižice'' is another outcome of Kožičić's thoroughly established literary and linguistic concept that includes explicit correlation of Church-Slavonic and Old-Croatian linguistic features.
The paper reviews the development of Quebec lexicography in the 18th-19th centuries analyzing Quebec lexicographic works which are practically unknown to Russian scientists in lexical variation studies. The article offers a general linguistic characteristic of the works, namely description of their volume and structure, analysis of the methods to present word entries, comparison of registered examples of the French language contacts with the languages of the autochthonic American Indian population and with the English language. Special attention is paid to the approaches of various Quebec authors to the description of Quebec word usage of the considered period (prescriptive/descriptive ideologies). The author of the article characterizes the social context of the discussed lexicographic works. The article outlines scientific reflection formation in the field of language norm in Quebec variant of the French language.
We present a Bayesian nonparametric model for estimating tree insertion grammars (TIG), building upon recent work in Bayesian inference of tree substitution grammars (TSG) via Dirichlet processes. Under our general variant of TIG, grammars are estimated via the Metropolis-Hastings algorithm that uses a context free grammar transformation as a proposal, which allows for cubic-time string parsing as well as tree-wide joint sampling of derivations in the spirit of Cohn and Blunsom (2010). We use the Penn treebank for our experiments and find that our proposal Bayesian TIG model not only has competitive parsing performance but also finds compact yet linguistically rich TIG representations of the data. 1
The volume Semantic Processing of Legal Texts contains a total of thirteen papers that share the common theme of processing legal documents. One undisputable merit of the book is that of being the first collection to focus specifically on computational linguistic aspects of this task. Otherwise the papers in the collection represent a variety of topics as distinct as ontology engineering, multi-label classification, and translation quality assurance. They deal with theoretical foundations as well as commercial applications, and the authors’ affiliations range from universities to industry. The book is based on selected papers presented at the first workshop on Semantic Processing of Legal Texts (held at LREC 2008 in Marrakech) but comprises further, invited contributions.
Non-verbal communication enables efficient transfer of information among people. In this context, classic orchestras are a remarkable instance of interaction and communication aimed at a common aesthetic goal: musicians train for years in order to acquire and share a non-linguistic framework for sensorimotor communication. To this end, we recorded violinists' and conductors' movement kinematics during execution of Mozart pieces, searching for causal relationships among musicians by using the Granger Causality method (GC). We show that the increase of conductor-to-musicians influence, together with the reduction of musician-to-musician coordination (an index of successful leadership) goes in parallel with quality of execution, as assessed by musical experts' judgments. Rigorous quantification of sensorimotor communication efficacy has always been complicated and affected by rather vague qualitative methodologies. Here we propose that the analysis of motor behavior provides a potentiall)
Humans make systematic errors in the 3D interpretation of the optic flow in both passive and active vision. These systematic distortions can be predicted by a biologically-inspired model which disregards self-motion information resulting from head movements (Caudek, Fantoni, & Domini 2011). Here, we tested two predictions of this model: (1) A plane that is stationary in an earth-fixed reference frame will be perceived as changing its slant if the movement of the observer's head causes a variation of the optic flow; (2) a surface that rotates in an earth-fixed reference frame will be perceived to be stationary, if the surface rotation is appropriately yoked to the head movement so as to generate a variation of the surface slant but not of the optic flow. Both predictions were corroborated by two experiments in which observers judged the perceived slant of a random-dot planar surface during egomotion. We found qualitatively similar biases for monocular and binocular viewing of the simul)
Amazon’s Mechanical Turk is an online labor market where requesters post jobs and workers choose which jobs to do for pay. The central purpose of this article is to demonstrate how to use this Web site for conducting behavioral research and to lower the barrier to entry for researchers who could benefit from this platform. We describe general techniques that apply to a variety of types of research and experiments across disciplines. We begin by discussing some of the advantages of doing experiments on Mechanical Turk, such as easy access to a large, stable, and diverse subject pool, the low cost of doing experiments, and faster iteration between developing theory and executing experiments. While other methods of conducting behavioral research may be comparable to or even better than Mechanical Turk on one or more of the axes outlined above, we will show that when taken as a whole Mechanical Turk can be a useful tool for many researchers. We will discuss how the behavior of workers compares with that of experts and laboratory subjects. Then we will illustrate the mechanics of putting a task on Mechanical Turk, including recruiting subjects, executing the task, and reviewing the work that was submitted. We also provide solutions to common problems that a researcher might face when executing their research on this platform, including techniques for conducting synchronous experiments, methods for ensuring high-quality work, how to keep data private, and how to maintain code security.
Quantitative linguistics (QL) is a discipline of linguistics, that, using real texts, studies languages with quantitative mathematical approaches, aiming to precisely describe and explain, with a system of mathematical laws, the operation and development of language systems. Later in this review, we will address the relationship between QL and computational linguistics. Quantitative Syntax Analysis is a recent work on QL by Reinhard Köhler that not only provides a comprehensive introduction to the work of QL on the syntactic level, but also sketches the theoretical grounds, the research paradigm, and the ultimate goals of quantitative linguistics in general.In the first chapter, Köhler points to the vital role of syntax in language: Syntax enables language users to code structures instead of ideas as wholes. A text embodies a complex cognitive formation and meets several basic requirements in human communication, which implies that language is not autonomous, but a dynamic communicative system used by human beings. Hence, the ultimate understanding and explanation of syntax (and the whole language system) depends on usage-based investigation of the cognitive basis and the functional requirements of language, which is somewhat neglected in many mainstream syntactic studies. Even in those cases where explanatory power is acknowledged as the ultimate goal of linguistic investigation, the necessary knowledge is still required as to what a scientific linguistic theory is and how such a theory may be built. So far, it is rare for quantitative means to be used in syntactic study. One reason is that many syntacticians are too addicted to the enshrined traditional paradigms that have been proven to be somewhat inadequate when it comes to processing real texts. And that is why in computational linguistics, “devout executors of the belief in strictly formal methods as opposed to statistical ones do not have any chance to succeed” (page 4).The second chapter, entitled “The Quantitative Analysis of Language and Text,” begins with an explanation of the difference between quantitative linguistics and the formal branches of linguistics that have once been widely used in computational linguistics: QL is concerned with the quantitative properties important for understanding the development and the operation of linguistic systems, whereas the formal branches of linguistics use only qualitative mathematical means and formal logics to model structural properties of language, overlooking, in most cases, the aspects of systems that exceed structure, viz., functions, dynamics, and processes. Köhler points out that the successes of modern natural sciences (the exact, testable statements, the precise predictions, and the copious applications) all derive from their instruments and their advanced models. This implies that these instruments and models, for which the quantitative parts of mathematics (probability theory and statistics, function theory, differential equations) are indispensable ingredients, are worth integrating into linguistics, which is the aim of QL.Chapter 3, entitled “Empirical Analysis and Mathematical Modeling,” reviews the important works of quantitative syntactic analysis. In the first section, Köhler gives a long list of the important syntactic units and properties defined within the frameworks of both phrase structure syntax and dependency syntax; this reflects the fact that researchers in both fields have been engaged in some fruitful quantitative studies. In Section 3.2, he defines quantitation of syntactic concepts as counting the objects under study, because syntactic analysis investigates only discrete objects. Section 3.4 is a detailed review of the important works on various syntactic phenomena within the frameworks of both phrase structure syntax and dependency syntax, including sentence length, probabilistic grammars and probabilistic parsing, Markov chains, Frumkina’s law on the syntactic level, distribution of dependency distance, and distribution of dependency types, and so on. These quantitative models, which have been empirically corroborated with real texts, or sometimes treebanks and dictionaries (of various languages), can be linguistically, cognitively, or functionally interpreted—a rare achievement in the past statistical investigations of language. Apart from the models concerning probabilistic grammars and Markov chains, which have already been widely used in computational linguistics, there are some other works that may also have practical applications in various fields. For example, the mathematical model of sentence length, which describes the probability of neighboring length classes as a function of the probability of the first of the two given classes, may contribute to practical applications such as text classification and the measurement of text comprehensibility, and so forth. The frequency studies of word and syntactic constructions have obtained many results useful for language teaching, the construction of parsing algorithms, and the estimation of effort of (automatic) rule learning, and more. The syntactic studies on Frumkina’s law, which is concerned with the number of text blocks with x occurrences of a given syntactic element or category, may benefit certain types of computational text processing if specific constructions or categories can be differentiated and found automatically by their particular distributions. One advantage of QL is that all its findings are mathematically formulated and linguistically interpreted, which at least makes it possible to be used in constructing models necessary for computational linguistics.In Chapter 4, Köhler introduces his efforts to build a real “linguistic theory,” for he believes that “there is not yet any elaborated linguistic theory in the sense of the philosophy of science” (page 21). Building such a theory begins with “plausible hypotheses,” which may become laws when sufficiently attested and may then be further integrated into a coherent system. This is the process of setting up a scientific theory, as succinctly summarized in the title of this chapter: “Hypotheses, Laws, and Theories.”The first section of this chapter shows the first step toward a scientific linguistic theory, the process in which “plausible hypotheses” are deduced, interpreted, and empirically attested before finally becoming laws. In Section 2, Köhler introduces the foundation of his synergetic linguistics, which views language as a dynamic, self-organizing, and self-regulating system where the so-called enslaving principle and order parameters are the crucial elements. On this basis, the author builds a synergetic syntactic model in Section 4.2.7 with certain modeling principles. Within the framework of phrase-structure syntax, eight properties of syntactic constructions and four inventories are chosen to build this model, which are linked together by laws resulting from the verified hypotheses and subject to the regulation of some order parameters.Quantitative linguistics, which is an unfamiliar field of study for many linguists, depends heavily on real texts and mathematical tools. Therefore we believe it is worthwhile to briefly clarify the differences and the relations between QL and corpus linguistics, on one hand, and between QL and computational linguistics, on the other hand. In comparison with QL, corpus linguistics is in fact more of a research methodology rather than an independent linguistic discipline, reflecting a shift of focus from competence to performance, from introspection to empirical study. This makes both the common ground and the difference between corpus linguistics and QL, which aims to quantitatively and mathematically explore, on the basis of real texts and treebanks, the fundamental laws governing the structure and evolution of language, and integrate them into a systematic theory capable of explanation and prediction.Computational linguistics (CL) is an interdisciplinary field investigating the structure of natural languages from a formal, mathematical, and computational point of view. Compared with QL, CL seems to be more interested in research that can have direct applications in such fields as language understanding, language generation, machine translation, and so forth, than in the explanation of language structure, operation, and evolution. Traditionally, the mathematical models used in CL are derived from the linguistic theories via the process of formalization. Due to the limitations of the qualitative mathematical models, however, the statistical models, most of which so far fail linguistic interpretations, have recently become dominant in the field of computational linguistics. That is perhaps why Shuly Wintner (2009, page 641), in a “Last Words” article in this journal, called for “the return of linguistics to computational linguistics,” implying that the advances in the field of CL may ultimately lie in the advances in the understanding of language itself.Wintner’s appeal does reflect the present situation of computational linguistics: a discipline in which many works are heavily oriented towards engineering and weakly grounded in linguistics. The formal linguistic models seem now outshone by the purely statistical paradigms, as the result of their inadequacy in processing real-world languages. Of course, this means no renouncement of the value of the traditional formal models. But it is obvious that considerable updating and enrichment are necessary for these formal models if they are to play significant roles in the future. In this regard, QL, which has a solid linguistic foundation, may help by providing quantitative cues (which are linguistically interpretable) to improve the performance in NLP, as has been illustrated in the case of probabilistic grammar that ingeniously integrates quantitative, statistical devices into qualitative linguistic models. QL has provided some useful models for CL. And it is reasonable to believe that it will continue to do so in the future, though some of its achievements seem now not ready to be directly used in CL.But perhaps QL can do more than that. Though QL has not yet drawn much attention from computational linguistics, it is potentially a proper answer to Wintner’s call for the return of linguistics to computational linguistics. CL aims to replicate in computers the patterns of human language behavior, whereas QL endeavors to mathematically and quantitatively reveal the laws and the principles that govern human language behavior—this is a relation between theory and practice. The success of a scientific discipline is usually based on precise models that are well-grounded in profound understanding of the object of study. Computational linguistics is no exception. Two features of QL are hence noteworthy. One is that it aims to explain, within a certain linguistic framework, the operation and the evolution of language by uncovering the systematic cognitive and functional regulations that underlie human languages. The other is that it tries to mathematically model, with systematic and precise quantitative laws, these regulations and the resulting mechanism of language. In view of these two features, we believe that the success of QL will somehow and somewhat boost the studies in the field of CL and that the communication between QL and CL is and will be not only possible but also mutually beneficial. This is why we hold that it is worthwhile to recommend this book to researchers in the field of computational linguistics, a book presenting a panorama of QL in general and discoveries on the syntactic level in particular.
We present FinnTreeBank 3, a treebank and parsebank for Finnish, and focus on its design and development. First, we describe our method of specifying the core linguistic representation with descriptive grammars (rather than with text corpus samples) and conclude with a reference to empirical experiments on the specifiability of Finnish dependency syntax. We outline the linguistic representation used in describing Finnish morphology and syntax. We describe one use of a systematically specified linguistic representation and “grammar definition corpus”: as a task specification for a subcontractor to create a parser engine and a parsebank for the language resource service.
Abstract This study examines the role of the voseo, tuteo, and ustedeo in written advertising in Montevideo, Uruguay. Analysis of 133 samples shows a preference for voseo in publicity, although tú and usted are also present. While written voseo - with or without a tonic pronoun - has traditionally been stigmatized in public education (Bertolotti & Coll 2003 and Gabbiani 2000), its predominant role in advertising suggests an increased acceptance of that written form in public venues. The study examines the language of publicity within the framework of politeness strategies, which explains that written publicity uses linguistic forms that lessen social distance and power [-D, -P] for the purpose of persuading consumers to buy. As such, voseo is the preferred form of address in commercial advertising in Montevideo, appearing in 83% of the samples. Voseo also has a strong presence in non-commercial advertising, where it appears in 44% of the examples in this study, irrespective of any stigma attached to its use.
The paper presents an integral framework for multilingual lexical databases (henceforth MLLD) based on Compreno technology. It differs from the existing approaches to MLLD in the following aspects: 1) it is based on a universal semantic hierarchy (SH) of thesaurus type filled with language-specific lexicon; 2) the position in the SH generally determines semantic and syntactic model of a word; 3) this model proposes a suite of elaborate tools to determine universal and language-specific semantic and syntactic properties and deals efficiently with problems of cross-lingual lexical, semantic and syntactic asymmetry. Currently, it includes English, Russian, German, French and Chinese and proves to be a compatible MLLD for typologically different languages that can be used as a comprehensive lexical-semantic database for various NLP applications.
Finding coordinations provides useful infor-mation for many NLP endeavors. However, the task has not received much attention in the literature. A major reason for that is that the annotation of major treebanks does not re-liably annotate coordination. This makes it virtually impossible to detect coordinations in which two conjuncts are separated by punctu-ation rather than by a coordinating conjunc-tion. In this paper, we present an annotation scheme for the Penn Treebank which intro-duces a distinction between coordinating from non-coordinating punctuation. We discuss the general annotation guidelines as well as prob-lematic cases. Eventually, we show that this additional annotation allows the retrieval of a considerable number of coordinate structures beyond the ones having a coordinating con-junction.
We describe our method of traditional Phrase Structure Grammar (PSG) parsing in CIPS-Bakeoff2012 Task3. First, bagging is proposed to enhance the baseline performance of PSG parsing. Then we suggest exploiting another TreeBank (CTB7.0) to improve the performance further. Experimental results on the development data set demonstrate that bagging can boost the baseline F1 score from 81.33 % to 84.41%. After exploiting the data of CTB7.0, the F1 score reaches 85.03%. Our final results on the official test data set show that the baseline closed system using bagging gets the F1 score of 80.17%. It outperforms the best closed system by nearly 4 % which uses a single model. After exploiting the CTB7.0 data, the F1 score reaches 81.16%, demonstrating further increases of about 1%. 1
Minimal hepatic encephalopathy (MHE) encompasses a number of neuropsychological and neurophysiological disorders in patients suffering from liver cirrhosis, who do not display abnormalities during a medical interview or physical examination. A negative influence of MHE on the quality of life of patients suffering from liver cirrhosis was confirmed, which include retardation of ability of operating motor vehicles and disruption of multiple health-related areas, as well as functioning in the society. The data on frequency of traffic offences and accidents amongst patients diagnosed with MHE in comparison to patients diagnosed with liver cirrhosis without MHE, as well as healthy persons is alarming. Those patients are unaware of their disorder and retardation of their ability to operate vehicles, therefore it is of utmost importance to define this group. The term minimal hepatic encephalopathy (formerly "subclinical" encephalopathy) erroneously suggested the unnecessity of diagnostic and therapeutic procedures in patients with liver cirrhosis. Diagnosing MHE is an important predictive factor for occurrence of overt encephalopathy - more than 50% of patients with this diagnosis develop overt encephalopathy during a period of 30 months after. Early diagnosing MHE gives a chance to implement proper treatment which can be a prevention of overt encephalopathy. Due to continuing lack of clinical research there exist no commonly agreed-upon standards for definition, diagnostics, classification and treatment of hepatic encephalopathy. This article introduces the newest findings regarding the importance of MHE, scientific recommendations and provides detailed descriptions of the most valuable diagnostic methods.
International audience
We present a number of experiments on parsing the Ancient Greek Dependency Treebank (AGDT), i.e. the largest syntactically annotated corpus of Ancient Greek currently available (350k words ca). Although the AGDT is rather unbalanced and far from being representative of all genres and periods of Ancient Greek, no attempt has been made so far to perform automatic dependency parsing of Ancient Greek texts. By testing and evaluating one probabilistic dependency parser (MaltParser), we focus on how to improve the parsing accuracy and how to customize a feature model that fits the distinctive properties of Ancient Greek syntax. Also, we prove the impact of genre and author diversity on parsing performances.
Sanskrit since many thousands of years has been the oriental language of India. It is the base for most of the Indian Languages. Ambiguity is inherent in the Natural Language sentences. Here, one word can be used in multiple senses. Morphology process takes word in isolation and fails to disambiguate correct sense of a word. Part-Of-Speech Tagging (POST) takes word sequences in to consideration to resolve the correct sense of a word present in the given sentence. Efficient POST have been developed for processing of English, Japanese, and Chinese languages but it is lacking for Indian languages. In this paper our work present simple rule-based POST for Sanskrit language. It uses rule based approach to tag each word of the sentence. These rules are stored in the database. It parses the given Sanskrit sentence and assigns suitable tag to each word automatically. We have tested this approach for 15 tags and 100 words of the language this rule based tagger gives correct tags for all the inflected words in the given sentence.
We describe a shared task on parsing web text from the Google Web Treebank. Participants were to build a single parsing system that is robust to domain changes and can handle noisy text that is commonly encountered on the web. There was a constituency and a dependency parsing track and 11 sites submitted a total of 20 systems. System combination approaches achieved the best results, however, falling short of newswire accuracies by a large margin. The best accuracies were in the 80-84% range for F1 and LAS; even part-ofspeech accuracies were just above 90%.
We have created a bilingual treebank for 99% of the sentences in the FraCaS test suite. The treebank is built together with an associated bilingual English-Swedish lexicon written in the Grammatical Framework Resource Grammar. The original FraCaS sentences are English, and we have tested the multilinguality of the Resource Grammar by analysing the grammaticality and naturalness of the Swedish translations. 86% of the sentences are grammatically and semantically correct and sound natural. About 10% can probably be fixed by adding new lexical items or grammatical rules, and only a small amount are considered to be difficult to cure.\n
To facilitate future research in unsupervised induction of syntactic structure and to standardize best-practices, we propose a tagset that consists of twelve universal part-of-speech categories. In addition to the tagset, we develop a mapping from 25 different treebank tagsets to this universal set. As a result, when combined with the original treebank data, this universal tagset and mapping produce a dataset consisting of common parts-of-speech for 22 different languages. We highlight the use of this resource via two experiments, including one that reports competitive accuracies for unsupervised grammar induction without gold standard part-of-speech tags.
We present two approaches (rule-based and statistical) for automatically annotating intra-chunk dependencies in Hindi. The intra-chunk dependencies are added to the dependency trees for Hindi which are already annotated with inter-chunk dependencies. Thus, the intra-chunk annotator finally provides a fully parsed dependency tree for a Hindi sentence. In this paper, we first describe the guidelines for marking intra-chunk dependency relations. Although the guidelines are for Hindi, they can easily be extended to other Indian languages. These guidelines are used for framing the rules in the rule-based approach. For the statistical approach, we use MaltParser, a data driven parser. A part of the ICON 2010 tools contest data for Hindi is used for training and testing the MaltParser. The same set is used for testing the rule-based approach.
International audience
A method is presented for transferring dependency treebanks between similar languages by using a bilingual lexicon, aiming to improve dependency parsing accuracy on the target language. It is illustrated by transferring the Slovene Dependency Treebank to Croatian by using a GIZA++ bilingual lexicon constructed from the Croatian-Slovene 1984 parallel corpus from the Multext East project. The transferred treebank is merged with the Croatian Dependency Treebank and the merged treebank is used to train and test two graph-based dependency parsers. MSTParser and CroDep accuracy on parsing the 1984 fictional text shows a statistically significant increase and a similar decrease on parsing the Croatian Dependency Treebank newspaper text.
A treebank may contain the annotation of different phenomena such as word order, morphological features, syntactic and semantic relations, etc., which are rather different in their nature. Quite often, the annotation of these phenomena is combined in a single structure, which leads to low-quality training results and is verifiably deficient from a theoretical (linguistic) perspective. We argue that the annotation of corpora requires a well-defined linguistic model which supports multi-level annotation, with one type of phenomenon per level. Our experience with dependency treebanks created or adjusted for surface-oriented natural language generation and based on the Meaning-Text Theory, a multi-level linguistic model, supports this argumentation.
There is evidence that women may be less successful when attempting to quit smoking than men. One potential contributory cause of this gender difference is differential craving and stress reactivity to smoking- and negative affect/stress-related cues. The present human laboratory study investigated the effects of gender on reactivity to smoking and negative affect/stress cues by exposing nicotine dependent women (n = 37) and men (n = 53) smokers to two active cue types, each with an associated control cue: (1) in vivo smoking cues and in vivo neutral control cues, and (2) imagery-based negative affect/stress script and a neutral/relaxing control script. Both before and after each cue/script, participants provided subjective reports of smoking-related craving and affective reactions. Heart rate (HR) and skin conductance (SC) responses were also measured. Results indicated that participants reported greater craving and SC in response to smoking versus neutral cues and greater subjective stress in response to the negative affect/stress versus neutral/relaxing script. With respect to gender differences, women evidenced greater craving, stress and arousal ratings and lower valence ratings (greater negative emotion) in response to the negative affect/stressful script. While there were no gender differences in responses to smoking cues, women trended towards higher arousal ratings. Implications of the findings for treatment and tobacco-related morbidity and mortality are discussed.