Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
The list of entries of the Pralex lexical database (LDB) was at first built according to the Frequency Dictionary of Czech, later enriched by other entries from the Dictionary of the Standard Czech Language and continually enriched from the corpus SYN, too (so far optionally). To make the database an inventory of lexical units of contemporary Czech, as well as a base for a new middle-sized monolingual dictionary or other special dictionaries, we would like to expand it by choosing relevant lexical items from both newer reference resources (dictionaries) and material resources, e.g. other corpuses of written texts made by the Institute of the Czech National Corpus, various other databases (e.g. text archive Newton) or our own databases (excerption database of neologisms and special excerption database of various branches).
Cet article presente un corpus parallele francais-allemand de plus de 4 millions de mots issu de la numerisation d’un corpus alpin multilingue. Ce corpus est une precieuse ressource pour de nombreuses etudes de linguistique comparee et du patrimoine culturel ainsi que pour le developpement d’un systeme statistique de traduction automatique dans un domaine specifique. Nous avons annote un echantillon de ce corpus parallele et aligne les structures arborees au niveau des mots, des constituants et des phrases. Cet “alpine treebank” est le premier corpus arbore parallele francais-allemand de haute qualite (manuellement controle), de libre acces et dans un domaine et un genre nouveau: le recit d’alpinisme.
This article reports on a practical, semi-automated procedure towards creating a clean, morphologically annotated Zulu corpus of tractable size that could eventually serve both as a gold standard for Zulu computational morphology and as basis for further linguistic annotation. A corpus development architecture is proposed which includes the corpus in various stages of development, a pre-processing module, the Zulu morphological analyser and its guesser variant, the machine-readable lexicon that serves as comprehensive lexical database for Zulu, and a human elicitation function for ensuring the integrity of the lexical database. The approach is novel in the sense that an existing rule-based, finitestate Zulu computational morphological analyser is used as a core technology in this procedure to facilitate the complex, agglutinative nature of Zulu morphology. The corpus, at present consisting of the Zulu version of the South African Constitution, will have morphological analysis and tagging as a first level of annotation.
The Copenhagen Dependency Treebanks (CDT) are a set of parallel treebanks for Danish, English, German, Italian, and Spanish. One of the main objectives of the CDT is to arrive at a unified description and annotation system for syntax, morphology, discourse, and anaphora. The treebanks are currently in the process of being annotated for these levels in all five languages. After a brief discussion of the subdivisions of the so-called bridging anaphors proposed by different scholars, we describe the classification and terminology adopted in the CDT. The main distinction here is the very common one between coreferential and associative anaphors, special attention being given to the latter group. Resumptive and evolving anaphors are treated as special subgroups of the coreferential anaphors. A list of the associative relations proposed by the CDT with authentic examples concludes the paper.
Reconstructing the History of Marriage and Residence Strategies in Indo-European—Speaking Societies Laura Fortunato Keywords Indo-European, Cultural Phylogenetics, Marriage, Monogamy, Polygyny, Affinal Terminology, Residence, Neolocality, Uxorilocality, Virilocality This file provides additional information on the data and methods used in Fortunato (2011a,b), and discussion of the results of the fossilization of nodes Proto-Indo-Hittite (PIH) and Proto-Indo-European (PIE) for marriage and residence strategies. Data and Methods Below I provide details on the criteria used to collate the cross-cultural sample, with the cross-cultural data in table form, and information on the procedure used by Pagel et al. (2007) to infer the posterior probability distribution of trees on which I mapped the cross-cultural data. Finally, I provide a detailed description of the method used for the comparative analyses. Cross-Cultural Data. Variable identifiers in this section follow Gray's (1999) Ethnographic Atlas (EA) codebook. I collated the cross-cultural sample by matching societies scored as speaking Indo-European (IE) languages (based on EA variable 98) with speech varieties in Dyen et al.'s (1992) IE basic vocabulary database, where needed using information from additional ethnographic and linguistic sources (e.g., Gordon 2005; Levinson 1991-1996; Price 1989; Ruhlen 1991). I also checked for correspondence between speech varieties in the linguistic database and the 62 societies in the EA with linguistic affiliation unknown and located in East Eurasia or in the Circum-Mediterranean region (based on EA variable 91). In some cases, more than one speech variety in the linguistic database could be matched with the same society in the EA. For example, Dyen et al. (1992) include five entries for Greek: three for dialectal forms (Greek D, Greek K, Greek ML), one for modern Greek (Greek Mod), and one for modern spoken [End Page 129] Greek (Greek MD), the latter compiled from dictionary data. In these cases, where available I selected the variety derived from dictionary data, which is likely to be less specific than other entries; alternatively, I selected the variety with data for the greatest number of meanings, or the first variety listed in Dyen et al. (1992, pp. 99-101). The phylogenetic tree model used to represent how societies are related captures the process of diversification of taxa from a common ancestor; therefore, I included in the sample only societies located in Eurasia, corresponding to the geographic range of IE languages before 1492 CE (Diamond and Bellwood 2003). I excluded the Icelanders because the EA description for this society refers to 1100 CE, while the descriptions for the 27 societies included in the sample refer to the "ethnographic present," with dates ranging from 1880 to 1960 CE, and median 1945 CE (Murdock 1967). Table 1 includes the recoded data on marriage strategy and residence strategy (prevailing and alternative modes) for the 27 societies. Tree Sample. Pagel et al. (2007) inferred the posterior probability distribution of trees from Dyen et al.'s (1992) IE basic vocabulary database, using the Bayesian Markov chain Monte Carlo (MCMC) phylogenetic tree-building method developed by Pagel and Meade (2004). The linguistic database includes word forms and cognacy judgments for 95 modern IE speech varieties (languages, dialects, and creoles) across the Swadesh 200-word list of items of basic vocabulary; two or more word forms are cognate if they share a common origin. Swadesh lists consist of cross-culturally universal items of vocabulary such as pronouns, body parts, and numerals, which are less prone to innovation and borrowing (i.e., horizontal transmission) than other meanings (Swadesh 1952). The tree-building analysis was performed on a data matrix obtained from the linguistic database as follows. First, Pagel et al. (2007) excluded eleven speech varieties suspected of methodological bias by Dyen et al. (1992) and added data for three extinct varieties (Hittite, Tocharian A, Tocharian B) to be used as "outgroup" taxa. Outgroups provide information on the direction of change in the data by virtue of being distantly related to the groups under investigation, the "ingroup" taxa; they are used in tree-building for determining ancestor-descendant relationships (Felsenstein 2004, p. 6). As discussed in Fortunato (2011a), Hittite belongs to the extinct sister-group to the IE languages, the...
“Handling emotions in human–computer dialogues”, written by Pittermann, Pittermann and Minker, is a complete and interesting book about affective computing in spoken dialogue systems. Dialogue systems are an integrated part of our daily life. They generally mean simplicity, time saving and safety. For example, when driving, hand-free operations are necessary and therefore the possibility of giving commands through speech is a necessity. However, to implement more flexible dialogue systems, it is important that these systems can adapt to the speaker. In particular, it seems very important that dialogue systems should be able to recognize and cope with our emotions. To human, the emotions recognition process occurs in an automatic, unconscious, and effortless fashion. Emotions explicitly affect our autonomic nervous system (e.g., cardiovascular and skin conductance changes) and our somatic nervous system (motor expression in face, voice and body). We usually don’t have many...
Background: Diversity patterns of livestock species are informative to the history of agriculture and indicate uniqueness of breeds as relevant for conservation. So far, most studies on cattle have focused on mitochondrial and autosomal DNA variation. Previous studies of Y-chromosomal variation, with limited breed panels, identified two Bos taurus (taurine) haplogroups (Y1 and Y2; both composed of several haplotypes) and one Bos indicus (indicine/zebu) haplogroup (Y3), as well as a strong phylogeographic structuring of paternal lineages. Methodology and Principal Findings: Haplogroup data were collected for 2087 animals from 138 breeds. For 111 breeds, these were resolved further by genotyping microsatellites INRA189 (10 alleles) and BM861 (2 alleles). European cattle carry exclusively taurine haplotypes, with the zebu Y-chromosomes having appreciable frequencies in Southwest Asian populations. Y1 is predominant in northern and north-western Europe, but is also observed in several Ibe)
Humans reached present-day Island Southeast Asia (ISEA) in one of the first major human migrations out of Africa. Population movements in the millennia following this initial settlement are thought to have greatly influenced the genetic makeup of current inhabitants, yet the extent attributed to different events is not clear. Recent studies suggest that south-to-north gene flow largely influenced present-day patterns of genetic variation in Southeast Asian populations and that late Pleistocene and early Holocene migrations from Southeast Asia are responsible for a substantial proportion of ISEA ancestry. Archaeological and linguistic evidence suggests that the ancestors of present-day inhabitants came mainly from north-to-south migrations from Taiwan and throughout ISEA approximately 4,000 years ago. We report a large-scale genetic analysis of human variation in the Iban population from the Malaysian state of Sarawak in northwestern Borneo, located in the center of ISEA. Genome-wide s)
Understanding the role that social cues have on interpersonal choice, and their susceptibility to contextual effects, is of core importance to models of social decision-making. Language, on the other hand, is one of the main means of communication during social interactions in our culture. The present experiments tested whether positive and negative linguistic descriptions of alleged partners in a modified Ultimatum Game biased decisions made to the same set of offers, and whether the contextual uncertainty of the game modulated this biasing effect. The results showed that in an uncertain context, the same offers were accepted with higher probability when they were preceded by positive rather than by negative valenced trait-words. Participants also accepted fair offers with higher probability than unfair offers, but this effect did not interact with the valence of the social descriptive words. In addition, the speed of the decision was affected by valence: acceptance choices were fast)
Personal naming practices exist in all human groups and are far from random. Rather, they continue to reflect social norms and ethno-cultural customs that have developed over generations. As a consequence, contemporary name frequency distributions retain distinct geographic, social and ethno-cultural patterning that can be exploited to understand population structure in human biology, public health and social science. Previous attempts to detect and delineate such structure in large populations have entailed extensive empirical analysis of naming conventions in different parts of the world without seeking any general or automated methods of population classification by ethno-cultural origin. Here we show how 'naming networks', constructed from forename-surname pairs of a large sample of the contemporary human population in 17 countries, provide a valuable representation of cultural, ethnic and linguistic population structure around the world. This innovative approach enriches and adds)
Reviewed by: From poets to padonki: Linguistic authority and norm negotiation in modern Russian culture Anastassia Zabrodskaja Ingunn Lunde and Martin Paulsen, eds. From poets to padonki: Linguistic authority and norm negotiation in modern Russian culture. Bergen: University of Bergen, 2009. [Slavica Bergensia, 9.] The issues discussed in From poets to padonki became a part of my life in 1999 when I began my university studies and using Russian and Estonian became my everyday reality. Choosing different languages for different purposes, code-switching, and multilingual language play have all been part of my daily language use for the last eleven years. This collection challenges definitions of what can be meant by "language", "standard" language, and the linguistic "norm". It acquaints readers with a wide range of linguistic phenomena in modern Russian culture. Padonki refers to a subculture within the Russian-speaking Internet (Runet), whose representatives use erratic spellings for words, aiming at creating a comic effect. Their nickname padonki itself illustrates such a trend-it is an alteration of podonki 'dregs'. Padonki is also characterized by gratuitous use of profanity and a penchant for obscene subjects. During the last decade, a body of literature has emerged proposing that (socio)linguists direct their attention away from the traditional focus of linguistics, i.e., language as a bounded system, towards broader semiotic resources, to see what is really going on when people use "language" (Stroud 2003, Jacquemet 2005, Shohamy 2006, Makoni and Pennycook 2007, Blommaert 2010). The notion of "language" becomes especially questionable in cases of (multilingual) computer-mediated communication. The last decade has also witnessed rising scholarly interest in language on and of the Internet in general and in e-mails and postings on Internet discussion forums or message boards in particular (e.g., Koutsogiannis and Mitsikopoulou 2003, Palfreyman and al Khalil 2003, Hinrichs 2006, Dorleijn and Nortier 2009, Androutsopoulos 2006, 2009). García (2009: 32) instead of "language" offers a more suitable term for the multiple discursive practices-languaging, i.e., "social practices that are actions performed by our meaning-making selves." For her, dialects, pidgins, creoles, and academic language [End Page 153] are all examples of languaging, as there are differences between language practices at home, in communities, and in academic contexts. I would argue that standard languages are idealized constructs, and none can remain unaffected by language contacts during its entire history. While studying language use by individuals, it is important to shift "from focus on structure to focus on function-from focus on linguistic form in isolation to linguistic form in human context" (Hymes 1974: 77). The volume under review offers fascinating reading for (socio)linguists, who work with different manifestations of real language practices rather than seeking for linguistic norm descriptions. Comprising 17 contributions written in English and Russian, the book summarizes analyses of the norm in modern Russian language culture based on data from multiple sources: "literary fiction, Internet slang, literary criticism and aesthetics, writers' blogs, linguistic play, and various arenas for 'talk about talk,' such as the classroom, blogs, the media, the courtroom, etc." (11). The book opens with Ingunn Lunde and Martin Paulsen's excellent general introduction, an insightful synthesis of how various approaches to the standard Russian language, its norms, and linguistic standards see the history, development, and future of the Russian language culture. The first paper, "Living norms" by Henning Andersen, gives a comprehensive overview of the history and development of the notion of language norms. Andersen reviews historical contributions to the understanding of norms. Making a distinction between declarative and deontic norms and dividing them further into explicit and implicit norms, he highlights the leading ideas in the field of (socio)linguistic studies that concern the position of norms. On the example of excerpts from Soviet grammars and dictionaries, he illustrates how norms can be governed from above by agencies of the state. "The standard norms include only prescribed and permitted forms", and all other forms must be avoided (25). Andersen also comments on spoken language standards, noting that "the [Russian] language continues to be spoken effectively in its numerous variants all over the inherited Russian language area as well as in the diasporas, old and new" (32). Martin Paulsen...
Music processing may be preserved in subjects with Alzheimer disease (AD). It is not known which neural substrates are engaged in music processing, and how music familiarity moderates the engagement of these substrates in AD. We investigated fMRI patterns of brain activation during listening to familiar and non-familiar classical music excerpts in subjects with mild to moderate AD and healthy age-matched controls. We related these patterns to behavioral data on familiarity ratings and musical abilities. Five subjects with AD (age M = 76.2, SD = 6.6, MMSE M = 18.6, SD = 7.7, range 9-26) and five healthy controls (age M72.8, SD = 8.0, MMSE M = 29.6, SD =.6, range 29-30) underwent fMRI with a block design paradigm consisting of 75s of familiar music excerpts followed by 30s white noise, vs. 75s of unfamiliar music excerpts. Participants were instructed to just listen. They received a battery of behavioral tests including repeated familiarity ratings of the presented music excerpts (1=very familiar to 5=very unfamiliar), the Montreal Battery for Evaluation of Amusia (MBEA), and the Seashore test of musical abilities. For fMRI a mixed-model 2x2 ANOVA was used to examine effect of group AD vs. controls) and stimulus (familiar vs. unfamiliar) using a p-value <.01 and a minimum cluster size of 200ÂμL. For behavioural data, independent- and paired-sample t-tests were used with p-value <.05. We found a significant group by stimulus effect in fMRI activation patterns. When familiar to unfamiliar music activations were compared, the following regions showed increases in AD and decreases in controls: left/right lingual gyrus, left inferior parietal, left superior, middle and inferior temporal gyri, left precuneus, left culmen, left/right striatum. Ratings for familiar and unfamiliar excerpts did not differ by group (AD M = 1.2, SD =.2 and M = 1.9, SD =.6; Controls M = 1.2, SD =.3, and M = 1.9, SD =.3). Performance on music ability tests also did not differ by group except for Seashore loudness and rhythm (p <.05). Subjects with AD appear to respond more intensely to non-familiar than familiar music by activating regions associated with recognizing familiar patterns and emotions. They do not differ from controls on behavioral measures. This finding suggests differential neural recruitment with respect to preserved music recognition, and its potential application in diagnosis and treatment.
This paper is to analyze curricular changes of Chongryon Korean schools in Japan. Chongryon Korean schools belong to the category of miscellaneous schools in the classification by the Ministry of Education in Japan. They neither need to follow curricular set by the Japanese government nor receive subsidies from it. They have their own curricula and textbooks. Historically, there have been 6 curricular reforms in Chongryon Korean schools. After characterizing those reforms, this paper compares the ``Korean`` textbooks of the 1993 and 2003 reforms. In general, the political color of advocating socialism and the Juche idea and anti-American and anti-Seoul propoganda gets thinner in new textbooks. The 2003 Korean textbooks emphasize speaking practice and adapt dialogues rich with story-telling. This paper also examines linguistic norms followed in the Korean textbooks. (Osaka University of Economics and Law)
We describe a new interactive annotation scheme between a human annotator who carries out simplified annotations on CFG trees, and a statistical parser that converts the human annotations automatically into a richly annotated HPSG treebank. In order to check the proposed scheme's effectiveness, we performed automatic pseudo-annotations that emulate the system's idealized behavior and measured the performance of the parser trained on those annotations. In addition, we implemented a prototype system and conducted manual annotation experiments on a small test set.
The study is focused on how to make use of the lexical database Pralex for a processing of nominal entries, how to describe a particular specific phenomenon or a partial issue in an entry form as well as which nominal data are included (mandatory items). To clarify this, the authors follow the order of individual parts of an entry form. Special attention is paid to the issue of homonymy, variation, explanation of a meaning and division of polysemantic entries.
The reconstruction of standardized texts in the Prague Dependency Treebank of Spoken Czech enables the comparison of authentic spoken utterances and standardized texts The authors concentrate on the questions: What does the syntactic identity of the Czech spoken and written texts consist of? What syntactic constructions are „natural“ in the spoken and in the written text? What is the difference in the density of the cohesive links, in the explicit and implicite relations between units?
Treebanks are a necessary prerequisite for many NLP tasks, including, but not limited to, semantic role labeling. For many languages, however, treebanks are either nonexistent or too small to be useful. Time-critical applications may require rapid deployment of natural language software for a new critical language—much faster than the development time of a traditional treebank. This dissertation describes a method for generating a treebank and training syntactic and semantic models using only semantic training information—that is, no human-annotated syntactic training data whatsoever. This will greatly increase the speed of development of natural language tools for new critical languages in exchange for a modest drop in overall accuracy. Using Combinatory Categorial Grammar (CCG) in concert with Propbank semantic role annotations allows us to accurately predict lexical categories in combination with a partially hidden Markov model. By training the Berkeley parser on our generated syntactic data, we can achieve SRL performance of 65.5% without using a treebank, as opposed to 74% using the same feature set with gold-standard data.
This study has a dual focus in that it aims to develop a viable methodology for elicitation experiments in English linguistics, while simultaneously applying the proposed methods to investigate an actual subject, the distribtion of the additive particles 'also' and 'too'. Traditionally, data for linguistic research is gained by sampling natural language corpora. Although this approach is valid and, indeed, has been applied here, elicitation experiments can gain in validity and informative value by additionally introducing questionnaires to accompany corpus research. Online questionnaires particularly are a cost-effective and highly customizable tool to create a linguistic database against which existing data can be tested. For the purpose of this study, I have created six online questionnaires to test three hypotheses about the distribution of 'also' and 'too'. Two interdependent hypotheses assume that the use of the two particles is sensitive to structural properties of the `added constituent' while the third one, the information-structural hypothesis, argues that the use of 'also' and 'too' is controlled by the information structure of the sentence. In addition to the questionnaires, a balanced sample was extracted from the "British National Corpus" and tested against corpus data from previous studies as well as the data elicited online. In the course of this study, the additive particles will firstly be defined in terms of their structural properties, and the hypotheses about their use introduced and explicated. Furthermore, the data elicitation process will be detailed, as well as results from previous studies be taken into account. The hypotheses will subsequently be tested against the data from both corpus research and elicitation per questionnaires, and the outcome discussed. Concluding the study, I will focus on the results of the distribution analysis as well as evaluate the introduction of the online questionnaires and their application in the context of testing the hypotheses against empirical linguistic data.
In this paper, we expand Morzycki (2009)’s claims that degree readings of size adjectives are attributed to syntax. We introduce a corpus-based analysis in Dutch to verify and extend his claim into the semantic domain. Using the LASSY Treebank, we extract syntactic and semantic properties of noun phrases consisting of the adjectives “gigantisch”, “kolossaal”, and “reusachtig ” and manually annotate each adjective-noun pair with a gradable or nongradable label. Using these features, we construct a statistical model based on logistic regression and find that the grammatical role, definiteness, and particular semantic noun groups derived from Cornetto (a Dutch WordNet with referential relations) have a significant effect on the likelihood that an adjective-noun pair is interpreted by the reader to have a degree reading. 1.
This paper briefly depicts major achievements in the Chinese language information processing and roughly reviews the computational linguistic research for the recent 20 years in China.The author questions the current methodologies such as POS tagging and treebank for Chinese.The paper presents some new ideas about the construction of Chinese data resources.The authors suggest that for Chinese we should address deep and semantic annotation instead of current shallow and syntactic one.The future annotation will include targeted tackling,diversified content,stepwise procedure,and non-professional annotators.The paper predicts some eye-catching features: merge of technologies and human-centered computing.
This paper proposes a method for shallow parsing on the basis of CRF and transformation-based error-driven learning.The method is applied to Penn Chinese Treebank and gets a good performance of chunking identification.First,CRF model is used to identify chunks to acquire candidate transformation rules by error-driven learning.Then,an evaluation function is used to filter candidate transformation rules.And last,transformation rules are used to revise the chunking results of CRF.The experimental results show that this approach is effective,and outperforms the single CRF-based approach in shallow parsing.Precision,recall and F-values are improved respectively.
While some visual objects prompt strong affective responses (e.g., guns and ice cream), most objects are thought to be affectively neutral. Last year we reported evidence for the existence of “micro-valences” (Lebrecht & Tarr, VSS, 2010): that nominally neutral objects actually possess subtle valences that we hypothesize form an integral part of object perception. In the current experiment we used fMRI to investigate: a) the extent to which micro-valences are coded within the extended visual object recognition network (Bar, 2007); b) how micro-valences are neurally instantiated with respect to valence strength and direction. Using slow event-related fMRI, participants viewed an object picture for 500ms and evaluated the object's “pleasantness” on each trial. Participants were shown 120 everyday, nominally neutral objects (e.g., teapots and clocks) and 120 strongly valenced objects (e.g., gold and a skull). Objects were assigned to these conditions based on mean valence ratings acquired in a prior experiment with a different population of participants. Individualized ratings for all objects were also acquired for our fMRI participants during a post-scan session. Regions of interest for further analysis were identified using two independent localizers: a) objects versus scrambled objects; b) strongly valenced objects versus minimally valenced objects (e.g., paperclips). Two results stand out. First, somewhat consistent with previous findings, lateral regions of PFC and regions of medial OFC are selective to a positive versus negative comparison for strongly valenced objects. Second, and intriguingly, almost all participants show selectivity for micro-valence objects, comparing positive to negative, in a region adjacent to the region for strongly valenced objects. We posit that intrinsic to visual object perception, object valence – for all objects – is evaluated in PFC. This valence metric forms one of many associated object properties that can influence subsequent perceptual and non-perceptual object-related processing.
Background: The well-established left hemisphere specialisation for language processing has long been claimed to be based on a low-level auditory specialization for specific acoustic features in speech, particularly regarding 'rapid temporal processing'. Methodology: A novel analysis/synthesis technique was used to construct a variety of sounds based on simple sentences which could be manipulated in spectro-temporal complexity, and whether they were intelligible or not. All sounds consisted of two noise-excited spectral prominences (based on the lower two formants in the original speech) which could be static or varying in frequency and/or amplitude independently. Dynamically varying both acoustic features based on the same sentence led to intelligible speech but when either or both acoustic features were static, the stimuli were not intelligible. Using the frequency dynamics from one sentence with the amplitude dynamics of another led to unintelligible sounds of comparable spectro-te)
We employ syntactic parsing to describe and to discover lexico-grammatical features of English regional varieties. In the absence of suitable Treebanks, automatically parsed corpora (tree jungles) can be used. As an example we focus on Indian English, using the International Corpus of English (ICE), and the British National Corpus (BNC). We use a largely corpus-driven method. There are few differences in frequencies of syntactic relations between the corpora, but considerable differences when taking the intricate relations between grammar and lexis into account. We describe differences in the use of zero articles, verb-preposition constructions, and ditransitive verbs. We show that relatively small corpora can be used to discover subtle lexico-grammatical differences.
Background: During sentence processing we decode the sequential combination of words, phrases or sentences according to previously learned rules. The computational mechanisms and neural correlates of these rules are still much debated. Other key issue is whether sentence processing solely relies on language-specific mechanisms or is it also governed by domain-general principles. Methodology/Principal Findings: In the present study, we investigated the relationship between sentence processing and implicit sequence learning in a dual-task paradigm in which the primary task was a non-linguistic task (Alternating Serial Reaction Time Task for measuring probabilistic implicit sequence learning), while the secondary task were a sentence comprehension task relying on syntactic processing. We used two control conditions: a non-linguistic one (math condition) and a linguistic task (word processing task). Here we show that the sentence processing interfered with the probabilistic implicit seque)
We want to demonstrate on some selected linguistic issues that classical structural and functional linguistics even with its seemingly traditional approaches has something to offer to a formal description of language and its applications in natural language processing and to illustrate by a brief reference to Functional Generative Grammar (on the theoretical side of CL) and Prague Dependency Treebank (on the applicational side) a possible interaction between linguistics and CL.
Prepositional phrase (PP) consists of two parts which are a preposition as the leading part and a word or phrase as the tail part. In accordance with this fact, this paper proposes a new approach for identifying PP. In this method, PP identification is transformed into the collocation identification of preposition itself and the right boundary word. The Cascaded Conditional Random Fields (CCRFs) is used in this approach. With the Penn Chinese Treebank 5.1 as our experiment corpus, the F <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sub> rises to 94.63%. This approach obtains breakthrough in this specific field as the current F <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sub> is about 8.6% higher than any publicly published paper.
Human languages evolve continuously, and a puzzling problem is how to reconcile the apparent robustness of most of the deep linguistic structures we use with the evidence that they undergo possibly slow, yet ceaseless, changes. Is the state in which we observe languages today closer to what would be a dynamical attractor with statistically stationary properties or rather closer to a non-steady state slowly evolving in time? Here we address this question in the framework of the emergence of shared linguistic categories in a population of individuals interacting through language games. The observed emerging asymptotic categorization, which has been previously tested - with success - against experimental data from human languages, corresponds to a metastable state where global shifts are always possible but progressively more unlikely and the response properties depend on the age of the system. This aging mechanism exhibits striking quantitative analogies to what is observed in the stati)
We present an approach of identifying the most prominent text/sentences using various shallow linguistic features, taking degree of connectiveness among the text units into consideration so as to minimize the poorly linked sentences in the resulting summary. As per the limitations of the current summarizing systems, the summary generated by those systems contains poorly linked sentences and are not topically salient. Thus, the paper aims at highlighting the effect of lexical chain scoring after the nouns and compound nouns are chained by searching for lexical cohesive relationships between words in the text using WordNet and using lexicographical relationships such as synonymy and hyponyms. In this paper, our algorithm ranks sentences based on the sum of the scores of the words in each sentence involving approaches like term frequencies, location of sentence in the text, cue words and phrases, word occurrences, and measuring lexical similarity(measuring chain score, word score and finally sentence score) for ranking the text units. We then identified and extracted high scored sentences and then the Vector Space approach is used to measure the relatedness/similarity between the extracted sentence and the topic words involving again the WordNet lexical database relationships to prioritise the topically related sentences. A threshold angle between the two vectors is predefined experimentally to which the ranked/scored sentences to be dropped and which the significant sentences with ranking/scores higher than threshold to be extracted. Note that the value of threshold is predetermined based on the percentage of output summary required to be generated.
In this paper, we describe and compare two statistical parsing approaches for the hybrid dependency-constituency syntactic representation used in the Quranic Arabic Treebank (Dukes and Buckwalter, 2010). In our first approach, we apply a multi-step process in which we use a shift-reduce algorithm trained on a pure dependency preprocessed version of the treebank. After parsing, the dependency output is converted into the hybrid representation. This is compared to a novel one-step parser that is able to learn the hybrid representation without preprocessing. We define an extended labelled attachment score (ELAS) as our performance metric for hybrid parsing, and report 87.47 % (F1 score) for the multi-step approach, and 89.03 % (F1 score) for the onestep integrated algorithm. We also consider the effect of using different sets of morphological features for parsing the Quran, comparing our results to recent work on Modern Standard Arabic.
This paper introduces Chart Inference (CI), an algorithm for deriving a CCG category for an unknown word from a partial parse chart. It is shown to be faster and more precise than a baseline brute-force method, and to achieve wider coverage than a rule-based system. In addition, we show the application of CI to a domain adaptation task for question words, which are largely missing in the Penn Treebank. When used in combination with self-training, CI increases the precision of the baseline StatCCG parser over subjectextraction questions by 50%. An error analysis shows that CI contributes to the increase by expanding the number of category types available to the parser, while self-training adjusts the counts. 1
BACKGROUND: The neuroanatomic basis of affective processing deficits in Huntington disease is insufficiently understood. We investigated whether Huntington disease-related deficits in emotion recognition and experience are associated with specific changes in grey matter volume. METHOD: We assessed grey matter volume in symptomatic patients with Huntington disease and healthy controls using voxel-based morphometry, and we correlated regional grey matter volume with participants' affective ratings. RESULTS: We enrolled 18 patients with Huntington disease and 18 healthy controls in our study. Patients with Huntington disease showed normal affective experience but impaired recognition of negative emotions (disgust, anger, sadness). The patients perceived the emotions as less intense and made more classification errors than controls. These deficits were correlated with regional atrophy in emotion-relevant areas (insula, orbitofrontal cortex) and in memory-relevant areas (dorsolateral prefrontal cortex, hippocampus). LIMITATIONS: Our study was limited by the small sample size and the resulting modest statistical power relative to the number of tests. CONCLUSION: Our study sheds new light on the importance of a cognitive-affective brain circuit involved in the affect recognition impairment in patients with Huntington disease.
The primary sensory cortices are characterized by a topographical mapping of basic sensory features which is considered to deteriorate in higher-order areas in favor of complex sensory features. Recently, however, retinotopic maps were also discovered in the higher-order visual, parietal and prefrontal cortices. The discovery of these maps enabled the distinction between visual regions, clarified their function and hierarchical processing. Could such extension of topographical mapping to high-order processing regions apply to the auditory modality as well? This question has been studied previously in animal models but only sporadically in humans, whose anatomical and functional organization may differ from that of animals (e.g. unique verbal functions and Heschl's gyrus curvature). Here we applied fMRI spectral analysis to investigate the cochleotopic organization of the human cerebral cortex. We found multiple mirror-symmetric novel cochleotopic maps covering most of the core and hig)
Emotion recognition algorithms for spoken dialogue applications typically employ lexical models that are trained on labeled in-domain data. In this paper, we propose a domainindependent approach to affective text modeling that is based on the creation of an affective lexicon. Starting from a small set of manually annotated seed words, continuous valence ratings for new words are estimated using semantic similarity scores and a kernel model. The parameters of the model are trained using least mean squares estimation. Word level scores are combined to produce sentence-level scores via simple linear and non-linear fusion. The proposed method is evaluated on the SemEval news headline polarity task and on the ChIMP politeness and frustration detection dialogue task, achieving state-of-theart results on both. For politeness detection, best results are obtained when the affective model is adapted using in domain data. For frustration detection, the domain-independent model and non-linear fusion achieve the best performance. Index Terms: language understanding, emotion, affect, affective lexicon
OBJECTIVES: This research compared sensory processing and personality traits involved in deciding to try a novel fruit (guava) in adults and children. DESIGN: The research employed an age, sex, and food neophobia matched between-participant design to examine sensory decision making in choosing to eat a novel fruit. METHODS: Forty-four adults (Study 1) and 68 children (Study 2) took part. In each study, participants were separated into two groups to investigate whether prior assessment of a familiar and liked fruit (apple) that shares similar visual characteristics to the target novel fruit (guava) increased the likelihood that an individual would decide to try it. All participants completed appetitive and familiarity ratings by sensory stages: vision, smell, and touch, prior to trying (tasting) the fruit. Participants (or their parents) also completed the general and food neophobia scales and adults also completed the sensation-seeking scale. RESULTS: Twenty-eight adults (64%) tried the guava and 16 did not (36%). In the second study, 22 children decided not to try the novel fruit (32%). Significant predictors of whether the adult tried the target fruit were Thrill and Adventure Seeking, Experience Seeking, General Neophobia, and 'appealing to touch'. In children, Food Neophobia, concurrent presentation of a familiar fruit alongside the target and visual assessment of the target predicted decision to try the novel fruit. CONCLUSIONS: This study suggests that touch is pertinent to adults' decision to try a novel fruit, whereas visual cues appear to be more important for children.
Fake content is flourishing on the Internet, ranging from basic random word salads to web scraping. Most of this fake content is generated for the purpose of nourishing fake web sites aimed at biasing search engine indexes: at the scale of a search engine, using automatically generated texts render such sites harder to detect than using copies of existing pages. In this paper, we present three methods aimed at distinguishing natural texts from artificially generated ones: the first method uses basic lexicometric features, the second one uses standard language models and the third one is based on a relative entropy measure which captures short range dependencies between words. Our experiments show that lexicometric features and language models are efficient to detect most generated texts, but fail to detect texts that are generated with high order Markov models. By comparison our relative entropy scoring algorithm, especially when trained on a large corpus, allows us to detect these “hard” text generators with a high degree of accuracy.
Participants read aloud swear words, euphemisms of the swear words, and neutral stimuli while their autonomic activity was measured by electrodermal activity. The key finding was that autonomic responses to swear words were larger than to euphemisms and neutral stimuli. It is argued that the heightened response to swear words reflects a form of verbal conditioning in which the phonological form of the word is directly associated with an affective response. Euphemisms are effective because they replace the trigger (the offending word form) by another word form that expresses a similar idea. That is, word forms exert some control on affect and cognition in turn. We relate these findings to the linguistic relativity hypothesis, and suggest a simple mechanistic account of how language may influence thinking in this context. [ABSTRACT FROM AUTHOR], Copyright of PLoS ONE is the property of Public Library of Science and its content may not be copied or emailed to multiple sites or posted to )
This study investigated cognitive and emotional effects of syncopation, a feature of musical rhythm that produces expectancy violations in the listener by emphasising weak temporal locations and de-emphasising strong locations in metric structure. Stimuli consisting of pairs of unsyncopated and syncopated musical phrases were rated by 35 musicians for perceived complexity, enjoyment, happiness, arousal, and tension. Overall, syncopated patterns were more enjoyed, and rated as happier, than unsyncopated patterns, while differences in perceived tension were unreliable. Complexity and arousal ratings were asymmetric by serial order, increasing when patterns moved from unsyncopated to syncopated, but not significantly changing when order was reversed. These results suggest that syncopation influences emotional valence (positively), and that while syncopated rhythms are objectively more complex than unsyncopated rhythms, this difference is more salient when complexity increases than when it decreases. It is proposed that composers and improvisers may exploit this asymmetry in perceived complexity by favoring formal structures that progress from rhythmically simple to complex, as can be observed in the initial sections of musical forms such as theme and variations.
Part-of-speech (POS) is an indispensable feature in dependency parsing. Current research usually models POS tagging and dependency parsing independently. This may suffer from error propagation problem. Our experiments show that parsing accuracy drops by about 6 % when using automatic POS tags instead of gold ones. To solve this issue, this paper proposes a solution by jointly optimizing POS tagging and dependency parsing in a unique model. We design several joint models and their corresponding decoding algorithms to incorporate different feature sets. We further present an effective pruning strategy to reduce the search space of candidate POS tags, leading to significant improvement of parsing speed. Experimental results on Chinese Penn Treebank 5 show that our joint models significantly improve the state-of-the-art parsing accuracy by about 1.5%. Detailed analysis shows that the joint method is able to choose such POS tags that are more helpful and discriminative from parsing viewpoint. This is the fundamental reason of parsing accuracy improvement. 1
Background: The language faculty is probably the most distinctive feature of our species, and endows us with a unique ability to exchange highly structured information. In written language, information is encoded by the concatenation of basic symbols under grammatical and semantic constraints. As is also the case in other natural information carriers, the resulting symbolic sequences show a delicate balance between order and disorder. That balance is determined by the interplay between the diversity of symbols and by their specific ordering in the sequences. Here we used entropy to quantify the contribution of different organizational levels to the overall statistical structure of language. Methodology/Principal Findings: We computed a relative entropy measure to quantify the degree of ordering in word sequences from languages belonging to several linguistic families. While a direct estimation of the overall entropy of language yielded values that varied for the different families con)
Previous research suggests that neural and behavioral responses to surprised faces are modulated by explicit contexts (e.g., "He just found $500"). Here, we examined the effect of implicit contexts (i.e., valence of other frequently presented faces) on both valence ratings and ability to detect surprised faces (i.e., the infrequent target). In Experiment 1, we demonstrate that participants interpret surprised faces more positively when they are presented within a context of happy faces, as compared to a context of angry faces. In Experiments 2 and 3, we used the oddball paradigm to evaluate the effects of clearly valenced facial expressions (i.e., happy and angry) on default valence interpretations of surprised faces. We offer evidence that the default interpretation of surprise is negative, as participants were faster to detect surprised faces when presented within a happy context (Exp. 2). Finally, we kept the valence of the contexts constant (i.e., surprised faces) and showed that participants were faster to detect happy than angry faces (Exp. 3). Together, these experiments demonstrate the utility of the oddball paradigm to explore the default valence interpretation of presented facial expressions, particularly the ambiguously valenced facial expression of surprise.
There has been a rapid increase in the volume of research on data-driven dependency parsers in the past five years. This increase has been driven by the availability of treebanks in a wide variety of languages—due in large part to the CoNLL shared tasks—as well as the straightforward mechanisms by which dependency theories of syntax can encode complex phenomena in free word order languages. In this article, our aim is to take a step back and analyze the progress that has been made through an analysis of the two predominant paradigms for data-driven dependency parsing, which are often called graph-based and transition-based dependency parsing. Our analysis covers both theoretical and empirical aspects and sheds light on the kinds of errors each type of parser makes and how they relate to theoretical expectations. Using these observations, we present an integrated system based on a stacking learning framework and show that such a system can learn to overcome the shortcomings of each non-integrated system.
This paper proposes a method to improve the accuracy of bilingual texts (bitexts) dependency parsing by using an auto-generated bilingual treebank created with the help of statistical machine translation (SMT) systems. Previous bitext parsing methods use human-annotated bilingual treebanks that are costly and troublesome to obtain. In the proposed method, we use an auto-generated bilingual treebank to train the parsing models. First, an SMT system is used to translate a monolingual treebank into the target language; then, a monolingual parser for the target language is used to parse the translated sentences. Since the auto-translated sentences and auto-parsed trees in the auto-generated bilingual treebank are far from perfect, the bilingual constraints are not sufficiently reliable. To overcome this problem, we propose a method to verify the reliability of the constraints using a large amount of target monolingual and bilingual unannotated data. Finally, we design a set of effective bilingual features for parsing models on the basis of the verified constraints. We conduct the experiments using a standard test data. The experimental results show that our bitext parser significantly outperforms monolingual parsers. Moreover, our method is still able to provide improvement when we use a larger monolingual treebank containing over 50 000 sentences. We also test the proposed method with different SMT systems and the results show that our method is very robust to the noise. In particular, the proposed method can be used in a purely monolingual setting with the help of SMT. That is, it does not need the human translation of the test set as previous methods do.
Transition-based dependency parsers generally use heuristic decoding algorithms but can accommodate arbitrarily rich feature representations. In this paper, we show that we can improve the accuracy of such parsers by considering even richer feature sets than those employed in previous systems. In the standard Penn Treebank setup, our novel features improve attachment score form 91.4 % to 92.9%, giving the best results so far for transitionbased parsing and rivaling the best results overall. For the Chinese Treebank, they give a signficant improvement of the state of the art. An open source release of our parser is freely available.
Text clustering is of substantial importance to information retrieval.The method of applying the information of syntactic distribution to text clustering is presented,in order to avoid the complex clustering algorithm whileenabling the linguistic interpretation of clustering features and the results of clustering.According to the dependency Treebank,ten dependency relations are suggested with distinctive distribution between oral and written Chinese By using five of them as clustering feature,the similarity of spoken and written classes achieves 71.98% and 83.13%,respectively.The experiment result shows that the proposed method of applying dependency relations to text clustering is feasible and effective.
For the task of automatic treebank conversion, this paper presents a feature-based approach which encodes bracketing structures in a treebank into features to guide the conversion of this treebank to a different standard. Experiments on two Chinese treebanks show that our approach improves conversion accuracy by 1.31 % over a strong baseline. 1
International audience
We propose a method for capturing vocalizations that is designed to avoid some of the limiting factors found in traditional bioacoustical methods, such as the impossibility of obtaining continuous long-term registers or analyzing amplitude due to the continuous change of distance between the subject and the position of the recording system. Using Bluetooth technology, vocalizations are captured and transmitted wirelessly into a receiving system without affecting the quality of the signal. The recordings of the proposed system were compared to those obtained as a reference, which were based on the coding of the signal with the so-called pulse-code modulation technique in WAV audio format without any compressing process. The evaluation showed p < .05 for the measured quantitative and qualitative parameters. We also describe how the transmitting system is encapsulated and fixed on the animal and a way to video record a spider monkey’s behavior simultaneously with the audio recordings.
Documentation of medical records requires professional knowledge,which is also the basis for medical translation.Centering on medical terminology and anatomical logic,this article,by way of examples,discusses Chinese to English translation skills at lexical,syntactical,grammatical and textual levels.It is hoped that these skills will help translators produce English versions of medical records up to the standards and norms of this profession.