Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
OBJECTIVE: To explore the possible differences in subjective analysis of the emotional stimuli from the International Affective Picture System between elderly and young samples. METHOD: 187 elderly subjects ranked the International Affective Picture System images according to the directions from the Manual of Affective Ratings. Their scores were compared to those obtained from International Affective Picture System studies with young people. RESULT: There is an age-related difference in arousal and valence in the International Affective Picture System rating. The correlation between affective valence and arousal is strong, and negative for the elderly. The expected versus the observed frequency of International Affective Picture System images between elderly and young samples show a statistical difference. CONCLUSION: This study shows an inter-age statistical dichotomy in how elderly and young people subjectively evaluate International Affective Picture System images.
In this paper, we propose two different language modeling approaches, namely skip trigram and across sentence boundary, to capture the long range dependencies. The skip trigram model is able to cover more predecessor words of the present word compared to the normal trigram while the same memory space is required. The across sentence boundary model uses the word distribution of the previous sentences to calculate the unigram probability which is applied as the emission probability in the word and the class model frameworks. Our experiments on the Penn Treebank [1] show that each of our proposed models and also their combination significantly outperform the baseline for both the word and the class models and their linear interpolation. The linear interpolation of the word and the class models with the proposed skip trigram and across sentence boundary models achieves 118.4 perplexity while the best state-of-the-art language model has a perplexity of 137.2 on the same dataset. 1.
Message Oriented Middleware (MOM) is getting popular along with the development of heterogeneous platforms and applications. Most Message Oriented Middleware supports Publish/Subscribe scheme for message interoperation and it usually matches topics directly by matching keywords. Recent Semantic Message Oriented Middleware base on ontology which requires knowledge to be predefined. In this paper, we provide a semantic method for matching topics (keywords) by using a lexical database (WordNet). By expanding the publishing queries among topics (keywords), highly related topics can also be subscribed all at once. As a result, we propose a Semantic Publish/Subscribe Message Oriented Middleware which is suitable for many general purposes without special configurations.
This dissertation considers the language socialization of law students. One message that the law students encounter is that legal Swedish is an entirely new language. The main aim is to investigate what linguistic norms are conveyed to the students through the teachers’ comments on the students’ texts and through various forms of writing instructions. The material consists of student texts with teacher comments and documentation on various phases of instruction with a focus on writing. Teacher comments on texts written during the first year of the law programme are analyzed and categorized. The analysis stems from two models. The first model is based on different text levels, like formal conventions of writing, sentence construction, text structure, word choice and style, and content. The second model distinguishes different linguistic norms based on three layers: The first layer consists of written language norms in general language practice, the second of academic language norms and the third of norms that are specific to the use of legal language. The results show that word choice and style is the most common category for the teachers’ comments in the first term of the law programme and content is the most common in the second term (with word choice and style the second most common). Formal conventions of writing, sentence structure and different types of grammatical constructions are some of the things the teachers criticize. Surprisingly few of the teachers’ comments concern more overarching aspects such as text structure or the aim and genre of the text. Comments are made on local features in the text, but rarely on more global features. The teaching practice that the writing of law students belongs to entails, among other things, that the students’ texts are assessed anonymously for the sake of fairness. This means that there is not much opportunity for a student to discuss the text with the teacher who commented on and assessed it. The construction of the teachers’ text comments is particularly important when dialogue between student and teacher on the text draft and final version is not an integral part of instruction. The teachers’ written comments are usually brief and do not allow much space for a consideration of linguistic norms and text patterns, which reduces the opportunities for the teachers and the law programme to contribute to a deeper linguistic awareness in the law students.
Proceedings of the Ninth International Workshop \non Treebanks and Linguistic Theories. \nEditors: Markus Dickinson, Kaili Müürisep and Marco Passarotti. \nNEALT Proceedings Series, Vol. 9 (2010), 31-42. \n© 2010 The editors and contributors. \nPublished by \nNorthern European Association for Language \nTechnology (NEALT) \nhttp://omilia.uio.no/nealt. \nElectronically published at \nTartu University Library (Estonia) \nhttp://hdl.handle.net/10062/15891.
Nivre’s method was improved by enhancing deterministic dependency parsing through application of a tree-based model. The model considers all words necessary for selection of parsing actions by including words in the form of trees. It chooses the most probable head candidate from among the trees and uses this candidate to select a parsing action. In an evaluation experiment using the Penn Treebank (WSJ section), the proposed model achieved higher accuracy than did previous deterministic models. Although the proposed model’s worst-case time complexity is O(n 2), the experimental results demonstrated an average parsing time not much slower than O(n). 1
Language is not a system of signs that allow the exchange of information among individuals but also the means to engage in intersubjective relationship as well as to mark the identity of the speaker: the connection between language and identity is often so strong that a single feature of linguistic usage can be enough to identity someone’s belonging in a certain group. The linguistic matrix, that consists of being aware of how linguistic register expectations become linguistic norms, constitutes the basis of the community of speakers and society. Linguistic behaviour is exposed to distinct social dynamics and new linguistic habits and can be influenced by the sense of belonging, as perceived by the speakers, and the strength of the language (vitality and prestige). Pluralinguistic formations make allowance for the different ways of communication, adapting them to local customs, in this way avoiding harm, as well as introducing a positive bias, in populations so far relegated to minority status.
Objective: To investigate whether interviewer personality, sex or being of the same sex as the interviewee, and training account for variance between interviewers’ ratings in a medical student selection interview. Design, setting and participants: In 2006 and 2007, data were collected from cohorts of each year's interviewers (by survey) and interviewees (by interview) participating in a multiple mini-interview (MMI) process to select students for an undergraduate medical degree in Australia. MMI scores were analysed and, to account for the nested nature of the data, multilevel modelling was used. Main outcome measures: Interviewer ratings; variance in interviewee scores. Results: In 2006, 153 interviewers (94% response rate) and 268 interviewees (78%) participated in the study. In 2007, 139 interviewers (86%) and 238 interviewees (74%) participated. Interviewers with high levels of agreeableness gave higher interview ratings (correlation coefficient [r] = 0.26 in 2006; r = 0.24 in 2007) and, in 2007, those with high levels of neuroticism gave lower ratings (r = − 0.25). In 2006 but not 2007, female interviewers gave higher overall ratings to male and female interviewees (t = 2.99, P = 0.003 in 2006; t = 2.16, P = 0.03 in 2007) but interviewer and interviewee being of the same sex did not affect ratings in either year. The amount of variance in interviewee scores attributable to differences between interviewers ranged from 3.1% to 24.8%, with the mean variance reducing after skills-based training (20.2% to 7.0%; t = 4.42, P = 0.004). Conclusion: This study indicates that rating leniency is associated with personality and sex of interviewers, but the effect is small. Random allocation of interviewers, similar proportions of male and female interviewers across applicant interview groups, use of the MMI format, and skills-based interviewer training are all likely to reduce the effect of variance between interviewers.
The paper presents an approach to valency frame extraction for Croatian verbs on basis of morphological and syntactic features of wordforms from syntactically annotated sentences. We have used a gold standard sample of approximately 1200 sentences and 30.000 tokens from the Croatian Dependency Treebank and a frame instance extraction algorithm. We extracted 936 verb frame instances for 424 different verbs – consisting of lemmas, morphosyntactic tags and syntactic functions of the encountered wordforms – and manually assigned tectogrammatical functors to their elements. Distributional properties are given in terms of co-occurrences for each of these features. The obtained results will serve for further development of valency frame extraction procedures.
Cytotoxic T cell (CTL) covers several subtypes, which are CD8+, CD4 and CD4-CD8-. CTL derives from T cell repertoire in lymphoid hematopoietic stem cells. It matures in thymus and is activated in peripheral lymphoid tissues. Effector CTL kills the target cells by 2 ways. One is apoptotic effect mediated by FasL-Fas pathway and the other one is cytolytic effect mediated by granzymes. CTL has aroused great attention due to its significance in anti-tumor and anti-virus.
In this paper, we argue for and demonstrate the use of Prolog as a tool to query annotated corpora. We present a case study based on the German TüBa-D/Z Treebank to show that flexible and efficient corpus querying can be started with a minimal amount of effort. We end this paper with a brief discussion of performance, that suggests that the approach is both fast enough and scalable. 1
This study explored the conceptual framework of dieticians' intentions to recommend functional food and the mediating role of consumption frequency. A web-based survey was designed using a self-administered questionnaire. A sample of Korean dieticians (N=233) responded to the questionnaire that included response efficacy, risk perception, consumption frequency, and recommendation intention for functional foods. A structural equation model was constructed to analyze the data. We found that response efficacy was positively related to frequency of consumption of functional foods and to recommendation intention. Consumption frequency also positively influenced recommendation intention. Risk perception had no direct influence on recommendation intention; however, the relationship was mediated completely by consumption frequency. Dieticians' consumption frequency and response efficacy were the crucial factors in recommending functional foods. Dieticians may perceive risks arising from the use of functional foods in general, but the perceived risks do not affect ratings describing dieticians' intentions to recommend them. The results also indicated that when dieticians more frequently consume functional foods, the expression of an intention to recommend functional foods may be controlled by the salience of past behaviors rather than by attitudes.
The novel anthocyanins, malvidin 3-O-(6-O-(4-O-malonyl-alpha-rhamnopyranosyl)-beta-glucopyranoside)-5-O-beta-glucopyranoside (2), malvidin 3-O-(6-O-alpha-rhamnopyranosyl-beta-glucopyranoside)-5-O-(6-O-malonyl-beta-glucopyranoside) (3), malvidin 3-O-(6-O-(4-O-malonyl-alpha-rhamnopyranosyl)-beta-glucopyranoside)-5-O-(6-O-malonyl-beta-glucopyranoside) (4), malvidin 3-O-(6-O-(4-O-malonyl-alpha-rhamnopyranosyl)-beta-glucopyranoside) (5) and malvidin 3-O-(6-O-(Z)-p-coumaroyl-beta-glucopyranoside)-5-O-beta-glucopyranoside (6), in addition to the 3-O-(6-O-alpha-rhamnopyranosyl-beta-glucopyranoside)-5-O-beta-glucopyranoside (1) and the 3-O-(6-O-(E)-p-coumaroyl-beta-glucopyranoside)-5-O-beta-glucopyranoside (7) of malvidin have been isolated from purple leaves of Oxalis triangularis A. St.-Hil. In pigments 2, 4 and 5 a malonyl unit is linked to the rhamnose 4-position, which has not been reported previously for any anthocyanin before. The identifications were mainly based on 2D NMR spectroscopy and electrospray MS.
Abstract Experiences with racism and age negatively affect how Afro-Brazilians in Salvador and São Paulo rate democracy. Older cohorts are more likely to rate democracy high compared to younger cohorts who rate it as low. Respondents in Salvador tend to rate democracy lower than respondents in São Paulo. Moreover, interviews reveal that as citizens believe they are not accorded full rights, they do not agree that Brazil's political system is fully democratic. Studies examining democracy in Brazil and racial politics throughout the diaspora would benefit from examining racialized experiences of citizens, rather than simply including the demographic variable of race. It is these experiences that affect rating of democracy rather than ascribed notions of race.
In this methodological investigation, we examined the influence of cultural background on viewers' interpretations of visual stimuli and verbs elicited by these materials. French and Mandarin native speakers' interpretations of seventeen short movies, produced by French speakers, depicting various state-changing actions were collected by a 25-item cultural protocol. A slight difference in the familiarity rating of movies is found between French and Mandarin participants. We also found that Mandarin speakers used more general verbs when describing actions depicted by movies with low familiarity rating and children used more conventional forms with movies of higher familiarity. Hierarchical cluster analyses were conducted in selecting movies that were matched in action-interpretations by both language groups.
Two of the main corpora available for training discourse relation classifiers are the RST Discourse Treebank (RST-DT) and the Penn Discourse Treebank (PDTB), which are both based on the Wall Street Journal corpus. Most recent work using discourse relation classifiers have employed fully-supervised methods on these corpora. However, certain discourse relations have little labeled data, causing low classification performance for their associated classes. In this paper, we attempt to tackle this problem by employing a semi-supervised method for discourse relation classification. The proposed method is based on the analysis of feature cooccurrences in unlabeled data. This information is then used as a basis to extend the feature vectors during training. The proposed method is evaluated on both RST-DT and PDTB, where it significantly outperformed baseline classifiers. We believe that the proposed method is a first step towards improving classification performance, particularly for discourse relations lacking annotated data.
We present algorithms for higher-order dependency parsing that are “third-order” in the sense that they can evaluate substructures containing three dependencies, and “efficient ” in the sense that they require only O(n4) time. Importantly, our new parsers can utilize both sibling-style and grandchild-style interactions. We evaluate our parsers on the Penn Treebank and Prague Dependency Treebank, achieving unlabeled attachment scores of
Dietary interventions with a household support component show promise for improving household social support and may impact magnitude of dietary change.
For centuries, scholars have explored the deep links among human languages. In this paper, we present a class of probabilistic models that use these links as a form of naturally occurring supervision. These models allow us to substantially improve performance for core text processing tasks, such as morphological segmentation, part-of-speech tagging, and syntactic parsing. Besides these traditional NLP tasks, we also present a multilingual model for the computational decipherment of lost languages. 1. Overview Electronic text is currently being produced at a vast and unprecedented scale across the languages of the world. Natural Language Processing (NLP) holds out the promise of automatically analyzing this growing body of text. However, over the last several decades, NLP research efforts have focused on the English language, often neglecting the thousands of other languages of the world (Bender, 2009). Most of these languages are currently beyond the reach of NLP technology due to several factors. One of these is simply the lack of the kinds of hand-annotated linguistic resources that have helped propel the performance of English language systems. For complex tasks of linguistic analysis, hand-annotated corpora can be prohibitively time-consuming and expensive to produce. For example, the most widely used annotated corpus in the English language, the Penn Treebank (Marcus et al., 1994), took years for a team of professional linguists to produce. It is unrealistic to expect such resources to ever exist for the majority of the world’s languages.
Recognition of special linguistic patterns in a certain language is very helpful for many NLP applications such as information extraction, machine translation and parsing. State-of-the-arts syntax parsers are based on given grammar. The used grammar is context free and cannot discover complex patterns which contain multiple linguistic units. We propose an unsupervised method to automatically discover the complex linguistic patterns from a classically parsed corpus. A specialized and efficient algorithm is applied to mine the frequent subtrees in the forest and the found subtrees are formalized as the linguistic patterns. The approach is validated on the Penn Chinese Treebank with found linguistic patterns.
This paper proposes a new cascade algorithm based on conditional random fields. The algorithm is applied to automatic recognition of Chinese verb-object collocation, and combined with a new sequence labeling of “ONIY”. Experiments compare identified results under two segmentations and part-of-speech tag sets. The comprehensive experimental results show that the best performance is 90.65% in F-score over Tsinghua Treebank, and 82.00% in F-score over the segmentation and part-of-speech tagging scheme of Peking University. Our experiments show that the proposed algorithm can greatly improve recognition accuracy of multi-nested collocation, and play a positive role on long distance collocation.
We describe a process for converting the Penn Arabic Treebank into the CCG formalism. Previous efforts have yielded CCGbanks in English, German, and Turkish, thus opening these languages to the sophisticated computational tools developed for CCG and enabling further cross-linguistic development. Conversion from a context free grammar treebank to a CCGbank is a four stage process: head finding, argument classification, binarization, and category conversion. In the process of implementing a basic CCGbank conversion algorithm, we reveal properties of Arabic grammar that interfere with conversion, such as subject topicalization, genitive constructions, relative clauses, and optional pronominal subjects. All of these problematic phenomena can be resolved in a variety of ways- we discuss advantages and disadvantages of each in their respective sections. We detail these and describe our categorial analysis of each of these Arabic grammatical phenomena in depth, as well as technical details on their integration into the conversion algorithm. 1.
Semantic dependency analysis is practicable way to semantic analysis. This paper describes a Chinese semantic dependency analysis system using HowNet. The system takes sentences with phrase syntactic information as input. First, it determines the headword of each phrase to get the dependency structure of the sentence. Second, it takes the syntactic constituent as the basic unit of semantic labeling and determines the semantic relation through searching HowNet and Semantic Information Structure Library.We randomly extract 100 sentences from Penn Chinese Treebank as test data. There are totally 2783 pairs of word. The system determines the semantic relations of 2546 pairs.The labeling ratio is 91.5%.
Large-scale phrase structure treebank and dependency structure treebank are developed and interconverted for the purpose of syntactic analysis on true corpus. The head percolation table is constructed on modern Chinese dependency grammar by discussing the relationship between phrase structure and dependency structure based on Penn Chinese Treebank (CTB),and CTB from phrase structure is converted to dependency structure treebank using the head percolation table. 200 sentences are chosen from CTB randomly to evaluate the conversion performance. Precision of the conversion has attained 99.50%. The achieved dependency structure treebank can be used to analyze Chinese dependency relation.
Megastudies with processing efficiency measures for thousands of words allow researchers to assess the quality of the word features they are using. In this article, we analyse reading aloud and lexical decision reaction times and accuracy rates for 2,336 words to assess the influence of subjective frequency and age of acquisition on performance. Specifically, we compare newly presented word frequency measures with the existing frequency norms of Kucera and Francis (1967), HAL (Burgess & Livesay, 1998), Brysbaert and New (2009), and Zeno, Ivens, Millard, and Duvvuri (1995). We show that the use of the Kucera and Francis word frequency measure accounts for much less variance than the other word frequencies, which leaves more variance to be "explained" by familiarity ratings and age-of-acquisition ratings. We argue that subjective frequency ratings are no longer needed if researchers have good objective word frequency counts. The effect of age of acquisition remains significant and has an effect size that is of practical relevance, although it is substantially smaller than that of the first phoneme in naming and the objective word frequency in lexical decision. Thus, our results suggest that models of word processing need to utilize these recently developed frequency estimates during training or setting baseline activation levels in the lexicon.
This paper describes the transfer component of a syntax-based Example-based Machine Translation system. The source sentence parse tree is matched in a bottom-up fashion with the source language side of a parallel example treebank, which results in a target forest which is sent to the target language generation component. The results on a 500 sentences test set are compared with a top-down approach to transfer of the same system, with the bottom-up approach yielding much better results. 1
The article analyzes 97 elementary schoolbooks in Buenos Aires to determine which social representations about linguistic norm underlie in these school materials. The paper reviews -especially in the defi nitions of categories, and exercises and activities- the concepts of linguistic variety, standard language and español neutro. Based on these variables, this article sees the possible repercussions in social representations that students and teachers can develop from point of view of the publishing companies.
Lexicon-Grammar tables are a very rich syntactic lexicon for the French language. This linguistic database is nevertheless not directly suitable for use by computer programs, as it is incomplete and lacks consistency. Tables are defined on the basis of features which are not explicitly recorded in the lexicon. These features are only described in literature. Our aim is to define for each tables these essential properties to make them usable in various Natural Language Processing (NLP) applications, such as parsing.
uni-tuebingen.de This paper describes a CoNLL-style chunk representation for the Tübingen Treebank of Written German, which assumes a flat chunk structure so that each word belongs to at most one chunk. For German, such a chunk definition causes problems in cases of complex prenominal modification. We introduce a flat annotation that can handle these structures via a stranded noun chunk. 1
Previous studies demonstrate that lexical coding of colour influences categorical perception of colour, such that participants are more likely to rate two colours to be more similar if they belong to the same linguistic category (Roberson et al., 2000, 2005). Recent work shows changes in Greek–English bilinguals' perception of within and cross-category stimulus pairs as a function of the availability of the relevant colour terms in semantic memory, and the amount of time spent in the L2-speaking country (Athanasopoulos, 2009). The present paper extends Athanasopoulos' (2009) investigation by looking at cognitive processing of colour in Japanese–English bilinguals. Like Greek, Japanese contrasts with English in that it has an additional monolexemic term for ‘light blue’ (mizuiro). The aim of the paper is to examine to what degree linguistic and extralinguistic variables modulate Japanese–English bilinguals' sensitivity to the blue/light blue distinction. Results showed that those bilinguals who used English more frequently distinguished blue and light blue stimulus pairs less well than those who used Japanese more frequently. These results suggest that bilingual cognition may be dynamic and flexible, as the degree to which it resembles that of either monolingual norm is, in this case, fundamentally a matter of frequency of language use.
This article discusses the treatment of collocations in the context of along-term project on the development of multilingual NLP tools. Besides“classical” two-word collocations, we will focus on the case of complexcollocations (3 words or more) for which a recursive design is presented in theform of collocation of collocations. Although comparatively less numerous thantwo-word collocations, the complex collocations pose important challenges forNLP. The article discusses how these collocations are retrieved from corpora,inserted and stored in a lexical database, how the parser uses such knowledgeand what are the advantages offered by a recursive approach to complexcollocations.
In this paper, we offer broad insight into the underperformance of Arabic constituency parsing by analyzing the interplay of linguistic phenomena, annotation choices, and model design. First, we identify sources of syntactic ambiguity understudied in the existing parsing literature. Second, we show that although the Penn Arabic Treebank is similar to other treebanks in gross statistical terms, annotation consistency remains problematic. Third, we develop a human interpretable grammar that is competitive with a latent variable PCFG. Fourth, we show how to build better models for three different parsers. Finally, we show that in application settings, the absence of gold segmentation lowers parsing performance by 2–5 % F1. 1
The aim of this article is to analyze the formation of Old English adverbs (A-Y) as retrieved from the lexical database of Old English Nerthus within the theoretical framework of the Layered Structure of the Word. Firstly, a critical review of the literature on Old English adverb formation is offered in order to emphasize the necessity of an exhaustive and theoretically up-to-date study that distinguishes clearly synchronic from diachronic aspects on the one hand, and inflectional from derivational aspects on the other. Secondly, an exhaustive analysis of the derivation of adverbs in Old English by means of different word-formation processes (zero derivation, conversion, affixation and compounding) is given. In the theoretical part, the conclusion reached is that conversion requires a Complex Word structure and that a distinction has to be drawn between syntactic exocentricity and morphological exocentricity.
Nouns are generally easier to learn than verbs (e.g., Bornstein, 2005; Bornstein et al., 2004; Gentner, 1982; Maguire, Hirsh-Pasek, & Golinkoff, 2006). Yet, verbs appear in children's earliest vocabularies, creating a seeming paradox. This paper examines one hypothesis about the difference between noun and verb acquisition. Perhaps the advantage nouns have is not a function of grammatical form class but rather related to a word's imageability. Here, word imageability ratings and form class (nouns and verbs) were correlated with age of acquisition according to the MacArthur-Bates Communicative Development Inventory (CDI) (Fenson et al., 1994). CDI age of acquisition was negatively correlated with words' imageability ratings. Further, a word's imageability contributes to the variance of the word's age of acquisition above and beyond form class, suggesting that at the beginning of word learning, imageability might be a driving factor.
BACKGROUND: In this controlled postdiagnosis study, the authors examined various aspects of body image of breast cancer survivors in cross-sectional and longitudinal designs. METHODS: In 2004 and 2007 the Body Image Scale (BIS) was completed by the same 248 disease-free women who had been treated for stage II and III breast cancer between 1998 and 2002. "Poorer" body image was defined as greater than the 70th percentile (N=76 women) of the BIS scores in contrast to "better" body image (N=172 women). Breast cancer survivors were examined clinically in 2004, and their BIS scores were compared with the scores from an age-matched group of women from the general population. RESULTS: In this cross-sectional study, poorer body image in 2004 was associated significantly with modified radical mastectomy, undergoing or planning to undergo breast-reconstructive surgery, a change in clothing, poor physical and mental health, chronic fatigue, and reduced quality of life (QoL). In univariate analyses, most of these factors and manually planned radiotherapy were significant predictors of poorer body image in 2007. In multivariate analyses, manually planned radiotherapy, poor physical QoL and high BIS score in 2004 remained independent predictors of a poorer body image in 2007. Body image ratings were relatively stable from 2004 to 2007. Twenty-one percent of breast cancer survivors reported body image dissatisfaction, similar to the proportion of dissatisfaction in controls. CONCLUSIONS: In this cross-sectional analysis, body image in breast cancer survivors was associated with the types of surgery and radiotherapy and with mental distress, reduced health, and impaired QoL. Body image ratings were relatively stable over time, and the antecedent body image score was a strong predictor of body image at follow-up. Body image in breast cancer survivors differed very little from that in controls.
The late positive potential (LPP) depicts brain electrical activity during both automatic and controlled sustained attentional processing of emotional stimuli. We investigated in a sample of 18 healthy women how the LPP is modulated by facial expression during an explicit valence rating task and an implicit sex classification task. Midline LPP amplitudes were significantly larger for valence rating than for sex classification. During valence rating, faces with a positive valence resulted in larger LPP amplitudes at centrofrontal electrodes than faces with a negative valence. During sex classification, a similar valence effect was observed at midline parietal electrodes. This implicit LPP valence effect appears to depend on higher visual processing, as during an additional sex classification task with blurred faces no such implicit valence effect was found.
Treebank is a text corpus with syntactic annotation. It records the syntactic tree, i.e. the syntactic parse, of every sentence in running texts. Since 1990s, automatic parsing of natural languages has again become the focus of the international community of computational linguistics, and one of the crucial reasons is the realization of the Penn Treebank (PTB). The performances of statistical parsers, which are based on automatically induced Probabilistic Context-Free Grammar (PCFG), outperform significantly those rule-and unification-based parsers. A Treebank of any language in the world, represented with either phrase structures or dependency structures, takes sentence as its basic description unit. Dependency Grammar is a lexicalized grammar; it denies the notion of phrase structures and describes only the various word-word relations in a sentence, in which the head-word is the dominant of a given relation, and the other word of the word-pair at stake is the dependent of the head. Dependency Grammar creates a transparent interface between the dependency syntax and semantics of a language. This paper highly estimates the life force of the sentence-based syntax and the head-driven sentence analyzing method advocated by Jinxi Li, because they have not only dominated grammar teaching in middle schools more than half century before and after the foundation of the People’s Republic of China, but also guides the treebanking practice today.
In this paper we present an experimental toolbox for automatic tree-to-tree alignment based on local classification and alignment inference. The aligner implements a recurrent architecture for structural prediction using history features and a sequential classification procedure. The discriminative base classifier uses a log-linear model which enables simple integration of various features extracted from the data. The Lingua-Align toolbox provides a flexible framework for feature extraction including contextual properties and implements several alignment inference procedures. Various settings and constraints can be controlled via a simple frontend or called from external scripts. Lingua-Align supports different treebank formats and includes additional tools for conversion and evaluation. In our experiments we can show that our tree aligner produces results with high quality and outperforms unsupervised techniques proposed otherwise. It also integrates well with another existing tool for manual tree alignment which makes it possible to quickly integrate additional training material and to run semi-automatic alignment strategies. 1.
French, like most Romance languages, displays both prenominal and postnominal placement of at-tributive adjectives. The fact that choice in position is not random led many linguists (Abeille and Godard (1999); Forsgren (1978); Wilmet (1981)) to propose constraints based on different dimensions of language (syntax, semantics, phonology, morphology, pragmatics...), but most of them are only trends and it is very difficult to draw a general picture of the phenomenon. We thus proposed in a previous work (Anonymous (tted)) a prediction model built on data from the French Treebank corpus (FTB) (Abeille et al. (2003)), along the lines of Bresnan et al. (2007), based on most of the proposed constraints, to test the contribution weight of each of them and the impact of their interaction. Even if results were encouraging, several facts were outlined: first, if introspection leads to the conclusion that most adjectives alternate, usage shows much more fixity: less than 10% of the adjectives in our data are actually in both positions. This suggests that locutors' mental representations for every item are in fact much stronger for a given position compared to its counterpart. Second, some contrainsts dont have a significant effect in our model. Yet, a qualitative exam showed that they are greatly correlated to a position for specific adjectives. For instance, different 'different', which appears equally in both positions, displays for each position a pattern linked to the nature of the determiner, whereas the type of determiner is not relevant at a more general level. More precisely, we observe a strong cooccurrence of a definite determiner with different in anteposition and a similar pattern between the indefinite determiner and the postposed adjective. This indicates that some constraints are relevant, but only for specific adjectives. Like (Bybee and Mcclelland (2005); Goldberg (2006); Croft (2001)) we believe that locutors' knowledge is based on much more specific information on the item, but also on the context in which it appears: formal characteristics of a particular sequence, frequency of use, distributionality... This work focuses on a qualitative study, on another corpus (ER (2010) 147,934,722 tokens), of the adjectives identified as displaying this alternation in the FTB, the aim being to better understand their functioning, and to propose a model that will better handle actual usage in the FTB. Our methodology is inspired by Gries and Stefanowitsch (2004): we search which lexical elements occur in a particular pattern and identify their attraction strength to the construction by means of a statistical analysis. Results show different types of behaviour on a continuum going from very fixed general patterns to more alternation. The first class of adjectives is very close to the fixed position adjectives: they massively prefer a given position, except in a few cases of idiomatic/collocational sequences. For instance, majeur 'major' is always postposed to the noun, unless it is in the sequence majeure partie 'most part'. In a second class, we still see a great preference for a given position but the alternating cases show less fixity. Two patterns appear within this class: the alternate order either corresponds to a use driven by one major constraint, or to a cumulation/interaction of different constraints. For the first pattern, the constraint involved is not necessarilly the same for every adjective. It can be semantically grounded, which usually leads to distinct nominal paradigms combining with the adjective given its position (e.g. ancien: 'ancien+N' means 'former', 'N+ancien' means 'old'), or it can be based on other devices, as illustrated by different. The data of the ER corpus differs from the FTB by the fact that it shows a strong preference for anteposition. The findings concerning the definite/indefinite nature of the NP were however confirmed, with 97,5% cases of definite in anteposition and 96% of indefinite in postposition. The adjective nouveau 'new' illustrates the second pattern: it prefers anteposition, but postposition is favoured when combined with a concrete noun, or ehanced when the NP is complement of a preposition. The third class concerns adjectives for which the general pattern shows a much weaker preference, if any, for one position over the other. There are however differences in usage for each position. The two patterns outlined for the preceeding class may also apply here: for instance, the placement of principal 'main' depends on a cumulation of information based on different grounds, whereas the semantics of pauvre (pitiful + N/N + not rich) clearly separates the uses. To sum up, the problem of adjective alternation does not depend on general principles, it is tightly linked to the item and to the NP within which it appears. Our study shows that the constraints previously proposed play a role in the placement of adjectives, but on a more specific level than the broad NP. In other words, locutors speak according to more specific patterns present in their linguistic knowledge.