Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Collins’ widely-used parsing models treat noun phrases (NPs) in a different manner to other constituents. We investigate these differences, using the recently released internal NP bracketing data (Vadas and Curran, 2007a). Altering the structure of the Treebank, as this data does, has a number of consequences, as parsers built using Collins’ models assume that their training and test data will have structure similar to the Penn Treebank’s. Our results demonstrate that it is difficult for Collins’ models to adapt to this new NP structure, and that parsers using these models make mistakes as a result. This emphasises how important treebank structure itself is, and the large amount of influence it can have.
Both classroom instruction and lexical database development stand to benefit from applied research on sign language, which takes into consideration American Sign Language rules, pedagogical issues, and teacher characteristics. In this study of technical science signs, teachers' experience with signing and, especially, knowledge of content, were found to be essential for the identification of signs appropriate for instruction. The results of this study also indicate a need for a systematic approach to examine both sign selection and its impact on learning by deaf students. Recommendations are made for the development of lexical databases and areas of research for optimizing the use of sign language in instruction. (PsycINFO Database Record (c) 2016 APA, all rights reserved)
Collaborative and content-based filtering are the recommendation techniques most widely adopted to date. Traditional collaborative approaches compute a similarity value between the current user and each other user by taking into account their rating style, that is the set of ratings given on the same items. Based on the ratings of the most similar users, commonly referred to as neighbors, collaborative algorithms compute recommendations for the current user. The problem with this approach is that the similarity value is only computable if users have common rated items. The main contribution of this work is a possible solution to overcome this limitation. We propose a new content-collaborative hybrid recommender which computes similarities between users relying on their content-based profiles, in which user preferences are stored, instead of comparing their rating styles. In more detail, user profiles are clustered to discover current user neighbors. Content-based user profiles play a key role in the proposed hybrid recommender. Traditional keyword-based approaches to user profiling are unable to capture the semantics of user interests. A distinctive feature of our work is the integration of linguistic knowledge in the process of learning semantic user profiles representing user interests in a more effective way, compared to classical keyword-based profiles, due to a sense-based indexing. Semantic profiles are obtained by integrating machine learning algorithms for text categorization, namely a naïve Bayes approach and a relevance feedback method, with a word sense disambiguation strategy based exclusively on the lexical knowledge stored in the WordNet lexical database. Experiments carried out on a content-based extension of the EachMovie dataset show an improvement of the accuracy of sense-based profiles with respect to keyword-based ones, when coping with the task of classifying movies as interesting (or not) for the current user. An experimental session has been also performed in order to evaluate the proposed hybrid recommender system. The results highlight the improvement in the predictive accuracy of collaborative recommendations obtained by selecting like-minded users according to user profiles. (PsycINFO Database Record (c) 2016 APA, all rights reserved)
This paper presents a scheme for ranking of spelling error corrections for Urdu. Conventionally spell-checking techniques do not provide any explicit ranking mechanism. Ranking is either implicit in the correction algorithm or corrections are not ranked at all. The research presented in this paper shows that for Urdu, phonetic similarity between the corrections and the erroneous word can serve as a useful parameter for ranking the corrections. This combined with a new technique Shapex that uses visual similarity of characters for ranking gives an improvement of 23% in the accuracy of the one-best match compared to the result obtained when the ranking is done on the basis of word frequencies only.
This paper presents a new research and development project called Papillon [Planas00]. It is a French-Japanese cooperation between laboratories GETA/CLIPS (Grenoble, France) and NII (Tokyo, Japan). Its goal is to build a French-English-Japanese multilingual lexical database by using interlingual links and to extract from it digital bilingual French-Japanese and Japanese-French dictionaries. These dictionaries will be available under the terms of an open source license. This project, initiated by some computational linguists, aims at being useful and open to all those who are interested in Japanese and French. A seminar was organized on the 10-12 of August 2000 in Tokyo [Planas00]. It was devoted to discussions aiming at reaching a general consensus on the structure and content of the database, and to decide some technical aspects of database development, i.e. database configuration, contents of the entries and link between the entries. Introduction There are few French-Japanese usa...
The lexicon of the modern Norwegian bokmal standard needs a better description and documentation than what is the situation today. The article presents a plan for building a modern lexical database based on a balanced corpus of 40 million words of modern bokmal. This base should serve as a source for a traditional scientific dictionary as well as a dictionary for language technological applications.
Vernacularisation et traduction des textes pragmatiques en Afrique — La traduction des textes comportant des lacunes d'ordre grammatical, lexical, stylistique ou idiomatique présente habituellement des difficultés particulières, lesquelles sont amplifiées lorsqu'elles sont attribuables à la vernacularisation d'une langue étrangère. Dans les sociétés postcoloniales, l'absence ou la non-disponibilité des études linguistiques sur la plupart des langues locales rend ardue l'analyse des interférences entre ces dernières et les langues officielles étrangères. Cette situation, ajoutée à la grande diversité ethnolinguistique ambiante, ne facilite pas l'interprétation des textes produits par les personnes semi-lettrées. Le traducteur de ces textes se présente davantage comme un rédacteur qui, à partir de l'idée globale qui se dégage de l'original, conçoit et produit un texte répondant aux normes de la langue cible. L'évaluation d'un tel travail ne peut se faire qu'en comparant la finalité des deux textes.
espanolEn este estudio hemos querido trazar un primer acercamiento al estudio de la lengua poetica del autor italiano Tommaso Stigliani y lo hemos hecho analizando su mayor obra, Il Mondo Nuovo (1628), poema epico sobre el descubrimiento de America. Hemos analizado los principales fenomenos de la ortografia, la fonetica, la morfologia, la sintaxis y el lexico que aparecen en el texto, teniendo en cuenta la norma linguistica del siglo XVII y las opiniones y teorias de los principales linguistas de la epoca. Con este estudio, por tanto, ofrecemos un acercamiento a la concepcion de lengua poetica que Tommaso Stigliani, a traves de su poema y de sus escritos teoricos, puso de manifiesto a lo largo de su carrera literaria. EnglishIn this study, we have attempted to make inroads into the study of the poetic language of the Italian author, Tommaso Stigliani, analysing his principal work, Il Mondo Nuovo (1628), an epic poem on the discovery of America. We have examined the main phenomena of orthography, phonetics, morphology, syntax and lexicon that appear in the text, taking into account the linguistic norms of the 17th century, as well as the opinions and theories of the major linguistic experts of that period. Our study, therefore, gives an insight into Tommaso Stigliani?s conception of poetic language, which, through his poem and his theoretical writings, he maintained for the duration of his literary career
This paper addresses the question of how to obtain consistent semantic annotation on the basis of a set of noisy texts. Many potential real-world applications of semantic computing are faced with the need to handle texts which are not well-edited, and for which a resource-intensive treebanking effort is not feasible. Student-produced short answers contain many grammatical and lexical errors, making consistent annotation a challenge. Nevertheless, this paper demonstrates that semantic role annotation can be done in a consistent and useful manner even under these constraints.
Youngster violence/violence on youngsters. Crossed views on the perception of verbal violence among immigrant school populations Among the different forms of violence, from and addressed to the youth, those exerted within the school framework are many and various. The form that has more particularly caught our attention, is the so-called normative violence connected to the language practices among immigrant school populations.Treating the issue of violence in relation to the norm, whether it be of a linguistic kind or another, is legitimate as far as the former constitutes and destroys the latter. Indeed, the school presents itself as « the place favouring a struggle to impose linguistic norms ». The way of speaking of school actors – a.o. learners and teachers – is an indicator of the relationship that the latter have with social rules in general and school norms in particular. Beyond the description of obscenities, language coarseness and vulgarity of the former, and of the French language « in the manner of Charles-Henri » of the others, these forms of verbal violence in the school framework are worth thinking about.Moreover, the social imagery related to immigration is so virulent that the equation « practice of languages different from "correct" French (the so-called "bon usage"), immigration and delinquency » may seem astonishing to some while it is admitted by others. The practice of the language would explain acts of delinquency: to act on the language would be preventive. « Young immigrants had better behave themselves », or even « they’d better talk correctly »! The eradication of violence seems to be at the cost of this simple solution.It is precisely beyond this naive optimism that we would like to go in this article, centred on the notion of youngster verbal violence, in the way it is perceived and lived in the school framework. To do so, the contribution of sociolinguistics can help in order to delimit and redefine language practices and to supply didactics of French with useful marks for efficient school norm teaching.
Abstract. We describe a lexical database consisting of morphologically and phonetically tagged words that occur in the texts primarily used for
Ratings of familiarity and pronounceability were obtained for a sample of 199 names and 199 nouns. Frequency and familiarity were more closely related in the proper name pool than the word pool, although the correlation was modest in both cases. Familiarity and pronounceability were highly related for both names and nouns. Although word-level models of speech recognition have become very successful in the past several years due to the great increase in computer capacity and processing speed, even the most successful models generally require a great deal of specific training in order to reach high levels of recognition. For items such as proper names, which may number in the tens of thousands and have multiple pronunciations, it is impractical, if not impossible, to train on the entire set. Name recognition is of considerable practical interest given the possibilities of building acoustic interfaces to telephone directories or library catalogs, and shows promise as an area of great overlap between human and machine word recognition. Names occupy a unique position in lexical access: they have no inherent meaning thus they require phonological (sound-based) recognition; on the other hand, proper names have aspects of frequency and familiarity that may allow them to act like words already in the lexicon. A promising immediate approach to designing a name recognition system is to incorporate statistical aspects of proper names (frequency and familiarity) directly. There exists relatively little data on the distribution of proper names in the language (see however (l)), and there are a number of reasons to suppose that words and names will be rated differently on frequency and/or familiarity. We expect that those names that are familiar will also be easy to pronounce and, importantly for the computational aspect, ones that will lead to relatively little variability in pronunciation. Less familiar names may be more difficult to pronounce and result in more varied pronunciations. The data below are a first effort at obtaining reliable familiarity ratings for a random sample of surnames.
We argue in favor of the the use of labeled directed graph to represent various types of linguistic structures, and illustrate how this allows one to view NLP tasks as graph transformations. We present a general method for learning such transformations from an annotated corpus and describe experiments with two applications of the method: identification of non-local depenencies (using Penn Treebank data) and semantic role labeling (using Proposition
In this paper, we describe the application of a bidirectional dependency parser trained on the Turin University Treebank.
This paper describes the mutually beneficial relationship between a cultural heritage digital library and a historical treebank: an established digital library can provide the resources and structure necessary for efficiently building a treebank, while a treebank, as a language resource, is a valuable tool for audiences traditionally served by such libraries. 1
This paper addresses the question of how to obtain consistent semantic annotation on the basis of a set of noisy texts. Many potential real-world applications of semantic computing are faced with the need to handle texts which are not well-edited, and for which a resource-intensive treebanking effort is not feasible. Student-produced short answers contain many grammatical and lexical errors, making consistent annotation a challenge. Nevertheless, this paper demonstrates that semantic role annotation can be done in a consistent and useful manner even under these constraints.
This was the first article explicitly on the theory of agency published in a regular, i.e., nonproceedings, issue of a journal in social science. The paper presents a fiduciary function model of policing in agency, with an application to attempts to influence regulatory performance by policing the behavior of regulators. Four types of agents - the pure fiduciary, lexical fiduciary, lexical self-interest agent, and pure self-interest agents - are identified. The paper notes that the rational principal would not police his agent if he did not expect a net gain from the attempt; this is one of the key logics of agency theory. The paper notes the effects of the fiduciary norm in economizing on specification and policing (agency) costs. An apparent paradox can occur when policing the agent appears to lower rather than increase the return to the principal. In other words, agent fidelity does not necessarily correlate with the level of principal return. In the context of public regulation, this can take the form of producing a more honest or better-behaved regulatory agent in a government that produces a poorer return to the public interest. The posted version of this article contains some corrections of errors/omissions introduced by the journal's publisher in the publication process.
The interpretation of nominal compounds is one of the most difficult problems in natural language processing. This paper proposes a new model for the automatic classification of four coarse-grained semantic relations involved in Chinese compound nominalizations. In such a model, for a compound nominalization, its paraphrased syntactic role occurrences (PSRO) in a treebank are exploited to form feature vectors for supervised classifiers. To solve the problem of data sparseness, the World Wide Web is used to discover relational clusters and such clusters are employed to produce smoothed PSRO feature vectors for the compound nominalizations. The experimental results show that such a method is very effective.
The purpose of this chapter is to provide a two-dimensional approach to language documentation (Hi mmelman 1998). In addition to building a database, we also conducted a sociolinguistic survey des igned to document the state of health of a language in a particular spatio-temporal frame. Our goa l is to share our fieldwork experience of documenting Kavalan, a seriously endangered language in sou theastern Taiwan now spoken by fewer than just a few dozen speakers. We first discuss our field exp eriences in working with speakers of Kavalan in Sinshe village, the only significant Kavalan set tlement left in Taiwan, and the state of the Kavalan language, based in part on Huang and Cha ng’s (19 95) earlier sociolinguistic survey, and in part on a recent more in-depth village-wide survey of lan guage use in the community. Next, we introduce the NTU Corpus of Formosan Languages, part of which incorporates our corpus data in Kavalan. The NTU Corpus of Formosan Languages aims to establish a standard for the creation of linguistic corpus databases through the application of information technology to linguistic research. The creation of this linguistic database enables us both to preserve valuable linguistic data and to provide a systematic recording of these languages, for the benefit of future linguistic research.
An ontology is a main component of an evolving knowledge base that caters to multiple clients. Consider a scenario where an automated procedure (a computer vision algorithm) used in an analyst tool detects different kinds of “roads” in images, and features in the ontology are used to distinguish a “paved ” road from a “dirt road”. In another scenario, the ontology enables reasoning about “locations”, supporting analysts' geospatial information processing tasks. In this paper, we describe the creation of a multiuse geospatial and visual information ontology, GVIO 1, building on and integrating with the lexical database, WordNet. To ensure that GVIO can interoperate with other ontologies in useful ways, we inherit as much of the WordNet structure and content as is relevant for the domain of aerial surveillance and link in new content/structure as necessary. 1.
Men receive conflicting messages about their sexual roles in heterosexual relationships. Men are socialized to initiate and direct sexual activities with women; yet societal norms also proscribe the sexual domination and coercion of women. The authors test these competing hypotheses by assessing whether men inhibit the link between sex and dominance. In Studies 1a and b, using a subliminal priming procedure embedded in a lexical decision task, the authors demonstrate that men automatically suppress the concept of dominance following exposure to subliminal sex primes relative to neutral primes. In Studies 2 and 3, the authors show that men who are less likely to perceive sexual assertiveness as necessary, to refrain from dominant sexual behavior, and who do not invest in masculine gender ideals are more likely to inhibit dominant thoughts following sex primes. Implications for theories of automatic cognitive networks and gender-based sexual roles are discussed.
Emotions are an important factor in human-computer interaction. One of the challenges in building emotionally intelligent systems is the automatic recognition of affective states. We are developing and evaluating a method for measuring user affect that incorporates psychological, behavioral, and physiological measures. During affective stimulation, breathing parameters, skin conductance level (SCL) and corrugator EMG activity correlate with self-reported levels of valence and arousal. Valence, at the level of subjective experience, summarizes how well one is doing, whereas arousal refers to a sense of energy. A stimulus activates appetitive or defensive motivation (the valence dimension) with some degree of energy mobilization (the arousal dimension). In the laboratory, moods are induced using different procedures. Only few studies have investigated the critical question of how long induced moods actually last. Further, knowledge concerning the persistence of physiological responding, when the stimuli are withdrawn, remains spare. The goals of this study were first, to assess the somato-physiological activity during affective stimulation (film clip viewing) and its relation to valence and arousal, and second, to determine if the response patterns persist, dissipate or otherwise change when the stimuli are withdrawn and the subjects perform a computer task. Seventy-six participants viewed a neutral film clip (an educational program) and completed immediately afterward the task (control condition). Then, they viewed either a positive high-arousal clip (sport scenes), a positive low-arousal clip (takes from landscapes), a negative high-arousal clip (scene depicting captives forced to play Russian roulette) or a negative low-arousal clip (a documentary about a slum in Belgium) and completed the task a second time (experimental condition). The task required participants to shop on an e-commerce website for office supplies. After each clip and each task, the participants rated their current mood. We tested valence and arousal effects during the last 90 s of the films, the first and last 90 s of the task of the experimental condition. Viewing of the selected film clips resulted in increasing defensive and appetitive activation in the expected ways both subjectively and physiologically. Corrugator EMG activity was higher at the end of the negative clips than the positive clips, and minute ventilation and SCL were higher for the arousing clips than for the less arousing clips. After the approximately 9-minute task, people who viewed the negative clips still reported more negative valence than those who viewed the positive clips. On the contrary, there were no differences in the arousal ratings. The valence effect in the mood state was paralleled by valence effects in the somato-physiological measures during the task. Increased facial frowning at the end of the negative clips was maintained during the task indicating persistence of defensive activation. SCL was lower for the negative film groups, especially at the end of the task, suggesting that sympathetic activation was lowered in subjects in the negative mood as compared with subjects in the positive mood. The findings of this study have several implications. First, they enrich our knowledge concerning the relationships between subjective feelings and their physiological correlates. Second, they inform us about the effectiveness of film clips as a mood induction instrument. Third, they suggest that induced changes in arousal are quickly overridden by the degree of activation "imposed" by the cognitive task, whereas induced changes in valence are more resistant and thus likely to impact the execution of the task. Finally, they show which physiological measures may be useful in tracking mood states during human-computer interaction (i.e., corrugator EMG activity and SCL). Feedback from these parameters could be used by computers to recognize mood states in the user.
Assessments of the quality of parts of syntactic grammars of natural languages are useful for the validation of their construction. We extended a grammar of French determiners that takes the form of a recursive transition network and evaluated its quality. The result of the application of this local grammar gives deeper syntactic information than chunking or information available in treebanks. We performed the evaluation by comparison with a corpus independently annotated with information on determiners. We obtained 85 % precision and 93 % recall on text not tagged for parts of speech.
An increased interest in body image (more specifically, becoming or staying thin) is a common trend that is increasing among children and adolescents. Previous research shows that the interest to become or remain thin peaks during early adolescence, particularly among females. PURPOSE: Because research on this topic is limited in younger populations, the purpose of this study was to examine the interest in weight control for a population of preadolescent children. METHODS: Subjects included 261 (122 female and 139 male) children from third (n=92), fourth (n=80), and fifth (n=89) grades. Average age of the participants was 9.5 years. The primary investigator met one-on-one with each child. Height was measured to the nearest centimeter and mass to the nearest 1/2 kilogram. Each child was asked if they judged themselves to be overweight (fat), underweight (skinny) or in-between. Also, each child was asked if they would like to lose weight, gain weight, or stay the same. Categorical data were separated through cross tabulation and significant differences were assessed with Chi-Square analyses. RESULTS: Average body mass index (BMI) for girls was 18.2 (just below the 75th percentile for age and gender) and for boys was 18.7 (the 75th percentile for age and gender). Thirty nine percent of all participants wanted to lose weight or remain thin. For boys and girls respectively, significant percentages 23.7% and 30.3% wanted to loose weight. Moreover, 33.8% and 43.4% of boys and girls respectively wanted to remain or become thin in spite of “in-between” self image ratings and normal BMIs (X2=19.742, p=0.001 and X2=16.418, p=0.003 for boys and girls respectively). No significant differences were noted between grades 3, 4, and 5 for boys or girls. CONCLUSIONS: The results of this study suggest that preadolescent children are focusing on body image and specifically on wanting to become or remain thin. This is especially of interest as these children had normal BMIs and generally viewed themselves as having “in-between” body images. The data also suggest that while interest in body image may peak in early adolescence, it clearly begins for both boys and girls in early prepubescent years.
Reinforcing value of a behavior refers to the motivation to engage in the behavior. A reinforcing behavior will support more work to obtain the behavior. Individual differences in the reinforcing value of physical activity predict the usual physical activity in children. Another factor that may influence physical activity is liking of physical activity. Liking or hedonics refers to an affective rating associated with the behavior, and people are more likely to engage in physical activities that they like than ones that they do not like. Liking correlates with physical activity in youth. Although the independence of reinforcing value and liking of physical activity has not yet been tested, the motivation to gain access to a behavior and liking for that behavior are likely different constructs. PURPOSE: To determine whether liking and relative reinforcing value (RRV) of physical activity independently predict time youth spend in moderate-to-vigorous physical activity (MVPA). METHODS: Boys (n = 21) and girls (n = 15) age 8 to 12 years were measured for height, weight, aerobic fitness, liking and RRV of physical activity, and minutes in MVPA using accelerometers. RESULTS: Using multiple regression to control for individual differences in age, sex, BMI percentile, aerobic fitness, and time the accelerometer was worn, liking (P < 0.05) and RRV (P < 0.01) of physical activity independently predicted time in MVPA. When using median splits of the RRV and liking data to form subject groups, the group of children with both a high liking and RRV of physical activity participated in greater (P < 0.05) minutes per week of MVPA (1340 + 72 min) than groups with high RRV-low liking (1040 + 95 min), low RRV-high liking (978 + 89 min), or low RRV-low liking (1007 + 70 min) of physical activity. CONCLUSIONS: RRV and liking of MVPA are separate constructs as they independently predict MVPA of children. Those children who find physical activity the most reinforcing and also have a high liking of physical activity engage in 33% more MVPA than children who either find physical activity highly reinforcing or have a high liking of physical activity. Interventions that concurrently increase the reinforcing value and liking of physical activity may be the most effective for increasing youth participation in free-living MVPA. Supported by NIH Grant RO1 HD42766.
This paper reports on experiments in frame-semantic annotation of a parallel treebank. Selected English and Swedish sentences that contained verbs of motion and communication were annotated independently by two annotators. We found that they assigned the same frame to corresponding sentences in 52% of the cases. This leads us to the conclusion that parallel treebanks can save considerable effort when building semantically annotated resources.
This article presents an algorithm for translating the Penn Treebank into a corpus of Combinatory Categorial Grammar (CCG) derivations augmented with local and long-range word-word dependencies. The resulting corpus, CCGbank, includes 99.4% of the sentences in the Penn Treebank. It is available from the Linguistic Data Consortium, and has been used to train wide-coverage statistical parsers that obtain state-of-the-art rates of dependency recovery. In order to obtain linguistically adequate CCG analyses, and to eliminate noise and inconsistencies in the original annotation, an extensive analysis of the constructions and annotations in the Penn Treebank was called for, and a substantial number of changes to the Treebank were necessary. We discuss the implications of our findings for the extraction of other linguistically expressive grammars from the Treebank, and for the design of future treebanks.
A real-time user-independent emotion detection system using physiological signals has been developed. The system has the ability to classify affective states into 2-dimensions using valence and arousal. Each dimension ranges from 1 to 5 giving a total of 25 possible affective regions. Physiological signals were measured using 3 biometric sensors for Blood Volume Pulse (BVP), Skin Conductance (SC) and Respiration (RESP). Two emotion inducing experiments were conducted to acquire physiological data from 13 subjects. The data from 10 of these subjects were used to train the system, while the remaining 3 datasets were used to test the performance of the system. A recognition rate of 62% for valence and 67% for arousal was achieved within +/- 1 units of the valence and arousal rating.
Proceedings of the 16th Nordic Conference \nof Computational Linguistics NODALIDA-2007. \nEditors: Joakim Nivre, Heiki-Jaan Kaalep, Kadri Muischnek and Mare Koit. \nUniversity of Tartu, Tartu, 2007. \nISBN 978-9985-4-0513-0 (online) \nISBN 978-9985-4-0514-7 (CD-ROM) \npp. 152-159.
We present a simple history-based model for sentence generation from LFG f-structures, which improves on the accuracy of previous models by breaking down PCFG independence assumptions so that more f-structure conditioning context is used in the prediction of grammar rule expansions. In addition, we present work on experiments with named entities and other multi-word units, \nshowing a statistically significant improvement of generation accuracy. Tested on section 23 of the PennWall Street Journal Treebank, the techniques described in this paper improve BLEU scores from 66.52 to 68.82, and coverage from 98.18% to 99.96%.
The Penn Treebank does not annotate within base noun phrases (NPs), committing only to flat structures that ignore the complexity of English NPs. This means that tools trained on Treebank data cannot learn the correct internal structure of NPs. This paper details the process of adding gold-standard bracketing within each noun phrase in the Penn Treebank. We then examine the consistency and reliability of our annotations. Finally, we use this resource to determine NP structure using several statistical approaches, thus demonstrating the utility of the corpus. This adds detail to the Penn Treebank that is necessary for many NLP applications.
Lists of names are an important knowledge source for many systems which carry out named entity recognition. It is shown that augmenting hand-crafted lists with those derived from corpora can improve their performance. Two methods for improving automatically acquired lists are presented. The best corpus-derived lists are shown to out-perform the hand-crafted ones by 4%. 1. Introduction Named entity (NE) recognition is the process of identifying and categorising names in text. NE recognition and corpora research can be mutually benficial. Corpora are often more valuable when linguistic information has been added to them. For example the SUZANNE and Penn TreeBank corpora contain texts which have been parsed and as a consequence these corpora are widely used resources in NLP research. In a similar fashion NE recognition can be used to annotate the names in texts and thereby produce a richer corpus. Conversely, the information in annotated corpora can be very useful in the development of...
In morphologically rich languages, should morphological and syntactic disambiguation be treated sequentially or as a single problem? We describe several efficient, probabilisticallyinterpretable ways to apply joint inference to morphological and syntactic disambiguation using lattice parsing. Joint inference is shown to compare favorably to pipeline parsing methods across a variety of component models. State-of-the-art performance on Hebrew Treebank parsing is demonstrated using the new method. The benefits of joint inference are modest with the current component models, but appear to increase as components themselves improve. 1
BACKGROUND: Emotion theory holds that unpleasant events prime withdrawal actions, whereas pleasant events prime approach actions. Recent studies have suggested that passive viewing of emotion eliciting images results in postural adjustments, which become manifest as changes in body center of pressure (COP) trajectories. From those studies it appears that posture is modulated most when viewing pictures with negative valence. The present experiment was conducted to test the hypothesis that pictures with negative valence have a greater impact on postural control than neutral or positive ones. Thirty-four healthy subjects passively viewed a series of emotion eliciting images, while standing either in a bipedal or unipedal stance on a force plate. The images were adopted from the International Affective Picture System (IAPS). We analysed mean and variability of the COP and the length of the associated sway path as a function of emotion. RESULTS: The mean position of the COP was unaffected by emotion, but unipedal stance resulted in overall greater body sway than bipedal stance. We found a modest effect of emotion on COP: viewing pictures of mutilation resulted in a smaller sway path, but only in unipedal stance. We obtained valence and arousal ratings of the images with an independent sample of viewers. These subjects rated the unpleasant images as significantly less pleasant than neutral images, and the pleasant images as significantly more pleasant than neutral images. However, the subjects rated the images as overall less pleasant and less arousing than viewers in a closely comparable American study, pointing to unknown differences in viewer characteristics. CONCLUSION: Overall, viewing emotion eliciting images had little effect on body sway. Our finding of a reduction in sway path length when viewing pictures of mutilation was indicative of a freezing strategy, i.e. fear bradycardia. The results are consistent with current knowledge about the neuroanatomical organization of the emotion system and the neural control of behavior.
As an extension of decades of syntactic theorizing, treebanks have inherited a small set of phrasal categories, which abstract over the environments that the categories can occur in. Extending ideas from Johnson (1998), we explore encoding information from the local tree context in each category. We then discuss two clustering techniques which preserve the distributionally relevant category distinction, forming linguistically relevant generalizations and improving PCFG parsing performance. 1
In this paper I discuss some phenomena of lexical semantics using data from D.Dobrovol’skij’s approaches and the results of the analysis of the combinatorial profile of Russian звать. This word makes up the core of a lexical class ( вызвать, позвать, etc.) and I analyse their use in various contexts. The comparison of the combinatorial profiles of these words proves to be an efficient instrument for defining the combinatorial norms and helps in the lexicographic description of these words.
The purpose of this study is to show the reasons translators have problems related to lexical choices when translating euphemisms and dysphemisms. As euphemistic expression is used to make a concept less offensive and more acceptable and avoid possible loss of face, it tends to have an ambiguous meaning. This means that translating euphemisms and dysphemisms is not a matter of just the accuracy of translation. Therefore, these pragmatic factors such as face saving, the cooperative principle, situational context and politeness, each play a crucial role in lexical choice. Figurative expressions, circumlocutions, general-for-specific substitutions and part-for-whole substitutions are widely used in news media on purpose. In this particular text, which is read by various people and races, translators need to be careful when translating. Every culture has different norms and face-work strategies. Non-native speakers are often unaware of these differences and because of this may unintentionally cause offense. Since the choice of euphemism and dysphemism is determined within a given context, translating these expressions is not always successful. Consequently, we try to find the most desirable way of translating to eliminate strange meanings caused by a literal translation and convey figurative senses which are peculiar to SL. In addition, many more alternative expressions to euphemism and dysphemism need to be added to the dictionary.
Language deviation is a linguistic device of purposeful violation of language norms,which serves not only as the necessity but also as the inevitability of language development.There are various forms of language deviation,including phonological deviation,lexical deviation,grammatical deviation,semantic deviation,graphological deviation,deviation of register and figurative deviation.It is the reflection of variety of language and its social nature.Language deviation does bring richer cultural connotation to every language.
This article examines the corpus of multinationals’ codes of conduct on CSR issues which has been collated by the ILO. Through lexical software analysis we identify three main points of reference in CSR codes of conduct: respect for ILO norms, discussion of the company’s relationship to society, and reinforcement of its internal discipline and organisation. Surprisingly, the issue of corporate responsibility itself constitutes a small part of the text of the codes. Their main targets are employees, who are charged with a dual task: to ensure the implementation of the principles stated in the codes, and to protect the assets of the company. In a reflexive dimension, codes of conduct help us to understand the key characteristics of the companies which made them.
Legal English is a special mode of expression and norm developed by a long practice of judicature in common law countries and has its own special features.This paper discusses lexical and grammatical features of legal English.Legal English is difficult to understand not only in mixture of daily words,professional words,borrowed words and ancient words,and accumulation of synonyms and antonyms but also in redundancy of sentence and complication of conception.
XARA is a rule-based PropBank labeler for Alpino XML files, written in Java. I used XARA in my research on semantic role labeling in a Dutch corpus to bootstrap a dependency treebank with semantic roles. Rules in XARA are based on XPath expressions, which makes it a versatile tool that is applicable to other treebanks as well.