Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
Social relations between humans critically depend on our affective experiences of others. Oxytocin enhances prosocial behavior, but its effect on humans' affective experience of others is not known. We tested whether oxytocin influences affective ratings, and underlying brain activity, of faces that have been aversively conditioned. Using a standard conditioning procedure, we induced differential negative affective ratings in faces exposed to an aversive conditioning compared with nonconditioning manipulation. This differential negative evaluative effect was abolished by treatment with oxytocin, an effect associated with an attenuation of activity in anterior medial temporal and anterior cingulate cortices. In amygdala and fusiform gyrus, this modulation was stronger for faces with direct gaze, relative to averted gaze, consistent with a relative specificity for socially relevant cues. The data suggest that oxytocin modulates the expression of evaluative conditioning for socially relevant faces via influences on amygdala and fusiform gyrus, an effect that may explain its prosocial effects.
Two Languages - One Annotation Scenario? Experience from the Prague Dependency Treebank This paper compares the two FGD-based annotation scenarios for Czech and for English, with the Czech as the basis. We discuss the secondary predication expressed by infinitive and its functions in Czech and English, respectively. We give a few examples of English constructions that do not have direct counterparts in Czech (e.g., tough movement and causative constructions with make, get, and have ), as well as some phenomena central in English but much less employed in Czech (object raising or control in adjectives as nominal predicates), and, last, structures more or less parallel both in their function and distribution, whose respective annotation differs due to significant differences in the respective linguistic traditions (verbs of perception).
The paper describes an approach to automati-cally annotate a Hindi Treebank using Pan-inian dependency framework. The annotator is a rule based system and the rules use certain syntactic cues available in a sentence. This automated annotation scheme aims at facilitat-ing manual annotation by reducing time and effort of manual annotators. Also, the aim of automatic annotation, among other things, is to increase the efficiency of a broad coverage constraint based Hindi parser. We also evalu-ate this tool and show its accuracy and cover-age. 1
In this paper, we will give an overview of the reconstruction process of the Swedish treebank Talbanken, created in the first half of the 70’s. Talbanken contains both written and spoken material, both encoded in the MAMBA-format. The goal has been to construct two new versions of the original data, one based on phrase structure and one on dependency structure. The outcome of the reconstruction, i.e. different versions of Talbanken, is available for non-commercial research and educational purposes. 1
Periods in the development of the lexical database in the Czech Language Institute, programmes.
Purpose Customer satisfaction is seen to be one of the main determinants of loyalty. However, the relationship between customer satisfaction and loyalty does not seem to be linear, many researchers have reported doubts about the predictability of loyalty solely due to customer satisfaction ratings which ignore image as predictor of loyalty. This paper aims to address the issues. Design/methodology/approach The authors report a study of ski resorts where they first established a causal model of customer satisfaction and image predicting customer loyalty, and then map the scores in a four‐fields‐grid. Additionally the authors conducted a moderator analysis to assess the relative importance of image and satisfaction for loyalty intentions between two different groups (first‐time‐visitors, and regular guests). Findings The results show that those ski resorts with the highest satisfaction ratings and the highest image ratings have the highest loyalty scores. Among first‐time‐visitors overall satisfaction is more important than image, with increasing number of repeat visits the importance of overall satisfaction declines and that of image relatively augments. Practical implications Besides measuring customer satisfaction, managers must assess also image ratings in order to get a realistic view of the loyalty intentions of their customer base. The scores can than be mapped together with the ratings of other ski resorts, and serve as a benchmark study. Originality/value Second order analysis of image (comprising three different dimensions), the image‐satisfaction‐grid, moderating effect of experience to relative importance of satisfaction and image on loyalty.
We present a new English→Czech machine translation system combining linguistically motivated layers of language description (as defined in the Prague Dependency Treebank annotation scenario) with statistical NLP approaches.
CONTEXT: Cognitive decline, mood, behavioral and sleep disturbances, and limitations of activities of daily living commonly burden elderly patients with dementia and their caregivers. Circadian rhythm disturbances have been associated with these symptoms. OBJECTIVE: To determine whether the progression of cognitive and noncognitive symptoms may be ameliorated by individual or combined long-term application of the 2 major synchronizers of the circadian timing system: bright light and melatonin. DESIGN, SETTING, AND PARTICIPANTS: A long-term, double-blind, placebo-controlled, 2 x 2 factorial randomized trial performed from 1999 to 2004 with 189 residents of 12 group care facilities in the Netherlands; mean (SD) age, 85.8 (5.5) years; 90% were female and 87% had dementia. INTERVENTIONS: Random assignment by facility to long-term daily treatment with whole-day bright (+/- 1000 lux) or dim (+/- 300 lux) light and by participant to evening melatonin (2.5 mg) or placebo for a mean (SD) of 15 (12) months (maximum period of 3.5 years). MAIN OUTCOME MEASURES: Standardized scales for cognitive and noncognitive symptoms, limitations of activities of daily living, and adverse effects assessed every 6 months. RESULTS: Light attenuated cognitive deterioration by a mean of 0.9 points (95% confidence interval [CI], 0.04-1.71) on the Mini-Mental State Examination or a relative 5%. Light also ameliorated depressive symptoms by 1.5 points (95% CI, 0.24-2.70) on the Cornell Scale for Depression in Dementia or a relative 19%, and attenuated the increase in functional limitations over time by 1.8 points per year (95% CI, 0.61-2.92) on the nurse-informant activities of daily living scale or a relative 53% difference. Melatonin shortened sleep onset latency by 8.2 minutes (95% CI, 1.08-15.38) or 19% and increased sleep duration by 27 minutes (95% CI, 9-46) or 6%. However, melatonin adversely affected scores on the Philadelphia Geriatric Centre Affect Rating Scale, both for positive affect (-0.5 points; 95% CI, -0.10 to -1.00) and negative affect (0.8 points; 95% CI, 0.20-1.44). Melatonin also increased withdrawn behavior by 1.02 points (95% CI, 0.18-1.86) on the Multi Observational Scale for Elderly Subjects scale, although this effect was not seen if given in combination with light. Combined treatment also attenuated aggressive behavior by 3.9 points (95% CI, 0.88-6.92) on the Cohen-Mansfield Agitation Index or 9%, increased sleep efficiency by 3.5% (95% CI, 0.8%-6.1%), and improved nocturnal restlessness by 1.00 minute per hour each year (95% CI, 0.26-1.78) or 9% (treatment x time effect). CONCLUSIONS: Light has a modest benefit in improving some cognitive and noncognitive symptoms of dementia. To counteract the adverse effect of melatonin on mood, it is recommended only in combination with light. TRIAL REGISTRATION: controlled-trials.com/isrctn Identifier: ISRCTN93133646.
Parser self-training is the technique of taking an existing parser, parsing extra data and then creating a second parser by treating the extra data as further training data. Here we apply this technique to parser adaptation. In particular, we self-train the standard Charniak/Johnson Penn-Treebank parser using unlabeled biomedical abstracts. This achieves an f-score of 84.3% on a standard test set of biomedical abstracts from the Genia corpus. This is a 20% error reduction over the best previous result on biomedical data (80.2% on the same test set).
We describe a parsing approach that makes use of the perceptron algorithm, in conjunction with dynamic programming methods, to recover full constituent-based parse trees. The formalism allows a rich set of parse-tree features, including PCFG-based features, bigram and trigram dependency features, and surface features. A severe challenge in applying such an approach to full syntactic parsing is the efficiency of the parsing algorithms involved. We show that efficient training is feasible, using a Tree Adjoining Grammar (TAG) based parsing formalism. A lower-order dependency parsing model is used to restrict the search space of the full model, thereby making it efficient. Experiments on the Penn WSJ treebank show that the model achieves state-of-the-art performance, for both constituent and dependency accuracy.
A nonhuman primate model was used to evaluate the value of the dexamethasone suppression test as an index of hypothalamic-pituitary-adrenal responsiveness to arousal. In 8 rhesus monkeys plasma cortisol was suppressed by dexamethasone in a dose-dependent fashion at doses between 0.75 and 33 microgram/kg. A replication study was performed 5 months later using a single dexamethasone dose (17 microgram/kg) known to produce maximal plasma cortisol suppression. This yielded highly correlated results (r = 0.91, p less than 0.005) suggesting that dexamethasone suppressibility may be a stable characteristic of individual animals. In 9 other animals whose arousal responses to a stressful procedure (nasogastric tube insertion) had been rated daily over a previous 3-month period, baseline plasma cortisol levels and the percent suppression of plasma cortisol by dexamethasone were evaluated. Baseline plasma cortisol levels did not significantly correlate with the degree of dexamethasone-suppression and the mean arousal ratings within animals. However, the postdexamethasone percent of baseline cortisol did correlate significantly (r = 0.75, p less than 0.025) with individual mean arousal ratings. These preliminary results suggest that assessment of the sensitivity of an individual's hypothalamic-pituitary glucocorticoid feedback system may be a better predictor than its baseline cortisol concentrations of its degree of behavioral arousal to stress.
Unlike previous emotional studies using functional neuroimaging that have focused on either locating discrete emotions in the brain or linking emotional response to an external behavior, this study investigated brain regions in order to validate a three-dimensional construct--namely pleasure, arousal, and dominance (PAD) of emotion induced by marketing communication. Emotional responses to five television commercials were measured with Advertisement Self-Assessment Manikins (AdSAM) for PAD and with functional magnetic resonance imaging (fMRI) to identify corresponding patterns of brain activation. We found significant differences in the AdSAM scores on the pleasure and arousal rating scales among the stimuli. Using the AdSAM response as a model for the fMRI image analysis, we showed bilateral activations in the inferior frontal gyri and middle temporal gyri associated with the difference on the pleasure dimension, and activations in the right superior temporal gyrus and right middle frontal gyrus associated with the difference on the arousal dimension. These findings suggest a dimensional approach of constructing emotional changes in the brain and provide a better understanding of human behavior in response to advertising stimuli.
We describe experiments on learning latent variable grammars for various German tree-banks, using a language-agnostic statistical approach. In our method, a minimal initial grammar is hierarchically refined using an adaptive split-and-merge EM procedure, giving compact, accurate grammars. The learning procedure directly maximizes the likelihood of the training treebank, without the use of any language specific or linguistically constrained features. Nonetheless, the resulting grammars encode many linguistically interpretable patterns and give the best published parsing accuracies on three German treebanks.
The suitability of different parsing methods for different languages is an important topic in syntactic parsing. Especially lesser-studied languages, typologically different from the languages for which methods have originally been developed, pose interesting challenges in this respect. This article presents an investigation of data-driven dependency parsing of Turkish, an agglutinative, free constituent order language that can be seen as the representative of a wider class of languages of similar type. Our investigations show that morphological structure plays an essential role in finding syntactic relations in such a language. In particular, we show that employing sublexical units called inflectional groups, rather than word forms, as the basic parsing units improves parsing accuracy. We test our claim on two different parsing methods, one based on a probabilistic model with beam search and the other based on discriminative classifiers and a deterministic parsing strategy, and show that the usefulness of sublexical units holds regardless of the parsing method. We examine the impact of morphological and lexical information in detail and show that, properly used, this kind of information can improve parsing accuracy substantially. Applying the techniques presented in this article, we achieve the highest reported accuracy for parsing the Turkish Treebank.
A number of researchers have recently conducted experiments comparing “deep” hand-crafted wide-coverage with “shallow” treebank- and machine-learning-based parsers at the level of dependencies, using simple and automatic methods to convert tree output generated by the shallow parsers into dependencies. In this article, we revisit such experiments, this time using sophisticated automatic LFG f-structure annotation methodologies with surprising results. We compare various PCFG and history-based parsers to find a baseline parsing system that fits best into our automatic dependency structure annotation technique. This combined system of syntactic parser and dependency structure annotation is compared to two hand-crafted, deep constraint-based parsers, RASP and XLE. We evaluate using dependency-based gold standards and use the Approximate Randomization Test to test the statistical significance of the results. Our experiments show that machine-learning-based shallow grammars augmented with sophisticated automatic dependency annotation technology outperform hand-crafted, deep, wide-coverage constraint grammars. Currently our best system achieves an f-score of 82.73% against the PARC 700 Dependency Bank, a statistically significant improvement of 2.18% over the most recent results of 80.55% for the hand-crafted LFG grammar and XLE parsing system and an f-score of 80.23% against the CBS 500 Dependency Bank, a statistically significant 3.66% improvement over the 76.57% achieved by the hand-crafted RASP grammar and parsing system.
We present a dependency-driven parser that parses both dependency structures and constituent structures. Constituency representations are automatically transformed into dependency representations with complex arc labels, which makes it possible to recover the constituent structure with both constituent labels and grammatical functions. We report a labeled attachment score close to 90% for dependency versions of the TIGER and TúBa-D/Z treebanks. Moreover, the parser is able to recover both constituent labels and grammatical functions with an F-Score over 75% for TüBa-D/Z and over 65% for TIGER.
Problematic types of prepositions, conjuctions, and particles and possible ways of their treatment in the lexical database.
This paper examines whether a learningbased coreference resolver can be improved using semantic class knowledge that is automatically acquired from a version of the Penn Treebank in which the noun phrases are labeled with their semantic classes. Experiments on the ACE test data show that a resolver that employs such induced semantic class knowledge yields a statistically significant improvement of 2 % in F-measure over one that exploits heuristically computed semantic class knowledge. In addition, the induced knowledge improves the accuracy of common noun resolution by 2-6%. 1
Abstract The area of probabilistic phrase structure parsing has been a central and active field in computational linguistics. Stochastic methods in natural language processing, in general, have become very popular as more and more resources become available. One of the main advantages of probabilistic parsing is in disambiguation: it is useful for a parsing system to return a ranked list of potential syntactic analyses for a string. In this article, we introduce probabilistic context‐free grammars (PCFGs) and outline some of their strengths and weaknesses. We concentrate on the automatic extraction of stochastic grammars from treebanks (large collections of hand‐corrected syntactic structures). We describe the current state of the field and the current research on improving the basic PCFG model. This includes lexicalized, history‐based and generative models. Finally, we briefly mention some research into probabilistic phrase structure parsing for domains other than traditional treebank text and languages other than English (Chinese, Arabic, German and French).
International audience
The purpose of this study is to show the reasons translators have problems related to lexical choices when translating euphemisms and dysphemisms. As euphemistic expression is used to make a concept less offensive and more acceptable and avoid possible loss of face, it tends to have an ambiguous meaning. This means that translating euphemisms and dysphemisms is not a matter of just the accuracy of translation. Therefore, these pragmatic factors such as face saving, the cooperative principle, situational context and politeness, each play a crucial role in lexical choice. Figurative expressions, circumlocutions, general-for-specific substitutions and part-for-whole substitutions are widely used in news media on purpose. In this particular text, which is read by various people and races, translators need to be careful when translating. Every culture has different norms and face-work strategies. Non-native speakers are often unaware of these differences and because of this may unintentionally cause offense. Since the choice of euphemism and dysphemism is determined within a given context, translating these expressions is not always successful. Consequently, we try to find the most desirable way of translating to eliminate strange meanings caused by a literal translation and convey figurative senses which are peculiar to SL. In addition, many more alternative expressions to euphemism and dysphemism need to be added to the dictionary.
07–484 Aceto, Michael (East Carolina U, USA; acetom@ecu.edu ), Statian Creole English: An English-derived language emerges in the Dutch Antilles. World Englishes (Blackwell) 25.3 & 4 (2006), 411–435. 07–485 Anchimbe, Eric A. (U Munich, Germany), World Englishes and the American tongue. English Today (Cambridge University Press) 22.4 (2006), 3–9. 07–486 Bartha, Csilla & Anna Borbély (Hungarian Academy of Sciences, Budapest, Hungary; bartha@nytud.hu ), Dimensions of linguistic otherness: Prospects of minority language maintenance in Hungary. Language Policy (Springer) 5.3 (2006), 337–365. 07–487 Coetzee-Van Rooy, Susan (North-West U, Potchefstroom, South Africa; basascvr@puk.ac.za ), Integrativeness: Untenable for world Englishes learners? World Englishes (Blackwell) 25.3 & 4 (2006), 437–450. 07–488 Gooskens, Charlotte (U Groningen, The Netherlands; c.s.gooskens@rug.nl ) & Renée van Bezooijen, Mutual comprehensibility of written Afrikaans and Dutch: Symmetrical or asymmetrical? Literary and Linguistic Computing (Oxford University Press) 21.4 (2006), 543–557. 07–489 Gooskens, Charlotte & Wilbert Heeringa (U Groningen, The Netherlands; c.s.gooskens@rug.nl ), The relative contribution of pronunciational, lexical, and prosodic differences to the perceived distances between Norwegian dialects. Literary and Linguistic Computing (Oxford University Press) 21.4 (2006), 477–492. 07–490 Guilherme, Manuela (U De Coimbra, Portgual), English as a Global language and education for cosmopolitan citizenship. Language and International Communication (Multilingual Matters) 7.1 (2007), 72–90. 07–491 Koscielecki, Marek (The Open U, Hongk Kong, China). Japanized English, its context and socio-historical background. English Today (Cambridge University Press) 22.4 (2006), 25–31. 07–492 Meilin, Chen (Three Gorges University, China) & Hu Xiaoqiong, Towards the acceptability of China English at home and abroad. English Today (Cambridge University Press) 22.4 (2006), 44–52. 07–493 Mesthrie, Rajend (U Cape Town, South Africa; raj@humanities.uct.ac.za ), World Englishes and the multilingual history of English. World Englishes (Blackwell) 25.3 & 4 (2006), 381–390. 07–494 Poole, Brian (Ministry of Manpower, Muscat, the Sultanate of Oman), Some effects of Indian English on the language as it is used in Oman. English Today (Cambridge University Press) 22.4 (2006), 21–24. 07–495 Robinson, Ian (U Calabria, Italy), Genre and loans: English words in an Italian newspaper. English Today (Cambridge University Press) 22.4 (2006), 9–20. 07–496 Ross, Kathryn (U Oxford, UK; kathryn.ross@trinity.ox.ac.uk ), Status of women in highly literate societies: The case of Kerala and Finland. Literacy (Blackwell) 40.3 (2006), 171–178. 07–497 Sala, Bonaventure M. (Cameroon), Does Cameroonian English have grammatical norms? English Today (Cambridge University Press) 22.4 (2006), 59–64. 07–498 Wei-Yu Chen, Cheryl (National Taiwan Normal U, Taiwan; wychen66@hotmail.com ), The mixing of English in magazine advertisements in Taiwan. World Englishes (Blackwell) 25.3 & 4 (2006), 467–478. 07–499 Wong, Jock (National U Singapore, Singapore; jockonn@hotmail.com ), Contextualizing aunty in Singaporean English. World Englishes (Blackwell) 25.3 & 4 (2006), 451–466. 07–500 Xiaoxia, Cui (Yunnan U, China), An understanding of ‘China English’ and the learning and use of the English language in China. English Today (Cambridge University Press) 22.4 (2006), 40–43. 07–501 Young, Ming Yee Carissa (Macao U Science & Technology, Macau; myyoung@must.edu.mo ), Macao students' attitudes toward English: A post-1999 survey. World Englishes (Blackwell) 25.3 & 4 (2006), 479–490.
Students learned teaching principles either with or without (control group) the presentation of a classroom exemplar in video or text format. Across 2 experiments, the video group produced higher transfer scores and affective ratings than the other groups. Four weeks later, the video group recalled more information about the exemplar than the text group, but no treatment effects were found on transfer. Qualitative analyses (Experiment 2) showed that the video group produced a significantly larger number of modeled behaviors in the transfer test than the text (immediate) and control (immediate and delayed) groups. Results encourage using classroom video exemplars to promote students ’ affect and retention, but suggest that additional pedagogies are needed to promote longer term transfer of theory into practice.
782 SEER, 85, 4, OCTOBER 2007 television is largely Prague-based may have a more profound bearing on the use of language than has generally been appreciated. Not surprisingly, this study has many of the strengths and some of the weaknesses of a typical doctoral thesis. It offers a comprehensive summary and evaluation of existing research and provides very useful cross-references. It also highlights the complexity of language usage in a linguistic settingwhere stylisticallyand functionally divergent forms coexist, and where theprestigious 'standard' variant is not the spoken norm. Most importantly, it offers new statistical information to add to the existing body of data on morphological, phonological and lexical variation, and to substantiate claims that language choice always depends to a significant extent on the purpose of the dialogue and the formality of the situation.However, minor problems with editing and proof-reading detract from the overall quality of thework. Furthermore, the selection of television broadcasts inevitably contains a degree of subjectivity and is not indicative of the speech of the population as a whole. Finally, itwould appear that a lack of space may have prevented the author from developing some of hermore interesting ideas, such as the notion thatwomen may be treated differendy tomen in the television studio, and that this may be reflected in theiruse of language. In summary, despite some shortcomings, this is an original and stimulating study,which is of relevance to all scholars of language variation and change, and presents considerable scope for further research. School of Humanities, Languages and Social Sciences Tom Dickins Universityof Wolverhampton Pushkin, Alexander. 'TheGypsies' and Other NarrativePoems. Translated, with an introduction and notes, by Antony Wood. Engravings by Simon Brett. Angel Books, London, 2006. xl + 116pp. Notes.?14.95. Anyone who has ever attempted to translate nineteenth-century Russian verse into English should make a point of turning to the Afterword of Antony Wood's new book. Subtided 'Pushkin's Voice inEnglish', it is a pithy credo from one of the UK's leading translators of verse. One statement in particular should be writ large above any translator's desk: 'the rise of translation theory in recent decades has not been accompanied by the emergence of any sub stantial body of translation of Pushkin's verse that has impressed as verse in English' (p. 114). Wood makes clear how he intends to remedy thisdeficiency. To begin with he is uncontroversial. He will eschew alternating masculine and feminine rhymes as being too difficult to achieve inRussian. He will have recourse to half-rhymes, since rhymes are far easier to find in an inflected language than in an uninflected language. His other points, however, are more contentious. He is clearly no enthusiast for translations which reproduce exactly themetre and rhyme scheme of the original, considering that their effect 'tends to be self-conscious, self-satisfied,unengaged and disembodied' (p. in). Nor does he think that the number of lines of the original should necessarily be maintained. reviews 783 The Afterword is one of the items added toWood's earlier work 'The Bridegroom', with 'Count Nulin} and 'TheTale of the Golden CockereT,published by Angel Books in 2002 and reviewed in this journal (vol. 82, July 2004, no. 3). These three poems are reproduced here with slight amendments, one of which, fromThe Bridegroom,isparticularly felicitous.Whereas in 2002 we find in the tenth stanza of the poem 'and then a jet/Over Natasha's head', the revised translation reads 'then splash a/Dash of iton Natasha. The new book is some twice the length of the earlier book. The new translations are Pushkin's first 'problem' poema,The Gypsiesand the skazka,The Tale of the Dead Princess and theSevenChampions.True to his credo,Wood does not attempt to replicate Pushkin's iambic tetrameter throughout his transla tion of The Gypsies. His favoured departure from this involves removing the initial unstressed syllable and turning the line into trochaic tetrameter. There are numerous examples of the type 'Life resounds on every side' (p. 3 ). In addition there are variants of this variant, all scrupulously noted in the Afterword. These departures from Pushkin's metre are clearly no oversight and Wood shows considerable expertise in producing...
This paper investigates probability distributions of dependency distances in six texts ex- tracted from a Chinese dependency treebank. The fitting results reveal that the investigated distribu- tion can be well captured by the right truncated Zeta distribution. In order to restrict the model only to natural language, two samples with randomly generated governors are investigated. One of them can be described e.g. by the Hyperpoisson distribution, the other satisfies the Zeta distribution. The paper also presents a study on sequential plot and mean dependency distance of six texts with three analyses (syntactic, and two random). Of these three analyses, syntactic analysis has a minimum (mean) dependency distance.
Proceedings of the Sixth International Workshop on Treebanks and \nLinguistic Theories. \nEditors: Koenraad De Smedt, Jan Hajič and Sandra Kübler. \nNEALT Proceedings Series, Vol. 1 (2007), 1-6. \n© 2007 The editors and contributors. \nPublished by \nNorthern European Association for Language \nTechnology (NEALT) \nhttp://omilia.uio.no/nealt. \nElectronically published at \nTartu University Library (Estonia) \nhttp://hdl.handle.net/10062/4476.
We present the Modified French Treebank (MFT), a completely revamped French Treebank, derived from the Paris 7 Tree-bank (P7T), which is cleaner, more co-herent, has several transformed structures, and introduces new linguistic analyses. To determine the effect of these changes, we investigate how theMFT fares in statistical parsing. Probabilistic parsers trained on the MFT training set (currently 3800 trees) already perform better than their counter-parts trained on five times the P7T data (18,548 trees), providing an extreme ex-ample of the importance of data quality over quantity in statistical parsing. More-over, regression analysis on the learning curve of parsers trained on the MFT lead to the prediction that parsers trained on the full projected 18,548 tree MFT training set will far outscore their counterparts trained on the full P7T. These analyses also show how problematic data can lead to problem-atic conclusions–in particular, we find that lexicalisation in the probabilistic parsing of French is probably not as crucial as was once thought (Arun and Keller (2005)).
In this paper we present several use cases for the Stockholm TreeAligner, a software tool originally designed for annotating the alignments in a parallel treebank. The tool has been extended and improved to the point that it can now also serve as a general tool for browsing and searching monolingual and parallel treebanks. Among the use cases presented are: building a parallel treebank, browsing mono- and bilingual treebanks, consistency checking using the search function, comparing PP-attachment in different languages, and viewing different versions of the same treebank. A demonstration of the software will be held during the workshop.
This paper presents recent extensions to Poliqarp, an open source tool for indexing and searching morphosyntactically annotated corpora, which turn it into a tool for indexing and searching certain kinds of treebanks, complementary to existing treebank search engines. In particular, the paper discusses the motivation for such a new tool, the extended query syntax of Poliqarp and implementation and efficiency issues.
Proceedings of the Sixth International Workshop on Treebanks and \nLinguistic Theories. \nEditors: Koenraad De Smedt, Jan Hajič and Sandra Kübler. \nNEALT Proceedings Series, Vol. 1 (2007), 189-200. \n© 2007 The editors and contributors. \nPublished by \nNorthern European Association for Language \nTechnology (NEALT) \nhttp://omilia.uio.no/nealt. \nElectronically published at \nTartu University Library (Estonia) \nhttp://hdl.handle.net/10062/4476.
Functional Arabic Morphology is a formulation of the Arabic inflectional system seeking the working interface between morphology and syntax. ElixirFM is its high-level implementation that reuses and extends the Functional Morphology library for Haskell. Inflection and derivation are modeled in terms of paradigms, grammatical categories, lexemes and word classes. The computation of analysis or generation is conceptually distinguished from the general-purpose linguistic model. The lexicon of ElixirFM is designed with respect to abstraction, yet is no more complicated than printed dictionaries. It is derived from the open-source Buckwalter lexicon and is enhanced with information sourcing from the syntactic annotations of the Prague Arabic Dependency Treebank.
The interpretation of nominal compounds is one of the most difficult problems in natural language processing. This paper proposes a new model for the automatic classification of four coarse-grained semantic relations involved in Chinese compound nominalizations. In such a model, for a compound nominalization, its paraphrased syntactic role occurrences (PSRO) in a treebank are exploited to form feature vectors for supervised classifiers. To solve the problem of data sparseness, the World Wide Web is used to discover relational clusters and such clusters are employed to produce smoothed PSRO feature vectors for the compound nominalizations. The experimental results show that such a method is very effective.
Abstract Measurement of an individual’s subjective experience of emotion has long been a key component of emotion research, but it presents some unique challenges. Researchers have developed a number of different methods to assess the subjective emotional experiences of study participants, each of which has its strengths and limitations. Self-report measures such as the Positive and Negative Affect Schedule (PANAS; Watson, Clark, & Tellegen, 1988) are well established and easy to complete and provide useful information; however, administration of any written measure necessitates an interruption in the flow of an experiment and does not allow frequent or continuous sampling of affective states. More involved methods, such as interviewing, give a detailed and comprehensive picture of a person’s emotions, but they are time-consuming and can provide only a retrospective report of affect. The interactive computer version of the Self-Assessment Manikin (SAM) Scales (Bradley & Lang, 1994) allows online assessment of both emotional valence and arousal levels, but it, too, does not provide a continuous record of affect.
%XLOGLQJ WKH &URDWLDQ 'HSHQGHQF\
The Referentiebestand Nederlands (RBN) is a lexical database for the Dutch language. Although its main objective is to function as a lexical resource for the production of bilingual dictionaries with Dutch as either source or target language, the RBN is designed as a multi-purpose lexical database. As such, the RBN is successfully used for the production of bilingual dictionaries as well as in the domain of language technology. Being a multi-purpose lexical database the RBN has a flexible structure and it contains more information about the entries than is strictly needed for the production of bilingual dictionaries. In this paper we describe the lexical information in the database and some of the principles and choices underlying its design. We will discuss some aspects of the macrostructure and we will elaborate on the microstructure by describing the lexical information about nouns and verbs. Special attention will be given to some of the characteristic properties of the RBN: the handling (of polysemy by using meaning shifts, the minute description of complementation patterns for verbs and the description of the many examples of combinations, like idioms and collocations.
This article examines the corpus of multinationals’ codes of conduct on CSR issues which has been collated by the ILO. Through lexical software analysis we identify three main points of reference in CSR codes of conduct: respect for ILO norms, discussion of the company’s relationship to society, and reinforcement of its internal discipline and organisation. Surprisingly, the issue of corporate responsibility itself constitutes a small part of the text of the codes. Their main targets are employees, who are charged with a dual task: to ensure the implementation of the principles stated in the codes, and to protect the assets of the company. In a reflexive dimension, codes of conduct help us to understand the key characteristics of the companies which made them.
Databases of hierarchically annotated text occupy a central place in linguistic research and language technology development. We describe a new approach to tree query which we call "Query by Annotation". Users express a query by annotating a tree, and the annotation is compiled into an expression in a path language. The result trees are overlaid with the original query, permitting the user to see why they match. Since queries and results are annotated trees, users can easily refine and resubmit their queries. The approach to Query by Annotation is motivated and exemplified using databases of linguistic trees, or treebanks.
While heavy lexical borrowing can pose a problem to any approach to linguistic prehistory, it has often been regarded as an especially difficult problem for lexicostatistics, especially in such areas as Australia, where some believe that extensive borrowing is the norm. The present paper applies lexicostatistics to what is arguably the most massive case of borrowing known for Australia, namely between the Jingulu and Mudburra languages of the Northern Territory, and finds that it actually leads to what is generally considered the correct genetic classification of these languages. This result is then shown to depend on certain relationships among the lexicostatistical percentages that may not always obtain in other cases of heavy borrowing.