Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
We present a novel method for evaluating the output of Machine Translation (MT), based on comparing the dependency structures of the translation and reference rather than their surface string forms. Our method uses a treebank-based, widecoverage, probabilistic Lexical-Functional Grammar (LFG) parser to produce a set of structural dependencies for each translation-reference sentence pair, and then calculates the precision and recall for these dependencies. Our dependency-based evaluation, in contrast to most popular string-based evaluation metrics, will not unfairly penalize perfectly valid syntactic variations in the translation. In addition to allowing for legitimate syntactic differences, we use paraphrases in the evaluation process to account for lexical variation. In comparison with other metrics on 16,800 sentences of Chinese-English newswire text, our method reaches high correlation with human scores. An experiment with two translations of 4,000 sentences from Spanish-English Europarl shows that, in contrast to most other metrics, our method does not display a high bias towards statistical models of translation.
Recent work identifies two properties that appear particularly relevant to the characterization of graph-based dependency models of syntactic structure: the absence of interleaving substructures (well-nestedness) and a bound on a type of discontinuity (gap-degree ≤ 1) successfully describe more than 99% of the structures in two dependency treebanks (Kuhlmann and Nivre 2006). Bodirsky et al. (2005) establish that every dependency structure with these two properties can be recast as a lexicalized Tree Adjoining Grammar (LTAG) derivation and vice versa. However, multi-component extensions of TAG (MC-TAG), argued to be necessary on linguistic grounds, induce dependency structures that do not conform to these two properties (Kuhlmann and Möhl 2006). In this paper, we observe that several types of MC-TAG as used for linguistic analysis are more restrictive than the formal system is in principle. In particular, tree-local MC-TAG, tree-local MC-TAG with flexible composition (Kallmeyer and Joshi 2003), and special cases of set-local TAG as used to describe certain linguistic phenomena satisfy the well-nested and gap degree ≤ 1 criteria. We also observe that gap degree can distinguish between prohibited and allowed wh-extractions in English, and report some preliminary work comparing the predictions of the graph approach and the MC-TAG approach to scrambling.
Background Irritable bowel syndrome (IBS) is conceptualized as a syndrome of enhanced central stress circuit responsiveness, likely associated with altered adrenergic and autonomic responses. Aims (1) To determine if baseline autonomic nervous system (ANS) differences exist in IBS versus control subjects; (2) to determine group differences in ANS response to yohimbine (YOH) and clonidine (CLO); and (3) to determine group differences in the effects of YOH and CLO on affective states. Methods IBS and control subjects were enrolled. ANS and affective measures were taken before and after drug ingestion. ANS was measured by HRV (high frequency, HF, indicating cardiovagal and low/high frequency ratio, LF/HF, indicating sympathetic) and systolic BP. Affective ratings were made with the Stress Symptom Rating Questionnaire. Results Mean baseline HF was lower in IBS versus controls (34.29 nu vs 62.88 nu, p =.013). Mean baseline LF/HF was higher in IBS versus controls (2.59 vs 0.91, p =.057). No significant HRV changes were seen in response to CLO or YOH in either group. YOH significantly increased BP (p =.01) and CLO significantly reduced BP (p =.02) in pooled subjects; no group difference in BP was seen. No group differences in affective ratings were seen. In combined groups, YOH increased anxiety (p =.05) and CLO led to increased fatigue (p = 0.01) with decreased arousal (p =.02). Conclusion These findings confirm that there is greater baseline sympathetic and lower parasympathetic activity in IBS patients compared with controls. A larger sample of patients is likely needed to elucidate group differences in drug responses.
Our paper reports an attempt to apply an unsupervised clustering algorithm to a Hungarian treebank in order to obtain semantic verb classes. Starting from the hypothesis that semantic metapredicates underlie verbs' syntactic realization, we investigate how one can obtain semantically motivated verb classes by automatic means. The 150 most frequent Hungarian verbs were clustered on the basis of their complementation patterns, yielding a set of basic classes and hints about the features that determine verbal subcategorization. The resulting classes serve as a basis for the subsequent analysis of their alternation behavior.
Proceedings of the Sixth International Workshop on Treebanks and \nLinguistic Theories. \nEditors: Koenraad De Smedt, Jan Hajič and Sandra Kübler. \nNEALT Proceedings Series, Vol. 1 (2007), 31-42. \n© 2007 The editors and contributors. \nPublished by \nNorthern European Association for Language \nTechnology (NEALT) \nhttp://omilia.uio.no/nealt. \nElectronically published at \nTartu University Library (Estonia) \nhttp://hdl.handle.net/10062/4476.
This paper reports about our efforts in creating a tri-lingual parallel treebank. The focal points are consistency checking and all aspects of sub-sentential alignment. We discuss the alignment guidelines, the importance of quality checks, and special alignment problems. Then we look at alignment algorithms and alignment visualization tools and we compare our own TreeAligner with other alignment tools. Our constituent structure treebanks contain just over 1,000 sentences and around 18,000 tokens in each language.
The correct attachment of prepositional phrases (PPs) is a central disambiguation problem when parsing natural languages. This paper compares the baseline situation for French as exemplified in the Le Monde treebank with earlier findings for English, German and Swedish. We perform uniform treebank queries and show that the noun attachment rate for French prepositions is strongly influenced by the preposition de which is by far the most frequent preposition and has a strong tendency for noun attachment. We therefore also compute the noun attachment rate for the other prepositions separately as well as for the many complex prepositions that are explicitly marked in this treebank.
Insights from corpus linguistics have come to be seen as having a significant impact in second language pedagogy. Learner corpora, or collections of texts spoken or written by non-native speakers (NNS) of a language, are now being used for the purposes of enhancing language teaching. Specifically, by comparing the corpus of NNS with native speakers (NS) it is possible to identify instances of learners' under- or over-use of spoken vocabulary, as well as to investigate how far, and in what ways learners deviate from NS's norms. This preliminary study compares a small NNS corpus of 43,651 words with an established NS corpus. Results revealed that the Japanese NNS differed markedly in many areas, especially in their underuse of certain lexical items such as discourse markers, modal items, adjectives for specific evaluations, some interactive words, delexical verbs and terms for marking vagueness and hedges. On the other hand, NNS overused some high frequency and auxiliary verbs and some common adjectives. From this data, we suggest that knowledge from learner corpora has important pedagogical implications which include giving higher priority to certain classes of vocabulary including multi-word clusters that appear to be underused among Japanese learners.
Proceedings of the Sixth International Workshop on Treebanks and \nLinguistic Theories. \nEditors: Koenraad De Smedt, Jan Hajič and Sandra Kübler. \nNEALT Proceedings Series, Vol. 1 (2007), 189-200. \n© 2007 The editors and contributors. \nPublished by \nNorthern European Association for Language \nTechnology (NEALT) \nhttp://omilia.uio.no/nealt. \nElectronically published at \nTartu University Library (Estonia) \nhttp://hdl.handle.net/10062/4476.
The amount of information available on the Web and in Digital Libraries is increasing over time. In this context, the role of user modeling and personalized information access is becoming crucial: Users need a personalized support in sifting through large amounts of retrieved information according to their interests. Information filtering and retrieval systems relying on this idea adapt their behavior to individual users by learning their preferences during the interaction in order to construct a profile of the user that can be later exploited in the search process. We propose a novel technique to learn user profiles which exploits word sense disambiguation based on the WordNet lexical database, in an attempt to produce semantic user profiles that might discover topics semantically closer to the user interests. Semantic profiles are used in the definition of a retrieval model that turns the traditional document-query search paradigm into a novel document-query-profile paradigm. As an example of this paradigm, we present an extension of the vector space model in which profiles are used to modify the ranking of search results obtained in response to a query, hopefully putting personally relevant items on the top of the result list. Experimental results in a movie retrieval scenario indicate that the proposed model to personalize Web search is effective.
Abstract This paper considers lexical combinations of choice functions where at least one is interpreted as arising from a norm. It is shown that in for all possibilities in which a norm is present, in general final choice may be consistent with preference optimization, but that it need not be so. It is concluded therefore that a fruitful approach to understanding the effect of norms on choice is to consider particular classes of norms rather than norms in general as in the work by Wulf Gaertner among others.
We describe some challenges of adaptation in the 2007 CoNLL Shared Task on Domain Adaptation. Our error analysis for this task suggests that a primary source of error is differences in annotation guidelines between treebanks. Our suspicions are supported by the observation that no team was able to improve target domain performance substantially over a state of the art baseline. 1
This paper presents a uniform approach to data extraction from syntactically annotated corpora encoded in XML. XQuery, which incorporates XPath, has been designed as a query language for XML. The combination of XPath and XQuery offers flexibility and expressive power, while corpus specific functions can be added to reduce the complexity of individual extraction tasks. We illustrate our approach using examples from dependency treebanks for Dutch.
The ability to detect similarity in conjunct heads is potentially a useful tool in helping to disambiguate coordination structures - a difficult task for parsers. We propose a distributional measure of similarity designed for such a task. We then compare several different measures of word similarity by testing whether they can empirically detect similarity in the head nouns of noun phrase conjuncts in the Wall Street Journal (WSJ) treebank. We demonstrate that several measures of word similarity can successfully detect conjunct head similarity and suggest that the measure proposed in this paper is the most appropriate for this task.
We study the correlations in the connectivity patterns of large scale syntactic dependency networks. These networks are induced from treebanks: their vertices denote word forms which occur as nuclei of dependency trees. Their edges connect pairs of vertices if at least two instance nuclei of these vertices are linked in the dependency structure of a sentence. We examine the syntactic dependency networks of seven languages. In all these cases, we consistently obtain three findings. Firstly, clustering, i.e., the probability that two vertices which are linked to a common vertex are linked on their part, is much higher than expected by chance. Secondly, the mean clustering of vertices decreases with their degree — this finding suggests the presence of a hierarchical network organization. Thirdly, the mean degree of the nearest neighbors of a vertex x tends to decrease as the degree of x grows—this finding indicates disassortative mixing in the sense that links tend to connect vertices of dissimilar degrees. Our results indicate the existence of common patterns in the large scale organization of syntactic dependency networks.
We present an unsupervised linguistically-based approach to discourse relations recognition, which uses publicly available resources like manually annotated corpora (Discourse Graph Bank, Penn Discourse TreeBank, RST-DT), as well as empirically derived data from “causally” annotated lexica like LCS, to produce a rule-based algorithm. In our approach we use the subdivision of Discourse Relations into four subsets – CONTRAST, CAUSE, CONDITION, ELABORATION, proposed by [1] in their paper where they report results obtained with a machine-learning approach from a similar experiment against which we compare our results. Our approach is fully symbolic and is partially derived from the system called GETARUNS, for text understanding, adapted to a specific task: recognition of Discourse Causal Relations in free text. We show that in order to achieve better accuracy both in the general task and in the specific one, semantic information needs to be used besides syntactic structural information. Our approach outperforms results reported in previous papers [2].
Based on the language of 17th century Bosnian Franciscan literature and enriched with features of the Neo-Štokavian folklore koine, the language of the 18th century writers represents a consistent system. Although it was not subject to willful codification, the language of the 18th century writers has codification elements. This is primarily implied by functional distribution, pronounced independence from common speech, especially on the syntactic level, compulsory use for all users and specific prescriptiveness in grammar handbooks of the time.
The purpose of the Chinese PropBank (CPB) project is to add a layer of annotation to the hand-parsed sentences in the Chinese Treebank (CTB) (Xue et al., 2005). This layer of annotation assigns predicatespecific argument labels to the constituents in a parse tree. The arguments of each predicate in the sentence, which are limited to verbs and their nominalizations in the work we report here, receive an argument label in the form of ArgN, where N is an integer between 0 and 5. These numbered arguments represent core arguments that are defined in relation to the predicate, which is labeled as Rel. Each core argument plays a unique role with regard to the predicate and generally the total number of core arguments for each predicate does not exceed 6. The core arguments annotated for the verb N (”investigate”) in Example (1) are the NPs (”the police”) and (”accident”) I(”cause”), which are labeled as Arg0 and Arg1 respectively. The semantic role labels added to the parse tree are in bold.
Functional Arabic Morphology is a formulation of the Arabic inflectional system seeking the working interface between morphology and syntax. ElixirFM is its high-level implementation that reuses and extends the Functional Morphology library for Haskell. Inflection and derivation are modeled in terms of paradigms, grammatical categories, lexemes and word classes. The computation of analysis or generation is conceptually distinguished from the general-purpose linguistic model. The lexicon of ElixirFM is designed with respect to abstraction, yet is no more complicated than printed dictionaries. It is derived from the open-source Buckwalter lexicon and is enhanced with information sourcing from the syntactic annotations of the Prague Arabic Dependency Treebank. MorphoTrees is the idea of building effective and intuitive hierarchies over the information provided by computational morphological systems. MorphoTrees are implemented for Arabic as an extension to the TrEd annotation environment based on Perl. Encode Arabic libraries for Haskell and Perl serve for processing the non-trivial and multi-purpose ArabTEX notation that encodes Arabic orthographies and phonetic transcriptions in parallel.
We present in this paper methods to improve HMM-based part-of-speech (POS) tagging of Mandarin. We model the emission probability of an unknown word using all the characters in the word, and enrich the standard left-to-right trigram estimation of word emission probabilities with a right-to-left prediction of the word by making use of the current and next tags. In addition, we utilize the RankBoost-based reranking algorithm to rerank the N-best outputs of the HMMbased tagger using various n-gram, morphological, and dependency features. Two methods are proposed to improve the generalization performance of the reranking algorithm. Our reranking model achieves an accuracy of 94.68 % using n-gram and morphological features on the Penn Chinese Treebank 5.2, and is able to further improve the accuracy to 95.11 % with the addition of dependency features. 1
Collins’ widely-used parsing models treat noun phrases (NPs) in a different manner to other constituents. We investigate these differences, using the recently released internal NP bracketing data (Vadas and Curran, 2007a). Altering the structure of the Treebank, as this data does, has a number of consequences, as parsers built using Collins’ models assume that their training and test data will have structure similar to the Penn Treebank’s. Our results demonstrate that it is difficult for Collins’ models to adapt to this new NP structure, and that parsers using these models make mistakes as a result. This emphasises how important treebank structure itself is, and the large amount of influence it can have.
In 1993 the Ministers of Education in the Netherlands and Flanders decided to install a binational committee of experts in order to co-ordinate, streamline, improve and stimulate the production of bilingual dictionaries and lexical databases with Dutch as a source or target language. This committee, called Commissie voor Lexicografische Vertaalvoorzieningen (Committee for Interlingual Lexicographical Resources) or CLVV, has, under the presidency of W. Martin, set up Action Plans involving some twenty dictionary projects which have been finished or are nearly finished by now. In this article the general policy lines of the CLVV are presented next to the criteria for the selection of language pairs, the infrastructure used, the results obtained and the lessons to be drawn from this ‘Dutch’ approach. The article also serves as a framework in which to situate the articles that follow.
We present two methods to address the problem of sparsity in the FrameNet lexical database. The first method is based on the idea that a word that belongs to a frame is ``similar'' to the other words in that frame. We measure the similarity using a WordNet-based variant of the Lesk metric. The second method uses the sequence of synsets in WordNet hypernym trees as feature vectors that can be used to train a classifier to determine whether a word belongs to a frame or not. The extended dictionary produced by the second method was used in a system for FrameNet-based semantic analysis and gave an improvement in recall. We believe that the methods are useful for bootstrapping FrameNets for new languages.
Vernacularisation et traduction des textes pragmatiques en Afrique — La traduction des textes comportant des lacunes d'ordre grammatical, lexical, stylistique ou idiomatique présente habituellement des difficultés particulières, lesquelles sont amplifiées lorsqu'elles sont attribuables à la vernacularisation d'une langue étrangère. Dans les sociétés postcoloniales, l'absence ou la non-disponibilité des études linguistiques sur la plupart des langues locales rend ardue l'analyse des interférences entre ces dernières et les langues officielles étrangères. Cette situation, ajoutée à la grande diversité ethnolinguistique ambiante, ne facilite pas l'interprétation des textes produits par les personnes semi-lettrées. Le traducteur de ces textes se présente davantage comme un rédacteur qui, à partir de l'idée globale qui se dégage de l'original, conçoit et produit un texte répondant aux normes de la langue cible. L'évaluation d'un tel travail ne peut se faire qu'en comparant la finalité des deux textes.
Iraj Mirza’s poetry occupies a special place in Persian literature as compared to the works of other poets of his time, due to his almost unrivalled use of language and rhetorics. Deviating from syntactic, semantic and pragmatic norms, he creates a new atmosphere with the simple language he uses, which draws his poetry close to the language of nature. The present paper examines Iraj Mirza’s poetry in terms of language function and his expert play with language. He deviates from the accepted linguistic norms of syntax, semantics and pragmatics, with an artistic courage, creating a new atmosphere in literary language: It is worth mentioning that he does so with such a simple language that one can claim, without unnecessary exaggeration that his poems are closer to the language of the nature than those of his contemporary poets. This article studies some of the language functions of Iraj's poems, revealing a small part of his skill in playing with the language.
BACKGROUND: Approved for treatment of treatment-resistant depression and for epilepsy, vagus nerve stimulation (VNS) therapy involves stimulation of the vagus nerve, affecting both mood and appetite regulating systems. VNS is associated with changes in food intake and weight loss in animals. Studies of its impact on food intake and weight with humans are limited. It is not known whether or how VNS influences emotional response to food, but vagus afferents project to regions in the insula involving satiety and taste. METHOD: Thirty-three participants were recruited for three groups: depressed patients undergoing VNS therapy, depressed patients not undergoing VNS therapy, and healthy controls. All participants viewed images of foods twice in random order. When applicable, VNS devices were turned on for one viewing and off for the other. Participants were instructed to rate immediately after the viewings how each picture made them feel on a visual analog on three dimensions (unhappy to happy, calm to aroused, and small/submissive to big/domineering). RESULTS: Controlling for time since last meal, a significant main effect was found for arousal ratings in response to sweet food images. Post-hoc analyses indicated that the VNS group demonstrated significant changes in arousal ratings between paired food image viewings compared to controls. Sixty-four percent of VNS participants demonstrated increases and 36% demonstrated decreases in arousal. Higher body mass indexes and greater levels of self-reported sweet cravings were associated with increased arousal during VNS activation. CONCLUSIONS: This study was the first to examine the effects of acute left cervical VNS on emotional ratings of food in adults with major depression. Results suggest that VNS device activation may be associated with acute alteration in arousal response to sweet foods among depressed patients. Future research is needed to replicate these findings and to assess how activation of the vagus nerve affects eating and weight.
One of the goals of natural language processing (NLP) systems is determining the meaning of what is being transmitted. Although much work has been accomplished in traditional written and spoken language domains, little has been performed in the newer computer-mediated communication domain enabled by the Internet, to include text-based chat. This is due in part to the fact that there are no annotated chat corpora available to the broader research community. The purpose of our research is to build a chat corpus, initially tagged with lexical and discourse information. Such a corpus could be used to develop stochastic NLP applications that perform tasks such as conversation thread topic detection, author profiling, entity identification, and social network analysis. During the course of our research, we preserved 477,835 chat posts and associated user profiles in an XML format for future investigation. We privacy-masked 10,567 of those posts and part-of-speech tagged a total of 45,068 tokens. Using the Penn Treebank and annotated chat data, we achieved part-of-speech tagging accuracy of 90.8%. We also annotated each of the privacy-masked corpus's 10,567 posts with a chat dialog act. Using a neural network with 23 input features, we achieved 83.2% dialog act classification accuracy.
Presentation of new tools used for processing of the Czech Lexical Database as a main source for preparation of various dictinaries.
The traditional English text chunking approach identifies phrases by using only one model and phrases with the same types of features. It has been shown that the limitations of using only one model are that: the use of the same types of features is not suitable for all phrases, and data sparseness may also result. In this paper, a divide-conquer strategy is proposed and applied in the identification of English phrases. And then, this strategy is rapid transplanted to Chinese text chunking. This strategy divides the task of chunking into several sub-tasks according to sensitive features of each phrase and identifies different phrases in parallel. Then, a two-stage decreasing conflict strategy is used to synthesize each sub-task's answer, where the main features are: one, each phrase uses its own sensitive features; two, avoidance of data sparseness. Through testing on public corpus (English) and Chinese Penn Treebank (Chinese), F score of English chunking achieves to 95.14% and that of Chinese chunking is 95.23%. These results are state of the art with the best results that have been reported..
We compare the accuracy of a statistical parse ranking model trained from a fully-annotated portion of the Susanne treebank with one trained from unlabeled partially-bracketed sentences derived from this treebank and from the Penn Treebank. We demonstrate that confidence-based semi-supervised techniques similar to self-training outperform expectation maximization when both are constrained by partial bracketing. Both methods based on partially-bracketed training data outperform the fully supervised technique, and both can, in principle, be applied to any statistical parser whose output is consistent with such partial-bracketing. We also explore tuning the model to a dieren t domain and the eect of in-domain data in the semi-supervised training processes.
%XLOGLQJ WKH &URDWLDQ 'HSHQGHQF\
Proceedings of the 16th Nordic Conference \nof Computational Linguistics NODALIDA-2007. \nEditors: Joakim Nivre, Heiki-Jaan Kaalep, Kadri Muischnek and Mare Koit. \nUniversity of Tartu, Tartu, 2007. \nISBN 978-9985-4-0513-0 (online) \nISBN 978-9985-4-0514-7 (CD-ROM) \npp. 81-88.
Since noun phrases are the most popular phrases in texts, noun phrase identification is one of vital subtasks of natural language processing. Generally Chinese noun phrases have hierarchical inner structures. This paper proposes an approach of defining various levels of granularity for noun phrases, catering for different application demands. Three levels of granularity noun phrases are proposed, that is, concept noun phrase, base noun phrase and entire noun phrase. The task of noun phrase identification is to label word sequences with phrase tags. All granularity noun phrase identifications are cast as classification problem under certain encoding schemes. The experimental dataset is acquired empirically from Chinese Penn Treebank 5.1. F, measure of concept noun phrase, base noun phrase and entire noun phrase identification reaches 92.12%, 84.13% and 85.32% respectively.
There are many methods to improve performance of statistical parsers. Resolving structural ambiguities is a major task of these methods. In the proposed approach, the parser produces a set of n-best trees based on a feature-extended PCFG grammar and then selects the best tree structure based on association strengths of dependency word-pairs. However, there is no sufficiently large Treebank producing reliable statistical distributions of all word-pairs. This paper aims to provide a self-learning method to resolve the problems. The word association strengths were automatically extracted and learned by parsing a giga-word corpus. Although the automatically learned word associations were not perfect, the constructed structure evaluation model improved the bracketed f-score from 83.09% to 86.59%. We believe that the above iterative learning processes can improve parsing performances automatically by learning word-dependence information continuously from web.
Students learned teaching principles either with or without (control group) the presentation of a classroom exemplar in video or text format. Across 2 experiments, the video group produced higher transfer scores and affective ratings than the other groups. Four weeks later, the video group recalled more information about the exemplar than the text group, but no treatment effects were found on transfer. Qualitative analyses (Experiment 2) showed that the video group produced a significantly larger number of modeled behaviors in the transfer test than the text (immediate) and control (immediate and delayed) groups. Results encourage using classroom video exemplars to promote students ’ affect and retention, but suggest that additional pedagogies are needed to promote longer term transfer of theory into practice.
Abstract We introduce a robust and shallow approach for grammatical role labeling (dependency labeling) where data-driven and theory-driven aspects are combined in a principled way. A classifier provides empirically justified weights, linguistic theory contributes well-motivated global restrictions, both are combined under the regiment of optimization. The empirical results of our approach are promising. However, we have made idealized assumptions (small inventory of dependency relations and treebank-derived chunks) that clearly must be replaced by a realistic setting.
We aim to improve the performance of a syntactic parser that uses a part-of-speech (POS) tagger as a preprocessor. Pipelined parsers consisting of POS taggers and syntactic parsers have several advantages, such as the capability of domain adaptation. However the performance of such systems on raw texts tends to be disappointing as they are affected by the errors of automatic POS tagging. We attempt to compensate for the decrease in accuracy caused by automatic taggers by allowing the taggers to output multiple answers when the tags cannot be determined reliably enough. We empirically verify the effectiveness of the method using an HPSG parser trained on the Penn Treebank. Our results show that ambiguous POS tagging improves parsing if outputs of taggers are weighted by probability values, and the results support previous studies with similar intentions. We also examine the effectiveness of our method for adapting the parser to the GENIA corpus and show that the use of ambiguous POS taggers can help development of portable parsers while keeping accuracy high. 1
Visual stimuli are judged for their emotional significance based on two fundamental dimensions, valence and arousal, and may lead to changes in neural and body functions like attention, affect, memory and heart rate. Alterations in behaviour and mood have been encountered in patients with Parkinson's disease (PD) undergoing functional neurosurgery, suggesting that electrical high-frequency stimulation of the subthalamic nucleus (STN) may interfere with emotional information processing. Here, we use the opportunity to directly record neuronal activity from the STN macroelectrodes in patients with PD during presentation of emotionally laden and neutral pictures taken from the International Affective Picture System (IAPS) to further elucidate the role of the STN in emotional processing. We found a significant event-related desynchronization of STN alpha activity with pleasant stimuli that correlated with the individual valence rating of the pictures. Our findings suggest involvement of the human STN in valence-related emotional information processing that can potentially be altered during high-frequency stimulation of the STN in PD leading to behavioural complications.
One may need to build a statistical parser for a new language, using only a very small labeled treebank together with raw text. We argue that bootstrapping a parser is most promising when the model uses a rich set of redundant features, as in recent models for scoring dependency parses (McDonald et al., 2005). Drawing on Abney’s (2004) analysis of the Yarowsky algorithm, we perform bootstrapping by entropy regularization: we maximize a linear combination of conditional likelihood on labeled data and confidence (negative Rényi entropy) on unlabeled data. In initial experiments, this surpassed EM for training a simple feature-poor generative model, and also improved the performance of a feature-rich, conditionally estimated model where EM could not easily have been applied. For our models and training sets, more peaked measures of confidence, measured by Rényi entropy, outperformed smoother ones. We discuss how our feature set could be extended with cross-lingual or cross-domain features, to incorporate knowledge from parallel or comparable corpora during bootstrapping. 1
We examine the problem of choosing word order for a set of dependency trees so as to minimize total dependency length. We present an algorithm for computing the optimal layout of a single tree as well as a numerical method for optimizing a grammar of orderings over a set of dependency types. A grammar generated by minimizing dependency length in unordered trees from the Penn Treebank is found to agree surprisingly well with English word order, suggesting that dependency length minimization has influenced the evolution of English. 1
The authors examine personality variables and interview format as potential antecedents of impression management (IM) behaviors in simulated selection interviews. The means by which these variables affect ratings of interview performance is also investigated. The altruism facet of agreeableness predicted defensive IM behaviors, the vulnerability facet of emotional stability predicted self- and other-focused behaviors, and interview format (behavior description vs. situational questions) predicted self-focused and defensive behaviors. Consistent with theory and research on situational strength, antecedent—IM relations were consistently weaker in a strong situation in which interviewees had an incentive to manage their impressions. There was also evidence that IM partially mediated the effects of personality and interview format on interview performance in the weak situation.
Although using ontologies to assist information retrieval and text document processing has recently attracted more and more attention, existing ontology-based approaches have not shown advantages over the traditional keywords-based Latent Semantic Indexing (LSI) method. This paper proposes an algorithm to extract a concept forest (CF) from a document with the assistance of a natural language ontology, the WordNet lexical database. Using concept forests to represent the semantics of text documents, the semantic similarities of these documents are then measured as the commonalities of their concept forests. Performance studies of text document clustering based on different document similarity measurement methods show that the CF-based similarity measurement is an effective alternative to the existing keywords-based methods. Especially, this CF-based approach has obvious advantages over the existing keywords-based methods, including LSI, in dealing with text abstract databases, such as MEDLINE, or in P2P environments where it is impractical to collect the entire document corpus for analysis.
The paper outlines the scientific concepts of the pun phenomenon. The paper identifies the origins of viewpoints diversity and analyzes 2 scientific approaches: unification of "pun" and "play on words" concepts and separating them. Study of pun takes a special importance in the light of contemporary trends in mass media language: slang, language of advertising, destructive language play, the widespread disregard to literary and linguistic norms.