Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
We consider linguistic database summaries in the sense of Yager (1982), in an implementable form proposed by Kacprzyk & Yager (2001) and Kacprzyk, Yager & Zadrozny (2000), exemplified by, for a personnel database, “most employees are young and well paid” (with some degree of truth) and their extensions as a very general tool for a human consistent summarization of large data sets. We advocate the use of the concept of a protoform (prototypical form), vividly advocated by Zadeh and shown by Kacprzyk & Zadrozny (2005) as a general form of a linguistic data summary. Then, we present an extension of our interactive approach to fuzzy linguistic summaries, based on fuzzy logic and fuzzy database queries with linguistic quantifiers. We show how fuzzy queries are related to linguistic summaries, and that one can introduce a hierarchy of protoforms, or abstract summaries in the sense of latest Zadeh’s (2002) ideas meant mainly for increasing deduction capabilities of search engines. We show an implementation for the summarization of Web server logs.
One of the benefits of incremental sentence production is reduction of the working memory capacity needed for advance planning: The planning units can be considerably smaller (measured in terms of word length) than in case of non-incremental production. The same advantage has been claimed for the various forms of ellipsis, which preempt the need to plan the detailed shape of one or more constituents and thereby reduce the size of planning units. Because working memory load tends to be higher in spoken than in written language, one expects that speakers, in comparison with writers, will more frequently resort to the use of elliptical constructions. However, in two corpus studies into the incidence of Clausal Coordinate Ellipsis (CCE) in spoken and written English, Meyer (1995) and Greenbaum & Nelson (1999) obtained a data pattern opposite to this prediction: In written clausal coordinations, the proportion of CCE versions was about twice as high as in spoken coordinations. The pattern was explained in terms of audience design: Non-elliptical (unreduced) clauses include more repetition and thereby facilitate comprehension. Recent treebanks with large numbers of hand-parsed spoken (CGN2.0) and written (ALPINO) Dutch sentences, enabled us to verify the data pattern for another language: In written Dutch, the percentage of elliptical versions within the set of all clausal coordinations was even three times higher than in spoken Dutch: 34 % versus 11 % (Table 1).
A polyadic dynamic logic is introduced in which a model-theoretic version of nonlocal multicomponent tree-adjoining grammar can be formulated.It is shown to have a low polynomial time model checking procedure.This means that treebanks for nonlocal MCTAG, incl.all weaker extensions of TAG, can be efficiently corrected and queried.Our result is extended to HPSG treebanks (with some qualifications).The model checking procedures can also be used in heuristics-based parsing.* The model checking procedure described in this paper uses constructs from a model checking procedure introduced in joint work with Martin Lange.Thanks also to Laura Kallmeyer, Timm Lichte and Wolfgang Maier for introducing me to various extensions of tree-adjoining grammar, incl.nonlocal MCTAG.
Treebanks are used for various purposes in language technology, but the wealth of data they contain can also be put to good use for the purpose of linguistic description and linguistic theory. To demonstrate this I will show how the treebank of the Spoken Dutch Corpus can be exploited to improve our understanding of what it is that distinguishes predicative complements from other types of complements. Section 1 shows why this distinction matters, section 2 provides a brief presentation of the treebank, section 3 gives a comprehensive survey of the intransitive predicate selecting verbs, based on the treebank data, section 4 presents a number of factors which can be used to differentiate the predicate selecting uses of the relevant verbs from their other uses, and section 5 summarizes the results. 1 Predicative complements Distinguishing a predicative complement from an object complement is easy in pairs like (1). (1) a. Fred is a plumber. b. Fred knows a plumber. The complement of the copula denotes a property which is attributed to the referent of the subject, and the copula itself is little more than a carrier of mood and tense. By contrast, the complement of know denotes an entity and the verb denotes a binary relation between that entity and the referent of the subject. Since the verb is the only element that overtly distinguishes (1a) from (1b), it might seem sufficient to draw the distinction, but the matter is more complex, since many verbs are used either way. The second complement of call and make, for instance, is predicative in (2), but not in (3). (2) a. Don’t call me a liar. b. They will make you chairman. 1This work is part of a larger project on the syntax and semantics of clauses with predicative complements. So far, it has yielded an HPSG style analysis of such clauses (Van Eynde 2008) and a semantic analysis of the copula, presented at the HPSG-2009 conference. Proceedings of the 19th Meeting of Computational Linguistics in the Netherlands Edited by: Barbara Plank, Erik Tjong Kim Sang and Tim Van de Cruys. Copyright c ©2009 by the individual authors.
Discovering frequent structures within large natural language corpora is one of the core problems of corpus linguistics, but it is difficult to do for richly structured data. This paper describes a practical algorithm to extract frequent structures from treebanks or annotated corpora that can be represented as a tree structures. It extracts the most frequent structures first, so that not all structures have to be counted in order to find the most frequent ones. This algorithm assumes random constant-time access to all parts of the treebank and has space and time bounds broadly proportionate to the size of the output, which is not readily predictable in most cases. It is efficient enough to be usable with reasonable sized corpora using conventional desktop workstations.
I will cover some of the main stages of my projects in Computational Linguistics through 60 years, from their very beginning after World War II to the Internet era, including my most recent works on 'Lessico Tomistico Biculturale' and the syntactic annotation of data in the 'Index Thomisticus' Treebank project. The talk will describe how and why the idea of electronic data processing in literary and linguistic analysis came to my mind at first, remembering my meeting with Mr. Thomas Watson at IBM and his decision to fund the realization of 'Index Thomisticus' since 1949. I will describe how meticuolous is the research methodology in linguistics which is imposed by (and thanks to) the use of computers. Through the years, I applied this methodology to more than 20 different languages (starting from Latin) and to many fields of Computational Linguistics, such as semiautomatic lemmatization, machine dictionaries, textual typology, methods of treatment of different alphabets etc. 1
ABSTRACT: This paper argues that the ‘world Englishes paradigm’ and English as a lingua franca (ELF) research, despite important differences, have much in common. Both share the pluricentric assumption that ‘English’ belongs to all those who use it, and both are concerned with the sociolinguistic, socio‐psychological, and applied linguistic implications of this assumption. For example, issues of language contact, variation and change, linguistic norms and their acceptance, ownership of the language, and expression of social identities are central to both WE and ELF research. The growing body of descriptive ELF research that is now becoming available can thus add substance to work in the field as a whole. It can also offer fresh perspectives on several theoretical constructs central to WE, such as ‘community’, ‘variety’, ‘lingua franca’, even ‘language’.
Treebanks are used for various purposes in language technology, but the wealth of data they contain can also be put to good use for the purpose of linguistic description and linguistic theory. To demonstrate this I will show how the treebank of the Spoken Dutch Corpus can be exploited to improve our understanding of what it is that distinguishes predicative complements from other types of complements. Section 1 shows why this distinction matters, section 2 provides a brief presentation of the treebank, section 3 gives a comprehensive survey of the intransitive predicate selecting verbs, based on the treebank data, section 4 presents a number of factors which can be used to differentiate the predicate selecting uses of the relevant verbs from their other uses, and section 5 summarizes the results.
We present a novel transition system for dependency parsing, which constructs arcs only between adjacent words but can parse arbitrary non-projective trees by swapping the order of words in the input. Adding the swapping operation changes the time complexity for deterministic parsing from linear to quadratic in the worst case, but empirical estimates based on treebank data show that the expected running time is in fact linear for the range of data attested in the corpora. Evaluation on data from five languages shows state-of-the-art accuracy, with especially good results for the labeled exact match score.
The paper reports on the main findings of the LANCHART language attitudes studies. These studies were designed to falsify (or modify) the picture of adolescent language ideology – and its role in language change – that had emerged from previous sociolinguistic studies in Denmark. This picture is formulated as three hypotheses: (1) There are two value systems at two levels of consciousness, (2) Language change is governed by subconscious values, (3) Copenhagen is Denmark’s only linguistic norm centre. Following strict guidelines for data collection among 9th graders (aged 15–16) in Copenhagen, Næstved, Vissenbjerg, Odder, and Vinderup we obtained subconsciously offered attitudes that could be compared with consciously offered attitudes. The results neither falsify nor modify the established picture but strongly confirm it.
We describe a heuristics-based system for automatic measurement of syntactic complexity using the revised Developmental Level (D-Level) scale (Rosenberg & Abbeduto 1987; Covington et al. 2006). The system takes a raw sentence as input and assigns it to an appropriate developmental level on the scale. The system is designed with child language acquisition and psycholinguistic research in mind, and is therefore developed and evaluated using both written data from the Penn Treebank (Marcus et al. 1993) and spoken child language acquisition data from the CHILDES database (MacWhinney 2000). Experiment results show that the model achieves an accuracy of 94.0% and 93.2% on unseen test data from the Penn Treebank and the CHILDES database respectively. We illustrate how the system is used in an example application to investigate the correlation of average D-Level score and speaker age.
This paper proposes an approach to enhance dependency parsing in a language by using a translated treebank from another language. A simple statistical machine translation method, word-by-word decoding, where not a parallel corpus but a bilingual lexicon is necessary, is adopted for the treebank translation. Using an ensemble method, the key information extracted from word pairs with dependency relations in the translated text is effectively integrated into the parser for the target language. The proposed method is evaluated in English and Chinese treebanks. It is shown that a translated English treebank helps a Chinese parser obtain a state-of-the-art result.
Emotion plays a significant role in goal-directed behavior, yet its neural basis is yet poorly understood. In several psychological models the cardinal dimensions that characterize the emotion space are considered to be valence and arousal. Here 3T functional magnetic resonance imaging (fMRI) was used to reveal brain areas that show valence- and arousal-dependent blood oxygen level dependent (BOLD) signal responses. Seventeen healthy adults viewed pictures from the International Affective Picture System (IAPS) for brief 100 ms periods in a block design paradigm. In many brain regions BOLD signals correlated significantly positively with valence ratings of unpleasant pictures. Interestingly, partly in the same regions but also in several other regions BOLD signals correlated negatively with valence ratings of pleasant pictures. Therefore, there were several areas where the correlation across all pictures was of inverted U-shape. Such correlations were found bilaterally in the dorsolateral prefrontal cortex (DLPFC), dorsomedial prefrontal cortex (DMPFC) extending to anterior cingulate cortex (ACC), and insula. Self-rated arousal of those pictures which were evaluated to be unpleasant correlated with BOLD signal in the ACC, whereas for pleasant pictures arousal correlated positively with the BOLD signal strength in the right substantia innominata. We interpret our results to suggest a major division of brain mechanisms underlying affective behavior to those evaluating stimuli to be pleasant or unpleasant. This is consistent with the basic division of behavior to approach and withdrawal, where differentiation of hostile and hospitable stimuli is crucial.
While rules and exemplars are usually viewed as opposites, this paper argues that they form end points of the same distribution. By representing both rules and exemplars as (partial) trees, we can take into account the fluid middle ground between the two extremes. This insight is the starting point for a new theory of language learning that is based on the following idea: If a language learner does not know which phrase-structure trees should be assigned to initial sentences, s/he allows (implicitly) for all possible trees and lets linguistic experience decide which is the "best" tree for each sentence. The best tree is obtained by maximizing "structural analogy" between a sentence and previous sentences, which is formalized by the most probable shortest combination of subtrees from all trees of previous sentences. Corpus-based experiments with this model on the Penn Treebank and the Childes database indicate that it can learn both exemplar-based and rule-based aspects of language, ranging from phrasal verbs to auxiliary fronting. By having learned the syntactic structures of sentences, we have also learned the grammar implicit in these structures, which can in turn be used to produce new sentences. We show that our model mimicks children's language development from item-based constructions to abstract constructions, and that the model can simulate some of the errors made by children in producing complex questions.
Abstract Textual data is at the forefront of information management problems today. One response has been the development of visualizations of text data. These visualizations, commonly based on simple attributes such as relative word frequency, have become increasingly popular tools. We extend this direction, presenting the first visualization of document content which combines word frequency with the human‐created structure in lexical databases to create a visualization that also reflects semantic content. DocuBurst is a radial, space‐filling layout of hyponymy (the IS‐A relation), overlaid with occurrence counts of words in a document of interest to provide visual summaries at varying levels of granularity. Interactive document analysis is supported with geometric and semantic zoom, selectable focus on individual words, and linked access to source text.
Older adults' relatively better memory for positive over negative material (positivity effect) has been widely observed in Western samples. This study examined whether a relative preference for positive over negative material is also observed in older Koreans. Younger and older Korean participants viewed images from the International Affective Picture System (IAPS), were tested for recall and recognition of the images, and rated the images for valence. Cultural differences in the valence ratings of images emerged. Once considered, the relative preference for positive over negative material in memory observed in older Koreans was indistinguishable from that observed previously in older Americans.
Learning-based models of anxiety disorders emphasize the role of aversive conditioning and retarded extinction in the etiology and maintenance of anxiety disorders. Yet few studies have examined these underlying processes in children, despite that some anxiety disorders typically onset during childhood. The authors examined the acquisition and extinction of conditioned responses in 17 anxious children and 18 nonanxious control children between 8 and 12 years old using a discriminative Pavlovian conditioning procedure. One geometric shape conditional stimulus was paired with an unpleasant loud tone unconditional stimulus (CS+) whereas another geometric shape was presented alone (CS-). In the context of similar levels of discriminative conditioning in both groups, anxious children showed larger skin conductance responses to the CS+ and the CS- during acquisition and evaluated the CS+ as more arousing than the CS- compared with control children. They also showed greater resistance to extinction in skin conductance responses but not in arousal ratings to the CS+ vs. the CS- relative to control children. Results suggest that deficits in response inhibition to safety cues and retarded extinction may underlie learning processes involved in the pathogenesis of childhood anxiety disorders.
In this paper, we present a discriminative word-character hybrid model for joint Chinese word segmentation and POS tagging. Our word-character hybrid model offers high performance since it can handle both known and unknown words. We describe our strategies that yield good balance for learning the characteristics of known and unknown words and propose an error-driven policy that delivers such balance by acquiring examples of unknown words from particular errors in a training corpus. We describe an efficient framework for training our model based on the Margin Infused Relaxed Algorithm (MIRA), evaluate our approach on the Penn Chinese Treebank, and show that it achieves superior performance compared to the state-of-the-art approaches reported in the literature.
Enhanced sensitivity to information of negative (compared to positive) valence has an adaptive value, for example, by expediting the correct choice of avoidance behavior. However, previous evidence for such enhanced sensitivity has been inconclusive. Here we report a clear advantage for negative over positive words in categorizing them as emotional. In 3 experiments, participants classified briefly presented (33 ms or 22 ms) masked words as emotional or neutral. Categorization accuracy and valence-detection sensitivity were both higher for negative than for positive words. The results were not due to differences between emotion categories in either lexical frequency, extremeness of valence ratings, or arousal. These results conclusively establish enhanced sensitivity for negative over positive words, supporting the hypothesis that negative stimuli enjoy preferential access to perceptual processing.
In 2 cross-sectional studies, the authors examined age-related differences in the evaluation of emotional stimuli in 2 community samples, with participants ranging in age from young to older adulthood (18-81 years old). Pictures of the International Affective Picture System were used in Study 1, and written verbs were used in Study 2. Participants rated these stimuli along the 2 major affective dimensions of hedonic valence and emotional arousal, thus yielding a 2-dimensional affective space for each participant. Young adults showed the expected pattern of 2 distinct clusters of stimuli in this space, representing increasing pleasantness (appetitive activation) and unpleasantness (aversive activation) with increasing emotional arousal. In contrast, for older adults, emotional valence and arousal ratings were linearly related: Low-arousing stimuli were rated as most pleasant, and high-arousing stimuli were rated as most unpleasant. When regressed on age, these changes revealed a gradual decrease of appetitive activation (i.e., the relationship between pleasure and arousal) across adulthood and a linear increase in aversive activation (i.e., the relationship between displeasure and arousal). These results extend previous work on emotional development, adding information as to the role of emotional intensity for affective experience in different age groups.
The present study compared gender differences in directly reported and indirectly derived career preferences and tested the hypothesis that individuals' implicit preferences would show less gender-biased occupational choices than their directly elicited ones. Two hundred sixty-six visitors to a career-related Internet site were asked to (a) list 5 to 10 suitable occupations (the directly reported list) and (b) report their preferences in terms of 31 career-related aspects. The latter were used to produce a short list of promising occupational alternatives (the indirectly derived list), using the occupational database of an Internet-based career planning system. Each occupation in the database rated for sex dominance. The findings indicated that the sex dominance ratings of the occupations on the directly reported list accorded with the participants' gender for both men and women: Men's lists included mostly “masculine” occupations, whereas women's lists included mostly “feminine” occupations. This gender bias was significantly lower for the implicit lists. The difference between the directly reported and the indirectly derived lists was larger for women than for men, suggesting that the impact of stereotypes is more pronounced in women's than in men's directly reported career preferences.
Morphological segmentation breaks words into morphemes (the basic semantic units). It is a key component for natural language processing systems. Unsupervised morphological segmentation is attractive, because in every language there are virtually unlimited supplies of text, but very few labeled resources. However, most existing model-based systems for unsupervised morphological segmentation use directed generative models, making it difficult to leverage arbitrary overlapping features that are potentially helpful to learning. In this paper, we present the first log-linear model for unsupervised morphological segmentation. Our model uses overlapping features such as morphemes and their contexts, and incorporates exponential priors inspired by the minimum description length (MDL) principle. We present efficient algorithms for learning and inference by combining contrastive estimation with sampling. Our system, based on monolingual features only, outperforms a state-of-the-art system by a large margin, even when the latter uses bilingual information such as phrasal alignment and phonetic correspondence. On the Arabic Penn Treebank, our system reduces F1 error by 11% compared to Morfessor.
Broad-coverage annotated treebanks necessary to train parsers do not exist for many resource-poor languages. The wide availability of parallel text and accurate parsers in English has opened up the possibility of grammar induction through partial transfer across bitext. We consider generative and discriminative models for dependency grammar induction that use word-level alignments and a source language parser (English) to constrain the space of possible target trees. Unlike previous approaches, our framework does not require full projected parses, allowing partial, approximate transfer through linear expectation constraints on the space of distributions over trees. We consider several types of constraints that range from generic dependency conservation to language-specific annotation rules for auxiliary verb analysis. We evaluate our approach on Bulgarian and Spanish CoNLL shared task data and show that we consistently outperform unsupervised methods and can outperform supervised learning for limited training data.
DeSR is a statistical transition-based dependency parser which learns from annotated corpora which actions to perform for building parse trees while scanning a sentence. We describe recent improvements to the parser, in particular stacked parsing, exploiting a beam search strategy and using a Multilayer Perceptron classifier. For the Evalita 2009 Dependency Parsing task DesR was configured to use a combination of stacked parsers. The stacked combination achieved the best accuracy scores in both the main and pilot subtasks. The contribution to the result of various choices is analyzed, in particular for taking advantage of the peculiar features of the TUT Treebank.\n\nKeywords: parser, dependency parsing, perceptron, classifier, natural language.
CONTEXT: Although drug cues reliably activate the brain's reward system, studies rarely examine how the processing of drug stimuli compares with natural reinforcers or relates to clinical outcomes. OBJECTIVES: To determine hedonic responses to natural and drug reinforcers in long-term heroin users and to examine the utility of these responses in predicting future heroin use. DESIGN: Prospective design examining experiential, expressive, reflex modulation, and cortical/attentional responses to opiate-related and affective stimuli. The opiate-dependent group was reassessed a median of 6 months after testing to determine their level of heroin use during the intervening period. SETTING: Community drug and alcohol services and a clinical research facility. PARTICIPANTS: Thirty-three opiate-dependent individuals (mean age, 31.6 years) with stabilized opiate-substitution pharmacotherapy and 19 sex- and age-matched healthy non-drug users (mean age, 30 years). MAIN OUTCOME MEASURES: Self-ratings, facial electromyography, startle-elicited postauricular reflex, and event-related potentials combined with measures of heroin use at baseline and follow-up. RESULTS: Relative to the control group, the opiate-dependent group rated pleasant pictures as less arousing and showed increased corrugator activity, less postauricular potentiation, and decreased startle-elicited P300 attenuation while viewing pleasant pictures. The opiate-dependent group rated the drug-related pictures as more pleasant and arousing, and demonstrated greater startle-elicited P300 attenuation while viewing them. Although a startle-elicited P300 amplitude response to pleasant (relative to drug-related) pictures significantly predicted regular (at least weekly) heroin use at follow-up, subjective valence ratings of pleasant pictures remained the superior predictor of use after controlling for baseline craving and heroin use. CONCLUSIONS: Heroin users demonstrated reduced responsiveness to natural reinforcers across a range of psychophysiological measures. Subjective rating of pleasant pictures robustly predicted future heroin use. Our findings highlight the importance of targeting anhedonic symptoms within clinical treatment settings.
Cognitive factors such as catastrophic thoughts regarding pain, and conversely, one's acceptance of that pain, may affect emotional functioning among persons with chronic pain conditions. The aims of the present study were to examine the effects of both catastrophizing and acceptance on affective ratings of experimentally induced ischemic pain and also self-reports of depressive symptoms. Sixty-seven individuals with chronic back pain completed self-report measures of catastrophizing, acceptance, and depressive symptoms. In addition, participants underwent an ischemic pain induction procedure and were asked to rate the induced pain. Catastrophizing showed significant effects on sensory and intensity but not affective ratings of the induced pain. Acceptance did not show any significant associations, when catastrophizing was also in the model, with any form of ratings of the induced pain. Catastrophizing, but not acceptance, was also significantly associated with self-reported depressive symptoms when these two variables were both included in a regression model. Overall, results indicate negative thought patterns such as catastrophizing appear to be more closely related to outcomes of perceived pain severity and affect in persons with chronic pain exposed to an experimental laboratory pain stimulus than does more positive patterns as reflected in measures of acceptance.
BACKGROUND: There is an interest in investigating the relation between emotional memory impairments in schizophrenia and specific symptom dimensions. We explored potential links between emotional memory and social anhedonia severity in patients with schizophrenia and in healthy individuals. METHODS: Twenty-nine patients with schizophrenia and 27 matched healthy individuals completed the Chapman Revised Social Anhedonia Scale and then performed an emotional face recognition memory task involving happy, sad and neutral face expressions. We calculated emotional memory performance using 2 independent measures: the discrimination accuracy index Pr and the response bias Br. We also measured valence ratings of the face stimuli. We performed correlation analyses using the inter-individual variability in social anhedonia severity and the individual score obtained for each memory performance variable and for each face valence rating condition. RESULTS: Patients with schizophrenia reported higher levels of social anhedonia compared with healthy individuals. They also showed lower recognition accuracy for faces compared with healthy participants. We found no significant correlation between social anhedonia severity and any of the memory performance variables for both patients with schizophrenia and healthy individuals. Regarding potential links between social anhedonia severity and face valence ratings, we found that individuals with elevated social anhedonia had a tendency to rate the face stimuli as more negative. LIMITATIONS: Our negative finding may be partly explained by a lack of statistical power owing to our small patient sample. In addition, our patient sample had unusually high estimated IQ scores, which highlights potential issues regarding the generalization of our findings. Finally, we used a yes-no recognition memory task with a very short retention interval delay. CONCLUSION: Our results suggest that social anhedonia is not directly linked to emotional memory deficits and biases and does not interfere with the modulatory effect of positively valenced emotion on memory.
Query expansion is a widely studied technique for improving information retrieval effectiveness. In this paper we proposed a new query expansion technique using the comprehensive thesaurus WordNet and its semantic relatedness measure modules. Word sense disambiguation are performed on original query sentence, yielding the concept of each term in the query. Based on those recovered concepts, expanded query terms are generated from WordNet lexical database. The proposed method has been evaluated in document retrieval on the Web using query sentence. Our extensive experimental results demonstrate a 7% precision improvement over retrieval methods not employing query expansion techniques.
We present a semi-supervised method to improve statistical parsing performance. We focus on the well-known problem of lexical data sparseness and present experiments of word clustering prior to parsing. We use a combination of lexicon-aided morphological clustering that preserves tagging ambiguity, and unsupervised word clustering, trained on a large unannotated corpus. We apply these clusterings to the French Treebank, and we train a parser with the PCFG-LA unlexicalized algorithm of (Petrov et al., 2006). We find a gain in French parsing performance: from a baseline of F1=86.76% to F1=87.37% using morphological clustering, and up to F1=88.29% using further unsupervised clustering. This is the best known score for French probabilistic parsing. These preliminary results are encouraging for statistically parsing morphologically rich languages, and languages with small amount of annotated data.
This paper describes an empirical study of high-performance dependency parsers based on a semi-supervised learning approach. We describe an extension of semi-supervised structured conditional models (SS-SCMs) to the dependency parsing problem, whose framework is originally proposed in (Suzuki and Isozaki, 2008). Moreover, we introduce two extensions related to dependency parsing: The first extension is to combine SS-SCMs with another semi-supervised approach, described in (Koo et al., 2008). The second extension is to apply the approach to second-order parsing models, such as those described in (Carreras, 2007), using a two-stage semi-supervised learning approach. We demonstrate the effectiveness of our proposed methods on dependency parsing experiments using two widely used test collections: the Penn Treebank for English, and the Prague Dependency Tree-bank for Czech. Our best results on test data in the above datasets achieve 93.79% parent-prediction accuracy for English, and 88.05% for Czech.
Jointly parsing two languages has been shown to improve accuracies on either or both sides. However, its search space is much bigger than the monolingual case, forcing existing approaches to employ complicated modeling and crude approximations. Here we propose a much simpler alternative, bilingually-constrained monolingual parsing, where a source-language parser learns to exploit reorderings as additional observation, but not bothering to build the target-side tree as well. We show specifically how to enhance a shift-reduce dependency parser with alignment features to resolve shift-reduce conflicts. Experiments on the bilingual portion of Chinese Treebank show that, with just 3 bilingual features, we can improve parsing accuracies by 0.6% (absolute) for both English and Chinese over a state-of-the-art baseline, with negligible (~6%) efficiency overhead, thus much faster than biparsing.
This paper presents preliminary investigations on the statistical parsing of French by bringing a complete evaluation on French data of the main probabilistic lexicalized and unlexicalized parsers first designed on the Penn Treebank. We adapted the parsers on the two existing treebanks of French (Abeillé et al., 2003; Schluter and van Genabith, 2007). To our knowledge, mostly all of the results reported here are state-of-the-art for the constituent parsing of French on every available treebank. Regarding the algorithms, the comparisons show that lexicalized parsing models are outperformed by the unlexicalized Berkeley parser. Regarding the treebanks, we observe that, depending on the parsing model, a tag set with specific features has direct influence over evaluation results. We show that the adapted lexicalized parsers do not share the same sensitivity towards the amount of lexical material used for training, thus questioning the relevance of using only one lexicalized model to study the usefulness of lexicalization for the parsing of French.
In his later work in the philosophy of language Davidson analyses communicative exchanges and arrives at the startling and weighty conclusion that linguistic norms and conventions are entirely inessential to linguistic meaning. I argue that this inference is flawed: if we place the account of communication against the backdrop of Davidson's own views about radical interpretation, then it becomes evident that linguistic norms are an essential feature of the Davidsonian picture.
One of the biggest challenges in compiling a dictionary of a minority language is managing the large quantity of lexical data. Decisions about the format and content of the dictionary or the orthography typically evolve over the years that such projects usually take. This results in inconsistencies between older and newer entries. Revising the data for publication as a dictionary introduces further inconsistencies as does having multiple contributors and/or editors. Proofreading a lexical database takes a great deal of time and the richer its structure the more this is the case. The tools described in this presentation significantly reduce this effort. Tools developed for checking the consistency of the lexical database in the Iu Mien—Chinese—English dictionary project have proven extremely helpful. Two basic approaches are used: 1) use of a program written to check for likely errors that scans the lexical database and produces an error report that is used by a lexicographer to make appropriate corrections. 2) outputting the lexical data in alternate forms that make it easier for the lexicographer to spot problem areas. These alternative forms include the reverse indexes and views structured according to semantic domains. The Iu Mien—Chinese—English dictionary project, like many minority language dictionary projects, uses SIL's Toolbox software. It is very flexible software but its capabilities to enforce consistency are quite limited. Some parts of the approach described here are specific to MDF (Multi-Dictionary Formatter) lexical databases in Toolbox but will be equally useful for other MDF databases. Other parts are specific to each of the three languages involved but will be useful for non-Toolbox lexical databases. Every dictionary is unique and this applies not only to content of the entries but also the decisions about how entries should be arranged to suit the languages involved. Other decisions about the structure are likely to be made differently even in other dictionaries of the same languages. It is the way that each dictionary combines themes that are found in many dictionaries that makes them unique, e.g. to be root based or not, to have include subentries. Therefore our approach is to use a toolkit based approach to curating lexical databases. This allows checking techniques to be mixed and matched to suit the unique aspects of a lexical project. The checking software is written in Python and relies on the toolbox module in NLTK (The Natural Language Toolkit http://nltk.sourceforge.net).
Olfactory perception was examined in deficit syndrome (DS) and nondeficit syndrome (ND) schizophrenia patients. Participants included 22 controls (CN) and 41 patients with schizophrenia who were divided into DS (n = 15) and ND (n = 26) subtypes using the Schedule for the Deficit Syndrome (SDS). Olfactory perception for pleasant and unpleasant odors was assessed using the Brief Smell Identification Test. Participants were instructed to identifying each smell as well as provide hedonic judgment ratings of each smell on a 7-point scale (1 = extremely pleasant, 4 = neutral, and 7 = extremely unpleasant). Results indicated that when compared with the ND patients, the DS patients rated pleasant smells as being significantly less pleasant, although no difference between the groups was present for unpleasant smells, and both ND and DS groups significantly differed from CN on rating and identifying pleasant and unpleasant items. Additionally, lower smell identification accuracy was negatively correlated with SDS symptom severity, and valence ratings for pleasant odors were positively correlated with SDS diminished emotional range. Findings suggest that the DS is characterized by a unique pattern of olfactory valence judgment that is characterized by abnormalities in processing positively valenced stimuli.
Parkinson's disease (PD) involves facial masking, which may impair social interaction. Older adult observers who viewed segments of videotaped interviews of individuals with PD expressed less interest in relationships with women with higher masking and judged them as less supportive. Masking did not affect ratings of men in these domains, possibly because higher masking violates gender norms for expressivity in women but not in men. Observers formed less accurate ratings of the social supportiveness and social strain of women than men, and higher masking decreased accuracy for ratings of strain. Results suggest that some of the problems with social relationships in PD may be due to inaccurate impressions and reduced desire to interact with individuals with higher masking, especially women.
The paper first briefly outlines the differences between a lexical data base as it typically results from a language documentation project and the kind of dictionaries the speech community wants for educational purposes with respect to the choice of head words, grammatical information, definitions of meaning and translations, encyclopaedic information and the choice of examples. The second part of the paper then explores how in spite of limited resources in terms of time, money and man power the speech community and the linguists can develop a method of dictionary making that both satisfies the needs of the community and the interests of linguists. Since it is impossible to create a comprehensive dictionary in a language documentation project, we opted for the thematic approach in which lexicographers work on particular semantic domains such as body parts, architecture or fishing, and try to cover all those lexemes of the respective domain that seem to be important for the intended dictionary users. In our project the headwords of the lexical database were classified according to their domains, and then filtered and exported from the lexical database in order to produce a mini dictionary for each selected domain. In the case of body parts, for instance, we did not only select the nouns that signify the body parts, but also verbs that express bodily actions like ‘sweat’ and ‘comb your hair’ as well as speech formulas like ‘have a heavy heart’. The indigenous lexicographers then checked this preliminary mini-dictionary for missing head words and fixed multi-word expressions,and revised the examples which often were so context dependent that they did not make much sense in isolation. Focussing on one particular semantic domain at a time helps to easily identify various kinds of lexical relations such as taxonomies, meronymies and metonymies as well as metaphorical usages, collocational restrictions and grammatical constructions. As one and the same lexeme can belong to more than one semantic domain, the mini-dictionaries must be accompanied by an index. The paper concludes with a discussion of how this kind of practical lexicography is related to frame semantics and how it can be used for semantic typology.
This paper describes the MulTra project, aiming at the development of an efficient multilingual translation technology based on an abstract and generic linguistic model as well as on object-oriented software design. In particular, we will address the issue of the rapid growth both of the transfer modules and of the bilingual databases. For the latter, we will show that a significant part of bilingual lexical databases can be derived automatically through transitivity, with corpus validation.
This study evaluated evidence for 2 forms of emotional abnormality in posttraumatic stress disorder (PTSD): numbing and heightened negative emotionality. Forty-nine male veterans with PTSD and 75 without the disorder rated their emotional responses to photographs that depicted scenes of Vietnam combat or were drawn from the International Affective Picture System (Lang et al., 2005). Images varied in their trauma-relatedness and affective qualities. A series of repeated measures ANOVAs revealed that Vietnam combat veterans with PTSD responded to unpleasant images with greater negative emotionality (i.e., enhanced arousal and lower valence ratings) than those without the disorder and this effect was modified by the trauma-relatedness of the image with stronger effects for trauma-related images. In contrast, the 2 groups showed equivalent patterns of responses to pleasant images. Findings raise questions about the sensitivity of the International Affective Picture System rating protocol for the assessment of PTSD-related emotional numbing.
BACKGROUND: Time-limited group cognitive behavioral treatments (GCBT) for obsessive-compulsive disorder have demonstrated improvement in target symptoms. One small sample study of GCBT specifically for hoarding problems also showed benefit. This study examines the efficacy of a specialized GCBT for compulsive hoarding on a larger sample. METHODS: Thirty-two clients diagnosed with hoarding participated in five groups. Four groups met once weekly for 2 hour over 16 weeks (n=27) and one group met for 20 weeks (n=5). All participants had two individual 90-min home sessions. Self-report assessments were completed at baseline, mid-treatment, and post-treatment about hoarding behavior and related symptoms (e.g., depression). The sample was predominantly female, White, highly educated, unemployed, and not partnered/married; mean age was 53. A majority was diagnosed with major depressive disorder and obsessive-compulsive personality disorder. RESULTS: Participants showed significant improvement from pre- to post-treatment on the Saving Inventory Revised, Saving Cognitions Inventory, Clutter Image Rating, and Clinical Global Severity. The most recent group (n=8) that used a more formalized treatment and research protocol improved significantly more than did earlier members. CONCLUSION: This study demonstrates the feasibility and modest success of GCBT methods in improving hoarding symptoms. Group treatment may be especially valuable because of its cost-effectiveness, greater client access to trained clinicians, and reduction in social isolation and stigma linked to this problem. Further research is needed to improve the efficacy of GCBT methods for hoarding and to examine durability of change, predictors of outcomes, and processes that influence change.