Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
Aim: The present study aims to understand environmental problems perception related to affectivity. Introduction and Context: The seriousness of global environmental problems engenders a challenge to psychology regarding the need to induce environmentally responsible behavior and sustainable development. Environmental psychology research has classically been cognitively oriented. Recent studies on risk perception, however, highlight the important role that affect may be playing, and sustainable behavior seems to stem from emotions (feelings of connectedness with other beings or moral emotions such as indignation or guilt).Method: 170 undergraduates where asked to recall associations for three environmental problems selected from a previous survey: climate change, loss of biodiversity, and environmental unconsciousness. Affective images associated with these issues were collected through free association test. Affect was also measured using the validated Spanish version of The Positive and Negative Affect Schedule (PANAS). One assumption of the current research is that word association techniques may allow to explore possible links between imagery and behavior. Despite the fact that this measuring instrument has recently been applied to the analysis of environmental risks, our methodological contribution is its combination with one of the most widely used measure of affectivity. That is to say, the second assumption rests on evidence suggesting a two dimensional structure of affective experience. 1509 associations were generated through the imagery survey and the content of these associations was analysed for each inductive stimulus.Affective imagery: The primary image of the majority of respondents was a paraphrase of each environmental issue. Although very few associations to climate change denoted consequences on human health, and no one mentioned future generations, our findings indicate a strong tendency for respondents to specify effects such as disasters, ice melting or floods, rather than causes, such as greenhouse gas emissions. By the same token, loss of biodiversity was primarily associated with resource scarcity or devastation, not with excessive uses of ecosystems. A slight difference appeared in images associated with environmental unconsciousness: the second highest frequency of responses referred to a specific behavioral pattern (no recycling actions), and the following ones suggested a sense of personal responsibility for environmental outcomes. In addition to analyzing the content categories, the affective ratings of images were examined.Implications: The knowledge gained from this research may be applied to the construction of communication messages and although more research is needed to compare data from different socio-demographic groups, there is some evidence that affective dimensions may be predictive of participation in pro-environmental behaviors.
This paper describes the construction of a dependency bank gold standard for Arabic, DCU 250 Arabic Dependency Bank (DCU 250), based on the Arabic Penn Treebank Corpus (ATB) (Bies and Maamouri, 2003; Maamouri and Bies, 2004) within the theoretical framework of Lexical Functional Grammar (LFG). For parsing and automatically extracting grammatical and lexical resources from treebanks, it is necessary to evaluate against established gold standard resources. Gold standards for various languages have been developed, but to our knowledge, such a resource has not yet been constructed for Arabic. The construction of the DCU 250 marks the first step \ntowards the creation of an automatic LFG f-structure annotation algorithm for the ATB, \nand for the extraction of Arabic grammatical and lexical resources.
In this paper, we describe an annotation scheme for the attribution of abstract objects (propositions, facts, and eventualities) associated with discourse relations and their arguments annotated in the Penn Discourse TreeBank. The scheme aims to capture both the source and degrees of factuality of the abstract objects through the annotation of text spans signalling the attribution, and of features recording the source, type, scopal polarity, and determinacy of attribution. RESUME. Dans cet article, nous decrivons un schema d’annotation pour l’encodage des objets abstraits (propositions, faits et possibilites) associes aux relations de discours et a leurs arguments tels qu’annotes dans le Penn Discourse TreeBank. Ce schema a pour objet la capture de la source et du degre de factualite des objets abstraits. Les aspects cles de ce schema comprennent l’annotation des intervalles textuels signalant l’attribution, ainsi que l’annotation des proprietes caracterisant la source, le type, la polarite de la portee, et le degre de determination de l’attribution.
CIDOC-CRM is a new standard for encoding a wide range of information for Cultural Heritage (CH). At present, existing CH collections are stored using all sorts of formats, sometimes proprietary, often defined roughly, which \nmakes it difficult to share or access heterogeneous information among the CH community. There is a need for a tool to map diverse formats into CIDOC-CRM, assisted by another tool using intelligent language technology to help the mapping whenever fields are underspecified or loosely described, both tools being complementary. In some cases, it may even be better to build fragments of a CIDOC database directly from informal descriptions in natural \nlanguage only, as the CH community may be reluctant to switch to new formats of data entry. Therefore, this paper focus primarily on the mapping of CH data described in natural language into CIDOC-CRM triples, the building blocks of the full CIDOC-CRM ontology. The methods exploits the propositional nature of CIDOC-CRM triples. Using WordNet as a lexical database and the WEB as corpus, we first extract triples from examples provided in the CIDOC-CRM literature, and then from text describing the medieval city of Wolfenbüttel. We show the strong points of the system and suggest where and how it could be improved. Although the triples extracted automatically from texts do not provide a full picture of the CIDOC-CRM structure buried in the textual description, our results indicate that it provides a sound initial working basis for the mapping/translation process, saving time on what would otherwise have to be done by hand.
We present the implementation of a system which extracts not only lexicalized grammars but also feature-based lexicalized grammars from Korean Sejong Treebank. We report on some practical experiments where we extract TAG grammars and tree schemata. Above all, full-scale syntactic tags and well-formed morphological analysis in Sejong Treebank allow us to extract syntactic features. In addition, we modify Treebank for extracting lexicalized grammars and convert lexicalized grammars into tree schemata to resolve limited lexical coverage problem of extracted lexicalized grammars.
Solar-driven interfacial evaporation technology has attracted significant attention for water purification. However, design and fabrication of solar-driven evaporator with cost-effective, excellent capability and large-scale production remains challenging. In this study, inspired by plant transpiration, a tri-layered hierarchical nanofibrous photothermal membrane (HNPM) with a unidirectional water transport effect was designed and prepared via electrospinning for efficient solar-driven interfacial evaporation. The synergistic effect of the hierarchical hydrophilic-hydrophobic structure and the self-pumping effect endowed the HNPM with unidirectional water transport properties. The HNPM could unidirectionally drive water from the hydrophobic layer to the hydrophilic layer within 2.5 s and prevent reverse water penetration. With this unique property, the HNPM was coupled with a water supply component and thermal insulator to assemble a self-floating evaporator for water desalination. Under 1 sun illumination, the water evaporation rates of the designed evaporator with HNPM in pure water and dyed wastewater reached 1.44 and 1.78 kg·m<sup>-2</sup>·h<sup>-1</sup>, respectively. The evaporator could achieve evaporation of 11.04 kg·m<sup>-2</sup> in 10 h under outdoor solar conditions. Moreover, the tri-layered HNPM exhibited outstanding flexibility and recyclability. Our bionic hydrophobic-to-hydrophilic structure endowed the solar-driven evaporator with capillary wicking and transpiration effects, which provides a rational design and optimization for efficient solar-driven applications.
We describe several improvements to the method of treebank-based LFG induction for Spanish from the Cast3LB treebank (O’Donovan et al., 2005). We discuss the different categories of problems encountered and present the solutions adopted. Some of the problems involve a simple adoption of existing linguistic analyses, as in our treatment of clitic doubling and null subjects. In other cases there is no standard LFG account for the phenomenon \nwe wish to model and we adopt a compromise, conservative solution. This is exemplified by our treatment of Spanish periphrastic constructions. In yet another case, the less configurational nature of Spanish means that the LFG annotation algorithm has to rely mostly on Cast3LB function tags, and consequently a reliable method of adding those tags to parse trees had to be developed. This method achieves over 6% improvement over the baseline for the \nCast3LB-function-tag assignment task, and over 3% improvement over the baseline for LFG f-structure construction from function-tag-enriched trees.
This paper addresses a classical but important problem: The coupling of lexical tones and sentence intonation in tonal languages, such as Chinese, focusing particularly on voice fundamental frequency (F1) contours of speech. It is important because it forms the basis of speech synthesis technology and prosody analysis. We provide a solution to the problem with a constrained tone transformation technique based on structural modeling of the F1 contours. This consists of transforming target values in pairs from norms to variants. These targets are intended to sparsely specify the prosodic contributions to the F1 contours, while the alignment of target pairs between norms and variants is based on underlying lexical tone structures. When the norms take the citation forms of lexical tones, the technique makes it possible to separate sentence intonation from observed F0 contours. When the norms take normative F0 contours, it is possible to measure intonation variations from the norms to the variants, both having identical lexical tone structures. This paper explains the underlying scientific and linguistic principles and presents an algorithm that was implemented on computers. The method's capability of separating and combining tone and intonation is evaluated through analysis and re-synthesis of several hundred observed F0 contours.
Abstract We present a new method for learning to parse a bilingual sentence using Inversion Transduction Grammar trained on a parallel corpus and a monolingual treebank. The method produces a parse tree for a bilingual sentence, showing the shared syntactic structures of individual sentence and the differences of word order within a syntactic structure. The method involves estimating lexical translation probability based on a word-aligning strategy and inferring probabilities for CFG rules. At runtime, a bottom-up CYK-styled parser is employed to construct the most probable bilingual parse tree for any given sentence pair. We also describe an implementation of the proposed method. The experimental results indicate the proposed model produces word alignments better than those produced by Giza++, a state-of-the-art word alignment system, in terms of alignment error rate and F-measure. The bilingual parse trees produced for the parallel corpus can be exploited to extract bilingual phrases and train a decoder for statistical machine translation.
In the present paper, we examined the effects of autobiographically induced mood and music on emotional evaluations of and psychophysiological responses to music in 48 subjects. Participants listened to music after a mood induction. Both music and induction varied on the dimensions of valence (pleasant - unpleasant) and arousal (high - low). During mood induction and listening to music, psychophysiological responses were measured continuously to assess physiological arousal (indexed by electrodermal activity) and the valence of the emotional state (indexed by facial muscle activity) of the participant. After listening to music, participants evaluated the music using pictorial scales for valence and arousal. As expected, subjects were in a more positive emotional state during listening to pleasant than unpleasant music and also evaluated the music more positively after a pleasant compared to an unpleasant pre-existing mood. As also expected, high-arousal music and pre-existing mood generated both higher physiological arousal and higher arousal ratings compared to low-arousal pre-existing mood and music. We found no support for the principle of mood-congruency, which posits that individuals preferentially process emotional stimuli that are congruent in emotional tone with their current mood state.
Statistical parsers trained and tested on the Penn Wall Street Journal (WSJ) treebank have shown vast improvements over the last 10 years. Much of this improvement, however, is based upon an ever-increasing number of features to be trained on (typically) the WSJ treebank data. This has led to concern that such parsers may be too finely tuned to this corpus at the expense of portability to other genres. Such worries have merit. The standard "Charniak parser" checks in at a labeled precision-recall f-measure of 89.7% on the Penn WSJ test set, but only 82.9% on the test set from the Brown treebank corpus.This paper should allay these fears. In particular, we show that the reranking parser described in Charniak and Johnson (2005) improves performance of the parser on Brown to 85.2%. Furthermore, use of the self-training techniques described in (McClosky et al., 2006) raise this to 87.8% (an error reduction of 28%) again without any use of labeled Brown data. This is remarkable since training the parser and reranker on labeled Brown data achieves only 88.4%.
Recently, the Prague Dependency Treebank 2.0 (PDT 2.0) has emerged as the largest text corpora annotated on the level of tectogrammatical representation (“linguistic meaning”) described in Sgall et al. (2004) and containing about 0.8 milion words (see Hajič (2004)). We hope that this level of annotation is so close to the meaning of the utterances contained in the corpora that it should enable us to automatically transform texts contained in the corpora to the form of knowledge base, usable for information extraction, question answering, summarization, etc. We can use Multilayered Extended Semantic Networks (MultiNet) described in Helbig (2006) as the target formalism. In this paper we discuss the suitability of such approach and some of the main issues that will arise in the process. In section 1. we introduce formalisms underlying PDT 2.0 and MultiNet, in section 2. we describe the role MultiNet can play in the system of Functional Generative Description (FGD), section 3. discusses issues of automatic conversion to MultiNet and section 4. gives some conclusions. 1.
The current paper has a twofold objective. On the one hand, it describes the creation and the features of the Szeged Treebank, which is currently the largest manually processed Hungarian textual database serving as a reference material for research in natural language processing. On the other hand, detailed information is given about different experiments that aimed at the automatic recognition of syntactic structures with the use of machine learning algorithms. In order to provide comparable results, we applied methods of different categories, namely a rule-based, a logic and a numeric learner to pre-defined parsing problems. The aforementioned Szeged Treebank was used for the training and the testing of the algorithms.
The issue of information structure in language has been studied extensively both in the Prague School of Linguistics and in the Functional Generative Description [FGD, 11, 5]. This theory of representation of linguistic meaning is the framework for a family of multi-level annotation projects, in particular the Prague Dependency Treebank for Czech [PDT, 2, 3] and the Prague Arabic Dependency Treebank [PADT, 4, 12]. Information structure — the question of ‘the given’ and ‘the new’ in an utterance and how it is expressed — is considered to contribute to the linguistic meaning, and its annotation in PDT is part of the third, the most detailed and abstract level of linguistic description. Next to determining which elements in a sentence are context-bound and which are non-bound (the elementary distinctive feature from which the topic–focus dichotomy is derived), attention is also paid to the resolution of anaphoric relations and to capturing the communicative dynamism of a proposition (annotations of coreference and deep word order, respectively) [9]. In PADT, which now consists of the morphological and the analytical levels of description of Arabic, a similar annotation of information structure is being established. In our contribution, we would like to overview the theoretical concepts we work with, and present our formal treatment of a number of prototypical, yet corpus-based, instances of linguistic phenomena that have a principal impact on the structure of information in Arabic [6, 1]. The applicability of the general approach to written as well as spoken Arabic will be the main point of our account. In FGD, the description of information structure incorporates also the notions of intonation center and stress, contrast, subjective word order, or potential ellipsis, which are directly connected to prosody. We will provide references to other related computational research, too [cf. e.g. 8, 7, 10].
During auditory perception, neural oscillations are known to entrain to acoustic dynamics but their role in the processing of auditory information remains unclear. As a complex temporal structure that can be parameterized acoustically, music is particularly suited to address this issue. In a combined behavioral and EEG experiment in human participants, we investigated the relative contribution of temporal (acoustic dynamics) and nontemporal (melodic spectral complexity) dimensions of stimulation on neural entrainment, a stimulus-brain coupling phenomenon operationally defined here as the temporal coherence between acoustical and neural dynamics. We first highlight that low-frequency neural oscillations robustly entrain to complex acoustic temporal modulations, which underscores the fine-grained nature of this coupling mechanism. We also reveal that enhancing melodic spectral complexity, in terms of pitch, harmony, and pitch variation, increases neural entrainment. Importantly, this manipulation enhances activity in the theta (5 Hz) range, a frequency-selective effect independent of the note rate of the melodies, which may reflect internal temporal constraints of the neural processes involved. Moreover, while both emotional arousal ratings and neural entrainment were positively modulated by spectral complexity, no direct relationship between arousal and neural entrainment was observed. Overall, these results indicate that neural entrainment to music is sensitive to the spectral content of auditory information and indexes an auditory level of processing that should be distinguished from higher-order emotional processing stages.<b>NEW & NOTEWORTHY</b> Low-frequency (<10 Hz) cortical neural oscillations are known to entrain to acoustic dynamics, the so-called neural entrainment phenomenon, but their functional implication in the processing of auditory information remains unclear. In a behavioral and EEG experiment capitalizing on parameterized musical textures, we disentangle the contribution of stimulus dynamics, melodic spectral complexity, and emotional judgments on neural entrainment and highlight their respective spatial and spectral neural signature.
This paper presents a Constraint Grammar-inspired machine learner and parser, LingPars, that assigns dependencies to morphologically annotated treebanks in a function-centred way. The system not only bases attachment probabilities for PoS, case, mood, lemma on those features' function probabilities, but also uses topological features like function/PoS n-grams, barrier tags and daughter-sequences. In the CoNLL shared task, performance was below average on attachment scores, but a relatively higher score for function tags/deprels in isolation suggests that the system's strengths were not fully exploited in the current architecture.
Grammar induction is one of the most important research areas of the natural language processing. The lack of a large Treebank, which is required in supervised grammar induction, in some natural languages such as Persian encouraged us to focus on unsupervised methods. We have found the Inside-Outside algorithm, introduced by Lari and Young, as a suitable platform to work on, and augmented IO with a history notion. The result is an improved unsupervised grammar induction method called History-based IO (HIO). Applying HIO to two very divergent natural languages (i.e., English and Persian) indicates that inducing more conditioned grammars improves the quality of the resultant grammar. Besides, our experiments on ATIS and WSJ show that HIO outperforms most current unsupervised grammar induction methods.
Linguists use treebanks as resource for collecting evidence of phenomena which cannot be easily recovered from data that is annotated at word level only, this includes collecting quantitative data, getting non-categorical information such as heaviness or finding natural sounding counter examples 1 (e.g. Uszkoreit et al. (1998); Arnold et al. (2000); Bresnan et al. (to appear)) 2. Tools such as TIGERSearch allow us easy access to the encoded information. 3 This poster presents work on the Tübinger Baumbank deutscher Zeitungssprache (Tüba-D/Z). It describes the encoding of coordination phenomena in the treebank and gives a qualitative and quantitative survey. 2 The TüBa-D/Z Treebank It is a corpus of newspaper texts which currently comprises about 22 000 sentences (more than 381 000 tokens) taken from the Wissenschafts-CD of ’die tageszeitung ’ (taz). The annotation combines information on inflectional morphology, part of speech, phrase structure (or rather recursive chunking), grammatical dependencies and topological fields. In addition, it includes marking of named entities and annotation of anaphoric and coreference relations (cf. Hinrichs et al. (2004)).
In this paper, a new approach of using temporal information to assist in Mandarin speech recognition is discussed. It incorporates two types of temporal information into the recognition search. One is a statistical syllable duration model which considers the influences of 411 basesyllables, 5 tones, 4 position-in-word factors, and 3 positionin-sentence factors on syllable duration. Another is the timing information of modeling three types of inter-syllable boundary including intra-word, inter-word without punctuation mark (PM), and inter-word with PM. The uses of these two types of temporal information are expected to be useful for improving the segmentation accuracies in both acoustic decoding and linguistic decoding. Experimental results showed that the base-syllable/character/word recognition rates were slightly improved for both MATBN and Treebank datbase.
Each year the Conference on Computational Natural Language Learning (CoNLL) features a shared task, in which participants train and test their systems on exactly the same data sets, in order to better compare systems. The tenth CoNLL (CoNLL-X) saw a shared task on Multilingual Dependency Parsing. In this paper, we describe how treebanks for 13 languages were converted into the same dependency format and how parsing performance was measured. We also give an overview of the parsing approaches that participants took and the results that they achieved. Finally, we try to draw general conclusions about multi-lingual parsing: What makes a particular language, treebank or annotation scheme easier or harder to parse and which phenomena are challenging for any dependency parser?
Motivationally relevant stimuli have been shown to receive prioritized processing compared to neutral stimuli at distinct processing stages. This effect has been related to the evolutionary importance of rapidly detecting dangers and potential rewards and has been shown to be modulated by the distance between an organism and a faced stimulus. Similarly, recent studies showed degrees of emotional modulation of autonomic responses and subjective arousal ratings depending on stimulus size. In the present study, affective modulation of pictures presented in different sizes was investigated by measuring event-related potentials during a two-choice categorization task. Results showed significant emotional modulation across all sizes at both earlier and later stages of processing. Moreover, affective modulation of earlier processes was reduced in smaller compared to larger sizes, whereas no changes in affective modulation were observed at later stages.
In this paper, we attempt to automatically annotate the Penn Chinese Treebank with semantic dependency structure. Ini-tially a small portion of the Penn Chinese Treebank was man-ually annotated with headword and semantic dependency re-lations. An initial investigation is then done using a Naive Bayesian Classifier and some handcrafted rules. The results show that the algorithms and proposed approach are effective at determining semantic dependency structure automatically. The Naive Bayesian Classifier makes a good baseline algo-rithm for future research.
In data-oriented English-Chinese machine translation, knowledge source is the very important basis for translation processing. This paper presents a kind of construction strategy for knowledge source which contains affluent grammatical and syntactical information. Firstly, taking lexical function grammar as the theoretical basis, treebank including parse trees converted from every sentence in the source language corpus is acquired. Secondly, based on the decomposition algorithm, the corresponding fragment-bank composed of all the legal fragments extracted from the treebank is constructed. Finally, based on the combination algorithm, the fragment-combination-bank including all the possible fragment-combination forms of every parse tree in the treebank is built. Based on the successful construction of the knowledge source, the whole machine translation process can be implemented efficiently and accurately.
CT resulted in variable functional and structural changes in dementia, and conclusions are limited by heterogeneity and study quality. Larger, more robust studies are required to correlate these findings with clinical benefits from CT.
Abstract In the past decade the Albanian language has undergone a period of significant change in terms of lexical development. These developments are almost entirely attributable to extralinguistic factors. The turbulent transition from a centralized socialist system to a market economy in the 1990s was accompanied by an immediate opening of Albania to the outside world. This upheaval also brought about a social and intellectual liberation, a change in mental outlook, and freedom from the weight of dictatorship and the totalitarian mindset. The transformation from collectivized property to private property and the birth of political and cultural pluralism constituted a novel variable within Albanian society. The pluralism of press and broadcast media and the phenomenal growth in the number of political parties are two important constituents for change in the lexical norm of the Albanian language, both in a positive and negative sense. A predisposition towards direct contact with Western languages and cultures, interpreted as a sign of the “Europeanization” of Albania, led to the proliferation of many foreign words, primarily English, into the lexicon of Albanian, especially in the domains of economics, trade, jurisprudence, information technology, politics, and administration. This study concerns itself with interpretation, classification, and examples of these changes.
The amygdala is closely linked to basal ganglia circuitry and plays a key role in danger detection and fear-potentiated startle. Based on recent findings of amygdalar abnormalities in Parkinson's disease, we hypothesized that non-demented patients with this illness would show blunted reactivity during aversive/unpleasant events, as indexed by diminished emotional modulation of the startle eyeblink response. To test this hypothesis, 23 idiopathic patients with Parkinson's disease and 17 controls viewed standardized sets of aversive, pleasant and neutral pictures for 6 s each. During this time, white noise bursts (50 ms, 95 db) were binaurally presented to elicit startle eyeblink responses, measured from electrodes over the orbicularis oculi. After viewing each picture, subjects provided ratings of valence and arousal. The Parkinson's disease patients were in the early to middle stages of their disease, not demented or depressed, and were tested 'on' dopaminergic medication. The two groups were similar in age, education, gender and cognitive screening status. The control group had larger startle responses when viewing negative, aversive pictures than neutral or pleasant pictures. As predicted, startle enhancement during aversive pictures was significantly muted in the Parkinson's disease patients. This blunting was not due to abnormalities in the mechanics of the startle eyeblink per se. Nor was it related to depression symptoms, medications (psychotropics), or failure to perceive/appreciate the negative meaning of aversive pictures (i.e. normal valence ratings). Reduced startle reactivity in the disease group was related to disease severity (Hoehn-Yahr) and occurred in the context of reduced arousal ratings of aversive pictures. These findings of blunted startle reactivity add to the literature on emotional changes associated with Parkinson's disease. The basis for this muted reactivity is unknown but may involve an amygdala-based translational defect whereby the results of cognitive appraisal are not appropriately transcoded into somato-motor-arousal responses normally associated with an aversive motivational state. This may arise from faulty dopaminergic gating of the amygdala, resulting in 'inhibition' of the amygdala in the manner described by Marowsky et al. (Marowsky A, Yanagawa Y, Obata K, Vogt E. Neuron 2005; 48: 1025-37). More broadly, the findings of muted reactivity to aversive stimuli may reflect a 'bradylimbic' affective disturbance in patients with Parkinson's disease. Future studies are needed to address whether the physiologic blunting observed here might be a useful correlate of apathy.
The Ferret copy detector has been used since 2001 to find plagiarism in large collections of students’ coursework in English. This article reports on extending its application to Chinese, with experiments on corpora of coursework collected from two Chinese universities. Our experiments show that Ferret can find both artificially constructed plagiarism and actually occurring, previously undetected plagiarism. We discuss issues of representation, focus on the effectiveness of a sub-symbolic approach, and show that Ferret does not need to find word boundaries first.
Traditionally, orthographic variants have been modelled as different ways of spelling the same word - described at the level of the lexeme. But when inflection is taken into account, this runs into a problem: different citation forms have different inflectional paradigm - and orthographic variation does not merely affect the citation form, but the entire paradigm. The MorDebe database therefore models orthographic variation as a relation between distinct, yet still token-identical lexemes. This paper discusses the advantage of that approach, and the full set of practical problems that arose during the structural treatment of orthographic variation in the MorDebe database.
Syntactic parsing requires a fine balance between expressivity and complexity, so that naturally occurring structures can be accurately parsed without compromising efficiency. In dependency-based parsing, several constraints have been proposed that restrict the class of permissible structures, such as projectivity, planarity, multi-planarity, well-nestedness, gap degree, and edge degree. While projectivity is generally taken to be too restrictive for natural language syntax, it is not clear which of the other proposals strikes the best balance between expressivity and complexity. In this paper, we review and compare the different constraints theoretically, and provide an experimental evaluation using data from two treebanks, investigating how large a proportion of the structures found in the treebanks are permitted under different constraints. The results indicate that a combination of the well-nestedness constraint and a parametric constraint on discontinuity gives a very good fit with the linguistic data.
CONTEXT: There is extensive evidence implicating dysfunctions in stress responses and adaptation to stress in the pathophysiological mechanism of major depressive disorder (MDD) in humans. Endogenous opioid neurotransmission activating mu-opioid receptors is involved in stress and emotion regulatory processes and has been further implicated in MDD. OBJECTIVE: To examine the involvement of mu-opioid neurotransmission in the regulation of affective states in volunteers with MDD and its relationship with clinical response to antidepressant treatment. DESIGN: Measures of mu-opioid receptor availability in vivo (binding potential [BP]) were obtained with positron emission tomography and the mu-opioid receptor selective radiotracer carbon 11-labeled carfentanil during a neutral state. Changes in BP during a sustained sadness challenge were obtained by comparing it with the neutral state, reflecting changes in endogenous opioid neurotransmission during the experience of that emotion. SETTING: Clinics and neuroimaging facilities at a university medical center. PARTICIPANTS: Fourteen healthy female volunteers and 14 individually matched patient volunteers diagnosed with MDD were recruited via advertisement and through outpatient clinics. INTERVENTIONS: Sustained neutral and sadness states, randomized and counterbalanced in order, elicited by the cued recall of an autobiographical event associated with that emotion. Following imaging procedures, patients underwent a 10-week course of treatment with 20 to 40 mg of fluoxetine hydrochloride. MAIN OUTCOME MEASURES: Changes in mu-opioid receptor BP during neutral and sustained sadness states, negative and positive affect ratings, plasma cortisol and corticotropin levels, and clinical response to antidepressant administration. RESULTS: The sustained sadness condition was associated with a statistically significant decrease in mu-opioid receptor BP in the left inferior temporal cortex of patients with MDD and correlated with negative affect ratings experienced during the condition. Conversely, a significant increase in mu-opioid receptor BP was observed in healthy control subjects in the rostral region of the anterior cingulate. In this region, a significant decrease in mu-opioid receptor BP during sadness was observed in patients with MDD who did not respond to antidepressant treatment. Comparisons between patients with MDD and controls showed significantly lower neutral-state mu-opioid receptor BP in patients with MDD in the posterior thalamus, correlating with corticotropin and cortisol plasma levels. Larger reductions in mu-opioid system BP during sadness were obtained in patients with MDD in the anterior insular cortex, anterior and posterior thalamus, ventral basal ganglia, amygdala, and periamygdalar cortex. The same challenge elicited larger increases in the BP measure in the control group in the anterior cingulate, ventral basal ganglia, hypothalamus, amygdala, and periamygdalar cortex. CONCLUSIONS: The results demonstrate differences between women with MDD and control women in mu-opioid receptor availability during a neutral state, as well as opposite responses of this neurotransmitter system during the experimental induction of a sustained sadness state. These data demonstrate that endogenous opioid neurotransmission on mu-opioid receptors, a system implicated in stress responses and emotional regulation, is altered in patients diagnosed with MDD.
In this thesis lateralization of olfactory functions was investigated by both behavioral and electrophysiological assessment, the latter with the olfactory event-related potential (OERP) technique. The olfactory sense is primarily ipsilateral in that a stimulus that is presented to one nostril is initially processed in the same hemisphere. This makes it possible to observe differences between stimulated nostrils as an indication of hemispheric difference. Study I explored differences in olfactory cognitive functions with respect to side of rhinal stimulation and demonstrated that familiarity ratings are higher at right- compared to left-nostril stimulation. No differences were found in episodic recognition memory or free identification, possibly reflecting inter-hemispheric interactions in higher cognitive functions. Effects of repetition priming were present in odor identification and tended to be more pronounced when tested via left nostril. Study II further investigated the effect of previous exposure in odor identification by a different experimental set-up, and demonstrated effects of repetition priming when tested via left- but not right-nostril stimulation. This finding indicates the importance of reconsidering possible sequential effects in olfactory research. Study III examined methodological aspects of an OERP protocol with respect to stimulus duration, which was used in Study IV. No differences in amplitudes or latencies where found between the stimulus durations of 150, 200 and 250 ms, suggesting the commonly used duration of 200 ms in a standard protocol. Study IV investigated laterality effects in OERPs with respect to side of stimulation and electrode site. The results showed consistent amplitudes and latencies regardless of rhinal side of stimulation. Larger amplitudes were demonstrated on left hemisphere and midline compared to right hemisphere, possibly explained by smaller N1/P2 amplitudes at the right-hemisphere sites at left-nostril stimulation. Apart from a proposed OERP protocol, the findings support the notions of a right-hemisphere predominance in processes related to olfactory perception and indicate, in accordance with other findings, a left-side advantage in conceptual repetition priming.
Detecting idioms in a sentence is important to sentence understanding. This paper discusses the linguistic knowledge for idiom detection. The challenges are that idioms can be ambiguous between literal and idiomatic meanings, and that they can be “transformed” when expressed in a sentence. However, there has been little research on Japanese idiom detection with its ambiguity and transformations taken into account. We propose a set of linguistic knowledge for idiom detection that is implemented in an idiom dictionary. We evaluated the linguistic knowledge by measuring the performance of an idiom detector that exploits the dictionary. As a result, more than 90% of the idioms are detected with 90% accuracy.
Spoken dialogue systems (SDSs) can be used to operate devices, e.g. in the automotive environment. People using these systems usually have different levels of experience. However, most systems do not take this into account. In this paper, we present a method to build a dialogue system in an automotive environment that automatically adapts to the user’s experience with the system. We implemented the adaptation in a prototype and carried out exhaustive tests. Our usability tests show that adaptation increases both user performance and user satisfaction. We describe the tests that were performed, and the methods used to assess the test results. One of these methods is a modification of PARADISE, a framework for evaluating the performance of SDSs [Walker MA, Litman DJ, Kamm CA, Abella A (Comput Speech Lang 12(3):317–347, 1998)]. We discuss its drawbacks for the evaluation of SDSs like ours, the modifications we have carried out, and the test results.
The Dutch spelling system, like other European spelling systems, represents a certain balance between preserving the spelling of morphemes (the morphological principle) and obeying letter-to-sound regularities (the phonological principle). We present experimental results with artificial learners that show a competition effect between the two principles: adhering more to one principle leads to more violations of the other. The artificial learners, memory-based learning algorithms, are trained (1) to convert written words to their phonemic counterparts and (2) to analyze written words on their morphological composition, based on data extracted from the CELEX lexical database. As an exception to the competition effect we show that introducing the schwa as a letter in the spelling system causes both morphology and phonology to be learnt better by the artificial learners. In general we argue that artificial learning studies are a tool in obtaining objective measurements on a spelling system that may be of help in spelling reform processes. (PsycINFO Database Record (c) 2016 APA, all rights reserved)
With increasing numbers of Web users, there is a necessity to improve their Web site navigation experience over the Internet and a range of Web applications have emerged recently for this purpose. Many researchers have stressed the importance of identifying semantic relatedness of Web pages in such Web applications as Web site navigation, automatic tour generation and adaptive Web applications. One approach to identifying semantic relatedness between documents is to use lexical databases and lexical chains. For example, an approach using lexical chains has been proposed by Green for identifying paragraph similarity in a document [1]. However, due to the unacceptable length of time needed for lexical chaining and the difficulty of global representation of documents, Green used synset weight vectors to compare semantic relatedness between two documents. But his approach to identifying paragraph similarities can be extended to identify semantic similarities between documents. In this study, an approach to identifying semantic similarity between Web pages incorporating weighted lexical chains (SRWLC) and document properties based on reiteration, density, length and semantic distance is proposed. The two approaches (the proposed approach and Green’s approach - SR Green ) were empirically compared by determining the semantic relatedness of Web pages using human subjects. The research hypothesis of this research is that the proposed approach identifies significantly more semantically related pages with a higher precision than the approach that has been proposed by Green. The null hypothesis is that there is no significant different in identification of semantically related pages between the two approaches. In this context precision is defined as the proportion of retrieved pages that are relevant. Web pages belonging to the Department of Computer Science, Keele University are used for the empirical evaluation of the two methods. The semantic relatedness of all pages was identified using both approaches and a Web-based page categorisation exercise using human subjects was carried out for this empirical evaluation. The evaluation is Web based, and therefore can be carried out on the subjects’ preferred Web browser at his or her preferred time & place. Therefore the distortion effects are minimised and the results of the evaluation are realistic and also can be reliably generalised to some extent. An invitation e-mail was sent to twenty subjects during the first week of October 2004 giving a link and guidelines for them to start the experiment. Once the link on the email was clicked, subjects were shown the initial Web page, giving an introduction to the experiment and instructions on how to continue. Twelve out of twenty invited subjects completed the experiment. Two subjects attempted the experiments but couldn’t finish because of network problems; another three subjects couldn’t finish because of time restrictions, and the other three subjects did not respond at all. Therefore responses from only twelve subjects are used for experimental evaluation. The Wilcoxon signed ranks test returns a p value of 0.004, indicating that the null hypothesis can be rejected, and that there is evidence to suggest that the SR WLC approach identifies significantly more semantically-related pages with higher precision than the SR Green approach. The SRWLC approach should be evaluated further using different Web site contents and language styles (e.g. American/British English). It would be interesting to use more subjects from different backgrounds to do the evaluation. This would determine whether the results of the evaluation are influenced by the human subject’s background, such as their status (student, staff, or other) his familiarity with the pages of the test database, gender differences or the level of English knowledge. Most importantly the approach is believed to be valid for more general browsing environments than a computer science Website and a wider study is desirable.
This paper focuses on the electronic literacy practices of two Korean-American heritage language learners who manage Korean weblogs.Online users deliberately alter standard forms of written language and play with symbols, characters, and words to economize typing effort, mimic oral language, or convey qualities of their linguistic identity such as gender, age, and emotional states.However, little is known about the impact of computer-mediated nonstandard language use on heritage learners' linguistic development.Through in-depth case studies of two siblings, the study examines the linguistic and pragmatic practices of these learners online and the perceived effects of non-standard forms of computer-mediated language on their heritage language development and maintenance.The data show that electronic literacy practices provide authentic opportunities to use the language and support the development of a social network of Korean speakers, which results in greater sociopsychological attachment to the Korean language and culture.The informants report that the deviant language forms found in e-texts enable them to engage in online interactions without the pressures of having to spell the words correctly.However, they express frustrations in not being able to distinguish between correct and non-standard forms of the language, which appear to be affecting their offline language use. THE KOREAN CONTEXTThe Republic of Korea has one of the fastest-growing cybercommunities in the world.According to the Korea Network Information Center, over 63% of the entire South Korean population are Internet users, and 95% of individuals in the 6-29 age bracket report using it on a daily basis.Internet sites that enable users to create "personal spaces" to share and document their changing lives and keep connected with people they know are immensely popular among Koreans.A case in point is "Cyworld," an upgraded blog that features chatting, commentaries, pictures, music, a guest book, avatars and links to other homepages prompting users to network with their friends, family, and colleagues.As of August 2005, there are over 11 million Cyworld registered users.Participation in online forums such as Cyworld engages its members in a social process of learning through shared practices, internally constructed membership, and the formation of personal and group identities (Holmes & Jin Sook Lee Electronic literacy and heritage language maintenance Language Learning & Technology 94Myerhoff, 1999).Members are involved in a community of practice, where a group of people who come together around a joint enterprise develop common beliefs, values, and ways of doing things, which all influence the ways in which members communicate with one another (Eckert, 2000;Wenger, 1998).New forms of expression are constantly being negotiated and shared among online users, making it difficult to keep current with the changing face of electronic text.Computer-mediated communication is unique in that, despite its similarities to oral speech, it invites substantial deregulation effects on communication, which can foster the use of creative, non-standard language play (Sproull & Kiesler, 1986).Studies have documented non-standard 1 uses of language in online interactions (a) to mark certain individual characteristics such as provincial dialects, social class, gender, age, and/or personality traits, (b) to economize typing efforts, and/or (c) to mimic spoken language (Barnes, 2003;Herring, 2001;Song, 2002;Sproull & Kiesler, 1986).For example, Su ( 2004) found an emergent mock Taiwanese accent among Internet users as a form of language play to jointly construct "a young, lively, congenial, and witty presence" (p.61).Androutsopoulos ( 2000) also revealed that non-standard orthography in online fan media texts was representative of spoken language and purely graphemic modifications, which are used to serve as contextualization cues and cues of subcultural positioning.Although all natural languages inevitably change over time, drastic deviances from standard language ranging from non-standard orthography and incorrect grammar to unfamiliar lexical items and symbols have brought forth great concern about the preservation of standard orthography, grammar, and pragmatic uses of the Korean language (Choi, 2003;Kim, 2005;Park, 1989).Educators across grade levels in Korea are reporting that students display electronic textual features in their school work: they have difficulty with spelling and with the proper word spacing used to delineate word boundaries due to non-standard ways of Internet language use, which flout conventional norms of literacy practices (Ahn, 2000;Choi, 2003;Kim, 2005;Noh, 2000).For young children and Korean as foreign/second language learners who have not fully acquired literacy in the language, exposure to electronic texts may have adverse effects on their language development.However, Meskill, Mossop, and Bates (1999) state that "children in the age of electronic text are developing unique skills and strategies for inventing novel forms of understanding these texts that are quite often independent of formal instructional ('school') literacy training" (p.4), thus, highlighting the positive ways in which the development of electronic texts can benefit students' cognitive flexibility and skills.