Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
To test the effects of cortisol on affective experience, the authors orally administered a placebo, 20 mg cortisol, or 40 mg cortisol to 85 men. Participants' affective responses to negative and neutral stimuli were measured. Self-reported affective state was also assessed. Participants in the 40-mg group (showing extreme cortisol elevations within the physiological range) rated neutral stimuli as more highly arousing than did participants in the placebo and 20-mg groups. Furthermore, within the 20-mg group, individuals with higher cortisol elevations made higher arousal ratings of neutral stimuli. However, cortisol was unrelated to self-reported affective state. Thus, findings indicate that acute cortisol elevations cause heightened arousal in response to objectively nonarousing stimuli, in the absence of effects on mood.
The development of technologies for monitoring the welfare of crewmembers is a critical requirement for extended spaceflight. Behavior analytic methodologies provide a framework for studying the performance of individuals and groups, and brief computerized tests have been used successfully to examine the impairing effects of sleep, drug, and nutrition manipulations on human behavior. The purpose of the present study was to evaluate the feasibility and sensitivity of repeated performance testing during spaceflight. Four National Aeronautics and Space Administration crewmembers were trained to complete computerized questionnaires and performance tasks at repeated regular intervals before and after a 10-day shuttle mission and at times that interfered minimally with other mission activities during spaceflight. Two types of performance, Digit-Symbol Substitution trial completion rates and response times during the most complex Number Recognition trials, were altered slightly during spaceflight. All other dimensions of the performance tasks remained essentially unchanged over the course of the study. Verbal ratings of Fatigue increased slightly during spaceflight and decreased during the postflight test sessions. Arousal ratings increased during spaceflight and decreased postflight. No other consistent changes in rating-scale measures were observed over the course of the study. Crewmembers completed all mission requirements in an efficient manner with no indication of clinically significant behavioral impairment during the 10-day spaceflight. These results support the feasibility and utility of computerized task performances and questionnaire rating scales for repeated measurement of behavior during spaceflight.
The measure of lexical repetition constitutes one of the variables used to determine the lexical richness of literary texts, a value further employed in authorship attribution studies. Although most of the constants for lexical richness actually depend on text length, Yule’s characteristic is considered to be highly reliable for being text length independent. It is not the aim of this paper questioning the validity of K to measure the lexical repeat-rate, nor to evaluate its usefulness in authorship studies, but to review the most accurate procedure to calculate its value in the light of the lack of standardization found in the specific literature. At the same time, the peculiar calculation of Yule’s K by TACT is explained. Our study suggests that standardization will certainly help improve the studies where K is employed.
We investigated the performance efficacy of beam search parsing and deep parsing techniques in probabilistic HPSG parsing using the Penn treebank. We first tested the beam thresholding and iterative parsing developed for PCFG parsing with an HPSG. Next, we tested three techniques originally developed for deep parsing: quick check, large constituent inhibition, and hybrid parsing with a CFG chunk parser. The contributions of the large constituent inhibition and global thresholding were not significant, while the quick check and chunk parser greatly contributed to total parsing performance. The precision, recall and average parsing time for the Penn treebank (Section 23) were 87.85%, 86.85%, and 360 ms, respectively.
BACKGROUND: In the present study neurophysiological correlates related to mismatching information in lexical access were investigated with a fragment priming paradigm. Event-related brain potentials were recorded for written words following spoken word onsets that either matched (e.g., kan - Kante [Engl. edge]), partially mismatched (e.g., kan - Konto [Engl. account]), or were unrelated (e.g., kan - Zunge [Engl. tongue]). Previous psycholinguistic research postulated the activation of multiple words in the listeners' mental lexicon which compete for recognition. Accordingly, matching words were assumed to be strongly activated competitors, which inhibit less strongly activated partially mismatching words. RESULTS: ERPs for matching and unrelated control words differed between 300 and 400 ms. Difference waves (unrelated control words - matching words) replicate a left-hemispheric P350 effect in this time window. Although smaller than for matching words, a P350 effect and behavioural facilitation was also found for partially mismatching words. Minimum norm solutions point to a left hemispheric centro-temporal source of the P350 effect in both conditions. The P350 is interpreted as a neurophysiological index for the activation of matching words in the listeners' mental lexicon. In contrast to the P350 and the behavioural responses, a brain potential ranging between 350 and 500 ms (N400) was found to be equally reduced for matching and partially mismatching words as compared to unrelated control words. This latter effect might be related to strategic mechanisms in the priming situation. CONCLUSION: A left-hemispheric neuronal network engaged in lexical access appears to be gradually activated by matching and partially mismatching words. Results suggest that neural processing of matching words does not inhibit processing of partially mismatching words during early stages of lexical identification. Furthermore, the present results indicate that neurophysiological correlates observed in fragment priming reflect different aspects of target processing that are cumulated in behavioural responses. Particularly the left-hemispheric P350 difference potential appears to be closely related to fine-grained activation differences of modality-independent representations in the listeners' mental lexicon. This neurophysiological index might guide future studies aimed at investigating neural aspects of lexical access.
The aim of this paper is to investigate a case of transfer within the context of language death. By examining data from Jersey Norman French (known to its speakers as Jèrriais) it illustrates the difficulty in determining linguistic norms for this relatively undocumented variety and suggests possible strategies to overcome this problem. The study compares systematically the occurrence of overt and covert transfer in the speech of a sample of fifty native speakers of Jèrriais via the analysis of a number of linguistic variables. The extent to which transfer-induced changes are themselves becoming established as norms within this speech community will also be considered.
Broadly diffused languages, like Spanish, enjoy diverse variations. In the international space common to Spanish, such as that covered by radio and television, the most sought after variants are those with the largest audience, although criteria are often subjective. In order to make decisions in this matter, the article offers proposals based on dispersion criteria (e.g. the number of countries that use a given norm) and population criteria (number of speakers). To exemplify this, the phonetic norms most frequently heard internationally are described, and cases of lexical variation are presented. The conclusion is that the media, promote national and international standardization turning it is to their advantage.
This first issue of Language Resources and Evaluation is dedicated to the memory of Antonio Zampolli, whom few would dispute is the one person who has led the way in promoting and establishing the development of language resources (LR) of all kinds for the past four decades. In this inaugural issue, we have attempted to bring together articles by major figures in the field in order to provide an overview of the history, state of the art, and the future of the creation, annotation, exploitation, evaluation, and distribution of LR. Hopefully, this collection of articles will serve not only as a tribute to Antonio, but also as a framework out of which this journal – which almost certainly would not have existed were it not for him – can grow.
We report two studies in which the interplay between stimulus properties and perceiver characteristics in the appreciation car interiors was investigated. In Experiment 1 three design components, complexity, curvature and innovativeness, which are all thought to affect design appreciation were combined in a fully factorial design. All dimensions were confirmed to affect ratings, and curvature and innovativeness particularly affected the attractiveness ratings. Curved and non-innovative designs were generally preferred. Moreover, participants who were particularly interested in art were more sensitive to curvature and innovativeness. In Experiment 2 two dimensions of Experiment 1 were replicated using similar stimuli. Moreover, the specific effects of a design knowledge treatment were investigated. Results replicated the preference for curved and non-innovative (rather classic) designs. The treatment had only small effects, which support a general rather than dimension-specific effects of cognitive pre-information. Copyright © 2005 John Wiley & Sons, Ltd.
In this paper, we discuss the role that temporal information plays in natural language text, specifically in the context of question answering systems. We define a descriptive framework with which we can examine the temporally sensitive aspects of natural language queries. We then investigate broadly what properties a general specification language would need, in order to mark up temporal and event information in text. We present a language, TimeML, which attempts to capture the richness of temporal and event related information in language, while demonstrating how it can play an important part in the development of more robust question answering systems.
The paper presents WordNet and EuroWordNet, lexical databases in which concepts are linked with lexical and semantic relations. In the past decade, similar lexical databases for other languages have been based on them, for which two approaches have mostly been used: the creation of independent wordnets and their subsequent integration into multilingual lexical databases, and the expansion of existing lexical databases into the target language by translating the entries and adopting relations between them. Due to the great automation potential of the network creation process and a uniform structure of the resulting databases, the latter approach has become very popular. This is why the given model was tested in this paper for its ability to create a Slovene wordnet, and the results were compared to its language-motivated counterpart acquired from the Fida corpus as well as other mono- and multilingual language resources for Slovene.
This paper presents a lexicalized HMM-based approach to Chinese text chunking. To tackle the problem of unknown words, we formalize Chinese text chunking as a tagging task on a sequence of known words. To do this, we employ the uniformly lexicalized HMMs and develop a lattice-based tagger to assign each known word a proper hybrid tag, which involves four types of information: word boundary, POS, chunk boundary and chunk type. In comparison with most previous approaches, our approach is able to integrate different features such as part-of-speech information, chunk-internal cues and contextual information for text chunking under the framework of HMMs. As a result, the performance of the system can be improved without losing its efficiency in training and tagging. Our preliminary experiments on the PolyU Shallow Treebank show that the use of lexicalization technique can substantially improve the performance of a HMM-based chunking system.
We present a classifier-based parser that produces constituent trees in linear time. The parser uses a basic bottom-up shift-reduce algorithm, but employs a classifier to determine parser actions instead of a grammar. This can be seen as an extension of the deterministic dependency parser of Nivre and Scholz (2004) to full constituent parsing. We show that, with an appropriate feature set used in classification, a very simple one-path greedy parser can perform at the same level of accuracy as more complex parsers. We evaluate our parser on section 23 of the WSJ section of the Penn Treebank, and obtain precision and recall of 87.54% and 87.61%, respectively.
The annotations of the Penn Discourse Treebank (PDTB) include (1) discourse connectives and their arguments, and (2) attribution of each argument of each connective and of the relation it denotes. Because the PDTB covers the same text as the Penn TreeBank WSJ corpus, syntactic and discourse annotation can be compared. This has revealed significant differences between syntactic structure and discourse structure, in terms of the arguments of connectives, due in large part to attribution. We describe these differences, an algorithm for detecting them, and finally some experimental results. These results have implications for automating discourse annotation based on syntactic annotation.
The main objective of this study is to determine the impact of immigration interpreters on the testimony of Spanish-English bilingually conducted hearings in one U.S. immigration court. Specifically, I analyze the performance of nine immigration interpreters. I identify the precise linguistic strategies they employ when interpreting and, using conversational and other discourse analytical approaches, determine how they become active members of the proceedings. The immigration hearings I observed took place in one Federal immigration courtroom located in a large northeastern city. This research shows the extent to which interpreters play a pivotal role in controlling courtroom discourse—constructing courtroom reality and either mitigating or magnifying the culpability of defendants through a variety of linguistic mechanisms: a) inaccurate lexical choice, b) the use of source language rather than target language words and phrases, c) the use of definitions and calques, d) the improper addition or deletion of repair mechanisms and of hesitation forms such as pauses and fillers, and e) the addition of polite forms of address to convey solidarity, to adhere to Hispanic cultural norms, and to avoid face threatening acts. This study shows that the linguistic power interpreters wield exerts a coercive force, particularly on witnesses and defendants, and that such linguistic coerciveness on the part of interpreters influences other participants in the judicial proceeding. In this study, both judges and attorneys are shown to have been influenced by the lexical choices of interpreters. Finally, I show that the intrusiveness of interpreters changes the pragmatic force intended by the speakers, which constitutes a violation of the ethical standards set for interpreters in the United States by such authorities as the Federal Judicial Center.
Spoken language proficiency is intuitively related to effective and efficient communication in spoken interactions. However, it is difficult to derive a reliable estimate of spoken language proficiency by situated elicitation and evaluation of a person’s communicative behavior. This paper describes the task structure and scoring logic of a group of fully automatic spoken language proficiency tests (for English, Spanish and Dutch) that are delivered via telephone or Internet. Test items are presented in spoken form and require a spoken response. Each test is automatically-scored and primarily based on short, decontextualized tasks that elicit integrated listening and speaking performances. The tests present several types of tasks to candidates, including sentence repetition, question answering, sentence construction, and story retelling. The spoken responses are scored according to the lexical content of the response and a set of acoustic base measures on segments, words and phrases, which are scaled with IRT methods or parametrically combined to optimize fit to human listener judgments. Most responses are isolated spoken phrases and sentences that are scored according to their linguistic content, their latency, and their fluency and pronunciation. The item development procedures and item norming are described.
Most existing statistical surface realizers either make use of hand-crafted grammars to provide coverage or are tuned to specific applications. This paper describes an initial effort toward building a statistical surface realization model that provides both precision and coverage. We trained a Maximum Entropy model that given a predicate-argument semantic representation, predicts the surface form for realizing a semantic concept and the ordering of sibling semantic concepts and their parent, on the Penn TreeBank and Proposition Bank corpora. Initial results have shown that the precisions for predicting surface forms and orderings reached 80% and 90% respectively, on a held-out part of Penn TreeBank. We use the model to generate sentences from our domain representations. We are in the process of evaluating the model on a corpus collected for our in-car applications.
Norms may also be understood as social realization of correctness notions and linguistic norms as performance instructions. This article reports on a case study of the application of Toury' s norm theory, particularly his operative translation norms, in subtitle translation. The authors believe that is a norm-governed communicative activity between two or more languages, and that subtitle translation, governed by linguistic and textual norms, should seek invisibility of subtitling as the ultimate goal.
The Korean Treebank Annotations Version 2.0 is a second volume of The Korean Treebank Annotations (Palmer et al., 2002; Han et al., 2002). It contains new texts that are from the news domain: the original corpus for the Korean Treebank 2.0 was extracted from The Korean Newswire corpus published by LDC, catalog number LDC2000T45. The Korean Treebank Annotations Version 2.0 consists of 647 news articles in 112 files which contain 132,040 words and 5,010 sentences. There are 40,252 unique words and 13,844 unique morphemes (12,681 unique morphemes excluding foreign characters and arabic numbers). The annotated text measures about 8.5MB in size.\nWhile annotating the new texts, many new linguistic constructions and phenomena were encountered which called for setting additional guidelines. Furthermore, a few guidelines used for the first volume of the Korean Treebank were re-examined and modified in the second volume. This document outlines the guidelines that were newly introduced for the second volume of the Penn Korean Treebank, as well as the ones that have been revised since the publication of volume 1.0. Therefore, this is not a self-contained document, but is rather an addendum to the two previously published guidelines for the Penn Korean Treebank (Han and Han, 2001; Han et al., 2001).
Reviewed by: Medieval Cruelty: Changing Perceptions, Late Antiquity to the Early Modern Period Dianne Hall Baraz, Daniel, Medieval Cruelty: Changing Perceptions, Late Antiquity to the Early Modern Period ( Conjunctions of Religion and Power in the Medieval Past), Ithaca, Cornell University Press, 2003; cloth; pp. xi, 225; RRP US$36.50; ISBN 0801438179. By centring his analysis on the shifting meanings of cruelty as it appears in narrative texts over a wide range of geographical and chronological contexts, Baraz has written a closely argued and important book that deepens our understanding of medieval and early modern societies. He moves emphatically beyond assumptions that medieval societies were inherently violent or cruel and instead analyses how ancient, medieval and early modern thinkers and writers defined and used terms and descriptions of cruel or unusually violent events. After an introduction and a background chapter, there are five chapters organised chronologically with one on the late ancient period, three chapters on the medieval period – one each for early, central and late – and then a chapter for the Early Modern period. There are then six appendices with detailed background to the argument of the book itself. Baraz ranges across the ancient classical world, the medieval east and west, and early modern Europe with assurance and skill. He is at home with all the intricate, complex and very different sources and so can give the sort of wide-ranging overview required for analyses of shifts in meaning of concepts. His argument is that, while the meaning of violent acts is self-evident, the meanings attached to extraordinary violence or cruelty are dependent on the [End Page 191] historical and intellectual context and that this context changed dramatically during the antique and medieval periods. By the Early Modern period cruelty was seen as a defining feature of 'others', notably religious or ethnic groups different from the Christian western 'norm'. He argues that this transformation occurred primarily in the west, where the building blocks of the definitions of cruelty were in the Latin authors of the late antique period, and is not seen in Islamic or Eastern Christian cultures which reflect the divergent attitudes of the Greek from the Latin Christians in the formative late antique period. He uses different sources in different periods, concentrating on intellectual definitions and discussions of cruelty where they exist, particularly for the ancient period. For the medieval period he depends on narratives of events that included descriptions that were either labelled cruel or were described in such a way that cruelty was implied by reference to acts of violence which had been described as cruel in other contexts. He finds that there were multiple perceptions of cruelty which varied over time and depended on the specific contexts. In any study of changing perceptions in different contexts, definitions of terms are very important and Baraz gives this aspect of his study due weight in the introduction. He makes a careful distinction between violence and cruelty. He uses what he terms the modern sense of violence ('a quantifiable and comparable category' [p. 6]) and then gives a rough equation of medieval distinctions between violence and cruelty with modern differences between violence and crime. This is a thought-provoking split and one which he makes work for him in the context of the narratives he examines throughout the book. His distinction between violence and cruelty seems artificial and I am not totally convinced that the meaning of violence is so self evident and that it is cruelty which shifts in meaning in different contexts. Whether the distinctions between definitions of violence and cruelty apply in other contexts or with other medieval sources will emerge when his theories are tested by other scholars. Since he is dealing primarily with narrative texts, he gives a great deal of attention to the lexical differences that he encounters. In chapter 1, he analyses the use of words that denote cruelty by authors throughout his period. In this analysis he is attempting to answer the question 'Why is cruelty in general ignored in some periods and discussed in others?' (p. 13). After an interesting discussion ranging across many authors he turns to descriptions of violent events in each of his five...
OBJECTIVE: In multiple sclerosis (MS), magnetic resonance imaging (MRI) predictors of cognitive impairment are based on sophisticated computer-generated analyses that are difficult to apply in clinical settings. This study investigated the clinical usefulness of a new visual rating scale, the Cholinergic Pathways Hyperintensities Scale (CHIPS), in detecting cognitive dysfunction. METHODS: Forty clinically definite MS patients underwent a brain MRI. Based on the CHIPS, cholinergic pathway hyperintensities were rated in 10 regions on four axial slices. Computerized hyperintense lesion volumes were also obtained. For cognitive testing, The Neuropsychological Screening Battery for Multiple Sclerosis was used. "Low" and "High" lesion score groups were computed based on the mean of the total CHIPS score. Optimal sensitivity and specificity of the total CHIPS score in detecting cognitive impairment were determined using a receiver operator characteristic curve. RESULTS: Despite a similar demographic profile, subjects with a "High" lesion score performed significantly worse than the "Low" lesion score group on verbal (P =.007) and visuospatial (P =.02) memory, and on a global index of cognitive functioning (P =.001). Optimal sensitivity (82%) and specificity (83%) were reached with a threshold total CHIPS score of 18 points. Total CHIPS score and total hyperintense lesion load were correlated (sigma = 0.82, P <.0001). CONCLUSION: CHIPS is helpful in clinically predicting cognitive impairment in MS.
In this paper, we describe the automatic annotation of the Cast3LB Treebank with LFG f-structures for the subsequent extraction of Spanish probabilistic grammar and lexical resources. We adapt the approach and methodology of Cahill et al. (2004), O’Donovan et al. (2004) and elsewhere for English to Spanish and the Cast3LB treebank encoding. We report on the quality and coverage of the automatic f-structure annotation. Following the pipeline and integrated models of Cahill et al. (2004), we extract wide-coverage \nprobabilistic LFG approximations and parse unseen Spanish text into f-structures. We also extend Bikel’s (2002) Multilingual Parse Engine to include a Spanish language module. Using the retrained Bikel parser in the pipeline model gives the best results against a manually constructed gold standard (73.20% predsonly f-score). We also extract Spanish lexical resources: 4090 semantic form types with 98 frame types. Subcategorised prepositions and particles are included in the frames.
Modern Chinese is a highly developed language that has norms phonetically, lexically, and grammatically. The norms are neither invariable nor variable unconditionally outside language context or without the restrictions of communication or styles. Standardization of modern Chinese should be viewed in a multi - perspective dimension. The chimerical views of standardization that deviates from language context, ignores stylistic characteristics, overlooking pragmatic functions, and blindly pursues standardization for its own sake should be opposed to.
All studies that examined the co-occurrence of pleasure and displeasure revealed at least some reports of mixed feelings (i.e., reports of concurrent pleasure and displeasure). Some researchers attribute these reports to measurement error, whereas others regard them as valid. This article examined response latencies of affect ratings to test the validity of reported mixed feelings. First, I demonstrate that respondents need more time to indicate the presence of an affect than the absence of an affect on unipolar response formats. Then, I demonstrate that mixed feelings are reported even when response latencies show a unipolar pattern for displeasure ratings, while the results for pleasure ratings were more mixed, but additional evidence suggests that they also are valid reflections of pleasure. In addition, I demonstrate that pleasure and displeasure ratings are independent of item-order and item-spacing. These results provide further support for the validity of reported mixed feelings and two-dimensional representations of pleasure and displeasure.
Many recent annotation efforts for English have focused on pieces of the larger problem of semantic annotation, rather than initially producing a single unified representation. This paper discusses the issues involved in merging four of these efforts into a unified linguistic structure: PropBank, NomBank, the Discourse Treebank and Coreference Annotation undertaken at the University of Essex. We discuss resolving overlapping and conflicting annotation as well as how the various annotation schemes can reinforce each other to produce a representation that is greater than the sum of its parts.
There has been a contemporary surge of interest in the application of stochastic models of parsing. The use of tree-adjoining grammar (TAG) in this domain has been relatively limited due in part to the unavailability, until recently, of large-scale corpora hand-annotated with TAG structures. Our goals are to develop inexpensive means of generating such corpora and to demonstrate their applicability to stochastic modeling. We present a method for automatically extracting a linguistically plausible TAG from the Penn Treebank. Furthermore, we also introduce labor-inexpensive methods for inducing higher-level organization of TAGs. Empirically, we perform an evaluation of various automatically extracted TAGs and also demonstrate how our induced higher-level organization of TAGs can be used for smoothing stochastic TAG models.
INTRODUCTION: In this work, we introduce the concept of semantic role labeling to the medical domain. We report first results of porting and adapting an existing resource, Propbank, to the medical field. Propbank is an adjunct to Penn Treebank that provides semantic annotation of predicates and the roles played by their arguments. The main aim of this work is the applicability of the Propbank frame files to predicates typically encountered in the medical literature. METHODS: We analyzed a target corpus of 610,100 abstracts, which was selected by searching for publication type "case reports". From this target corpus, we randomly selected 10,000 sample abstracts to estimate the predicate distribution, and matched the predicates from this sample to the predicates in Propbank. RESULTS: Of the 1998 unique verbs in our sample, 76% were represented in Propbank. This included the 40 most frequent verbs, which represented 49% of all predicate instances in our sample and which matched the Propbank usage in a study of representative sentences. We propose extensions to Propbank that handle medical predicates, which are not adequately covered by Propbank. CONCLUSION: We believe that semantic role labeling using Propbank is a valid approach to capture predicate relations in the medical literature.
We describe a history-based generative parsing model which uses a k-nearest neighbour (k-NN) technique to estimate the model's parameters. Taking the output of a base n-best parser we use our model to re-estimate the log probability of each parse tree in the n-best list for sentences from the Penn Wall Street Journal treebank. By further decomposing the local probability distributions of the base model, enriching the set of conditioning features used to estimate the model's parameters, and using k-NN as opposed to the Witten-Bell estimation of the base model, we achieve an f-score of 89.2%, representing a 4% relative decrease in f-score error over the 1-best output of the base parser.
We present a heuristic technique for converting a constituency treebank into a dependency treebank. In particular, we comment on our experience in converting the Spanish treebank Cast3LB. We extract a context-free grammar from the treebank, automatically identify the head in each rule, and use this information for constructing the dependency tree. Our heuristics have 99 % precision and 80 % recall in identifying the head in the rules, which gives 92% accuracy in identifying dependencies between words.
Abstract This corpus-based contrastive study examines the thematic use of the semantic field of research and researchers in the Discussion section of biomedical reports in Spanish native texts and English-Spanish translations. This semantic field was divided into integral reference (specific named researchers), general nouns for researchers, and singular and plural nouns referring to research. Themes containing these lexical items were examined with regard to their syntactic manifestations and their lexicogrammatical relations with the main finite verb. Quantitative analysis was used to establish reference values for the native texts and to reveal differences between the two subcorpora. Qualitative contextual analysis then investigated how the data might be applied to the translated texts. The quantitative study showed that the Spanish texts had more integral references and more general researcher nouns in their themes whereas the translations had more singular research nouns, especially those referring to the current study. Singular research nouns were associated with more prepositional adjuncts in the Spanish texts but with more subject themes, either as head or as modifier, in the translations. The distribution of tenses was different in all categories except for plural research nouns, with a higher percentage of present and present perfect in the Spanish texts and more past indefinite in the translations. Differences were also found in the distribution of lexical verbs related to integral references and singular research nouns. The contextual analysis revealed that awareness of these differences and strategic choices based on them could lead to thematic and discourse patterns that come closer to the target-language norms for this genre.
We present a treebank conversion method by which we construct an RMRS bank for HPSG parser evaluation from the TIGER Dependency Bank. Our method effectively performs automatic RMRS semantics construction from functional dependencies, following the semantic algebra of (Copestake et al., 2001). We present the semantics construction mechanism, and focus on some special phenomena. Automatic conversion is followed by manual validation. First evaluation results yield high precision of the automatic semantics construction rules. 1
Incremental parsing gains its importance in natural language processing and psycholinguistics because of its cognitive plausibility. Modeling the associated cognitive data structures, and their dynamics, can lead to a better understanding of the human parser. In earlier work, we have introduced a recursive neural network (RNN) capable of performing syntactic ambiguity resolution in incremental parsing. In this paper, we report a systematic analysis of the behavior of the network that allows us to gain important insights about the kind of information that is exploited to resolve different forms of ambiguity. In attachment ambiguities, in which a new phrase can be attached at more than one point in the syntactic left context, we found that learning from examples allows us to predict the location of the attachment point with high accuracy, while the discrimination amongst alternative syntactic structures with the same attachment point is slightly better than making a decision purely based on frequencies. We also introduce several new ideas to enhance the architectural design, obtaining significant improvements of prediction accuracy, up to 25% error reduction on the same dataset used in previous work. Finally, we report large scale experiments on the entire Wall Street Journal section of the Penn Treebank. The best prediction accuracy of the model on this large dataset is 87.6%, a relative error reduction larger than 50% compared to previous results.
Prepositional Phrase-attachment is a common source of ambiguity in natural language. The previous approaches use limited information to solve the ambiguity -- four lexical heads -- although humans disambiguate much better when the full sentence is available. We propose to solve the PP-attachment ambiguity with a Support Vector Machines learning model that uses complex syntactic and semantic features as well as unsupervised information obtained from the World Wide Web. The system was tested on several datasets obtaining an accuracy of 93.62% on a Penn Treebank-II dataset; 91.79% on a FrameNet dataset when no manually-annotated semantic information is provided and 92.85% when semantic information is provided.
Dieser Beitrag beschäftigt sich mit sprachlichen Mischformen im frankophonen Kanada. Im Mittelpunkt steht die hybride Varietät des Chiac, eine in Moncton/Acadie gesprochene urbane Mischvarietät, die aus dem Sprachkontakt von Englisch und Französisch entstanden ist und auf einer französischen Grammatik mit englischen Lexikelementen basiert. Gegenstand der Betrachtung ist die in den späten 1960er Jahren entstandene littérature acadienne und ihr Umgang mit sprachlicher Variation im Spannungsfeld von gesellschaftlicher Norm und Formen ihrer Transgression. Am Beispiel der zeitgenössischen Autorin France Daigle und ihrem Umgang mit der Varietät des Chiac soll verdeutlicht werden, in welcher Weise hybride Formen von Sprache, die existente Sprach- und Kulturgrenzen in Frage stellen, dazu beitragen, individuelle und soziale Widersprüche zu erfassen und zu bearbeiten.
Contemporary neuropsychological studies have stressed the widely distributed and multicomponential nature of human affective processes. Here, we examined facial electromyographic (EMG) (zygomaticus and corrugator muscle activity), autonomic (skin conductance and heart rate) and subjective measures of affective valence and arousal in patient TG, a 30 year-old man with left anterior mediotemporal and left orbitofrontal lesions resulting from a traumatic brain injury. Both TG and a normal control group were exposed to hedonically valenced visual and olfactory stimuli. In contrast with control subjects, facial EMG and electrodermal activity in TG did not differentiate among pleasant, unpleasant and neutral pictures. In addition, the controls reacted spontaneously with larger corrugator EMG activity and higher skin conductance to unpleasant odors. By contrast, the subjective feeling states (pleasure and arousal ratings) remained preserved in TG. The covariation between facial and self-report measures of negative valence was also a function of the nature of the olfactory task in the patient only. Taken together, the data suggest a functional dissociation between brain substrates supporting generation of emotion and those supporting representation of emotion.
We introduce a method for transferring annotation from a syntactically annotated corpus in a source language to a target language. Our approach assumes only that an (unannotated) text corpus exists for the target language, and does not require that the parameters of the mapping between the two languages are known. We outline a general probabilistic approach based on Data Augmentation, discuss the algorithmic challenges, and present a novel algorithm for sampling from a posterior distribution over trees.
To address the task of detecting nonischemic motion abnormalities from animated displays of gated myocardial perfusion single photon emission computed tomography data, we performed an observer study to evaluate the difference in detection performance between gating to 8 and 16 frames. Images were created from the NCAT mathematical phantom with a realistic heart simulating hypokinetic motion in the left lateral wall. Realistic noise-free projection data were simulated for both normal and defective hearts to obtain 16 frames for the cardiac cycle. Poisson noise was then simulated for each frame to create 50 realizations of each heart, All datasets were processed in two ways: reconstructed as a 16-frame set, and collapsed to 8 frames and reconstructed. Ten observers viewed the cardiac images animated with a realistic real-time frame rate. Observers trained on 100 images and tested on 100 images, rating their confidence on the presence of a motion defect on a continuous scale. None of the observers showed a significant difference in performance between the two gating methods. The 95% confidence interval on the difference in areas under the ROC curve (Az8 - Az16) was -0.029-0.085. Our test did not find a significant difference in detection performance between 8-frame gating and 16-frame gating. We conclude that, for the task of detecting abnormal motion, increasing the number of gated frames from 8 to 16 offers no apparent advantage.
MONA is an automata toolkit providing a compiler for compiling formulae of monadic second order logic on strings or trees into string automata or tree automata. In this paper, we evaluate the option of using MONA as a treebank query tool. Unfortunately, we find that MONA is not an option. There are several reasons why the main being unsustainable query answer times. If the treebank contains larger trees with more than 100 nodes, then even the processing of simple queries may take hours.