Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
Named Entity Recognition (NER) is an important first step for BioNLP tasks, e.g., gene normalization and event extraction. Employing supervised machine learning techniques for achieving high performance recent NER systems require a manually annotated corpus in which every mention of the desired semantic types in a text is annotated. However, great amounts of human effort is necessary to build and maintain an annotated corpus. This study explores a method to build a high-performance NER without a manually annotated corpus, but using a comprehensible lexical database that stores numerous expressions of semantic types and with huge amount of unannotated texts. We underscore the effectiveness of our approach by comparing the performance of NERs trained on an automatically acquired training data and on a manually annotated corpus. 1
Gastel, A., S. Schulze, Y. Versley & E. Hinrichs (2011). Annotation of implicit discourse relations in the TüBa-D/Z treebank. In T. Hedeland, T. Schmidt & K. Wörner (eds.), Multilingual Resources and Multilingual Applications, Proceedings of the German Society of Computational Linguistics and Language Technology (GSCL) 2011, 99–104. Hamburg.
A method for deriving an approximately labeled dependency treebank from the Thai Categorial Grammar Treebank has been implemented. The method involves a lexical dictionary for assigning dependency directions to the CG types associated with the grammatical entities in the CG bank, falling back on a generic mapping of CG types in case of unknown words. Currently, all but a handful of the trees in the Thai CG bank can unambiguously be transformed into directed dependency trees. Dependency labels can optionally be assigned with a learned classifier, which in a preliminary evaluation with a very small training set achieves 76.5% label accuracy. In the process, a number of annotation errors in the CG bank were identified and corrected. Although rather limited in its coverage, excluding e.g. long-distance dependencies, topicalisations and longer sentences, the resulting treebank is believed to be sound in terms of structural annotational consistency and a valuable complement to the scarce Thai language resources in existence.
A method for deriving an approximately labeled dependency treebank from the Thai Categorial Grammar Treebank has been implemented. The method involves a lexical dictionary for assigning dependency directions to the CG types associated with the grammatical entities in the CG bank, falling back on a generic mapping of CG types in case of unknown words. Currently, all but a handful of the trees in the Thai CG bank can unambiguously be transformed into directed dependency trees. Dependency labels can optionally be assigned with a learned classifier, which in a preliminary evaluation with a very small training set achieves 76.5% label accuracy. In the process, a number of annotation errors in the CG bank were identified and corrected. Although rather limited in its coverage, excluding e.g. long-distance dependencies, topicalisations and longer sentences, the resulting treebank is believed to be sound in terms of structural annotational consistency and a valuable complement to the scarce Thai language resources in existence.
Manually performed treebanking is an expensive effort compared with automatic annotation.In return, manual treebanking is generally believed to provide higherquality/value syntactic annotation than automatic methods.Unfortunately, there is little or no empirical evidence for or against this belief, though arguments have been voiced for the high degree of subjectivity in other levels of linguistic analysis (e.g.morphological annotation).We report a double-blind annotation experiment at the level of dependency syntax, using a small Finnish corpus as the analysis data.The results suggest that an interannotator agreement can be reached as a result of reviews and negotiations that is much higher than the corresponding labelled attachment scores (LAS) reported for stateof-the-art dependency parsers.
This paper introduces Chart Inference (CI), an algorithm for deriving a CCG category for an unknown word from a partial parse chart. It is shown to be faster and more precise than a baseline brute-force method, and to achieve wider coverage than a rule-based system. In addition, we show the application of CI to a domain adaptation task for question words, which are largely missing in the Penn Treebank. When used in combination with self-training, CI increases the precision of the baseline StatCCG parser over subjectextraction questions by 50%. An error analysis shows that CI contributes to the increase by expanding the number of category types available to the parser, while self-training adjusts the counts. 1
The Internet has rarely been used in auditory perception studies due to concerns about standardisation and calibration across different systems and settings. However, not all auditory research is based on the investigation of fine-grained differences in auditory thresholds. Where meaningful ‘real-world’ listening, for instance the perception of speech, is concerned, the Internet may be a more appropriate and ecologically valid setting to collect data. This study compared affective ratings of low-pass-filtered infant-, foreigner- and British adult-directed speech obtained with traditional methods in the laboratory, with those obtained from an Internet sample. Dropout rates and demographic distribution of participants in the Internet condition were also assessed. The results show that affective ratings were similar for both the Internet and laboratory samples. These findings indicate the viability of Internet-based research into affective speech perception and suggest that precise acoustic environmental control may not always be necessary.
In this paper, we describe and compare two statistical parsing approaches for the hybrid dependency-constituency syntactic representation used in the Quranic Arabic Treebank (Dukes and Buckwalter, 2010). In our first approach, we apply a multi-step process in which we use a shift-reduce algorithm trained on a pure dependency preprocessed version of the treebank. After parsing, the dependency output is converted into the hybrid representation. This is compared to a novel one-step parser that is able to learn the hybrid representation without preprocessing. We define an extended labelled attachment score (ELAS) as our performance metric for hybrid parsing, and report 87.47 % (F1 score) for the multi-step approach, and 89.03 % (F1 score) for the onestep integrated algorithm. We also consider the effect of using different sets of morphological features for parsing the Quran, comparing our results to recent work on Modern Standard Arabic.
This chapter presents a meta-search approach, meant to deliver bibliography from the internet, according to trainees’ results obtained at an e-assessment task. The bibliography consists of web pages related to the knowledge gaps of the trainees. The meta-search engine is part of an education recommender system, attached to an e-assessment application for project management knowledge. Meta-search means that, for a specific query (or mistake made by the trainee), several search mechanisms for suitable bibliography (further reading) could be applied. The lists of results delivered by the standard search mechanisms are used to build thematically homogenous groups using an ontology-based clustering algorithm. The clustering process uses an educational ontology and WordNet lexical database to create its categories. The research is presented in the context of recommender systems and their various applications to the education domain.
Abstract. FrameNet frames have been used to develop lexical databases and annotated corpora for different languages. This paper analyses the use of FrameNet frames to build a legal ontology for the Brazilian Law. In order to discuss the problems of such approach to ontology development, the lexical units evoking the Criminal_process frame were contrasted in English and Portuguese. Frame divergence between languages has consequences not only for legal ontology development but also for the development of legal lexical resources, such as lexical databases and corpora annotation. 1.
We present an approach of identifying the most prominent text/sentences using various shallow linguistic features, taking degree of connectiveness among the text units into consideration so as to minimize the poorly linked sentences in the resulting summary. As per the limitations of the current summarizing systems, the summary generated by those systems contains poorly linked sentences and are not topically salient. Thus, the paper aims at highlighting the effect of lexical chain scoring after the nouns and compound nouns are chained by searching for lexical cohesive relationships between words in the text using WordNet and using lexicographical relationships such as synonymy and hyponyms. In this paper, our algorithm ranks sentences based on the sum of the scores of the words in each sentence involving approaches like term frequencies, location of sentence in the text, cue words and phrases, word occurrences, and measuring lexical similarity(measuring chain score, word score and finally sentence score) for ranking the text units. We then identified and extracted high scored sentences and then the Vector Space approach is used to measure the relatedness/similarity between the extracted sentence and the topic words involving again the WordNet lexical database relationships to prioritise the topically related sentences. A threshold angle between the two vectors is predefined experimentally to which the ranked/scored sentences to be dropped and which the significant sentences with ranking/scores higher than threshold to be extracted. Note that the value of threshold is predetermined based on the percentage of output summary required to be generated.
Therapist self-disclosure has been theorized and found to have both positive and negative effects. These effects depend, in part, on the nature of the disclosure. This study sought to examine the differential effects of therapist disclosures of more and less resolved countertransference issues on perceptions of therapists and therapy sessions. Using an analogue method, undergraduate participants (N = 116) were randomly assigned to watch one of two videos in which a therapist disclosed personal issues that were relatively resolved or relatively unresolved. As hypothesized, therapist disclosure of issues that were more resolved caused the therapist to be rated as more attractive and trustworthy and instilled greater hope than therapist disclosure of less resolved issues. The type of therapist disclosure, however, did not affect ratings of the expertness of the therapist, the depth or smoothness of the session, or the perceived universality between client and therapist. Implications of the results for the judicious use of self-disclosure are discussed.
This paper considers how implemented grammars can enhance descriptive ones. An implemented grammar encodes the analyses in machine readable form, facilitating automatic annotation of morphological, syntactic and semantic structures. A grammar augmented with a collection of such structures would enable, for example, a reader to search for items in which a PP argument fills the third most prominent semantic role.
Recent advances in parsing technology have made treebank parsing with discontinuous constituents possible, with parser output of competitive quality (Kallmeyer and Maier, 2010). We apply Data-Oriented Parsing (DOP) to a grammar formalism that allows for discontinuous trees (LCFRS). Decisions during parsing are conditioned on all possible fragments, resulting in improved performance. Despite the fact that both DOP and discontinuity present formidable challenges in terms of computational complexity, the model is reasonably efficient, and surpasses the state of the art in discontinuous parsing. 1
This article provides recent empirical evidence to support the argument that the everyday politics of race and fear of the non-desired Other still persist in Australia more than half a century after the official demise of the White Australia Policy. The article sheds so me insight into how Australian immigration policies are now deliberately designed to normalise and assimilate new migrants into narrow Anglo-Saxon cultural and linguistic norms, thereby inadvertently excluding people from culturally and linguistically diverse backgrounds who need Australian citizenship the most. The argument of the article is based on outcomes of a study on personal stories of African refugee background Australian citizens regarding their experiences with the Australian citizenship test; their opinions about the literacy-for-citizenship requirement; and their ideas about being and becoming Australian. The participants to the study expressed strong reservations with the idea of having to undertake a formal citizenship test that neither improves their understanding of the everyday way of life in Australia nor opens avenues for greater opportunities for socio-economic participation and recognition of the linguistic and cultural identities they bring to Australia.
OBJECTIVE: The purpose of this article is to evaluate the use of the periodically rotated overlapping parallel lines with enhanced reconstruction (PROPELLER) technique for artifact reduction and overall image quality improvement for intermediate-weighted and T2-weighted MRI of the shoulder. SUBJECTS AND METHODS: One hundred eleven patients undergoing MR arthrography of the shoulder were included. A coronal oblique intermediate-weighted turbo spin-echo (TSE) sequence with fat suppression and a sagittal oblique T2-weighted TSE sequence with fat suppression were obtained without (standard) and with the PROPELLER technique. Scanning time increased from 3 minutes 17 seconds to 4 minutes 17 seconds (coronal oblique plane) and from 2 minutes 52 seconds to 4 minutes 10 seconds (sagittal oblique) using PROPELLER. Two radiologists graded image artifacts, overall image quality, and delineation of several anatomic structures on a 5-point scale (5, no artifact, optimal diagnostic quality; and 1, severe artifacts, diagnostically not usable). The Wilcoxon signed rank test was used to compare the data of the standard and PROPELLER images. RESULTS: Motion artifacts were significantly reduced in PROPELLER images (p < 0.001). Observer 1 rated motion artifacts with diagnostic impairment in one patient on coronal oblique PROPELLER images compared with 33 patients on standard images. Ratings for the sequences with PROPELLER were significantly better for overall image quality (p < 0.001). Observer 1 noted an overall image quality with diagnostic impairment in nine patients on sagittal oblique PROPELLER images compared with 23 patients on standard MRI. CONCLUSION: The PROPELLER technique for MRI of the shoulder reduces the number of sequences with diagnostic impairment as a result of motion artifacts and increases image quality compared with standard TSE sequences. PROPELLER sequences increase the acquisition time.
This paper presents a taxonomy of cognitive functions that supports formal functional modeling of cognitive technical systems (CTSs) and cognitive products. To date, there is little support for functional modeling of such systems and products even though their interdisciplinary complexity exceeds that of electro-mechanical products and makes modeling support in conceptual design even more important. The taxonomy of cognitive functions is based on literature research and consists of a set of cognitive capabilities on three hierarchical levels as well as a defined set of flows. Relationships among cognitive capabilities have been identified using WordNet, a lexical database of English. The application of the taxonomy is demonstrated through the example of a coffee robot waiter, which has been designed and prototyped in the research group of the authors. Through defining a common taxonomy of cognitive functions and flows, a common practice for functional modeling of cognitive products is defined thus supporting re-use of functional models. This creates the foundation for creating model-based design repositories for CTSs and cognitive products to support their future development.
This paper examines the ways in which par-allelism can be used to speed the parsing of dense PCFGs. We focus on two kinds of parallelism here: Symmetric Multi-Processing (SMP) parallelism on shared-memory multi-core CPUs, and Single-Instruction Multiple-Thread (SIMT) parallelism on GPUs. We de-scribe how to achieve speed-ups over an al-ready very efficient baseline parser using both kinds of technology. For our dense PCFG parsing task we obtained a 60×speed-up us-ing SMP and SSE parallelism coupled with a cache-sensitive algorithm design, parsing sec-tion 24 of the Penn WSJ treebank in a little over 2 secs. 1
It is hypothesized that ratings of emotional stimuli are affected by a constant threat of traumatic events. Ratings of valence and arousal on the International Affective Picture System from young adults in the United States were compared to those of young Israeli adults. Israelis rated the pictures as less negative and less positive than did participants from the United States. Israeli women gave higher arousal ratings compared to the American women. These differences may be due to compulsory military service in Israel, during which exposure to traumatic events is more likely to occur, and to the timing of the study which followed a year of frequent suicide bomb attacks. The authors suggest that these findings may reflect mild symptoms of stress disorders.
Based on two syntactic dependency treebanks built with two different styles of Chinese, a statistical study is conducted regarding word-frequency and distributions. We extracted three grammatical words as the research objects and analyzed their network features, including all degree, out-degree, in-degree, all closeness, in-closeness, out-closeness and betweenness. Then these three nodes were removed from the networks. We recorded and compared the network features of the two original networks and the three networks from which one node is respectively removed, including the number of vertices, average degree, average path length, diameter, the number of isolated vertices, domain and density. The results show that all three function words are central nodes of the Chinese syntactic networks but have different status. Their influence to the overall structure is also quite different. The research not only provides a new method for the study about Chinese grammatical words but also provides a new way of thinking the node characteristics in the complex network.
While most dialectological research so far focuses on phonetic and lexical phenomena, we use recent fieldwork in the domain of dialect syntax to guide the development of multidialectal natural language processing tools. In particular, we develop a set of rules that transform Standard German sentence structures into syntactically valid Swiss German sentence structures. These rules are sensitive to the dialect area, so that the dialects of more than 300 towns are covered. We evaluate the transformation rules on a Standard German treebank and obtain accuracy figures of 85% and above for most rules. We analyze the most frequent errors and discuss the benefit of these transformations for various natural language processing tasks. 1
Good Dictionary Examples or GDEX is a tool in the Sketch Engine designed to help lexicographers with identifying dictionary examples by ranking sentences according to how likely they are to be good candidates. The ranking is done automatically using various syntactic and lexical features. So far, only GDEX for English has been available. This paper presents the design and evaluation of Slovene GDEX, which was used for finding good examples for the new lexical database of Slovene, one of the activities in the Communication in Slovene project. Several different GDEX configurations were designed, evaluated and compared. The evaluation involved examining sentences of lemmas belonging to different word classes. Good sentences were logged for subsequent analysis with external data-mining software, WEKA. The observed behaviour was then used to adjust the parameters of the GDEX classifiers. We believe that the procedure of identifying features of good examples and their values, described in this paper, can be used for the development of GDEX for any language.
We introduce synchronous tree adjoining grammars (TAG) into tree-to-string transla-tion, which converts a source tree to a target string. Without reconstructing TAG deriva-tions explicitly, our rule extraction algo-rithm directly learns tree-to-string rules from aligned Treebank-style trees. As tree-to-string translation casts decoding as a tree parsing problem rather than parsing, the decoder still runs fast when adjoining is included. Less than 2 times slower, the adjoining tree-to-string system improves translation quality by +0.7 BLEU over the baseline system only al-lowing for tree substitution on NIST Chinese-English test sets. 1
The dictionaries are among the wellknown tools for applications in everyday life, education, sciences, humanities, and human communication. The recent developments of information technologies contribute to the design and creation of new software tools with a wide range of applications, especially for natural language processing. The paper presents an online bilingual dictionary as a technological tool for applications in digital humanities, and describes the structure and content of the bilingual Lexical Database supporting a Bulgarian-Polish online dictionary. It focuses especially on the presentation of verbs, which form the richest from a specific characteristics viewpoint linguistic category in Bulgarian. The main software modules for webpresentation of this digital dictionary are also shortly described. 1
Abstract The aim of computational semantics is to capture the meaning of natural language expressions in representations suitable for performing inferences, in the service of understanding human language in written or spoken form. First‐order logic is a good starting point, both from the representation and inference point of view. But even if one makes the choice of first‐order logic as representation language, this is not enough: the computational semanticist needs to make further decisions on how to model events, tense, modal contexts, anaphora and plural entities. Semantic representations are usually built on top of a syntactic analysis, using unification, techniques from the lambda‐calculus or linear logic, to do the book‐keeping of variable naming. Inference has many potential applications in computational semantics. One way to implement inference is using algorithms from automated deduction dedicated to first‐order logic, such as theorem proving and model building. Theorem proving can help in finding contradictions or checking for new information. Finite model building can be seen as a complementary inference task to theorem proving, and it often makes sense to use both procedures in parallel. The models produced by model generators for texts not only show that the text is contradiction‐free; they also can be used for disambiguation tasks and linking interpretation with the real world. To make interesting inferences, often additional background knowledge is required (not expressed in the analysed text or speech parts). This can be derived (and turned into first‐order logic) from raw text, semi‐structured databases or large‐scale lexical databases such as WordNet. Promising future research directions of computational semantics are investigating alternative representation and inference methods (using weaker variants of first‐order logic, reasoning with defaults), and developing evaluation methods measuring the semantic adequacy of systems and formalisms.
Eighteenth-century language usage is markedly under-represented in the first two editions of the OED, whose quotations for this period were gathered almost entirely during the late nineteenth and early twentieth centuries. This article reviews some of the possible causes, characteristics and consequences of OED’s gap in eighteenth-century documentation and shows that female authors were particularly scanted. The role of quotations in the OED, as the evidential basis for the dictionary, is briefly considered, along with eighteenth-century (and Victorian/Edwardian) views on women and language, and the availability of female-authored texts for quotation by the lexicographers. The article reports sample reading in eighteenth-century female writers (especially Jean Adam, Penelope Aubin and Anna Seward), which shows that OED could easily have supplied its eighteenth-century deficiency from such authors, and that it often favoured distinctive usages in female-authored texts—innovative, eccentric or domestic vocabulary—rather than usage which exemplified linguistic norms (especially in poetry, where Seward’s case is examined). It also discusses revisions to the OED so far conducted in the third (ongoing) edition, and their implications for readers and editors of eighteenth-century texts.
Anaphora resolution is one of the most difficult tasks in NLP. The ability to identify non-referential pronouns before attempting an anaphora resolution task would be significant, since the system would not have to attempt resolving such pronouns and hence end up with fewer errors. In addition, the number of non-referential pronouns has been found to be non-trivial in many domains. The task of detecting non-referential pronouns could also be incorporated into a part-of-speech tagger or a parser, or treated as an initial step in semantic interpretation. In this article, I describe a machine learning method for identifying non-referential pronouns in an annotated subsegment of the Penn Arabic Treebank using three different feature settings. I achieve an accuracy of 97.22% with 52 different features extracted from a small window size of -5/+5 tokens surrounding each potentially non-referential pronoun.
Abstract The language of a speech community can only act as an identity marker for all of its speakers if linguistic norms are widely shared and if a minimal number of language varieties are spoken. This article examines briefly how a linguistic norm came to serve the whole of Iceland and how a situation of relative linguistic homogeneity was maintained for centuries. Sociolinguistic theory tells us that the speech community that we can reconstruct for early Iceland should lead to the establishment and maintenance of local norms. However, Iceland, arguably monodialectal, was certainly characterized by long-term linguistic homogeneity and remained a society where nucleated settlements barely formed over a thousand-year period. Scholars have argued that a mixture of dialects leveled shortly after the settlement of Iceland in the ninth century (Settlement). Studies show that dialect leveling requires dialect mixing, the convergence of people on one place, and sustained linguistic contact between the speakers. The settlement pattern of Iceland is indicative of population divergence (not convergence) and there is limited evidence of sustained contact. It is therefore proposed that the dialect leveling might be linked instead with significant population movements and social upheaval in mainland Scandinavia in the immediate pre-Viking period. The variety of Norse that was taken westward across the Atlantic might itself already have been the result of several earlier stages of mixing and koineization. It is only by combining linguistic, historical, and archaeological knowledge that this problem of how one linguistic norm came to serve the whole of Iceland can be understood.
This paper introduces, an XML format developed to serialise the object model defined by the ISO Syntactic Annotation Framework SynAF. Basing on widespread best practices we adapt a popular XML format for syntactic annotations, TigerXML, with additional features to support a variety of syntactic phenomena including constituent and dependency structures, binding, and different node types such as compounds or empty elements. We also define interfaces to other formats and standards including the Morpho-syntactic Annotation Framework MAF and the ISOCat Data Category Registry. Finally a case study of the German Treebank TueBa-D/Z is presented, showcasing the handling of constituent structures, topological fields and coreference annotation.
Psycholinguistic studies on whether classifiers facilitate processing object-extracted relative clauses (RC) in Mandarin have often made use of a classifier mismatch-match configuration, wherein a preceding classifier mismatches the following RC-subject but matches the modified head noun. However, an examination of the Chinese Treebank corpus 5.0 shows this configuration rarely occurs. None of the 10 tokens of pre-RC classifiers conforms to the mismatch-match configuration in a real sense. Instead, either a dropped RC-subject or some intervening item successfully avoids anticipated lexical disruption effects induced by a mismatching classifier. The results of analysis suggest that the constructed examples used in previous psycholinguistic studies may not realistically test natural language processing procedures.
OBJECTIVES: Dental phobia is currently classified as a specific phobia of the blood-injection-injury (BII) subtype. In another subtype, animal phobia, enhanced amplitudes of late event-related potentials have consistently been identified for patients during passive viewing of disorder-relevant pictures. However, this has not been shown for BII phobics, and studies with dental phobics are lacking. Findings on cardiac responses in BII phobia during exposure are heterogeneous, as some studies showed a diphasic pattern of heart rate acceleration and deceleration, whereas others observed pure acceleration. In contrast, heart rate increase has consistently been shown for dental phobics, resembling the reaction of animal phobics. Moreover, the BII subtype is characterized by elevated disgust reactivity whereas the role of habitual disgust proneness in dental phobia is unclear. METHODS: We recorded the electroencephalogram and the electrocardiogram from 18 dental phobic and 18 healthy women while they watched pictures depicting dental treatment, disgust, fear and neutral items. RESULTS: Phobics relative to controls showed an enhanced late positive potential (300-700 ms) and heart rate acceleration towards phobic material, reflecting motivated attention and fear. Affective ratings revealed that dental phobics experienced significantly higher levels of fear than disgust during exposure to phobia-relevant material. Patients' elevated habitual disgust proneness was restricted to specific domains, such as the oral incorporation of offensive objects. CONCLUSION: The psychophysiology of dental phobia resembles the fear-dominated subtypes of specific phobia reported in earlier studies. Future studies should continue to investigate whether the current classification of this disorder as BII phobia needs to be reconsidered.
Many recent experiments in the automatic classification of discourse relations have limited themselves to a small set of coarse categories. While there are eminent reasons to do so – annotation for a small sets of categories can be created more reliably, and possibly also be defined in a more clear – cut way than finer distinctions – it is an interesting question whether the finer-grained distinctions present in some annotated corpora can be reconstructed reliably. The present paper investigates the feasibility of such fine-grained tagging of discourse relations using data from the Penn Discourse Treebank. 1
This paper presents a probabilistic approach for POS tagging that combines HMMs and character language models being applied to Portuguese texts. In this approach, the emission probabilities for each hidden state in a HMM are estimated by a proper character language model. The tagger built has been trained and tested on Bosque, a subset of Floresta Sinta(c)tica treebank, reaching 96.2% accuracy with a 39-tag tagset and 92.0% with a 257-tag tagset extended with inflexion information.
OBJECTIVE: To determine the effectiveness of a modulated acoustic startle reflex paradigm with emotional imagery in studying physiological changes associated with emotional responses in persons with traumatic brain injury (TBI). SETTING: Outpatient rehabilitation hospital. PARTICIPANTS: Six individuals with moderate to severe TBI. Mean age was 32 years and mean years postinjury were 9.9. METHOD: The modulated acoustic startle reflex procedure involved imagery of emotional scripts (joy, anger, fear, and neutral) followed by a startle noise, versus startle noise alone (no script). MEASURES: Eyeblink and skin conductance response, subjective arousal and valence ratings of the scripts, and general anger questionnaire. RESULTS: Startle blink responses following anger imagery were significantly smaller than those following fear (P =.006) and neutral (P =.023) imagery. Skin conductance response did not change significantly based on the content of the scripts (P =.070). CONCLUSIONS: Large startle blink responses indicate avoidance of a stimulus. Our findings suggest that participants with TBI did not have an avoidant reaction to anger-inducing stimuli. Skin conductance response findings may imply arousal impairments. The modulated acoustic startle reflex was effective in measuring emotional responses; however, larger studies comparing persons with TBI with control groups are needed to further explore these findings.
The conventional sequence labeling methods for Chinese word segmentation do not fully utilize the linguistic information, which restricts further improvements of the performance. Chinese morphology intensively investigates the constructions and usages of Chinese words, which is helpful to Chinese word segmentation. Furthermore, some word segmentation ambiguities cannot be resolved only by means of the lexical information, and the final disambiguations take place in the parsing process. In this paper, we propose a parsing-based Chinese word segmentation model, which can fully utilize the morphological and syntactic information. Experiments on Penn Chinese Treebank(CTB) 5.0 show that the proposed model obtains competitive performances as the CRFs-based model. To investigate the relationship between our parsing-based model and the CRFs-based model, a maximum entropy model based framework for integrating different knowledge sources is employed. The integrating model obtains an F-measure of 97.9, 25% in segmentation error rate reduction relative to the CRFs-based model, which indicates that the two models are complementary to each other.
L’exploitation de corpus analyses syntaxiquement (ou corpus arbores) pour le public non specialiste n’est pas un probleme trivial. Si la communaute du TAL souhaite mettre a la disposition des chercheurs non-informaticiens des corpus comportant des annotations linguistiques complexes, elle doit imperativement developper des interfaces simples a manipuler mais permettant des recherches fines. Dans cette communication, nous presentons les modes de recherche « grand public » developpe(e)s dans le cadre du projet Scientext, qui met a disposition un corpus d’ecrits scientifiques interrogeable par partie textuelle, par partie du discours et par fonction syntaxique. Les modes simples sont decrits: un mode libre et guide, ou l’utilisateur selectionne lui-meme les elements de la requete, et un mode semantique, qui comporte des grammaires locales preetablies a l’aide des fonctions syntaxiques.
PURPOSE: To report how youths, both with and without idiopathic scoliosis (IS), respond to questions about their self-image and perceptions of body shape. An additional purpose is to describe themes that emerged as important to youths with IS to better understand scoliosis from their perspective. METHODS: Descriptive qualitative and quantitative methods were utilized. Subject interviews were conducted, as part of a larger cognitive interviewing study on the Spinal Appearance Questionnaire, using a cross-sectional sample of 76 females between 8 and 16 years of age with IS and who were typically developing (TD), without scoliosis. RESULTS: IS and TD subjects revealed similar findings when asked what makes them look good versus their peers; self-image ratings were also positive. Predominant themes from open-ended responses include physical appearance, feelings, brace wear, and discomfort. CONCLUSION: Self-image and body shape did not differ significantly between groups. The identified themes warrant further exploration as they are significant and important to youth with scoliosis.
Flat noun phrase structure was, up until recently, the standard in annotation for the Penn Treebanks. With the recent addition of internal noun phrase annotation, dependency parsing and applications down the NLP pipeline are likely affected. Some machine translation systems, such as TectoMT, use deep syntax as a language transfer layer. It is proposed that changes to the noun phrase dependency parse will have a cascading effect down the NLP pipeline and in the end, improve machine translation output, even with a reduction in parser accuracy that the noun phrase structure might cause. This paper examines this noun phrase structure’s effect on dependency parsing, in English, with a maximum spanning tree parser and shows a 2.43%, 0.23 Bleu score, improvement for English to Czech machine translation. 1
AIM: To evaluate a standardised MRI acquisition protocol and a new image rating scale for disease severity in patients with progressive supranuclear palsy (PSP) and multiple systems atrophy (MSA) in a large multicentre study. METHODS: The MRI protocol consisted of two-dimensional sagittal and axial T1, axial PD, and axial and coronal T2 weighted acquisitions. The 32 item ordinal scale evaluated abnormalities within the basal ganglia and posterior fossa, blind to diagnosis. Among 760 patients in the study population (PSP = 362, MSA = 398), 627 had per protocol images (PSP = 297, MSA = 330). Intra-rater (n = 60) and inter-rater (n = 555) reliability were assessed through Cohen's statistic, and scale structure through principal component analysis (PCA) (n = 441). Internal consistency and reliability were checked. Discriminant and predictive validity of extracted factors and total scores were tested for disease severity as per clinical diagnosis. RESULTS: Intra-rater and inter-rater reliability were acceptable for 25 (78%) of the items scored (≥ 0.41). PCA revealed four meaningful clusters of covarying parameters (factor (F) F1: brainstem and cerebellum; F2: midbrain; F3: putamen; F4: other basal ganglia) with good to excellent internal consistency (Cronbach α 0.75-0.93) and moderate to excellent reliability (intraclass coefficient: F1: 0.92; F2: 0.79; F3: 0.71; F4: 0.49). The total score significantly discriminated for disease severity or diagnosis; factorial scores differentially discriminated for disease severity according to diagnosis (PSP: F1-F2; MSA: F2-F3). The total score was significantly related to survival in PSP (p<0.0007) or MSA (p<0.0005), indicating good predictive validity. CONCLUSIONS: The scale is suitable for use in the context of multicentre studies and can reliably and consistently measure MRI abnormalities in PSP and MSA. Clinical Trial Registration Number The study protocol was filed in the open clinical trial registry (http://www.clinicaltrials.gov) with ID No NCT00211224.