Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Classifying patients' affect is a pivotal part of the mental status examination. However, this common practice is often widely inconsistent between raters. Recent advances in the field of Facial Action Recognition (FAR) have enabled the development of tools that can act to identify facial expressions from videos. In this study, we aimed to explore the potential of using machine learning techniques on FAR features extracted from videotaped semi-structured psychiatric interviews of 25 male schizophrenia inpatients (mean age 41.2 years, STD = 11.4). Five senior psychiatrists rated patients' affect based on the videos. Then, a novel computer vision algorithm and a machine learning method were used to predict affect classification based on each psychiatrist affect rating. The algorithm is shown to have a significant predictive power for each of the human raters. We also found that the eyes facial area contributed the most to the psychiatrists' evaluation of the patients' affect. This study serves as a proof-of-concept for the potential of using the machine learning FAR system as a clinician-supporting tool, in an attempt to improve the consistency and reliability of mental status examination.
Abstract This chapter poses the question of whether humans might be essentially normative animals, i.e. whether traditionally prominent specificities of the human life form—our linguistic, social, and moral “natures”—might ground in a basic susceptibility, or proclivity to the deontic regulation of thought and behaviour: the “normative animal thesis.” The chapter lays out the issues at stake in attempting to answer this question. It divides into two main parts. The first begins by clarifying the—norm-related—concept of normativity at issue, distinguishing it from the—reason-related—conceptualisation current in meta-ethics and theories of rationality. It then discusses the primary candidates for generic features of norms, before dividing the normative animal thesis into various sub-claims. The second part presents the key questions at issue in the discussion of social, moral, and linguistic norms, comparing ways of conceiving them and marking the significance of such conceptualisations for the normative animal thesis.
Purpose Verbs with low concreteness are frequent in discourse samples but rarely targeted in aphasia treatments for verbs. These verbs are an important part of functional communication, and recent studies have called for more research regarding aphasia and treatment stimuli with low concreteness. The aim of this study was to pilot the use of verbs with low concreteness in a novel sentence production intervention with persons with aphasia. Method The study took the form of a single-case experimental design with multiple baselines across behaviors and across participants. Three persons with chronic nonfluent aphasia and apraxia of speech participated in the study. Each participant received treatment designed to increase the semantic networks of verbs with high frequency and low concreteness. Sentence production was closely examined over the course of treatment for treated and untreated verbs of varying concreteness levels. Additional measures of language and cognitive functioning were also taken before and after treatment. Results Results indicated improved sentence production with target verbs attributable to the treatment in the 1st phase of 2 phases for 2 of the 3 participants. The increases corresponded with the application of treatment, despite the difference in number of baseline sessions for the participants. Where there were treatment effects, there was also considerable generalization to untreated sets of items during the 1st treatment phase. Word retrieval also improved for 2 participants. Conclusions The results suggest that the novel treatment may improve sentence production and word retrieval in persons with aphasia, even when using target verbs with low concreteness ratings. Future research is warranted into the use of low concreteness verbs. Supplemental Material https://doi.org/10.23641/asha.10870958.
Drawing on linguistic ethnographic data analysis, this article aims to expand the Rampton’s concept of ‘language crossing’ through integrating the notion of ‘linguistic racism’ experienced by Mongolian background immigrant women in Australia. These women encounter linguistic homogeneity, discrimination, and alienation in varied ways in their daily institutional and non-institutional settings based on how they speak English or their usage of heritage languages. As a result, they establish everyday linguistic resistance strategies to combat linguistic racism, which further add two new dimensions to the concept of language crossing – ‘crossing as a resistance strategy’ and ‘crossing as a passing strategy’. Adopting these crossing strategies allow these women to use their preferred forms of communication to resist dominant linguistic norms and standards in the dominant culture. These strategies further make it possible for these speakers to pass as the native speakers of that dominant language. Finally, the paper argues that it is almost impossible to understand ‘language crossing’ as a discrete understanding isolated from the concept of ‘linguistic racism’. It is better to examine these concepts together, as they seem to complement each other in terms of investigating the everyday linguistic practices, sociolinguistic realities and struggles that these immigrant women encounter.
The slowness of legal proceedings in the common law legal system is a widely known fact. Any tool which could help reduce the time taken for the resolution of a case is invaluable. Common legal systems place a great importance on precedents and retrieving the correct set of precedents is considerably time consuming. Hence, for any case whose proceedings are in progress, if there are suitable prior cases, then the court has to follow the same interpretations that were passed in the prior cases. This is to ensure that similar situations receive similar treatment, thus maintaining uniformity amongst the legal proceedings across all courts at all times. Hence, precedent cases are treated as important as any other written law (a statute) in this legal system. In this paper, we propose two new approaches to solve this information retrieval problem wherein the system accepts the current case document as the query and returns the relevant precedent cases as the result. The first approach is to calculate the document similarity using Wordnet, which is a lexical database that could be leveraged to quantify the semantic relatedness between two documents, using a semantic network. The second approach is the use of a Siamese Manhattan Long Short Term Memory network, which is a supervised model trained to understand the underlying similarity between two documents.
Many long short-term memory (LSTM) applications need fast yet compact models. Neural network compression approaches, such as the grow-and-prune paradigm, have proved to be promising for cutting down network complexity by skipping insignificant weights. However, current compression strategies are mostly hardware-agnostic and network complexity reduction does not always translate into execution efficiency. In this work, we propose a hardware-guided symbiotic training methodology for compact, accurate, yet execution-efficient inference models. It is based on our observation that hardware may introduce substantial non-monotonic behavior, which we call the latency hysteresis effect, when evaluating network size vs. inference latency. This observation raises question about the mainstream smaller-dimension-is-better compression strategy, which often leads to a sub-optimal model architecture. By leveraging the hardware-impacted hysteresis effect and sparsity, we are able to achieve the symbiosis of model compactness and accuracy with execution efficiency, thus reducing LSTM latency while increasing its accuracy. We have evaluated our algorithms on language modeling and speech recognition applications. Relative to the traditional stacked LSTM architecture obtained for the Penn Treebank dataset, we reduce the number of parameters by 18.0x (30.5x) and measured run-time latency by up to 2.4x (5.2x) on Nvidia GPUs (Intel Xeon CPUs) without any accuracy degradation. For the DeepSpeech2 architecture obtained for the AN4 dataset, we reduce the number of parameters by 7.0x (19.4x), word error rate from 12.9% to 9.9% (10.4%), and measured run-time latency by up to 1.7x (2.4x) on Nvidia GPUs (Intel Xeon CPUs). Thus, our method yields compact, accurate, yet execution-efficient inference models.
The paper presents a short introduction to several electronic resources for Ukrainian language, namely, two treebanks: the Gold standard (ab. 130 thousand tokens), manually annotated in the Universal Dependencies flavour (https://universaldependencies.org/), which comprises the training data for a machine-trained syntactic parser, and a big (near 3 billion tokens),
Dialectology and Dutch syntax from the perspective of linguistic (norm) change Jeroen van Craenenbroeck demonstrates in a highly convincing way that both synchronous descriptions and theoretical approaches of the syntax of Standard Dutch can be optimized by including analyses of dialect variation. Besides a minor reservation with respect to van Craenenbroeck’s interpretation of the conjugation of conjunctions, this response adds a complementary perspective to van Craenenbroeck’s overall argumentation by arguing that the inclusion of geolinguistic research on dialect variation is also indispensable for the diachronic study of Dutch syntax.
In the era of big data, mining the emotional tendency of opinions through natural language processing technology is meaningful for the timely understanding of human behavior data. Nowadays, Generalized Autoregressive Pre-training Language Modeling (XLNet) can not only capture bidirectional contextual knowledge but also learn the word dependency. However, its sentence-level representation didn't take broad features into account. In this paper, we design a novel architecture, called Broad Autoregressive Language Model (BroXLNet), to automatically process sentiment analysis task. BroXLNet integrates the advantage of generalized autoregressive language modeling and broad learning system, which has the ability of extracting deep contextual features and randomly searching high-level contextual representation in broad spaces. We evaluated our algorithm on binary Stanford Sentiment Treebank dataset. Compared with the state-of-the-art methods, e.g., BiLSTM, ELMo, OpenAI GPT, BERT and XLNet, BroXLNet achieved the best result of 94.0% in sentiment analysis task of binary Stanford Sentiment Treebank. The result demonstrated the excellent classifying ability of BroXLNet in sentiment analysis.
For scientists and researchers, it is very critical to ensure knowledge is accessible for re-use and development. Moreover, the way we store and manage scientific articles and their metadata in digital libraries determines the amount of relevant articles we can discover and access depending on what is actually meant in a search query. Yet, are we able to explore all semantically relevant scientific documents with the existing keyword-based search information retrieval systems? This is the primary question addressed in this thesis. Hence, the main purpose of our work is to broaden or expand the knowledge spectrum of researchers working in an interdisciplinary domain when they use the information retrieval systems of multidisciplinary digital libraries. However, the problem raises when such researchers use community-dependent search keywords while other scientific names given to relevant concepts are being used in a different research community.Towards proposing a solution to this semantic exploration task in multidisciplinary digital libraries, we applied several text mining approaches. First, we studied the semantic representation of words, sentences, paragraphs and documents for better semantic similarity estimation. In addition, we utilized the semantic information of words in lexical databases and knowledge graphs in order to enhance our semantic approach. Furthermore, the thesis presents a couple of use-case implementations of our proposed model
Neural networks have shown great potential in language modeling. Currently, the dominant approach to language modeling is based on recurrent neural networks (RNNs) and convolutional neural networks (CNNs). Nonetheless, it is not clear why RNNs and CNNs are suitable for the language modeling task since these neural models are lack of interpretability. The goal of this paper is to tailor an interpretable neural model as an alternative to RNNs and CNNs for the language modeling task. This paper proposes a unified framework for language modeling, which can partly interpret the rationales behind existing language models (LMs). Based on the proposed framework, an interpretable neural language model (INLM) is proposed, including a tailored architectural structure and a tailored learning method for the language modeling task. The proposed INLM can be approximated as a parameterized auto-regressive moving average model and provides interpretability in two aspects: component interpretability and prediction interpretability. Experiments demonstrate that the proposed INLM outperforms some typical neural LMs on several language modeling datasets and on the switchboard speech recognition task. Further experiments also show that the proposed INLM is competitive with the state-of-the-art long short-term memory LMs on the Penn Treebank and WikiText-2 datasets.
Most current language modeling techniques only exploit co-occurrence, semantic and syntactic information from the sequence of words. However, a range of information such as the state of the speaker and dynamics of the interaction might be useful. In this work we derive motivation from psycholinguistics and propose the addition of behavioral information into the context of language modeling. We propose the augmentation of language models with an additional module which analyzes the behavioral state of the current context. This behavioral information is used to gate the outputs of the language model before the final word prediction output. We show that the addition of behavioral context in language models achieves lower perplexities on behavior-rich datasets. We also confirm the validity of the proposed models on a variety of model architectures and improve on previous state-of-the-art models with generic domain Penn Treebank Corpus.
OBJECTIVE: Facial expressions communicate emotional states and regulate social bonds. An approach or avoidance-based valence might interact with direct or averted gaze to elicit different attentional allocation. These processes might be aberrant in major depression or first-episode psychosis and this requires empirical investigation. METHOD: This study examined higher order, controlled attentional processing of emotional facial expressions (happy, neutral, angry and fearful), with direct or averted gaze, using electroencephalogram (EEG) measures of the face-elicited Late Positive Potential (LPP), in young people diagnosed with major depression or first-episode psychosis, compared with a healthy control group. RESULTS: In the control group, there was no evidence of increased attentional allocation to emotional facial expressions, or to facial expressions with a matching emotional expression and gaze direction. There was no evidence, in the depression or first-episode psychosis groups, for a threat-based, attentional hypersensitivity to fearful or angry facial expressions, nor for this effect to be potentiated in response to angry direct or fearful averted gaze faces. However, the absence of such effects could not be concluded due to sample size and the absence of stimulus arousal and valence ratings. Importantly, there was significantly increased attentional allocation in the first-episode psychosis group to facial expressions regardless of emotional expression or gaze direction, compared to both the depression and control group. CONCLUSIONS: There might be an attentional hypersensitivity to facial expressions regardless of emotional expression or gaze direction in first-episode psychosis.
Dependency parser is one of the most important fundamental tools in the natural language processing, which extracts structure of sentences and determines the relations between words based on the dependency grammar. The dependency parser is proper for free order languages, such as Persian. In this paper, data-driven dependency parser has been developed with the help of phrase-structure parser for Persian. The defined feature space in each parser is one of the important factors in its success. Our goal is to generate and extract appropriate features to dependency parsing of Persian sentences. To achieve this goal, new semantic and syntactic features have been defined and added to the MSTParser by stacking method. Semantic features are obtained by using word clustering algorithms based on syntagmatic analysis and syntactic features are obtained by using the Persian phrase-structure parser and have been used as bit-string. Experiments have been done on the Persian Dependency Treebank (PerDT) and the Uppsala Persian Dependency Treebank (UPDT). The results indicate that the definition of new features improves the performance of the dependency parser for the Persian. The achieved unlabeled attachment score for PerDT and UPDT are 89.17% and 88.96% respectively.
Discourse relation classification has proven to be a hard task, with rather low performance on several corpora that notably differ on the relation set they use. We propose to decompose the task into smaller, mostly binary tasks corresponding to various primitive concepts encoded into the discourse relation definitions. More precisely, we translate the discourse relations into a set of values for attributes based on distinctions used in the mappings between discourse frameworks proposed by This arguably allows for a more robust representation of discourse relations, and enables us to address usually ignored aspects of discourse relation prediction, namely multiple labels and underspecified annotations. We study experimentally which of the conceptual primitives are harder to learn from the Penn Discourse Treebank English corpus, and propose a correspondence to predict the original labels, with preliminary empirical comparisons with a direct model.
This chapter engages with research in lingua franca scenarios, defining this term and arguing for the value of adopting a linguistic ethnographic approach in this area. The development of the field of English as a Lingua Franca (ELF) is outlined. Historical antecedents for lingua franca studies in interactional sociolinguistics are identified, including work on intercultural miscommunications, the study of English as an international language and the identification of pragmatic strategies in ELF scenarios such as the let-it-pass procedure. The chapter considers critical debates, including the need to challenge the norms of standard language pedagogies, and the importance of maintaining a critical perspective on the current global dominance of English. It provides a review of current areas of focus in this area, including studies of English as lingua franca in higher education in an increasingly internationalised university system, and in workplaces. The range of methods drawn on in ethnographic studies of lingua franca scenarios is described, and implications for practice are identified in relation to the teaching of English and language policy and planning. Directions for future research identified include the emergence of social and linguistic norms in interaction, and developing the focus on lingua francas other than English.
It is widely accepted that the human cognitive system organizes perceptual input into complex hierarchical descriptions which can be represented by tree structures. Tree structures have been used to describe linguistic, musical and visual perception. In this paper, we will investigate whether there exists an underlying model that governs perceptual organization in general. Our key idea is that the cognitive system strives for the simplest structure (the “simplicity principle”), but in doing so it is biased by the likelihood of previous experiences (the “likelihood principle”). We will present a model which combines these two principles by balancing the notion of most likely tree with the notion of shortest derivation. Experiments with linguistic and musical benchmarks (Penn Treebank and Essen Folksong Collection) show that such a combination outperforms models that are based on either simplicity or likelihood alone.
Counterconditioning (CC) is a form of retroactive interference that inhibits expression of learned behavior. But similar to extinction, CC can be a fairly weak and impermanent form of interference, and the original behavior is prone to relapse. Research on CC is limited, especially in humans, but prior studies suggest it is more effective than extinction at modifying some behaviors (e.g., preference or valence ratings) than others (e.g., physiological arousal). Here, we used a within-subjects design to compare the effects of aversive-to-appetitive CC versus standard extinction on two separate tests of long-term memory in human adults: implicit physiological arousal and explicit episodic memory. Participants underwent Pavlovian fear conditioning to two semantic categories (animals, tools) paired with an electric shock. Conditioned stimuli (i.e., category exemplars) from one category were then extinguished, while stimuli from the other category were paired with a positive outcome. Participants returned 24-h later for a test of skin conductance responses (SCR) to the conditioned exemplars, as well as a surprise recognition memory test for stimuli encoded the previous day. Results showed reduced SCRs at a test for unique stimuli from a category that had undergone CC, relative to stimuli from a category that had undergone standard extinction. Additionally, participants selectively remembered more stimuli encoded during CC than extinction. These results provide new evidence that aversive-to-appetitive CC, as compared to extinction, strengthens memory for items directly associated with a positive outcome, which may provide stronger retrieval competition against a fear memory at test to help diminish fear relapse.
The emergence of China as a global economic power in the 21st Century has brought about surging needs for cross-lingual and cross-cultural mediation, typically performed by translators. Advances in Artificial Intelligence and Language Engineering have been bolstered by Machine learning and suitable Big Data cultivation. They have helped to meet some of the translator's needs, though the technical specialists have not kept pace with the practical and expanding requirements in language mediation. One major technical and linguistic hurdle involves words outside the vocabulary of the translator or the lexical database he/she consults, especially Multi-Word Expressions (Compound Words) in technical subjects. A further problem lies in the multiplicity of renditions of a term in the target language. This paper discusses a proactive approach following the successful extraction and application of sizable bilingual Multi-Word Expressions (Compound Words) for language mediation in technical subjects, which do not fall within the expertise of typical translators, who have inadequate appreciation of the range of new technical tools available to help him/her. Our approach draws on the personal reflections of translators and teachers of translation and is based on the prior R&D efforts relating to 300,000 comparable Chinese-English patents. The subsequent protocol we have developed aims to be proactive in meeting four identified practical challenges in technical translation (e.g. patents). It has broader economic implication in the Age of Big Data (Tsou et al, 2015) and Trade War, as the workload, if not, the challenges, increasingly cannot be met by currently available front-line translators. We shall demonstrate how new tools can be harnessed to spearhead the application of language technology not only in language mediation but also in the “teaching” and “learning” of translation. It shows how a better appreciation of their needs may enhance the contributions of the technical specialists, and thus enhance the resultant synergetic benefits. © 2019 Incoma Ltd. All rights reserved.
Creating an opinionated lexicon is an important step towards a reliable social media analysis system. In this article we are proposing an approach and describing an experiment to build an Arabic polarised lexical database from analysing online implicitly and explicitly rated customer reviews. These reviews are written in modern standard Arabic and Palestinian/Jordanian dialect. Therefore, the produced lexicon contains casual slangs and dialectic entries used by the online community, which is useful for sentiment analysis of informal social media micro-blogs. We have extracted 28,000 entries from processing 15,100 reviews and by expanding the initial lexicon through Google translate. We calculated an implicit rating for every review driven by its text to address the problem of ambiguous opinions of certain online posts, where the text of the review does not match the given rating (the explicit rating). Each entry was given a polarity tag and a confidence score. High confidence scores have increased the precision of the polarisation process. Explicit rating has increased the coverage and confidence of polarity.
Previous studies have shown that hoarding behavior usually starts at a subclinical level in early adolescence and gradually worsens; however, a limited number of studies have examined the prevalence of hoarding behavior and its association with developmental disorders in young adults. The aims of this study were to estimate the prevalence of hoarding behavior and to identify correlations between hoarding behavior and developmental disorder traits in university students. The study participants included 801 university students (616 men, 185 women) who completed questionnaires (ASRS: Adult ADHD Self-Report Scale version 1.1, AQ16: Autism-Spectrum Quotient with 16 items, and CIR: Clutter Image Rating). Among 801 participants, 27 (3.4%) exceeded the CIR cut-off score. Moreover, the participants with hoarding behavior had a significantly higher percentage of ADHD traits compared to participants without hoarding behavior (HB(+) vs HB(−), 40.7% vs 21.7%). In addition, 7.4% of HB(+) participants had autism spectrum disorder (ASD) traits, compared to 4.1% of HB(−) participants. A correlation analysis revealed that the CIR composite score had a stronger correlation with the ASRS inattentive score than with the hyperactivity/impulsivity score (CIR composite vs ASRS IA, r = 0.283; CIR composite vs ASRS H/I, r = 0.147). The results showed a high prevalence of ADHD traits in the university students with hoarding behavior. Moreover, we found that the hoarding behavior was more strongly correlated with inattentive symptoms rather than with hyperactivity/impulsivity symptoms. Our results support the concept of a common pathophysiology behind hoarding behavior and ADHD in young adults.
Recurrent neural network language models (RNNLMs) have become an increasing popular choice for state-of-the-art speech recognition systems. RNNLMs are normally trained by minimizing the cross entropy (CE) using the stochastic gradient descent (SGD) algorithm. However, the SGD method doesn't consider the correlation between parameters and therefore can lead to unstable and slow convergence in training. Second-order optimization methods provide a possible solution to this issue. However these methods are either computationally heavy or do not have competitive performance. In this paper, a novel optimization method - stochastic natural gradient based on minimum variance assumption (SNGM) is proposed for training RNNLMs. It allows the natural gradient method to operate at a comparable training efficiency to the SGD method. By modifying the gradient according to the local curvature of the KL-divergence between current and updated probabilistic distributions, the proposed SNGM approach is shown to outperform both the SGD and limited memory BFGS methods across three tasks: Penn Treebank, Switchboard conversational speech recognition and AMI meeting room transcription in terms of both perplexity and word error rate.
espanolResumen en castellano: Esta tesis trata de la morfologia verbal del ingles antiguo para identificar y lematizar los verbos debiles de esta lengua en un corpus al que se accede a traves de una base de datos lexica. La lematizacion es una de las tareas mas importantes a la hora de construir un diccionario. Sin embargo, es una de las tareas pendientes en el campo de la linguistica historica debido a que no existen corpora exhaustivos y lematizados de esta lengua. El enfoque de esta tesis doctoral esta en la lematizacion de las tres clases de verbos debiles del ingles antiguo, aunque las areas de la Lexicografia y la Linguistica de Corpus son tambien relevantes para esta investigacion. Las fuentes principales de esta investigacion son las formas flexivas que estan atestiguadas en el Dictionary of Old English Corpus (DOEC) y que estan disponibles en el lematizador Norna, las fuentes lexicograficas que existen publicadas sobre esta lengua, principalmente el Dictionary of Old English (DOE), y otras fuentes textuales como el York-Toronto-Helsinki Parsed Corpus of Old English (YCOE) y una indexacion de fuentes secundarias del ingles antiguo. El objetivo principal supone la identificacion de las flexiones de los verbos debiles y de su lematizacion con uno de los lemas propuestos en las listas de referencia. Conseguir este objetivo implica manejar las fuentes disponibles en ingles antiguo para poder lematizar y validar los resultados del analisis y el diseno de un metodo que combine busquedas automaticas en la base de datos lexica Nerthus y la revision manual de los resultados. La metodologia incluye cuatro pasos sucesivos con diversas tareas en cada paso. El primero de estos pasos tiene como objetivo la lematizacion de las formas canonicas de los verbos debiles lanzando cadenas de busquedas especificas para cada clase de verbos debiles en el lematizador Norna, donde esta disponible un indice de tipos del DOEC, la fuente de informacion mas fiable de la que se dispone en ingles antiguo. Despues, los resultados se validan con el DOE y se anaden las formas no-canonicas de los verbos debiles entre las letras A y H. El tercer paso tiene como objetivo identificar las formas no-canonicas de las terminaciones flexivas y de las vocales de los radicales que aparecen con mas frecuencia en los verbos debiles para generar patrones de lematizacion. La busqueda de estos patrones y de la lista de prefijos no-canonicos que esta disponible en Norna culmina en la lematizacion de las formas flexivas no transparentes de los verbos debiles. La validacion de los resultados de las letras I a la Y supone el ultimo paso de la metodologia, donde se comparan los datos obtenidos con el analisis sintactico del YCOE y con los datos que se obtienen de una base de datos de indexacion de las fuentes secundarias del ingles antiguo. Los problemas que surgen a lo largo del proceso de lematizacion tienen que ver principalmente con las peculiaridades del ingles antiguo y las limitaciones de la lematizacion de tipos que esta investigacion sigue. La discusion de los resultados del analisis concluye esta tesis. Las principales aportaciones de esta tesis son las listas de lemas y sus formas flexivas, especialmente las de los verbos entre las letras I y la Y ya que no estan disponibles todavia, y el metodo que se ha disenado para identificar estas formas, incluyendo los patrones de lematizacion generados para lematizar las formas con terminaciones no comunes y vocales no canonicas en el radical. EnglishThis thesis deals with the verbal morphology of the Old English language in order to identify and lemmatise weak verbs in a corpus accessed through a lexical database. Lemmatisation is a pending task in the field of historical linguistics given the lack of comprehensive and lemmatised corpora in this language. The focus of this doctoral dissertation is on the lemmatisation of the three classes of weak verbs, although the linguistic fields of Lexicography and Corpus Linguistics are also relevant to this research. The main aim involves the identification of the canonical and non-canonical realisations of the Old English weak verbs and their lemmatisation with a lemma from a reference list of weak verbs. Achieving this goal involves, firstly, the use of the available sources of the Old English language in order to lemmatise and validate the results and, secondly, the design of a semi-automatic research methodology that combines automatic searches in the lexical database Nerthus and the manual revision of the results in order to achieve this task. The sources for this investigation are the inflectional forms that are attested in the Dictionary of Old English Corpus (DOEC) which are available in the lemmatiser Norna, the lexicographical sources published on the Old English language, mainly the Dictionary of Old English (DOE), and other textual sources such as the York-Toronto-Helsinki Parsed Corpus of Old English (YCOE) and an index of secondary sources of Old English. The methodology comprises four successive steps and several tasks within each step. The first step aims at the lemmatisation of the transparent forms of weak verbs with the search of specific query strings for each subclass of weak verbs in the lemmatiser Norna, where an index type of the DOEC, the most reliable source of information regarding the Old English language, is available. Then, the second step validates the results with the DOE and adds to the analysis the non-canonical attestations for the weak verbs from the letter A-H. Thirdly, the identification of the most recurrent non-canonical inflectional endings and stem vowels attested in weak verbs gives rise to lemmatisation patterns. The search of these sets of correspondences and the list of non-canonical prefixes that is available in Norna results in the lemmatisation of the non-canonical inflections of weak verbs. The validation of the results from the letter I-Y concludes the research methodology with the syntactic parsing provided by the YCOE and the data retrieved from the index of secondary sources of Old English Freya. The issues that arise throughout the lemmatisation process mainly concern the idiosyncrasy of the Old English language writing system and the limitations of the lemmatisation by type that this investigation follows. The quantitative and qualitative discussion of the results of the analysis concludes this thesis. The main contributions of this thesis are the lists of weak lemmas and their lemmatised inflectional forms, specially those of the verbs I-Y which are not available yet and the designed research methodology to identify these forms, including the sets of lemmatisation patterns of the non-canonical inflectional endings and stem vowels of weak verbs.
Recent research on discourse relations has found that they are cued not only by discourse markers (DMs) but also by other textual signals and that signaling information is indicative of genres. While several corpora exist with discourse relation signaling information such as the Penn Discourse Treebank (PDTB, Prasad et al. 2008) and the Rhetorical Structure Theory Signalling Corpus (RST-SC, Das and Taboada 2018), they both annotate the Wall Street Journal (WSJ) section of the Penn Treebank (PTB, Marcus et al. 1993), which is limited to the news domain. Thus, this paper adapts the signal identification and anchoring scheme (Liu and Zeldes, 2019) to three more genres, examines the distribution of signaling devices across relations and genres, and provides a taxonomy of indicative signals found in this dataset.
The Plain Meaning Rule is often assailed on the grounds that it is unprincipled—that it substitutes for careful analysis an interpreter’s ad hoc and impressionistic intuition about the meaning of legal texts. But what if judges and lawyers had the means to test their intuitions about plain meaning systematically? Then initial linguistic impressions about the meaning of a legal text might be viewed as hypotheses to be tested, rather than determinative criteria upon which to base important decisions. There exists very little legal scholarship on corpus linguistics—the study of language function and use through large, electronic linguistic databases called corpora—and the role that corpus methods might play in legal interpretation. This omission becomes more and more striking as scholars and jurists (and even the United States Supreme Court) have found themselves persuaded by corpus-based arguments. This Article argues that the plain or ordinary meaning of a given term in a given context is an empirical matter that may be quantified through corpus-based methods. These methods, when applied to questions of legal ambiguity, present significant advantages over existing empirical approaches to plain meaning and over the prevailing intuition-based interpretive approach of many courts. Because large, sophisticated linguistic corpora are widely available and easy to use, and because corpus methods offer a more principled and systematic alternative to the impressionistic interpretation of legal texts, corpus linguistics may one day revolutionize the process of legal interpretation.
Affect fluctuates in a moment-to-moment fashion, and it reflects the continuous relationship between the individual and the environment. Despite substantial research, there remain important open questions regarding how the continuous stream of sensory input is dynamically represented in experienced affect. Here, approaching affect as a temporally dependent process, we show that momentary affect is shaped by a combination of changes in recent stimuli (i.e. visually presented images for the current studies) and previously experienced affect. We also found that this temporally dependent relationship is influenced by context uncertainty. Participants, in each trial, viewed sequentially presented images and subsequently reported their affective experience, which was modeled based on images’ normative affect ratings and participants’ previously reported affect. Study 1 showed that self-reported valence and arousal in a given trial is partly shaped by the affective impact of the given images and previously experienced affect. In Study 2, we manipulated context uncertainty by controlling occurrence probabilities for normatively pleasant and unpleasant images in separate trials. Increasing context uncertainty (i.e. random occurrence of pleasant and unpleasant images) is associated with increased negative affect when the overall effect context is controlled. In addition, the relative contribution of the most recent image to momentary affect increased with increasing context uncertainty. Taken together, these findings provide clear behavioral evidence that affective experience fluctuates in a temporally dependent and continuous fashion based on recent changes in input variables and previous internal state, and that these fluctuations are sensitive to the affective context and its certainty.
Introduction: Hoarding behaviour is a common symptom seen in patients with Obsessive Compulsive Disorder (OCD). The phenomenology and prevalence of hoarding in OCD have been under studied in India and the phenomenon is less explored on routine clinical examination. Aim: To study the prevalence and phenomenology of hoarding as a symptom in patients with OCD and tried to elucidate some differences between OCD patients with and without hoarding symptoms. Materials and Methods: A total of 50 patients with OCD and 50 relatives of psychiatric patients were the subjects for the study. The OCD group was administered the Yale Brown Obsessive Compulsive Scale (YBOCS), the Hoarding Rating Scale and the Clutter Image Rating Scale. The 50 cases of OCD were further divided on the presence and absence of hoarding as a symptom into 2 groups and the scores on the scales used were statistically analysed using descriptive statistics like frequency and percentages, chi-square test and unpaired t-test. Results: The mean duration of illness was 8.01±5.17 years and the mean age of onset of the illness was 27.28±7.11 years for all patients with OCD. OCD patients with hoarding had a shorter total duration of illness than those without hoarding. Newspapers and scrap were hoarded the most with sentimental reasons along with importance of goods were cited as reasons for the behaviour. The two groups showed significant differences on compulsive sub scale of the YBOCS and no differences were noted in the other scales used. Conclusion: Patients having OCD with hoarding as one of the symptoms may differ from those not having hoarding. However larger studies across diverse groups are needed to corroborate these findings.
Lectometric approaches measure distances between language varieties (dialects, sociolects, registers etc.) by aggregating over observed differences in the realizations of a set of linguistic variables. In <em>lexical</em> lectometry, a variable consists of the alternative lexical expressions for one concept. In <em>corpus-based</em> lectometry, the observed realizations are culled from stratified corpora. Measuring semantically defined variables in corpora, and aggregating over them, poses specific methodological challenges that have been tackled in a number of studies (Heylen & Ruette 2013; Ruette et al. 2014; Ruette, Ehret & Szmrecsanyi 2016) with different statistical techniques, including Distributional Semantic Models. Yet so far, no general framework for corpus-based lexical lectometry has been formulated that systematically describes the issues and options in each step of the procedure so that it can be straightforwardly applied to new data and new languages, other than English (Ruette, Ehret & Szmrecsanyi 2016), Dutch (Geeraerts, Grondelaers & Speelman 1999) and Portuguese (Soares da Silva 2010). This paper can be characterized as a twofold extension of the previous studies. First, it aims to establish a general framework for lexical lectometry research that considers most if not all options for different steps. Second, we want to go beyond the Indo-European languages by extending the framework on a typologically unrelated language, i.e. Chinese varieties. For the general framework, we propose that a proper lexical lectometry research normally should involve the following steps: (1) compilation of a lectally stratified corpus; (2) sampling concepts as measuring points for lectometry; (3) identification of lexical expressions per concept; (4) disambiguation of lexical expressions in corpus data; (5) calculation of aggregated lexico-lectometric distances; (6) evaluation of measurement reliability and validity. For each step, we further provide possible options and caveats. For instance, step 2 and 3 can rely on existing concept-based lexical databases, like a synonym dictionary, or use corpus-driven keyword extraction and semantic vector space models. Step 4 can either make use of token-level distributional semantics models or rely on simpler n-gram language models. To assess the portability of the general framework, both in practical and linguistic-typological terms, we perform a lexical lectometric analysis for varieties of Chinese based on data from large-scale corpora of Mainland Chinese, Taiwan Chinese and Singapore Chinese.
BACKGROUND: People experiencing mental illness require services that provide them with a sense of personal safety, a place where they can experience a reduction to their distress and assistance in managing their feelings. Interventions need to explore therapies that enhance feelings of personal safety and comfort for consumers and within a forensic mental health service, therapies and support that can assist in combating the antecedents to violent offending. The practice of Qigong is reported to have numerous health benefits; however, little has been reported regarding the possible benefits of Qigong for people experiencing severe mental illness and, more specifically, for people experiencing severe mental illness who have serious offending histories such as forensic consumers. This study explores the possibility of using Qigong to reduce personal frustrations that can lead to violence. OBJECTIVES: The object of this study was to explore whether Qigong is an effective intervention on positive affect traits for forensic mental health consumers, and whether other benefits are experienced. METHODS: An exploratory design using quantitative and qualitative approaches was used. Consumers participated in weekly Qigong groups delivered for a 10-wk period. Data were collected using an adapted version of the positive affect rating scale measuring the degree to which people experience different positive emotions. Qualitative measures were added to the scale to obtain a deeper understanding of the consumer experience, with 67 scales completed. CONCLUSIONS: Consumers in a forensic hospital responded positively to participating in Qigong groups. Strategies such as Qigong are interventions that mental health clinicians can use to promote positive feelings of personal relaxation, peacefulness, and safety. Qigong can promote positive affective traits for consumers in forensic hospitals. These positive affective traits can act as protective factors to inpatient aggression and violence. Forensic consumers report that Qigong is easy to learn and helpful for them in managing their frustrations. The findings from this study may add to the paucity of data discussing the use of Qigong with consumers as an effective relaxation intervention and possibly as an intervention in reducing negative affective states by promoting positive affective states, thereby reducing aggression and possible violence occurring within the forensic inpatient environment.
This paper presents the submission by the CMU-01 team to the SIGMORPHON 2019\ntask 2 of Morphological Analysis and Lemmatization in Context. This task\nrequires us to produce the lemma and morpho-syntactic description of each token\nin a sequence, for 107 treebanks. We approach this task with a hierarchical\nneural conditional random field (CRF) model which predicts each coarse-grained\nfeature (eg. POS, Case, etc.) independently. However, most treebanks are\nunder-resourced, thus making it challenging to train deep neural models for\nthem. Hence, we propose a multi-lingual transfer training regime where we\ntransfer from multiple related languages that share similar typology.\n
This paper suggests annotation guidelines to build a Universal Dependencies (UD) treebank for Korean. We discuss the part-of-speech annotation of Korean specific-categories such as prenouns, numeral classifiers, and (pre)final endings, and propose how to implement UD scheme in Korean regarding selecting a head and assigning dependency relations to dependents. UD prioritizes content words over functional words since the former exhibits less cross-linguistic variations. In a noun phrase, for instance, a core noun is always a head of the entire noun phrase independently of a language. The rest are treated as a dependent: not only a modifier such as an adjective but also a functional category such as an article, numeral quantifier, demonstrative, and so on. However, when it comes to head-less constructions such as coordination or predicate ellipsis, UD firmly advocates the head-initial strategy. The present application of UD to Korean tries to follow UD’s principles as much as possible. Korean is a head-final language, so that headed constructions are analyzed head-finally. In contrast, head-less ones are tagged head-initially. This might disregard language-specific characteristics from a linguistic perspective, but the strategy allows us to build up a set of treebanks in a cross-linguistically consistent way (i.e., the fundamental purpose of UD).
This thesis studies the connections between parsing friendly representations and interlingua grammars developed for multilingual language generation. Parsing friendly representations refer to dependency tree representations that can be used for robust, accurate and scalable analysis of natural language text. Shared multilingual abstractions are central to both these representations. Universal Dependencies (UD) is a framework to develop cross-lingual representations, using dependency trees for multlingual representations. Similarly, Grammatical Framework (GF) is a framework for interlingual grammars, used to derive abstract syntax trees (ASTs) corresponding to sentences. The first half of this thesis explores the connections between the representations behind these two multilingual abstractions. The first study presents a conversion method from abstract syntax trees (ASTs) to dependency trees and present the mapping between the two abstractions – GF and UD – by applying the conversion from ASTs to UD. Experiments show that there is a lot of similarity behind these two abstractions and our method is used to bootstrap parallel UD treebanks for 31 languages. In the second study, we study the inverse problem i.e. converting UD trees to ASTs. This is motivated with the goal of helping GF-based interlingual translation by using dependency parsers as a robust front end instead of the parser used in GF. \n\nThe second half of this thesis focuses on the topic of data augmentation for parsing – specifically using grammar-based backends for aiding in dependency parsing. We propose a generic method to generate synthetic UD treebanks using interlingua grammars and the methods developed in the first half. Results show that these synthetic treebanks are an alternative to develop parsing models, especially for under-resourced languages without much resources. This study is followed up by another study on out-of-vocabulary words (OOVs) – a more focused problem in parsing. OOVs pose an interesting problem in parser development and the method we present in this paper is a generic simplification that can act as a drop-in replacement for any symbolic parser. Our idea of replacing unknown words with known, similar words results in small but significant improvements in experiments using two parsers and for a range of 7 languages.
Many text corpora exhibit socially problematic biases, which can be propagated or amplified in the models trained on such data. For example, doctor cooccurs more frequently with male pronouns than female pronouns. In this study we (i) propose a metric to measure gender bias; (ii) measure bias in a text corpus and the text generated from a recurrent neural network language model trained on the text corpus; (iii) propose a regularization loss term for the language model that minimizes the projection of encoder-trained embeddings onto an embedding subspace that encodes gender; (iv) finally, evaluate efficacy of our proposed method on reducing gender bias. We find this regularization method to be effective in reducing gender bias up to an optimal weight assigned to the loss term, beyond which the model becomes unstable as the perplexity increases. We replicate this study on three training corpora---Penn Treebank, WikiText-2, and CNN/Daily Mail---resulting in similar conclusions.
The emergence of deep learning as a commanding technique for learning heterogeneous layers of feature representations have consequently substituted traditional machine learning algorithms which are generally poor in analyzing compound sentences. Additionally, convolutional and recurrent neural networks have auspiciously yielded state-of-the-art results in sentiment classification and Natural Language Processing (NLP). In this paper, a deep sentiment representation model through the combination of multiple Convolutional Neural Networks (CNN) kernels with Long Short-Term Memory (LSTM) is proposed for sentiment classification. Our model gains word vector representation using pre-trained Global Vectors for Word Representation (GloVe) embeddings, thereafter used as input to the CNN layer which extracts higher local text representations. Finally, Bidirectional LSTM (biLSTM) generates sentiment classification of sentence representation based on context dependent features. Our combined approach of CNN and biLSTM was experimented using the Stanford Large Movie Review Dataset (IMDB) and Stanford Sentiment Treebank Dataset (SSTB) for binary classification. The evaluation achieves outstanding results in outperforming several existing approaches with 90.4% accuracy on the Stanford Sentiment Treebank dataset and 94.8% accuracy on the Stanford Large Movie Review dataset. These results are achieved with a drastic reduction of model parameters and without a pooling layer in the CNN architecture, helping to retain local and structural information in comparison to other existing deep neural network frameworks.
In sequence learning tasks such as language modelling, Recurrent Neural\nNetworks must learn relationships between input features separated by time.\nState of the art models such as LSTM and Transformer are trained by\nbackpropagation of losses into prior hidden states and inputs held in memory.\nThis allows gradients to flow from present to past and effectively learn with\nperfect hindsight, but at a significant memory cost. In this paper we show that\nit is possible to train high performance recurrent networks using information\nthat is local in time, and thereby achieve a significantly reduced memory\nfootprint. We describe a predictive autoencoder called bRSM featuring recurrent\nconnections, sparse activations, and a boosting rule for improved cell\nutilization. The architecture demonstrates near optimal performance on a\nnon-deterministic (stochastic) partially-observable sequence learning task\nconsisting of high-Markov-order sequences of MNIST digits. We find that this\nmodel learns these sequences faster and more completely than an LSTM, and offer\nseveral possible explanations why the LSTM architecture might struggle with the\npartially observable sequence structure in this task. We also apply our model\nto a next word prediction task on the Penn Treebank (PTB) dataset. We show that\na 'flattened' RSM network, when paired with a modern semantic word embedding\nand the addition of boosting, achieves 103.5 PPL (a 20-point improvement over\nthe best N-gram models), beating ordinary RNNs trained with BPTT and\napproaching the scores of early LSTM implementations. This work provides\nencouraging evidence that strong results on challenging tasks such as language\nmodelling may be possible using less memory intensive, biologically-plausible\ntraining regimes.\n
Throughout many studies which focus on brain laterality as a key component to the outcome of an experiment, or to a participants’ reaction to a stimulus, it can be noted that the different areas of the brain involved in a task response must work together to produce a viable outcome (i.e. lateralized brain processes). While there are laterality components in relation to the cognitive processes of handed and footed responses, it is still largely unknown how the different areas of the brain interpret the emotional stimuli to then affect these outcomes. The purpose of our study is to determine how emotional context affects a simple cognitive task that includes handed and footed responses, and if any observed differences can be traced back to the different systems at work within the brain. Subjects will be tested over two days for handed and footed responses in a cognitive Simon Task. Subjects will be tested with and without emotional context (i.e. a background image of a specific valence and arousal rating), and any resulting differences between non-emotional context and emotional context reaction times will be compared.
Do individual sounds carry meaning? The relationship between sound and meaning in human languages is typicallyassumed to be arbitrary, though recent research provides evidence for the existence of both iconicity and systematicitybetween word forms and their meaning. However, this research has not asked whether individual sounds in a languagecovary in systematic ways with aspects of meaning. In two analyses, we find evidence for more systematicity betweenthe initial phones of words and those words concreteness ratings than one would expect in a truly arbitrary lexicon. Thissuggests that initial phones may act as cues to aspects of word meaning, and raises questions about whether languagelearners detect and exploit these cues.
This talk will explore the role of individual social actors and their communities and networks in the formation, maintenance and dissolution of linguistic norms. It will consider several standardisation episodes in the history of Old and Early Middle English and connect them to other unification processes in the political and cultural history of England at the time. It will be suggested that the suppression of variability on the linguistic level often accompanies, or is a symptom of, a similar suppression on the ideological level. Political, religious, legal, and linguistic processes mingle in various ways in this period (as they do today) to replicate and enhance the social order, to support a reform movement, or to refute dissent. \nThree case studies will offer insights into these processes: 1) shire courts and the ‘standardised’ lexis of the Anglo-Saxon Chronicle in the reign of King Alfred and Edward the Elder; 2) chancery norms and charters of the eleventh century; and 3) religious reform and linguistic focusing in the thirteenth century.
Abstract This chapter is devoted to the presentation of the tools and methods used for the different steps of the semi-automatic syntactic annotation: automatic preprocessing; microsyntactic parsing with the FRMG tool, correction of the parsing with the Arborator tool, agreement analysis, post-validation correction, and development of the final format of the Rhapsodie syntactic treebank. As FRMG is a parser for written French that was not configured to analyze disfluencies and reformulation, we used our manual pile marking to unfold the piles and produce a series of simplified “sentences” with only government relations. Despite having two annotators plus a validator for the corrections, we found a substantial number of errors in the post-validation procedure by using a set of rules to determine the well-formedness of the trees.
Contextualized embeddings, which capture appropriate word meaning depending\non context, have recently been proposed. We evaluate two meth ods for\nprecomputing such embeddings, BERT and Flair, on four Czech text processing\ntasks: part-of-speech (POS) tagging, lemmatization, dependency pars ing and\nnamed entity recognition (NER). The first three tasks, POS tagging,\nlemmatization and dependency parsing, are evaluated on two corpora: the Prague\nDependency Treebank 3.5 and the Universal Dependencies 2.3. The named entity\nrecognition (NER) is evaluated on the Czech Named Entity Corpus 1.1 and 2.0. We\nreport state-of-the-art results for the above mentioned tasks and corpora.\n
We present a recurrent neural network memory that uses sparse coding to\ncreate a combinatoric encoding of sequential inputs. Using several examples, we\nshow that the network can associate distant causes and effects in a discrete\nstochastic process, predict partially-observable higher-order sequences, and\nenable a DQN agent to navigate a maze by giving it memory. The network uses\nonly biologically-plausible, local and immediate credit assignment. Memory\nrequirements are typically one order of magnitude less than existing LSTM, GRU\nand autoregressive feed-forward sequence learning models. The most significant\nlimitation of the memory is generalization to unseen input sequences. We\nexplore this limitation by measuring next-word prediction perplexity on the\nPenn Treebank dataset.\n
In this paper, we discuss constituent ordering generalizations in Japanese. Japanese has SOV as its basic order, but a significant range of argument order variations brought about by ‘scrambling’ is permitted. Although scrambling does not induce much in the way of semantic effects, it is conceivable that marked orders are derived from the unmarked order under some pragmatic or other motivations. The difference in the effect of basic and derived order is not reflected in native speaker’s grammaticality judgments, but we suggest that the intuition about the ordering of arguments may be attested in corpus data. By using the Keyaki treebank (a proper subset of which is NINJAL Parsed Corpus of Modern Japanese (NPCMJ)), it is shown that the naturally-occurring corpus data confirm that marked orderings of arguments are less frequent than their unmarked ordering counterparts. We suggest some possible motivations lying behind the argument order variations.
Classic natural language processing resources such as the Penn Treebank (Marcus et al. 1993) have long been used both as evaluation data for many linguistic tasks and as training data for a variety of off-the-shelf language processing tools. Recent work has highlighted a gender imbalance in the authors of this text data (Garimella et al. 2019) and hypothesized that tools created with such resources will privilege users from particular demographic groups (Hovy and Søgaard 2015). Domain adaptation is typically employed as a strategy in machine learning to adjust models trained and evaluated with data from different genres. However, the present work seeks to evaluate whether domain adaptation to demographic groups such as age or gender may be an effective strategy to ameliorate the effects of biased or outdated training corpora in linguistic preprocessing tasks. We find adaptation to demographic groups to be an effective strategy for improving preprocessing performance across all demographic groups.
Mirror-sensory synesthetes mirror the pain or touch that they observe in other people on their own bodies. This type of synesthesia has been associated with enhanced empathy. We investigated whether the enhanced empathy of people with mirror-sensory synesthesia influences the experience of situations involving touch or pain and whether it affects their prosocial decision making. Mirror-sensory synesthetes (<i>N</i> = 18, all female), verified with a touch-interference paradigm, were compared with a similar number of age-matched control individuals (all female). Participants viewed arousing images depicting pain or touch; we recorded subjective valence and arousal ratings, and physiological responses, hypothesizing more extreme reactions in synesthetes. The subjective impact of positive and negative images was stronger in synesthetes than in control participants; the stronger the reported synesthesia, the more extreme the picture ratings. However, there was no evidence for differential physiological or hormonal responses to arousing pictures. Prosocial decision making was assessed with an economic game assessing altruism, in which participants had to divide money between themselves and a second player. Mirror-sensory synesthetes donated more money than non-synesthetes, showing enhanced prosocial behaviour, and also scored higher on the Interpersonal Reactivity Index as a measure of empathy. Our study demonstrates the subjective impact of mirror-sensory synesthesia and its stimulating influence on prosocial behaviour.This article is part of the discussion meeting issue ‘Bridging senses: new developments in synaesthesia’.
Creating an opinionated lexicon is an important step towards a reliable social media analysis system. In this article we are proposing an approach and describing an experiment to build an Arabic polarised lexical database from analysing online implicitly and explicitly rated customer reviews. These reviews are written in modern standard Arabic and Palestinian/Jordanian dialect. Therefore, the produced lexicon contains casual slangs and dialectic entries used by the online community, which is useful for sentiment analysis of informal social media micro-blogs. We have extracted 28,000 entries from processing 15,100 reviews and by expanding the initial lexicon through Google translate. We calculated an implicit rating for every review driven by its text to address the problem of ambiguous opinions of certain online posts, where the text of the review does not match the given rating (the explicit rating). Each entry was given a polarity tag and a confidence score. High confidence scores have increased the precision of the polarisation process. Explicit rating has increased the coverage and confidence of polarity.
Like odor identification, remote odor memory, reflected in familiarity ratings, is impaired in AD (Murphy, Nature Reviews Neurology, 2019). We investigated the relative abilities of standard screening (MMSE), odor identification and remote odor memory to predict transitions from amnestic MCI (aMCI) to AD in a sample from the UCSD ADRC. The sample contained 170 controls, 210 AD, 26 aMCI converters to AD and 42 aMCI non-converters. A receiver operating characteristic (ROC) curve plots the trade-off between sensitivity and specificity. The area under the curve (AUC) indicates how well a marker discriminates patients from controls. Analyses showed higher predictive value for converting from aMCI to AD in ApoE ε4+ carriers for odor familiarity, odor identification and for the combination than for the MMSE. ROC/AUCs for the conversion from aMCI to AD have ranged from.63 -.67 for CSF biomarkers. Odor familiarity and odor identification had similar AUC values; however, combining odor familiarity and odor identification produced an ROC/AUC value of 1.0 in ε4 carriers, appreciably higher than for MMSE alone (.58). Olfactory biomarkers show real promise as early, non-invasive indicators of disease, particularly in samples enriched with ε4 carriers. Although odor identification has been the focus of olfactory biomarker work, the results suggest that other measures of olfactory function have the potential to enhance prediction. Combining odor familiarity and odor identification produced a predictive value of 1.0 in ε4 carriers. The results warrant further investigation into the potential for enhancing drug trials and clinical screening. Supported by NIH grants R01AG004085-26 (CM) and P50AG005131 (UCSD ADRC). We thank the UCSD ADRC and particularly Drs. Douglas Galasko and David Salmon.
Cet article presente la creation d’un treebank journalistique serbe, ParCoJour. Il est compose de 30K tokens et dote de trois couches d’annotation: etiquetage morphosyntaxique, lemmatisation et annotation syntaxique. Une fois construit, ParCoJour a ete utilise dans trois experiences afin d’evaluer l’impact du domaine textuel sur le parsing du serbe en comparant les performances de Talismane, un systeme par apprentissage automatique, sur deux types de corpus, journalistique et litteraire: 1) parsing du corpus journalistique avec un modele entraine sur le corpus journalistique; 2) parsing du corpus journalistique avec un modele entraine sur le corpus litteraire; 3) parsing du corpus litteraire avec un modele entraine sur le corpus journalistique. Les resultats sont compares a ceux ou les deux corpus relevaient du domaine litteraire. Le changement de domaine textuel dans la deuxieme et la troisieme experience entraine une baisse des performances, mais les resultats de parsing restent satisfaisants.