Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
In recent years, dependency parsing is a fascinating research topic and has a lot of applications in natural language processing. In this paper, we present an effective approach to improve dependency parsing by utilizing supertag features. We performed experiments with the transition-based dependency parsing approach because it can take advantage of rich features. Empirical evaluation on Vietnamese Dependency Treebank showed that, we achieved an improvement of 18.92% in labeled attachment score with gold supertags and an improvement of 3.57% with automatic supertags.
Text diacritization is a critical task which plays an important role for improving the performance of many NLP tasks for languages that include diacritics in their orthographies. In this paper, we handle the problem of Arabic text diacritization such that our system diacritize input Arabic sequence of words both morphologically and syntactically. The operation of the system is divided into three layers: the first layer uses HMM for the morphological diacritization of previously seen words, the second layer uses an external morphological analyzer for the morphological diacritization of OOV words, and the third layer uses CRF for the syntactic diacritization of all words. To evaluate the performance of the system, we used the benchmark LDC Arabic Treebank Part 3 datasets used by the state-of-the-art systems. The proposed system achieved a morphological WER of 4.3%, and a syntactic WER of 9.4%.
We train a language-universal dependency parser on a multilingual collection of treebanks. The parsing model uses multilingual word embeddings alongside learned and specified typological information, enabling generalization based on linguistic universals and based on typological similarities. We evaluate our parser's performance on languages in the training set as well as on the unsupervised scenario where the target language has no trees in the training data, and find that multilingual training outperforms standard supervised training on a single language, and that generalization to unseen languages is competitive with existing model-transfer approaches.
It is well recognised in psychology that music has affective connotations and that musical stimuli can modify affective states. The aim of this study was to assess the affective connotations of 120 fifteen-second musical excerpts, covering both modern musical genres such as pop, rock, jazz, rap/R&B and electronic music (5 x N = 20), and classical music ( N = 20). Expert judges used predetermined criteria to select excerpts with positive or negative valence that induced high arousal or low arousal. The excerpts were assessed by 50 undergraduate students (25 women) from different academic departments, aged between 18 and 28 years ( M = 21.46 years, SD = 1.85). They listened to all 120 fragments and rated them with respect to six dimensions: valence, arousal, dominance, origin, subjective significance and imageability. Analyses showed that ratings were reliable, with high split-half correlations and Cronbach’s alpha estimates. We did not identify any gender differences concerning affective reactions to the music. Some music genre specificity was found for all measures, and initial music preference appeared to shape affective ratings. The results presented here will be of interest to researchers working on musical perception and the influence of music on affective outcomes and emotional regulation.
PURPOSE: The focus of this study was to examine the influence of fundamental frequency (F0) and vocal tract length (VTL) modifications on speaker gender recognition in cochlear implant (CI) recipients for different stimulus types. METHOD: Single words and sentences were manipulated using isolated or combined F0 and VTL cues. Using an 11-point rating scale, CI recipients and listeners with normal hearing rated the maleness/femaleness of the corresponding voice. RESULTS: Speaker gender ratings for combined F0 and VTL modifications were similar across all stimulus types in both CI recipients and listeners with normal hearing, although the CI recipients showed a somewhat larger ambiguity. In contrast to listeners with normal hearing, F0-VTL and F0-only modifications revealed similar ratings in the CI recipients when using words as stimuli. However, when sentences were used, a difference was found between F0-VTL-based and F0-based ratings. Modifying VTL cues alone did not affect ratings in the CI group. CONCLUSIONS: Whereas speaker gender ratings by listeners with normal hearing relied on combined VTL and F0 cues, CI recipients made only limited use of VTL cues, which might be one reason behind problems with identifying the speaker on the basis of voice. However, use of the voice cues depended on stimulus type, with the greater information in sentences allowing a more detailed analysis than single words in both listener groups.
Drawing on the work of Lawrence Abu Hamdan, a British-Lebanese artist and researcher currently based in Beirut, this essay examines the juridical and conceptual field of critical forensis which is situated at the juncture of security studies, art, and architecture. Abu Hamdan extends forensics to the area of “new audibilities,” with a focus on the politics of juridical hearing in situations of legal-identity profiling and voice authentication (the “shibboleth test”). Abu Hamdan's projects investigate how accent monitoring and audio surveillance, voice recognition, translation technologies, sovereign acts of listening, and court determinations of linguistic norms emerge as so many technical constraints on “freedom of speech,” itself a malleable term ascribed to discrepant claims and principles, yet taking on performative force in site-specific situations.
Anxiety disorders may not only be characterized by specific symptomatology (e.g., tachycardia) in response to the fearful stimulus (primary problem or first-level emotion) but also by the tendency to negatively evaluate oneself for having those symptoms (secondary problem or negative meta-emotion). An exploratory study was conducted driven by the hypothesis that reducing the secondary or meta-emotional problem would also diminish the fear response to the phobic stimulus. Thirty-three phobic participants were exposed to the phobic target before and after undergoing a psychotherapeutic intervention addressed to reduce the meta-emotional problem or a control condition. The electrocardiogram was continuously recorded to derive heart rate (HR) and heart rate variability (HRV) and affect ratings were obtained. Addressing the meta-emotional problem had the effect of reducing the physiological but not the subjective symptoms of anxiety after phobic exposure. Preliminary findings support the role of the meta-emotional problem in the maintenance of response to the fearful stimulus (primary problem).
This paper illustrates the similarity between Thai and Laotian, and between Malay and Indonesian, based on an investigation on raw parallel data from Asian Language Treebank. The cross-lingual similarity is investigated and demonstrated on metrics of correspondence and order of tokens, based on several standard statistical machine translation techniques. The similarity shown in this study suggests a possibility on harmonious annotation and processing of the language pairs in future development.
The present study compared the effects of: (a) PETTLEP imagery (e.g. imaging in the environment), (b) prior-observation (i.e. observing prior to imaging), and (c) traditional imagery (e.g. imaging sat in a quiet room) on the ease and vividness of external visual imagery (EVI), internal visual imagery (IVI), and kinaesthetic imagery (KI) of movements. Fifty-two participants (28 female, 24 male, Mage = 19.60 years, SD = 1.59) imaged the movements described in the Vividness of Movement Imagery Questionnaire-2 under the three conditions in a counterbalanced order. Vividness and ease of imaging ratings were recorded for each movement. A repeated measure MANOVA revealed that ease and vividness ratings for EVI, IVI, and KI were higher during the PETTLEP imagery condition compared to the traditional imagery condition, and vividness of EVI was higher during the observation imagery condition compared to traditional imagery. Findings indicate that incorporating PETTLEP elements into the imagery instructions leads to easier and more vivid movement EVI, IVI, and KI imagery.
Abstract We present a work in progress aimed at extracting translation pairs of source and target dependency treelets to be used in a dependency-based machine translation system. We introduce a novel unsupervised method for parallel tree segmentation based on Gibbs sampling. Using the data from a Czech-English parallel treebank, we show that the procedure converges to a dictionary containing reasonably sized treelets; in some cases, the segmentation seems to have interesting linguistic interpretations.
This article addresses the question of the types of morphological relatedness that the lexical class of the adjective presents in Old English. After an exhaustive analysis of the derivational paradigms of the language based on data retrieved from the lexical database of Old English Nerthus, the following conclusions are reached. Two types of morphological relatedness are identifiable, namely explicit and implicit. Short distance and long distance relations overlap with explicit and implicit morphological relatedness. These relations involve four types of units, to wit lexical primes (the bases of lexical paradigms), derived adjectives (the input to recursive processes of word-formation), target adjectives (the output of processes that cannot be inputted to a recursive process) and morphologically unrelated adjectives (which are neither the input nor the output of a process of word-formation). © 2016 Society for Studia Neophilologica.
The object of study in the article is represented by orthographic norm as one of the varieties of linguistic norm. The aim was to identify the problems of realization of spelling rules associated with linguistic and extra-linguistic factors. Features of functioning of spelling rules before and after the publication of "The rules of Russian orthography and punctuation" in 1956 shows the understanding of the essence of spelling rules as a two-component phenomena (orthogrammes and spelling rules), taking into account the forms of its existence; stated dynamics of spelling rules. It was found that errors in implementing the orthographic norms in the practice of writing ordinary speakers have both intra-linguistic and extra-linguistic reasons. The main ones are the following: delayed change of fixing spelling rules as a result of the transformation of other linguistic norms (formative, morphological, orthoepic, etc...); a reflection on the letter settings "old" spelling rules as a result of ignorance of the writing content of the "new" rules; incorrectness of language game associated with the need to draw attention to the advertising text; poor reading of the younger generation; reading books and other printed products of inadequate quality editing and proofreading processing; impact of the language of Internet communication.
Light verb constructions (LVC) in Hindi are highly productive. If we can distinguish a case such as nirnay lenaa ‘decision take; decide’ from an ordinary verb-argument combination kaagaz lenaa ‘paper take; take (a) paper’,it has been shown to aid NLP applications such as parsing (Begum et al., 2011) and machine translation (Pal et al., 2011). In this paper, we propose an LVC identification system using language specific features for Hindi which shows an improvement over previous work(Begum et al., 2011). To build our system, we carry out a linguistic analysis of Hindi LVCs using Hindi Treebank annotations and propose two new features that are aimed at capturing the diversity of Hindi LVCs in the corpus. We find that our model performs robustly across a diverse range of LVCs and our results underscore the importance of semantic features, which is in keeping with the findings for English. Our error analysis also demonstrates that our classifier can be used to further refine LVC annotations in the Hindi Treebank and make them more consistent across the board.
Transient global amnesia (TGA) is a disorder with reversible anterograde disturbance of explicit memory, frequently preceded by an emotionally or physically stressful event. By using magnetic resonance imaging (MRI) following an episode of TGA, small hippocampal lesions have been observed. Hence it has been postulated that the disorder is caused by the stress-related transient inhibition of memory formation in the hippocampus. In experimental studies, stress has been shown to affect both explicit and implicit learning-the latter defined as learning and memory processes that lack conscious awareness of the information acquired. To test the hypothesis that impairment of implicit learning in TGA is present and related to stress, we determined the effect of experimental exposure to stress on hippocampal activation patterns during an implicit learning paradigm in patients who suffered a recent TGA and healthy matched control subjects. We used a hippocampus-dependent aversive learning procedure (context conditioning with the phases habituation, acquisition, and extinction) during functional MRI following experimental stress exposure (socially evaluated cold pressor test). After a control procedure, controls showed successful learning during the acquisition phase, indicated by increased valence, arousal and contingency ratings to the paired (CON+) vs. the non-paired (CON-) conditioned stimulus, and successful extinction of the conditioned responses. Following stress, acquisition was still successful, however extinction was impaired with persistently increased contingency ratings. In contrast, TGA patients showed impairment of conditioned responses and insufficient extinction after the control procedure, indicated by a lack of significant differences between CON+ and CON- for valence and arousal ratings after the acquisition phase and by significantly increased contingency ratings after the extinction. After stress, aversive learning was not successful with non-significant ratings of all parameters. Concerning brain activation patterns after the control procedure, controls showed increased hippocampal response during acquisition after the control procedure. This was not seen after stress exposure. In TGA patients, we observed an increased response in the right ventral striatum in the acquisition phase following stress. These findings suggest that alterations in implicit learning processes, including impaired hippocampal and increased striatal responses, might play a role in TGA pathophysiology, partly related to acute stress.
Research finds we make spontaneous trait inferences from facial appearance, even after brief exposures to a face (i.e., ≤ 100 ms). We examined spontaneous impressions of criminality from facial appearance, testing whether these impressions persist after repeated presentation (i.e., one to three exposures) and increased exposure duration (100, 500, or 1000 ms) to the face. Judgement confidence and response times were recorded. Other participants viewed the faces for an unlimited period of time, rating trustworthiness dominance, and criminal appearance. We found evidence that participants spontaneously make criminal appearance attributions. These inferences persisted with repeated presentation and increased exposure duration, were related to trustworthiness and dominance ratings, and were made with high confidence. Implications are discussed.
We developed a novel word sense disambiguation algorithm that uses the semantic relations of lexical database Poly-WordNet. The PolyWordNet is a lexical database that organizes multiple senses of a polysemy word in such a way that each sense of the polysemy word is linked with its related words by dividing these related words into verbs, nouns, adverbs and adjectives. Our algorithm does not count the overlap of words between the glosses of context and sense bags as in contextual overlap count knowledge-based word sense disambiguation algorithms. Instead, our algorithm searches the paths or links of context words with the senses of the target word. It keeps the track of each path or link that connects a context word and a sense of the target word. If the paths thus obtained connect only one sense of the target word, the algorithm output the linked sense as the correct sense of the target word for the given context. If there are paths that link more than one senses, then the algorithm counts the number of paths or links or connections for each linked sense. Then, the sense for which the number of connection paths is maximum is selected as a correct sense. The accuracy (96.11%)of our algorithm using PolyWordNet is found significantly higher than that of the accuracy (58.33%) of the other contextual overlap count Word Sense Disambiguation method that used the Princeton WordNet for sense disambiguation.
Abstract Gender neutral language has been one of the most hotly debated issues in Bible translation in recent decades, especially in translations into English. The article presents some aspects of this problem expanding the perspective and comparing gender neutral language usage in modern translations of Scripture into English and Polish: the New International Version and the Paulist Bible and the Poznan Bible, with occasional references to other English and Polish translations. Renditions of selected New Testament terms such as anthrōpos, anēr, adelphos/adelphoi and huioi are examined, as well as English and Polish translations of diakoneo when it describes women accompanying Jesus in the synoptic gospels. Translations of “Junia/Junius” (Rom 16:7) are also compared as well as the issue of Phoebe the “deaconess” in Rom 16:1. The author concludes that solutions concerning gender neutral language in English and Polish translations of the Bible, sometimes similar, are not identical due to differences between these languages, due to different socio-linguistic norms characterizing Polish and English audiences respectively and due to the fact that the English translation is addressed to the evangelical Christians, while the Polish ones to the Catholics.
The article deals with Finnish translations of varieties of spoken language in fiction from the late 19th century to the beginning of the 2000s. It presents the central findings of a comprehensive study on the changes and developments of translational norms in Finnish literature. The study is based on a corpus consisting of 200 literary works (the original and its translations are counted as one work), representing various genres: literary fiction, young-adult fiction, as well as genre fiction (romance and crime). During this 100-year period, the use of colloquial variants in translations has strongly increased, influenced by the changing literary and linguistic norms of original Finnish literature. The norms of different literary genres, however, vary, and rich, non-standard variation can be found in translated works from different periods.
Chinese word segmentation and Part-of-speech (POS) tagging have been studied for decades. However, most of the previous works mainly focus on pipeline method which will lead to error propagation. In order to make word segmentation and POS tagging jointly in one model, in this paper, we propose an effective neural network model to improve the accuracy of the segmentation and tagging. Our model works based on the hierarchical Long Short-Term Memory (LSTM) and trained jointly in one objective function. What's more, to better utilizing the transition features between tags, we further introduce the transition matrix which can help to search the best tagging sequence. Experiment on Chinese Treebank shows that our model achieves competitive accuracy on word segmentation and POS tagging.
We propose a method for improving the dependency parsing of complex sentences. This method assumes segmentation of input sentences into clauses and does not require to re-train a parser of one's choice. We represent a sentence clause structure using clause charts that provide a layer of embedding for each clause in the sentence. Then we formulate a parsing strategy as a two-stage process where (i) coordinated and subordinated clauses of the sentence are parsed separately with respect to the sentence clause chart and (ii) their dependency trees become subtrees of the final tree of the sentence. The object language is Czech and the parser used is a maximum spanning tree parser trained on the Prague Dependency Treebank. We have achieved an average 0.97% improvement in the unlabeled attachment score. Although the method has been designed for the dependency parsing of Czech, it is useful for other parsing techniques and languages.
During the first decade of this century, a new subculture emerged on the Runet (the Russian Internet). It promoted a version of the deliberately distorted Russian language, first of all by systematically altering orthographic and punctuation norms. This “padonkafskii” language (given the misspelled usage of the original word, it seems more accurate to translate it as “basterds,” as in in Tarantino’s movie) has been analyzed by linguists. Its political role and implicit message have not been scrutinized until now. In the article, Ilya Kukulin argues that this <i>basterd</i> language was used and promoted by pro-Kremlin political entrepreneurs to develop new strategies of communication: performances of cynical transgression and symbolic violence through the humiliation of counterparts. Cyberbullying became a distinctive style of seemingly nonconformist subculture, and already in this capacity was used by the regime. Central to this communication strategy was the cult of power, xenophobia, and discarding of any idealistic motivations of human behavior as hypocritical public relations technologies. At the same time, the <i>basterd</i> subculture itself claimed the status of nonconformist sincerity for its members. This image of a novel subculture helped to promote its aggressive xenophobia and support for the authorities among the most socially dynamic groups of Russia’s population. Political mobilization of the <i>basterds</i>’ language relied on the earlier aesthetic and political strategies of Russian mass culture of the 1990s, which were partially reconfigured or further developed. The radical and socially escapist irony underlying the cultural scene of the 1990s (the article uses Russian punk rock as an example) has been recast into the view of society as a total war, perceived from the position of total cynicism and nihilism. Cynicism has become the main mode of social thinking in modern Russia. Once popular with many bloggers, the <i>basterds</i>’ language is now out of vogue. Moreover, experiments with language transgressions seem to have lost their popularity. Instead, the role of transmitters of an ultra-cynical worldview and proponents of symbolic violence has been assumed by officialdom as represented by press secretaries of the president, key ministries, or MPs. They display the same conscious transgression of linguistic norm and moral standards as their <i>Basterd</i> predecessors, who pretended to be antiestablishment and nonconformist. The preponderance of <i>basterd</i> language during the previous decade cleared the ground for hate-speech and symbolic violence in the public sphere as an acceptable, attractive, and even necessary format of public communication.
108 Objectives To diagnose the dementia subtypes has significant information for determining the treatment strategies and predicting the clinical courses. Recently, many dementia patients undergo dopamine transporter (DAT) imaging, in addition to brain perfusion imaging to investigate the subtypes of dementia. However, patients must wait for tracer decay when using the conventional method, and this is burdensome for dementia patients and often delays the decision making for treatment. We developed a prototype CdTe SPECT system with 4-Pixel Matched Collimator for brain study. This system provides high energy resolution (6.6%), high sensitivity (220 cps/MBq/head) and provide high spatial resolution images of I-123 and Tc-99m simultaneously. The aim of this study was to evaluate findings and quantification on dual isotope study of cerebral blood flow (CBF) and Dopamine transporter (DAT) images with the new SPECT system. Methods We prospectively enrolled 21 patients with cognitive disorder. Every patient underwent imaging examinations on the same schedule. We used 3-head scanner (GCA-9300R, TOSHIBA) as a conventional scanner. Both 99mTc-ECD (CBF) and 123I-IFP (DAT) scans were independently performed with the conventional scanner on separate days, and simultaneously performed with the new system on another day. CBF and DAT images were both visually and quantitatively analyzed.For visual analyses, two nuclear medicine physicians visually interpreted both DAT and CBF images rating the severity into 4 grades (from 0 to 3). Totally 12 (6/hemisphere) supratentorial regions on the 99mTc-ECD images and 2 regions (left and right striatum) on the 123I-IFP were defined for each patient. The correlation of visual analysis results between GCA and SPICA was evaluated by weighted Kappa statistics. For quantitative analyses, we calculated the cerebrum-thalamus count ratio of 99mTc-ECD and the specific-to-background ratios of 123I-IFP. We assessed those regions of Pearson9s correlation coefficient and intraclass correlation coefficient (ICC) between the conventional scanner and the new system. Results The weighted Kappa statistics between GCA and SPICA were 0.67 and 0.68 for 99mTc-ECD and 123I-IFP, respectively, indicating high interrater reliability. The semi-quantitative analyses demonstrated that the 99mTc-ECD cerebrum-thalamus count ratio of SPICA was well correlated to GCA (R=0.81, p Conclusions The findings and quantitative results of both CBF and DAT images of dual isotope study with the new system had excellent correlation with those of single isotope studies with the conventional scanner with two separate days. In conclusion, our new SPECT scanner with semiconductor detectors enables quantitative dual tracer diagnostic imaging of CBF and Dopamine transporter imaging in patients with cognitive disorder. This technique may enable the one-stop imaging diagnosis for dementia and reduce the burden on patients.
Representation of syntactic structure is a core area of research in Computational Linguistics, disambiguating distinctions in meaning that are crucial for correct interpretation of language. Development of algorithms and statistical models over the past three decades has led to systems that are accurate enough to be deployed in industry, playing a key role in products such as Google Search and Apple Siri. However, syntactic parsers today are usually constrained to tree representations of language, and performance is interpreted through a single metric that conveys no linguistic information regarding remaining errors.In this dissertation, we present new algorithms for error analysis and parsing. The heart of our approach to error analysis is the use of structural transformations to identify more meaningful classes of errors, and to enable comparisons across formalisms. For parsing, we combine a novel dynamic program with careful choices in syntactic representation to create an efficient parser that produces graph structured output. Together, these developments allowed us to evaluate the outstanding challenges in parsing and to address a key weakness in current work.First, we present a search algorithm that, given two structures, finds a sequence of modifications leading from one structure to the other. We applied this algorithm to syntactic error analysis, where one structure is the output of a parser, the other is the correct parse, and each modification corresponds to fixing one error. We constructed a tool based on the algorithm and analyzed variations in behavior between parsers, types of text, and languages. Our observations shine light on several assumptions about syntactic errors, showing some to be true and others to be false. For example, prepositional phrase attachment errors are indeed a major issue, while coordination scope errors do not hurt performance as much as expected.Next, we describe an algorithm that builds a parse in one syntactic representation to match a parse in another representation. Specifically, we build phrase structure parses from Combinatory Categorial Grammar derivations. Our approach follows the philosophy of CCG, defining specific phrase structures for each lexical category and generic rules for combinatory steps. The new parse is built by following the CCG derivation bottom-up, gradually building the corresponding phrase structure parse. This produced significantly more accurate parses than past work, and enabled us to compare performance of several parsers across formalisms.Finally, we address a weakness we observed in phrase structure parsers: the exclusion of syntactic trace structures for computational convenience. We present an efficient dynamic programming algorithm that constructs the graph structure that has the highest score under an edge-factored scoring function. We define a parse representation compatible with the algorithm, and show how certain linguistic distinctions dramatically impact coverage. We also show various ways to modify the algorithm to improve performance by exploiting properties of observed linguistic structure. This approach to syntactic parsing is the first to cover virtually all structure encoded in the Penn Treebank.
The rapid accumulation of data in social media (in million and billion scales) has imposed great challenges in information extraction, knowledge discovery, and data mining, and texts bearing sentiment and opinions are one of the major categories of user generated data in social media. Sentiment analysis is the main technology to quickly capture what people think from these text data, and is a research direction with immediate practical value in ‘big data’ era. Learning such techniques will allow data miners to perform advanced mining tasks considering real sentiment and opinions expressed by users in additional to the statistics calculated from the physical actions (such as viewing or purchasing records) user perform, which facilitates the development of real-world applications. However, the situation that most tools are limited to the English language might stop academic or industrial people from doing research or products which cover a wider scope of data, retrieving information from people who speak different languages, or developing applications for worldwide users. More specifically, sentiment analysis determines the polarities and strength of the sentiment-bearing expressions, and it has been an important and attractive research area. In the past decade, resources and tools have been developed for sentiment analysis in order to provide subsequent vital applications, such as product reviews, reputation management, call center robots, automatic public survey, etc. However, most of these resources are for the English language. Being the key to the understanding of business and government issues, sentiment analysis resources and tools are required for other major languages, e.g., Chinese. In this tutorial, audience can learn the skills for retrieving sentiment from texts in another major language, Chinese, to overcome this obstacle. The goal of this tutorial is to introduce the proposed sentiment analysis technologies and datasets in the literature, and give the audience the opportunities to use resources and tools to process Chinese texts from the very basic preprocessing, i.e., word segmentation and part of speech tagging, to sentiment analysis, i.e., applying sentiment dictionaries and obtaining sentiment scores, through step-by-step instructions and a hand-on practice. The basic processing tools are from CKIP Participants can download these resources, use them and solve the problems they encounter in this tutorial. This tutorial will begin from some background knowledge of sentiment analysis, such as how sentiment are categorized, where to find available corpora and which models are commonly applied, especially for the Chinese language. Then a set of basic Chinese text processing tools for word segmentation, tagging and parsing will be introduced for the preparation of mining sentiment and opinions. After bringing the idea of how to pre-process the Chinese language to the audience, I will describe our work on compositional Chinese sentiment analysis from words to sentences, and an application on social media text (Facebook) as an example. All our involved and recently developed related resources, including Chinese Morphological Dataset, Augmented NTU Sentiment Dictionary (aug-NTUSD), E-hownet with sentiment information, Chinese Opinion Treebank, and the CopeOpi Sentiment Scorer, will also be introduced and distributed in this tutorial. The tutorial will end by a hands-on session of how to use these materials and tools to process Chinese sentiment. Content Details, Materials, and Program please refer to the tutorial URL: http://www.lunweiku.com/
My research focuses on the study of grammatical change in the recent history of the English language; in particular, I am currently writing my PhD dissertation on from Early Modern English to Present-Day English. In this PhD project, I analyse and compare the different factors that appear to influence in Present-Day English with earlier stages of the language (Late Modern English).The concept of ellipsis refers to a syntactic strategy in which expected elements have been left unpronounced in certain constructions. This omission triggers a mismatch between meaning (the intended message) and sound (what is in fact uttered). In particular, my research focuses on those examples of (Miller 2011, Miller and Pullum 2013), i.e. ellipsis types that occur after the following licensors (that is, those elements that license ellipsis): modal verbs, auxiliaries be, have and do, infinitival marker to and negator not. The main aim is to carry out an empirical analysis of from Late Modern English to Present-Day English (1700-1914), both quantitatively and qualitatively, by means of data retrieved from the Penn Corpora of Historical English. This project pays attention to syntactic variation, genre distribution and discourse variables (type of anaphora, mismatches in polarity, aspect, voice, modality, tense; comparison of clause types; distance, linking, type of focus).Esta tesis doctoral versa sobre la variacion diacronica de la elipsis sintactica desde ingles moderno temprano hasta la actualidad. La elipsis representa un desajuste entre el significado de lo que se dice (la intencion de un mensaje) y lo que se pronuncia en realidad. En el ambito de la linguistica moderna, la elipsis se estudia en el campo la semantica, la sintaxis, la pragmatica, la psicolinguistica, la linguistica de corpus, etc. y constituye la novedad de esta investigacion el tratamiento empirico de este fenomeno linguistico tratando de juntar variables procedentes de las distintas teorias consultadas. Cabe destacar que no existe un estudio multidisciplinar sobre este concepto sintactico en ingles moderno. Existen solo unos pocos estudios relativamente recientes pero se centran unicamente en el estudio del ingles actual. Por esta razon, se decidio llevar a cabo la investigacion tanto de los aspectos formales como de los funcionales de la elipsis y su evolucion diacronica teniendo en cuenta distintas variables discursivas. Para eso, se creo una base de datos en la que apareceria el ejemplo de Post-Auxiliary Ellipsis (elipsis despues de un auxiliar) con su numero identificador; el genero al que pertenece (de los dieciocho diferentes que existen en el corpus utilizado, el Penn Treebank); el licensor (aquel elemento que posibilita la elipsis); el tipo de union entre el antecedente de la elipse y la clausula eliptica (coordinacion, subordinacion, parataxis, etc.); el tipo de conector entre el antecedente y la parte la distancia existente entre el antecedente y la clausula eliptica (numero de clausulas); contexto sintactico en el que aparece la elipsis (clausulas principales, subordinadas, question-tags, etc); tipo de anafora (anaforico, cataforico, exoforico); categoria del antecedente (sintagma verbal, nominal, adjetival o no constituyente); categoria del material elidido (sintagma verbal, nominal, adjetival o no constituyente); presencia o ausencia de cambio de referente; comparacion del aspecto, la voz, la modalidad y el tiempo del antecedente con respecto a la parte presencia o ausencia de question-tag; comparacion entre el tipo de clausula del antecedente (declarativa, interrogativa o imperativa) y el de la clausula elidida; y tipo de foco de la parte elidida (auxiliary-choice (eleccion de auxiliar), subject-choice (eleccion de sujeto) o ambos).Esta tese de doutoramento versa sobre a variacion diacronica da elipse sintactica dende ingles moderno temperan ata a actualidade. A elipse representa un desaxuste entre o significado do que se di (a intencion dunha mensaxe) e o que se pronuncia en realidade. No ambito da linguistica moderna, a elipse estudase no campo a semantica, a sintaxe, a pragmatica, a psicolinguistica, a linguistica de corpus, etc. e constitue a novidade desta investigacion o tratamento empirico deste fenomeno linguistico tratando de xuntar variables procedentes das distintas teorias consultadas. Cabe destacar que non existe un estudo multidisciplinar sobre este concepto sintactico en ingles moderno. Existen so uns poucos estudos relativamente recentes pero centranse unicamente no estudo do ingles actual. Por esta razon, decidiuse levar a cabo a investigacion tanto dos aspectos formais coma dos funcionais da elipse e a sua evolucion diacronica tendo en conta distintas variables discursivas. Para iso, creouse unha base de datos na que apareceria o exemplo de Post-Auxiliary Ellipsis (elipse despois dun auxiliar) co seu numero identificador; o xenero ao que pertence (dos dezaoito diferentes que existen no corpus utilizado, o Penn Treebank); o licensor (aquel elemento que posibilita a elipse); o tipo de union entre o antecedente da elipse e a clausula eliptica (coordinacion, subordinacion, parataxe, etc.); o tipo de conector entre o antecedente e a parte a distancia existente entre o antecedente e a clausula eliptica (numero de clausulas); contexto sintactico no que aparece a elipse (clausulas principais, subordinadas, question-tags, etc); tipo de anafora (anaforico, cataforico, exoforico); categoria do antecedente (sintagma verbal, nominal, adxectival ou non constituinte); categoria do material elidido (sintagma verbal, nominal, adxectival ou non constituinte); presenza ou ausencia de cambio de referente; comparacion do aspecto, a voz, a modalidade e o tempo do antecedente con respecto a parte presenza ou ausencia de question-tag; comparacion entre o tipo de clausula do antecedente (declarativa, interrogativa ou imperativa) e o da clausula elidida; e tipo de foco da parte elidida (auxiliary-choice (eleccion de auxiliar), subject-choice (eleccion de suxeito) ou ambos).
Despite the proliferation of corpus-based studies focusing on the complex category of discourse markers in the recent years, consensus is yet to be found regarding the most reliable yet informative model to describe their behavior in authentic data. The major and most widely spread frameworks include the Penn Discourse TreeBank (Prasad et al. 2008), Rhetorical Structure Theory (Mann & Thompson 1988), Segmented Discourse Representation Theory (Asher & Lascarides 2003) and the Cognitive approach to Coherence Relations (Sanders et al. 1992). These models disagree both on the top levels (number and type of generic annotation levels, if any) and the specific relations included in them. It is precisely the relation between top levels and corresponding sublevels that will be discussed in this paper, starting from a recent proposal of functional taxonomy applied to the French-English spoken corpus DisFrEn (Crible, in press) and its revision in the framework of the LOCAS-F corpus (Degand, Martin & Simon 2014). In DisFrEn, four top-level functions or "domains" are distinguished, from the revision of existing proposals for both speech and writing (Cuenca 2013, González 2005, Halliday & Hasan 1976, Zufferey & Degand in press): ideational (objective relations), rhetorical (subjective, metadiscursive functions), sequential (structuring functions) and interpersonal (speaker-hearer relationship). In the original model, these four domains include a total of thirty functions, each of them belonging to one – and only one – domain. For instance, the function labeled "CAUSE" is always ideational, while "MOTIVATION" is always rhetorical, etc. This interdependent system is intended to maximize the informativity of the annotation labels (one label for one function, vs. combinations of labels) while giving the opportunity to filter the distribution from thirty to four values, hence more efficient for quantitative purposes. The annotation procedure doesn't specify which decision should be made first, the domain or the function, and it is assumed that both orders are possible although not equally relevant depending on the specific function at stake and/or the research question, annotator's expertise, etc. (see Crible & Degand 2015). An on-going project (Degand & Simon 2015) is currently working on a revision of the DisFrEn taxonomy aiming at reducing the number of options and enhancing the reliability and cognitive validity of the model. The main difference between this revision and the original is that domains and functions are no longer interdependent. On the contrary, it is assumed that most – if not all – functions can be assigned more than one domain, roughly following Sweetser’s (1990) discourse domains. For instance, a [CONTRAST] can be ideational (1), rhetorical (2), and even sequential (3) or interpersonal (4), a proposal which still needs careful investigation and discussion. (1) I wasn’t looking forward to doing it but I am now (DisFrEn EN-phon-01) (2) a rebate is when they send the money back // yes but how do you define it in economic terms (DisFrEn EN-clas-02) (3) (after a digression on the industries in Bristol area) but Bristol itself is a large metropolis (DisFrEn EN-intf-05) (4) I think the Marks is better // actually I’m not sure it is (DisFrEn EN-conv-01) In this presentation, we will address several methodological implications of the interdependence vs. independence of domains and functions, paying particular attention to issues of inter-rater reliability and overall validity and consistency of the taxonomy. In addition, quantitative results will be presented to illustrate the comparison of the two versions on samples of spoken French from the LOCAS-F corpus annotated with both systems.
Two fundamental components of causality are the Cause and the Result. In linguistic work the distinction between these aspects is commonly blurred, presumably because the primary research focus has been on describing how language encodes causality. The semantic nature of the component events and the constraints on their relationship are seldom discussed; however, the current work aims to shed light on a broader spectrum of features that underlie the concept. This is an essential foundation for understanding how language communicates Result. The present discussion explores and illuminates the nature of this concept focusing on a relatively open-ended set of linguistic elements that can play a role in shaping a discourse relation in addition to discourse connectives. This is in contrast to the majority of the previous research, which has been quite intensely concerned with investigating a limited collection of well-established causality markers. Also, despite the fact that English has been used in studies on causality both as a control language and a metalanguage, there is surprisingly little work on the semantics of the relations that occur specifically in English, let alone Result relations. By borrowing from several cognitively-oriented approaches and combining empirical data from two written corpora (British National Corpus and the Penn Discourse Treebank) with experimental work, the current study systematically investigates the conceptual and linguistic properties of several closely related Result relation types (including Purpose), along with the joint role of discourse connectives and other discourse elements in conveying the intended sense. The findings indicate that linguistic signals of the conceptual structure of the relation seem to play a more significant role in the interpretation than explicit marking. Two factors emerged as more vital cues than the presence of the ambiguous connective so. In Purpose relations, a modal auxiliary conveying an intended effect, and in Result relations the presence/absence of an intentionally acting actor are crucial for disambiguation. The multifunctional connective therefore seems to merely satisfy the mandatory marking requirement related to the intrinsically unrealized (‘nonveridical’) nature of Purpose. In Result the presence of an ambiguous marker is to a great extent optional in English. However, discourse markers can also reflect how language users categorize causal event types. This claim has been confirmed in several cross-linguistic analyses, but the lexicon of English connectives has not been systematically investigated from this vantage point. The few existing studies found that the uses of English connectives are quite unconstrained across causal categories. The present work contributes to this line of research and suggests that two unambiguous markers, as a result and for this reason, indeed cover a wide range of causal event types; however, they also exhibit significant tendencies to occur prototypically in certain relation types. The presence and role of an intentionally acting discourse participant behind both real-world and linguistic causally-related events contributes to these tendencies. The contexts that include such a participant are regarded as intrinsically subjective and have been found to manifest surface expressions of subjectivity in previous work on other languages. The current study confirms similar tendencies in the linguistic construal and marking of Result relations in English, which proves that certain language elements partake in establishing the intended interpretation on a par with discourse connectives. What emerges as a result of this discussion, is therefore an account on how English utilizes the broad category of Result and what linguistic elements are used to convey the array of resultative events.
El acto de destrucción de la lengua que lleva cabo uno de los sujetos de la escritura de Altazor abre el espacio y el tiempo para el surgimiento de nuevos significantes. Para este sujeto, el uso sistemático de la lengua -cumpliendo la norma lingüística- impide la representación (aparición de imágenes) y referencia de correlatos radicalmente nuevos. Este sujeto es discontinuo y coexiste con otros, en especial, con un sujeto voluntarioso, que sigue aferrándose a la tradicional concepción del mundo, fundada en la trascendencia divina. Las operaciones sobre la lengua de este poeta altazoriano se realizan, sobre todo, como utilización paródica de ella, como demolición intencional de la lengua, como transformación de las ruinas idiomáticas en significantes, como uso alegórico de la lengua (en el sentido sugerido por W. Benjamin). Este sujeto altazoriano propone orientar y sostener sentimentalmente la constitución de nuevas imágenes (en el sentido de los universales fantásticos de Vico). La poesía se hace, así, acontecimiento (Ereignis, según lo nombra Heidegger), actividad cuya plenitud se produce en su consumación y consunción, en la mostración de su temporalidad fundante. Pero la energía del poeta no alcanza para la continuidad del acontecimiento poético -en el caso de que sea posible-, para la urgencia y necesidad de su aparición. The act oflanguage destruction carried out by one ofthe subjects in the writing of "Altazor", opens both space and time to the birth ofnew signifiers. For such subject, the systematic use oflanguage sticking to the linguistic norm, hinders the representation (the creation o.fimages and the reference to radically new co-relatives. Such subject is discontinuous and coexists with others, in particular, with a willful subject that keeps its allegiance to the traditional world view, based on divine transcendence. The language operations ofthis altazorean poet are effected, above all, as parody; as transformation ofidiomatic debris into signifiers; and as an allegorical use of language (in the sense suggested by Walter Benjamín). The altarzorean subject propases to orientate and sustain sentimentally the creation of new images (in the sense of the "fantastic universals" of Vico). In this way, poetry becomes 'event' ("Ereignis", as named by Heidegger) an activity whosefullnes is achieved through its realization and consumption, in the exhibition of its founding tenporality.
INTRODUCTIONThe Strategic Integrated Management Seminar (SIMS) course is mandatory for every senior student in the school of business at a mid-size private university in the northeastern United States. The course allows students to integrate their accumulated knowledge and apply this knowledge to issues from a strategic perspective. It examines a firm from the position of top level management, focusing on the role of the general manager in formulating and implementing corporate and business level strategy. Strategic issues of an entire athletic (hereon, footwear company or company) and industry are analyzed. Students are expected to draw their accumulated knowledge of the functional areas of their majors into a homogenous team effort. Each individual student works on developing his/her ability to analyze information, draw logical conclusions, and offer sound supporting evidence for their arguments in written form and classroom discussions. The course is highly interactive with students taking the lead and the professors sharing knowledge and offering supplementary support.The SIMS course uses the Business Strategy Game (BSG) simulation to enable students to experience a top management team perspective in running a and experiencing competitive conditions in the athletic industry. In the BSG, students compete in teams (each team constitutes a company, hereon team/company will be synonymous) within a global arena that encompasses four regions - Europe-Africa, North America, Asia-Pacific, and Latin America (The Business Strategy Game, 2016). They compete against teams in their individual classes and compare/contrast data with teams/companies worldwide. Each competes head-to-head against companies run by other teams in the course, hence competition plays an important role in the experience. Each sells its brand of to retailers worldwide and to individuals buying online at the company's website.Competing in the BSG requires a series of complex decisions by the students, taking into account the team's strategy for their and the competitive conditions in the industry and the strategies of their competitors. The simulation allows for numerous decisions for each round, requiring students to choose which decisions are most important to implement their strategy and which areas of the business must receive attention in order for their firm to be its most competitive. Beyond overall strategy (corporate, competitive) are several key functional areas for decision making. Decision areas in operations include capacity planning (either adding to existing plants or building new plants in new geographic locations), production quality decisions for the athletic footwear, plant operations efficiency, and labor decisions. Footwear must be shipped to distribution centers around the world and students must choose where it is best to manufacture the and where to ship taking into consideration demand, shipping costs, tariffs, and exchange rates. Marketing decisions include pricing the product in a wholesale and a retail environment, advertising and use or non-use of celebrity endorsements, rebates, and incentives to retailers. Financial decisions include funding the capital structure of the firm using debt, equity, and/or cash. Dividend payouts and stock repurchases may be used by the companies.The simulation has students take control of an athletic that has been in operation for ten years. Teams make in total eight years of decisions (years 11 - 18), approximately one per week. Each decision rollover represents one year and includes many decisions within the decision. The simulation evaluates team performance based on five investor expectation performance targets: Earnings Per Share (EPS), Return on Equity (ROE), credit rating, image rating (a combination of market share and shoe quality), and stock price. Each measure of performance is equally weighted at 20% of the total score (The Business Strategy Game, 2016). …
This paper presents a novel high-order dependency parsing framework that targets non-projective treebanks. It imitates how a human parses sentences in an intuitive way. At every step of the parse, it determines which word is the easiest to process among all the remaining words, identifies its head word and then folds it under the head word. Further, this work is flexible enough to be augmented with other parsing techniques.
The aim of the current study was to investigate if Openness – to – Experience and Neuroticism personality traits are associated with curiosity. This will help us to estimate whether knowledge expansion is dependent on a person’s personality and which trait is more willing to invest time on learning. The experiment consisted of two different sessions. To estimate curiosity, 40 subjects first performed a word-synonymy task, where Shannon’s (1948) entropy was estimated and the result of which lead to the measurement of uncertainty. Then in a second session, participants had the option to request for feedback between a few alternative options at a cost (time), and they were also required to estimate their satisfaction about the answer on a valence rating scale. Finally, participants were screened for personality traits. Neurotic individuals appeared to be more willing in investing time on feedback request, in contrast to open individuals.
This paper aims at filling the gap between the accuracy of Italian and English constituency parsing: firstly, we adapt the Bllip parser, i.e., the most accurate constituency parser for English, also known as Charniak parser, for Italian and trained it on the Turin University Treebank (TUT). Secondly, we design a parse reranker based on Support Vector Machines using tree kernels, where the latter can effectively generalize syntactic patterns, requiring little training data for training the model. We show that our approach outperforms the state of the art achieved by the Berkeley parser, improving it from 84.54 to 86.81 in labeled F1.
espanolLa actitud que debemos tomar ante la norma linguistica no puede ser la misma en todos los ambitos profesionales. Un logopeda, profesional de la rehabilitacion del lenguaje, el habla y la voz alterados, debe tener una actitud ante la norma diferente a la que debe adoptar un maestro. En este trabajo hemos tomado como instrumento de investigacion una encuesta que plantea a los logopedas en activo una serie de cuestiones acerca de que norma linguistica toman como referencia en su labor rehabilitadora. Para ello partimos de las nociones de sistema linguistico frente a norma, con una metodologia basada en las encuestas de opinion seleccionadas con el objetivo de dilucidar la actitud que estos adoptan en su proceso de intervencion logopedica. Hemos recogido encuestas de logopedas de varias comunidades linguisticas que no comparten el mismo modelo de lengua para observar que variante de lengua toman como base en sus intervenciones. EnglishThe attitude that we should have towards linguistic norm should not be the same in all professional fields. A speech and language therapist (language rehabilitation, speech and altered voice professional), must have an attitude towards different norm which a teacher must adopt. In this study, we have used a survey as a research instrument to ask active speech and language therapists a series of questions on what linguistic norms do they use as reference in their rehabilitative work. Thus, we focus our study on linguistic system notions against standard norm, with a methodology based on 2 selected opinion surveys to elucidate the adopted attitude during speech therapy intervention. Consequently, we have collected speech therapists surveys from various linguistic communities that do not share the same language model to observe which language variant do they use as a basis of their interventions.
A morphological tagger is a computer program that provides complete morphological descriptions of sentences. Morphological taggers find applications in many NLP fields. For example, they can be used as a pre-processing step for syntactic parsers, in information retrieval and machine translation. The task of morphological tagging is closely related to POS tagging but morphological taggers provide more fine-grained morphological information than POS taggers. Therefore, they are often applied to morphologically complex languages, which extensively utilize inflection, derivation and compounding for encoding structural and semantic information. This thesis presents work on data-driven morphological tagging for Finnish and other morphologically complex languages. \n\nThere exists a very limited amount of previous work on data-driven morphological tagging for Finnish because of the lack of freely available manually prepared morphologically tagged corpora. The work presented in this thesis is made possible by the recently published Finnish dependency treebanks FinnTreeBank and Turku Dependency Treebank. Additionally, the Finnish open-source morphological analyzer OMorFi is extensively utilized in the experiments presented in the thesis. \n\nThe thesis presents methods for improving tagging accuracy, estimation speed and tagging speed in presence of large structured morphological label sets that are typical for morphologically complex languages. More specifically, it presents a novel formulation of generative morphological taggers using weighted finite-state machines and applies finite-state taggers to context sensitive spelling correction of Finnish. The thesis also explores discriminative morphological tagging. It presents structured sub-label dependencies that can be used for improving tagging accuracy. Additionally, the thesis presents a cascaded variant of the averaged perceptron tagger. In presence of large label sets, a cascaded design results in substantial reduction of estimation speed compared to a standard perceptron tagger. Moreover, the thesis explores pruning strategies for perceptron taggers. Finally, the thesis presents the FinnPos toolkit for morphological tagging. FinnPos is an open-source state-of-the-art averaged perceptron tagger implemented by the author.