Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
Creating an opinionated lexicon is an important step towards a reliable social media analysis system. In this article we are proposing an approach and describing an experiment to build an Arabic polarised lexical database from analysing online implicitly and explicitly rated customer reviews. These reviews are written in modern standard Arabic and Palestinian/Jordanian dialect. Therefore, the produced lexicon contains casual slangs and dialectic entries used by the online community, which is useful for sentiment analysis of informal social media micro-blogs. We have extracted 28,000 entries from processing 15,100 reviews and by expanding the initial lexicon through Google translate. We calculated an implicit rating for every review driven by its text to address the problem of ambiguous opinions of certain online posts, where the text of the review does not match the given rating (the explicit rating). Each entry was given a polarity tag and a confidence score. High confidence scores have increased the precision of the polarisation process. Explicit rating has increased the coverage and confidence of polarity.
Death is a vague, frightening and abstract concept, mostly considered as a taboo. The current study investigates the ways of conceptualizing death among Persian language speakers in a representative corpus of Persian texts. It seeks answer to the question: What kind of linguistic tools are used by Persian language speakers in order to conceptualize the phenomenon of death and what cultural elements are involved in this regard.To answer the research question, we used the Persian Linguistic Database (PLDB) and a selection of obituaries and epitaphs as data and attempted to identify and extract all the expressions which were directly or indirectly related to the concept of death. Using the conceptual metaphor theory (CMT), by Lakoff and Johnson (1987), and cultural conceptualization by Sharifian (2011), the ways death is conceptualized were identified on the basis of a cognitive-cultural approach.
Treebanks traditionally treat punctuation marks as ordinary words, but linguists have suggested that a tree’s “true” punctuation marks are not observed (Nunberg, 1990). These latent “underlying” marks serve to delimit or separate constituents in the syntax tree. When the tree’s yield is rendered as a written sentence, a string rewriting mechanism transduces the underlying marks into “surface” marks, which are part of the observed (surface) string but should not be regarded as part of the tree. We formalize this idea in a generative model of punctuation that admits efficient dynamic programming. We train it without observing the underlying marks, by locally maximizing the incomplete data likelihood (similarly to the EM algorithm). When we use the trained model to reconstruct the tree’s underlying punctuation, the results appear plausible across 5 languages, and in particular are consistent with Nunberg’s analysis of English. We show that our generative model can be used to beat baselines on punctuation restoration. Also, our reconstruction of a sentence’s underlying punctuation lets us appropriately render the surface punctuation (via our trained underlying-to-surface mechanism) when we syntactically transform the sentence.
Abstract: Temperatures above 20° Celsius have shown to adversely impact human behavior, leading to increased aggression and violence. Climate change will contribute to both the magnitude and severity of this pattern as temperatures continue their rise. Contributions to this field of research have only recently begun to analyze online behavior and language as a proxy for hedonic state, or well-being. From a development perspective this study is relevant since the poor tend to live in some of the warmest regions on earth, and would thus be disproportionately impacted by increased temperatures. We use several sources of data; U.S. based daily statewide temperature data from 2016 through 2017, as well as localized viewer chat data from a live video streaming website. We will sort chatting comments looking for key words (i.e. hate speech, swearing, etc.), and with the use of a word rating system we then assess the overall mood of the chatters contingent on high temperature readings on the precise day of the communications. After controlling for spatiotemporal fixed effects, we find strong evidence that hedonic state decreases above 20°c.
Abstract The topic of this paper is the interaction of aspectual verb coding, information content and lengths of verbs, as generally stated in Shannon’s source coding theorem on the interaction between the coding and length of a message. We hypothesize that, based on this interaction, lengths of aspectual verb forms can be predicted from both their aspectual coding and their information. The point of departure is the assumption that each verb has a default aspectual value and that this value can be estimated based on frequency – which has, according to Zipf’s law, a negative correlation with length. Employing a linear mixed-effects model fitted with a random effect for LEMMA, effects of the predictors’ DEFAULT – i.e. the default aspect value of verbs, the Zipfian predictor FREQUENCY and the entropy-based predictor AVERAGE INFORMATION CONTENT – are compared with average aspectual verb form lengths. Data resources are 18 UD treebanks. Significantly differing impacts of the predictors on verb lengths across our test set of languages have come to light and, in addition, the hypothesis of coding asymmetry does not turn out to be true for all languages in focus.
Antiaddictive social advertising is a special speech genre of modern communication with specific features determined by the target setting and the chosen strategy. Advertising can be considered as a special functional style, within which separate genres are distinguished, first of all commercial advertising and social advertising. Social advertising, which refers to ethical categories, has become especially popular because society is faced with such problems, the solution of which depends on mass behavior. Advertising has become an integral part of the daily life of a person. Since advertising is mass-replicated in the media, it can enter the consciousness of the addressee, even against his/her will and desire, no wonder advertising is sometimes defined as the “fifth power” after media, whose power is considered the “fourth power”. Therefore, it is so important for advertising to follow the moral principles of society, it is so important to comply with modern ethical and linguistic norms. If commercial advertising is widespread, the antiaddictive one is little known. Many specialists in advertising argue for the need of the strategy of “shock” advertising, a remarkable feature of which is hyperbolization, used to cause fear. In antiaddictive advertising there is often a morphological imperative in the meaning of categorical motivation. It is justified as such advertizing has to be as much as possible appellate, has to be understood unambiguously. Anti-drug social advertising should show positive, motivation to a healthy lifestyle.
In two experiments, the influence of inducing negative mood on cognitive performance was explored by analyzing physical arm reaching movements as indicators of mind wandering. Mood was induced by viewing a series of six photos per mood condition that were previously established for their emotionally valenced and arousal ratings. A reach tracking device recorded three metrics of arm movement that were expected to reflect instances of mind wandering: initiation latency, movement time, and arm curvature. In the first experiment, 29 participants were randomly assigned into one of two induced-mood groups, negative mood (n = 15) or neutral mood (n = 14). Participants performed a simple Go/No-go task in which arm movements were detected by the reach tracker. The first experiment indicated that the mood inducement was successful but the effect of negative mood on either self-reported mind wandering or variances in arm movement were not significant. Thus, the second experiment prompted the change to a visual-search target-selection task in which variances in initiation latency, movement time, and curvature were expected to be more pronounced. The second experiment consisted of 23 participants who were also randomly assigned to either negative (n = 12) or neutral (n = 11) mood condition. The second experiment revealed that the mood induction was still successful but that there were still no significant effects observed between mood and indicators of mind wandering. Though the results of this study did not reflect initial predictions, it may suggest that low-arousing negative moods in healthy individuals are not associated with increased mind wandering.
Abstract This article explores the continuing linguistic impact of the Mandarin Union Version by investigating and contrasting two Chinese translations of William Paul Young's global bestseller The Shack (2007): the Traditional Chinese version Xiaowu (《小屋》, 2009) and the Simplified Chinese version Pengwu (《棚屋》, 2010). Ever since its publication, the Mandarin Union Version has served as the predominant Bible within Mandarin-speaking Protestant communities across the world. This has brought about the standardisation of terminology in Chinese Protestantism. The Shack, though widely marked as a Christian novel, is also known for its unconventional fictional representations of Christianity that some Christians think depart from orthodoxy. Both Xiaowu and Pengwu were published by non-Christian publishing houses for a general readership. However, Xiaowu, translated by a Christian, exhibits a significant number of phrases that specifically belong to Chinese Christian terminology shaped by the Mandarin Union Version. Pengwu is a contrast in this regard. By comparing extracts from these two Chinese versions, this article highlights how far the Mandarin Union Version has contributed to the formation of the linguistic repertoire of Mandarin-speaking Christian translators as well as linguistic norms for translated Christian-themed texts into Chinese.
This study investigated the associations of imageability with fear reactivity. Imageability ratings of four word classes: positive and negative (i) emotional and (ii) propriosensitive, neutral and negative (iii) theoretical and (iv) neutral concrete filler, and fear reactivity scores – degree of fearfulness towards different situations (TF score) and total number of extreme fears and phobias (EF score), were obtained from 171 participants. Correlations between imageability, TF and EF scores were tested to analyze how word categories and their valence were associated with fear reactivity. Imageability ratings were submitted to recursive partitioning. Participants with high TF and EF scores had higher imageability for negative emotional and negative theoretical words. The correlations between imageability of negative emotional words and negative theoretical words for EF score were significant. Males showed stronger correlations for imageability of negative emotional words for EF and TF scores. High imageability for positive emotional words was associated with lower fear reactivity in females. These findings were discussed with regard to negative attentional bias theory of anxiety, influence on emotional systems, and gender-specific coping styles. This study provides insight into cognitive functions involved in mental imagery, semantic competence for mental imagery in relation to fear reactivity, and a potential psycholinguistic instrument assessing fear tendency.
We introduce a novel transition system for discontinuous constituency\nparsing. Instead of storing subtrees in a stack --i.e. a data structure with\nlinear-time sequential access-- the proposed system uses a set of parsing\nitems, with constant-time random access. This change makes it possible to\nconstruct any discontinuous constituency tree in exactly $4n - 2$ transitions\nfor a sentence of length $n$. At each parsing step, the parser considers every\nitem in the set to be combined with a focus item and to construct a new\nconstituent in a bottom-up fashion. The parsing strategy is based on the\nassumption that most syntactic structures can be parsed incrementally and that\nthe set --the memory of the parser-- remains reasonably small on average.\nMoreover, we introduce a provably correct dynamic oracle for the new transition\nsystem, and present the first experiments in discontinuous constituency parsing\nusing a dynamic oracle. Our parser obtains state-of-the-art results on three\nEnglish and German discontinuous treebanks.\n
Several studies have reported specific semantic dissociations in brain damaged patients. These semantic dissociations concern particularly the impairment of one domain of knowledge (living concepts) whilst the other is relatively spared (nonliving concepts). The origin of these dissociations remain controversial. A important number of semantic theories have advanced different explanations for this phenomenon. Among others we include: the sensory-functional theory (Warrington & McCarthy, 1983), the organized content hypothesis, the domain specific hypothesis (Caramazza & Shelton, 1998) and the conceptual structure account (Moss, Tyler, & Devlin, 2002). However, at first time some authors have suggested that category-specific semantic deficit for living things may be due to a lower familiarity and frequency and a more visual complexity of these concepts compared with nonliving concepts (Funnell & Sheridan, 1992). In addition, recently researches suggest that familiarity varies according to the gender and therefore it may to influence the emergence of category-specific semantic deficits. Thus, Laiacona, Barbarotto and Capitani (1998) showed that females with Alzheimer's disease were more impaired with nonliving categories, whereas males patients showed more deficit with living categories. Likewise, studies from normal subjects have shown a gender-category interaction in a semantic fluency task, females performed better with fruits and males with tools (Capitani, Laiacona, & Barbarotto, 1999). The goals of this study are a) to evaluate the familiarity of concepts according to domains (living and nonliving) and categories and b) to explore the gender effect on familiarity ratings.
ABSTRACT Issues surrounding English for Academic Purposes (EAP) and its use by English as an additional language (EAL) students in higher education have become increasingly significant in recent years, fueled both by increased international student mobility and increased linguistic and cultural diversity within and outside of the student body. As well as posing language-related challenges, the transfer of EAL students to an English-speaking foreign university also demands the negotiation of new university expectations, channeled through a new cultural environment. While Academic Literacies research has identified that concepts such as power, identity, and culture play a role in academic writing, students’ own perceptions remain relatively unexplored. Consequently, this study analyzes the ways in which EAL students articulate their relationship with academic writing at a tertiary institution in Ireland. Data for this study were gathered through questionnaires and interviews and analyzed through discourse analysis through a critical lens. The findings suggest that while participants generally positively reflect on their ability to negotiate academic writing through the English language, there is nonetheless a high level of conflict between dominant linguistic norms and the students’ expression of their identity and culture.
Mirror-sensory synesthetes mirror the pain or touch that they observe in other people on their own bodies. This type of synesthesia has been associated with enhanced empathy. We investigated whether the enhanced empathy of people with mirror-sensory synesthesia influences experience of situations involving touch or pain, and whether it affects their prosocial decision making. Mirror-sensory synesthetes (N=18, all female), verified with a touch-interference paradigm, were compared to a similar number of age-matched control individuals (all female). Participants viewed arousing images depicting pain or touch; we recorded subjective valence and arousal ratings, and physiological responses, hypothesizing more extreme reactions in synesthetes. The subjective impact of positive and negative images was stronger in synesthetes than in control participants; the stronger the reported synesthesia, the more extreme the picture ratings. However, there was no evidence for differential physiological or hormonal responses to arousing pictures. Prosocial decision making was assessed with an economic game assessing altruism, in which participants had to divide money between themselves and a second player. Mirror-sensory synesthetes donated more money than non-synesthetes, showing enhanced prosocial behaviour, and also scored higher on the Interpersonal Reactivity Index as a measure of empathy. Our study demonstrates the subjective impact of mirror-sensory synesthesia and its stimulating influence on prosocial behaviour.
Literary language is a style or form of language used in literary writing. The intent of this investigation is to disclose how and why literary writers foreground their texts and what meanings and effects are associated with foregrounding, deviation, creativity, Style and aesthetics on literature. This paper therefore appraised the characteristics of the language of literature, with a view to revealing the potency of creativity, style and aesthetics in some African and non-African poems and novels, which in turn portrays the skilfulness and dexterity of literary writers. Specifically, it examined the foregrounded parts of selected literary works; and to achieve this purpose, linguistic benchmarks were applied to these literary works. The descriptive system of data analysis, primary and secondary data collection methods and the foregrounding/deviation theory were employed. This survey therefore revealed that the literary genius contravenes the linguistic norms deliberately because he or she believes that the most proficient means of achieving distinction in writing is the use of distorted and strange forms. Thus, style heightens the language of literature to create a special effect and special meaning to the audience in order to arousing the interest and consciousness of the reader and society at large.Key Words: Deviation, Foregrounding, Literary Artist, Literature and Stylistics.
Mirror-sensory synesthetes mirror the pain or touch that they observe in other people on their own bodies. This type of synesthesia has been associated with enhanced empathy. We investigated whether the enhanced empathy of people with mirror-sensory synesthesia influences the experience of situations involving touch or pain and whether it affects their prosocial decision making. Mirror-sensory synesthetes (<i>N</i> = 18, all female), verified with a touch-interference paradigm, were compared with a similar number of age-matched control individuals (all female). Participants viewed arousing images depicting pain or touch; we recorded subjective valence and arousal ratings, and physiological responses, hypothesizing more extreme reactions in synesthetes. The subjective impact of positive and negative images was stronger in synesthetes than in control participants; the stronger the reported synesthesia, the more extreme the picture ratings. However, there was no evidence for differential physiological or hormonal responses to arousing pictures. Prosocial decision making was assessed with an economic game assessing altruism, in which participants had to divide money between themselves and a second player. Mirror-sensory synesthetes donated more money than non-synesthetes, showing enhanced prosocial behaviour, and also scored higher on the Interpersonal Reactivity Index as a measure of empathy. Our study demonstrates the subjective impact of mirror-sensory synesthesia and its stimulating influence on prosocial behaviour.This article is part of the discussion meeting issue ‘Bridging senses: new developments in synaesthesia’.
Classic natural language processing resources such as the Penn Treebank (Marcus et al. 1993) have long been used both as evaluation data for many linguistic tasks and as training data for a variety of off-the-shelf language processing tools. Recent work has highlighted a gender imbalance in the authors of this text data (Garimella et al. 2019) and hypothesized that tools created with such resources will privilege users from particular demographic groups (Hovy and Søgaard 2015). Domain adaptation is typically employed as a strategy in machine learning to adjust models trained and evaluated with data from different genres. However, the present work seeks to evaluate whether domain adaptation to demographic groups such as age or gender may be an effective strategy to ameliorate the effects of biased or outdated training corpora in linguistic preprocessing tasks. We find adaptation to demographic groups to be an effective strategy for improving preprocessing performance across all demographic groups.
Identity expression can be seen at either a personal or a social level. It can be shown in several ways, particularly through poetry. Therefore, this paper seeks to examine how identity is expressed modally in Darwish’s famous poem “Identity Card”. Leech’s (1974) theory of “seven types of meaning” is used as the theoretical framework for the study. As the study aims to figure out how modality is constructed when expressing identity in the poem, this linguistic norm has been used to track the mood and attitude of the poets in the composition of the stanzas. The findings of the study revealed that identity is expressed modally in the sensations that Mahmoud Darwish carries for the Palestinian, Arab, National, Cultural, Geographical and Historical identities. His language enacted the way he feels towards these six components of his belongings. Nationalism is seen as an important perspective in the affiliations of Darwish. The national identity is, however, the central concept of the poem. The modality provides a rough picture of what is going on in his mind as he experiencing the loss of land. It also comes in harmony with the general cultural context that contributes to his poetic experiences.
Identity expression can be seen at either a personal or a social level. It can be shown in several ways, particularly through poetry. Therefore, this paper seeks to examine how identity is expressed modally in Darwish’s famous poem “Identity Card”. Leech’s (1974) theory of “seven types of meaning” is used as the theoretical framework for the study. As the study aims to figure out how modality is constructed when expressing identity in the poem, this linguistic norm has been used to track the mood and attitude of the poets in the composition of the stanzas. The findings of the study revealed that identity is expressed modally in the sensations that Mahmoud Darwish carries for the Palestinian, Arab, National, Cultural, Geographical and Historical identities. His language enacted the way he feels towards these six components of his belongings. Nationalism is seen as an important perspective in the affiliations of Darwish. The national identity is, however, the central concept of the poem. The modality provides a rough picture of what is going on in his mind as he experiencing the loss of land. It also comes in harmony with the general cultural context that contributes to his poetic experiences.
In this paper, we discuss constituent ordering generalizations in Japanese. Japanese has SOV as its basic order, but a significant range of argument order variations brought about by ‘scrambling’ is permitted. Although scrambling does not induce much in the way of semantic effects, it is conceivable that marked orders are derived from the unmarked order under some pragmatic or other motivations. The difference in the effect of basic and derived order is not reflected in native speaker’s grammaticality judgments, but we suggest that the intuition about the ordering of arguments may be attested in corpus data. By using the Keyaki treebank (a proper subset of which is NINJAL Parsed Corpus of Modern Japanese (NPCMJ)), it is shown that the naturally-occurring corpus data confirm that marked orderings of arguments are less frequent than their unmarked ordering counterparts. We suggest some possible motivations lying behind the argument order variations.
We present a recurrent neural network memory that uses sparse coding to\ncreate a combinatoric encoding of sequential inputs. Using several examples, we\nshow that the network can associate distant causes and effects in a discrete\nstochastic process, predict partially-observable higher-order sequences, and\nenable a DQN agent to navigate a maze by giving it memory. The network uses\nonly biologically-plausible, local and immediate credit assignment. Memory\nrequirements are typically one order of magnitude less than existing LSTM, GRU\nand autoregressive feed-forward sequence learning models. The most significant\nlimitation of the memory is generalization to unseen input sequences. We\nexplore this limitation by measuring next-word prediction perplexity on the\nPenn Treebank dataset.\n
In this paper, a classifier of emails by level of urgency in service companies is presented, using a natural language processing algorithm to grammatically label the unstructured text of messages to give it meaning and structure before analyzing it. The proposed classifier uses a lexical database that is composed of unigrams classified into eight basic emotions (anger, fear, anticipation, confidence, surprise, sadness, joy and disgust) and two feelings (positive and negative). This allows to compare the text of the grammatically labeled messages with the unigrams of the database, in order to conduct the analysis of emotions and feelings. The analysis determines the percentage of each emotion and the polarity of feeling in order to classify the most negative emails as urgent and channel them to the corresponding departments to be attended to. The research has been implemented using a set of data taken from a company dedicated to electronic invoicing in Mexico.
Abstract The article presents empirical research of verbal prepositional “of“ structures, grammatical collocations of the verb and the preposition OF. The preposition OF belongs among the most frequent prepositions in the English language. The study is based on comparisons of English and Czech sentences containing verbs and prepositions that are followed by the object. Material was taken from the electronic data bank Prague Czech-English Dependency Treebank 2.0. The structures were examined and analyzed from morphological, syntactical and semantic points of view. The aim of the study is to create English-Czech verbal prepositional counterparts; to create verbal prepositional groups on the grounds of the similar semantic, syntactic features; to identify the features that are the same for each verb group and generalize them; to identify trends and tendencies for verbs when they collocate with a certain preposition. The findings are presented in several charts and tables.
<h3>Introduction</h3><br> BOLT Egyptian Arabic-English Word Alignment -- SMS/Chat Training was developed by the Linguistic Data Consortium (LDC) and consists of 349,414 words of Egyptian Arabic and English parallel text enhanced with linguistic tags to indicate word relations. <br> The DARPA <a href="https://www.ldc.upenn.edu/collaborations/current-projects/bolt">BOLT</a> (Broad Operational Language Translation) program developed machine translation and information retrieval for less formal genres, focusing particularly on user-generated content. LDC supported the BOLT program by collecting informal data sources -- discussion forums, text messaging and chat -- in Chinese, Egyptian Arabic and English. The collected data was translated and annotated for various tasks including word alignment, treebanking, propbanking and co-reference. <br> <h3>Data</h3><br> This release consists of Egyptian Arabic source text message and chat conversations collected using two methods: new collection via LDC's collection platform, and donation of SMS or chat archives from BOLT collection participants. The source data is released as BOLT Egyptian Arabic SMS/Chat and Transliteration (<a href="../../../LDC2017T07">LDC2017T07</a>). <br> The BOLT word alignment task was built on treebank annotation. Specifically, Egyptian Arabic source tree tokens were automatically extracted from tree files in LDC's BOLT Egyptian Arabic Treebank. Those tree files had been tagged for part-of-speech and syntactically annotated. That data was then aligned and annotated for the word alignment task. <br> The data profile broken down by character tokens, tree tokens and segments appears below: <br> <table border="1" cellpadding="5"><br> <tbody><br> <tr><br> <td>Language</td><br> <td>Genre</td><br> <td>Files</td><br> <td>Words</td><br> <td>Tree/POS-tokens</td><br> <td>Segments</td><br> </tr><br> <tr><br> <td>Egyptian Arabic</td><br> <td>SMS/Chat</td><br> <td>1367</td><br> <td>349,414</td><br> <td>475,665</td><br> <td>74,814</td><br> </tr><br> </tbody><br> </table><br> <h3>Acknowledgement</h3><br> This material is based upon work supported by the Defense Advanced Research Projects Agency (DARPA) under Contract No. HR0011-11-C-0145. The content does not necessarily reflect the position or the policy of the Government, and no official endorsement should be inferred. <br> <h3>Samples</h3><br> Please view the following samples: <br> <ul><br> <li><a href="desc/addenda/LDC2019T18.arz.tkn.txt">Egyptian Arabic Source</a></li><br> <li><a href="desc/addenda/LDC2019T18.eng.tkn.txt">English Translation</a></li><br> <li><a href="desc/addenda/LDC2019T18.wa.txt">Word Alignment</a></li><br> </ul><br> <h3>Updates</h3><br> None at this time. </br> Portions © 2019 Trustees of the University of Pennsylvania
We explore whether it is possible to leverage eye-tracking data in an RNN dependency parser (for English) when such information is only available during training, i.e., no aggregated or token-level gaze features are used at inference time. To do so, we train a multitask learning model that parses sentences as sequence labeling and leverages gaze features as auxiliary tasks. Our method also learns to train from disjoint datasets, i.e. it can be used to test whether already collected gaze features are useful to improve the performance on new non-gazed annotated treebanks. Accuracy gains are modest but positive, showing the feasibility of the approach. It can serve as a first step towards architectures that can better leverage eye-tracking data or other complementary information available only for training sentences, possibly leading to improvements in syntactic parsing.
Abstract This chapter describes the data structure of the Rhapsodie Treebank and discusses methodological issues stemming from the complexity of this structure, articulated around three independent, non-aligned, hierarchies: Microsyntactic, macrosyntactic and prosodic, and the challenging questions to be resolved in this context. It discusses the specific problems posed by the simultaneous processing of the phonological stream (prosodic level) and the orthographic stream (syntactic level), which are often far from being isomorphic in French, and the related problem of the processing of disfluent and/or overlapped strings, which have not the same representation in the syntactic and the prosodic hierarchy. Then, it presents the formats adopted to encode prosodic and syntactic annotations and query them simultaneously, given that the prosodic architecture is a non-recursive time-aligned representation while the syntactic one is a recursive tree-based representation.
This chapter provides a history of 'the King's English' as a context for an analysis of language, history and power in The Merry Wives of Windsor and the second tetralogy. The trope is used as a rhetorical and ideological tool in performatives. Associated with temperance and honesty, 'the King's English' belongs to a set of defining values of true Englishness. The project to produce this linguistic norm coincides with a homologous project to produce a stable, monetary system of 'good' coin through exclusion of 'bad', 'counterfeit' or 'clipped' coin. These projects testify to a shift of the centre of economic and cultural gravity from the court to the merchant citizen class. Shakespeare's one English comedy centred on English citizens which features his one use of 'the King's English' is shown to engage critically with this ideology, and to set against it an idea of 'our English' as an inclusive mix, the 'gallimaufry' loved by the linguistically extravagant gentleman John Falstaff. The comedy draws out the implications of the banishment of Falstaff in the second tetralogy, which sets history against the project of cultural reformation ideology to produce (the) 'true' English.
This paper examines the relation between gendered language and the processing of non-stereotypical gender representation through a psycholinguistic priming experiment consisting of a self-paced reading test. The experiment tests two things: The processing ease of 3rd person singular pronouns that either match or mismatch the stereotypical gender of their referents; and whether sentences with gendered language affect this. Processing ease is measured by reading time. Sentences with gendered language are used as priming; and pronouns with matching or mismatching stereotypicality as targets. The results showed no difference in the reading time of matching and mismatching pronouns, and no priming effect was found. This could point to the fact that gender stereotypical mismatch does not affect the informants; and that gendered language is so well-integrated in our language that it cannot prime for gender stereotypes. Moreover, the pronoun han was read significantly faster than the pronoun hun, which could indicate that the masculine is expected as a linguistic norm. This is considered in relation to the reflections on the missing priming effect as a result of gendered language being the norm.
This paper focuses on the process and principle of parsing, which is an essential task for machine to understand the syntactic, semantic structure of a sentence. First, a series of machine analysis procedures such as word segmentation, part-of-speech tagging and parsing of Chinese sentences are visually represented by using Cparser, a rule-based constituency parser developed by Peking University. Next, to better understand parsing mechanism, we explain in detail how the linguistic knowledge is embodied in the lexical, syntactic and semantic component of Cparser, showing their complex interplay that allows automatic parsing. As a practical example, a Chinese textbook treebank is also constructed using Cparser. According to the theoretical and practical discussion in this paper, Peking University Cparser, which is easy to reflect and modify linguistic knowledge, is expected to be widely used as an analysis and verification tool for Chinese grammar research.
This paper presents the first gold-standard resource for Russian annotated with compositionality information of noun compounds. The compound phrases are collected from the Universal Dependency treebanks according to part of speech patterns, such as ADJ+NOUN or NOUN+NOUN, using the gold-standard annotations. Each compound phrase is annotated by two experts and a moderator according to the following schema: the phrase can be either compositional, non-compositional, or ambiguous (i.e., depending on the context it can be interpreted both as compositional or noncompositional). We conduct an experimental evaluation of models and methods for predicting compositionality of noun compounds in unsupervised and supervised setups. We show that methods from previous work evaluated on the proposed Russian-language resource achieve the performance comparable with results on English corpora.
The article refers to the concept of intelligentsia as a social group which exerts significant influence on Polish standard patterns. Although the term intelligentsia is vague and questionable, it is well-established term in Polish linguistics, especially in sociolinguistics. Author argues that science communicators (young professional researchers, science journalists, PhD students) represent the young intelligentsia, because these well-educated people pursue their intellectual development and they have sense of public duty. The article examines standard of popular science texts in Internet, new tendencies in written Polish and attitude of young intelligentsia toward traditional linguistic norm. The errors (esp. punctuation and syntax) exemplify impact of technological changes and phenomenon of secondary orality. It would be useful for science communicators to edit carefully their texts. Both researchers and journalists need to improve their writing skills permanently. Nevertheless it must be emphasized that school education and competent teachers seem to have important influence on the linguistic patterns.
The communicative role of nonlinear vocal phenomena remains poorly understood since they are difficult to manipulate or even measure with conventional tools. In this study parametric voice synthesis was employed to add pitch jumps, subharmonics/sidebands, and chaos to synthetic human nonverbal vocalizations. In Experiment 1 (86 participants, 144 sounds), chaos was associated with lower valence, and subharmonics with higher dominance. Arousal ratings were not noticeably affected by any nonlinear effects, except for a marginal effect of subharmonics. These findings were extended in Experiment 2 (83 participants, 212 sounds) using ratings on discrete emotions. Listeners associated pitch jumps, subharmonics, and especially chaos with aversive states such as fear and pain. The effects of manipulations in both experiments were particularly strong for ambiguous vocalizations, such as moans and gasps, and could not be explained by a non-specific measure of spectral noise (harmonics-to-noise ratio) – that is, they would be missed by a conventional acoustic analysis. In conclusion, listeners interpret nonlinear vocal phenomena quite flexibly, depending on their type and the kind of vocalization in which they occur. These results showcase the utility of parametric voice synthesis and highlight the need for a more fine-grained analysis of voice quality in acoustic research.
Abstract This article presents results from a study on hybrid linguistic norms in translated articles from New York Times made available on UOL website. According to Faraco (2008) and Bagno (2012), there is a difference between norma padrão (a prescriptive norm, but not based on usage) and norma culta (an alternative, usage-based norm). The first one combines normative rules that determine correct linguistic forms, but generally hard to follow by most users, while the second one brings together a set of linguistic forms frequently employed by users, because they are more intuitively accessible, although not subscribed by the conservative standard norm (norma padrão) commonly taught in grammar books and in writing style manuals. The research was meant to verify if journalistic texts translated from English have been as permeable to linguistic forms not subscribed by the prescriptive standard norm, as the ones originally written in Portuguese have proven to be.
Location: Dewberry Hall With over 34,000 students representing 123 countries, George Mason University is a vastly diverse university with students bringing different learning experiences and skill sets with them into the classroom. The study that our team conducted analyzes the essays of native English speakers, as well as the essays of students for whom English is their second language. Our objective when conducting this research was to observe the essays for signs of syntactic complexity and patterns of language errors. Specifically, we looked for subordinating clauses, transitions, subject/verb agreement, article usage, run-on sentences, and fragments. We found that some errors in L1 and L2 populations were consistent with our expectations, but others reveled a more complex understanding of the linguistic norms of the groups studied. The results of these findings will give professors of all disciplines and modalities insight to the challenges that first-year L1 and L2 students confront when faced with a writing assignment.
Contextualized embeddings, which capture appropriate word meaning depending\non context, have recently been proposed. We evaluate two meth ods for\nprecomputing such embeddings, BERT and Flair, on four Czech text processing\ntasks: part-of-speech (POS) tagging, lemmatization, dependency pars ing and\nnamed entity recognition (NER). The first three tasks, POS tagging,\nlemmatization and dependency parsing, are evaluated on two corpora: the Prague\nDependency Treebank 3.5 and the Universal Dependencies 2.3. The named entity\nrecognition (NER) is evaluated on the Czech Named Entity Corpus 1.1 and 2.0. We\nreport state-of-the-art results for the above mentioned tasks and corpora.\n
This article proposes a character-level neural language model (NLM) that is based on quantum theory. The input of the model is the character-level coding represented by the quantum semantic space model. Our model integrates a convolutional neural network (CNN) that is based on network-in-network (NIN). We assessed the effectiveness of our model through extensive experiments based on the English-language Penn Treebank dataset. The experiments results confirm that the quantum semantic inputs work well for the language models. For example, the PPL of our model is 10%–30% less than the states of the arts, while it keeps the relatively smaller number of parameters (i.e., 6 m).
Abstract This chapter is devoted to the presentation of the tools and methods used for the different steps of the semi-automatic syntactic annotation: automatic preprocessing; microsyntactic parsing with the FRMG tool, correction of the parsing with the Arborator tool, agreement analysis, post-validation correction, and development of the final format of the Rhapsodie syntactic treebank. As FRMG is a parser for written French that was not configured to analyze disfluencies and reformulation, we used our manual pile marking to unfold the piles and produce a series of simplified “sentences” with only government relations. Despite having two annotators plus a validator for the corrections, we found a substantial number of errors in the post-validation procedure by using a set of rules to determine the well-formedness of the trees.
This talk will explore the role of individual social actors and their communities and networks in the formation, maintenance and dissolution of linguistic norms. It will consider several standardisation episodes in the history of Old and Early Middle English and connect them to other unification processes in the political and cultural history of England at the time. It will be suggested that the suppression of variability on the linguistic level often accompanies, or is a symptom of, a similar suppression on the ideological level. Political, religious, legal, and linguistic processes mingle in various ways in this period (as they do today) to replicate and enhance the social order, to support a reform movement, or to refute dissent. \nThree case studies will offer insights into these processes: 1) shire courts and the ‘standardised’ lexis of the Anglo-Saxon Chronicle in the reign of King Alfred and Edward the Elder; 2) chancery norms and charters of the eleventh century; and 3) religious reform and linguistic focusing in the thirteenth century.
This is a work-in-progress report, which aims to share preliminary results of a novel sequence-to-sequence schema for dependency parsing that relies on a combination of a BiLSTM and two Pointer Networks (Vinyals et al., 2015), in which the final softmax function has been replaced with the logistic regression. The two pointer networks co-operate to develop a latent syntactic knowledge, by learning the lexical properties of "selection" and the lexical properties of "selectability", respectively. At the moment and without fine-tuning, the parser implementation gets a UAS of 93.14% on the English Penn-treebank (Marcus et al., 1993) annotated with Stanford Dependencies: 2-3% under the SOTA but yet attractive as a baseline of the approach.
Do individual sounds carry meaning? The relationship between sound and meaning in human languages is typicallyassumed to be arbitrary, though recent research provides evidence for the existence of both iconicity and systematicitybetween word forms and their meaning. However, this research has not asked whether individual sounds in a languagecovary in systematic ways with aspects of meaning. In two analyses, we find evidence for more systematicity betweenthe initial phones of words and those words concreteness ratings than one would expect in a truly arbitrary lexicon. Thissuggests that initial phones may act as cues to aspects of word meaning, and raises questions about whether languagelearners detect and exploit these cues.
This paper is concerned with whether deep syntactic information can help surface parsing, with a particular focus on empty categories. We consider data-driven dependency parsing with both linear and neural disambiguation models. We find that the information about empty categories is helpful to reduce the approximation error in a structured prediction based parsing model, but increases the search space for inference and accordingly the estimation error. To deal with structure-based overfitting, we propose to integrate disambiguation models with and without empty elements. Experiments on English and Chinese TreeBanks indicate that incorporating empty elements consistently improves surface parsing.
Neural models have been investigated for sentiment classification over constituent trees. They learn phrase composition automatically by encoding tree structures but do not explicitly model sentiment composition, which requires to encode sentiment class labels. To this end, we investigate two formalisms with deep sentiment representations that capture sentiment subtype expressions by latent variables and Gaussian mixture vectors, respectively. Experiments on Stanford Sentiment Treebank (SST) show the effectiveness of sentiment grammar over vanilla neural encoders. Using ELMo embeddings, our method gives the best results on this benchmark.
Throughout many studies which focus on brain laterality as a key component to the outcome of an experiment, or to a participants’ reaction to a stimulus, it can be noted that the different areas of the brain involved in a task response must work together to produce a viable outcome (i.e. lateralized brain processes). While there are laterality components in relation to the cognitive processes of handed and footed responses, it is still largely unknown how the different areas of the brain interpret the emotional stimuli to then affect these outcomes. The purpose of our study is to determine how emotional context affects a simple cognitive task that includes handed and footed responses, and if any observed differences can be traced back to the different systems at work within the brain. Subjects will be tested over two days for handed and footed responses in a cognitive Simon Task. Subjects will be tested with and without emotional context (i.e. a background image of a specific valence and arousal rating), and any resulting differences between non-emotional context and emotional context reaction times will be compared.
In sequence learning tasks such as language modelling, Recurrent Neural\nNetworks must learn relationships between input features separated by time.\nState of the art models such as LSTM and Transformer are trained by\nbackpropagation of losses into prior hidden states and inputs held in memory.\nThis allows gradients to flow from present to past and effectively learn with\nperfect hindsight, but at a significant memory cost. In this paper we show that\nit is possible to train high performance recurrent networks using information\nthat is local in time, and thereby achieve a significantly reduced memory\nfootprint. We describe a predictive autoencoder called bRSM featuring recurrent\nconnections, sparse activations, and a boosting rule for improved cell\nutilization. The architecture demonstrates near optimal performance on a\nnon-deterministic (stochastic) partially-observable sequence learning task\nconsisting of high-Markov-order sequences of MNIST digits. We find that this\nmodel learns these sequences faster and more completely than an LSTM, and offer\nseveral possible explanations why the LSTM architecture might struggle with the\npartially observable sequence structure in this task. We also apply our model\nto a next word prediction task on the Penn Treebank (PTB) dataset. We show that\na 'flattened' RSM network, when paired with a modern semantic word embedding\nand the addition of boosting, achieves 103.5 PPL (a 20-point improvement over\nthe best N-gram models), beating ordinary RNNs trained with BPTT and\napproaching the scores of early LSTM implementations. This work provides\nencouraging evidence that strong results on challenging tasks such as language\nmodelling may be possible using less memory intensive, biologically-plausible\ntraining regimes.\n
The emergence of deep learning as a commanding technique for learning heterogeneous layers of feature representations have consequently substituted traditional machine learning algorithms which are generally poor in analyzing compound sentences. Additionally, convolutional and recurrent neural networks have auspiciously yielded state-of-the-art results in sentiment classification and Natural Language Processing (NLP). In this paper, a deep sentiment representation model through the combination of multiple Convolutional Neural Networks (CNN) kernels with Long Short-Term Memory (LSTM) is proposed for sentiment classification. Our model gains word vector representation using pre-trained Global Vectors for Word Representation (GloVe) embeddings, thereafter used as input to the CNN layer which extracts higher local text representations. Finally, Bidirectional LSTM (biLSTM) generates sentiment classification of sentence representation based on context dependent features. Our combined approach of CNN and biLSTM was experimented using the Stanford Large Movie Review Dataset (IMDB) and Stanford Sentiment Treebank Dataset (SSTB) for binary classification. The evaluation achieves outstanding results in outperforming several existing approaches with 90.4% accuracy on the Stanford Sentiment Treebank dataset and 94.8% accuracy on the Stanford Large Movie Review dataset. These results are achieved with a drastic reduction of model parameters and without a pooling layer in the CNN architecture, helping to retain local and structural information in comparison to other existing deep neural network frameworks.
This thesis studies the connections between parsing friendly representations and interlingua grammars developed for multilingual language generation. Parsing friendly representations refer to dependency tree representations that can be used for robust, accurate and scalable analysis of natural language text. Shared multilingual abstractions are central to both these representations. Universal Dependencies (UD) is a framework to develop cross-lingual representations, using dependency trees for multlingual representations. Similarly, Grammatical Framework (GF) is a framework for interlingual grammars, used to derive abstract syntax trees (ASTs) corresponding to sentences. The first half of this thesis explores the connections between the representations behind these two multilingual abstractions. The first study presents a conversion method from abstract syntax trees (ASTs) to dependency trees and present the mapping between the two abstractions – GF and UD – by applying the conversion from ASTs to UD. Experiments show that there is a lot of similarity behind these two abstractions and our method is used to bootstrap parallel UD treebanks for 31 languages. In the second study, we study the inverse problem i.e. converting UD trees to ASTs. This is motivated with the goal of helping GF-based interlingual translation by using dependency parsers as a robust front end instead of the parser used in GF. \n\nThe second half of this thesis focuses on the topic of data augmentation for parsing – specifically using grammar-based backends for aiding in dependency parsing. We propose a generic method to generate synthetic UD treebanks using interlingua grammars and the methods developed in the first half. Results show that these synthetic treebanks are an alternative to develop parsing models, especially for under-resourced languages without much resources. This study is followed up by another study on out-of-vocabulary words (OOVs) – a more focused problem in parsing. OOVs pose an interesting problem in parser development and the method we present in this paper is a generic simplification that can act as a drop-in replacement for any symbolic parser. Our idea of replacing unknown words with known, similar words results in small but significant improvements in experiments using two parsers and for a range of 7 languages.
Discourse Relations, also known as coherence or rhetorical relations, characterize the semantic or pragmatic relationships between clauses or sentences in discourse.Such relations are established in order to facilitate effective communication.In addition to the inventory of relations, previous research has also investigated how discourse relations are established or signaled.Discourse markers (DMs) are considered to be the most typical signals in discourse; however, focusing merely on DMs is inadequate as they can only account for a small number of relations in discourse.Thus, researchers have been exploring textual signals beyond DMs such as the Penn Discourse Treebank 2.0 (PDTB, Prasad et al. [22]) and the Rhetorical Structure Theory Signalling Corpus (RST-SC, Das and Taboada [5]).Despite their different theoretical groundings and approaches to relation signaling, both corpora annotated the Wall Street Journal (WSJ) section of the Penn Treebank (PTB, Marcus et al. [19]), i.e. the news articles.Nevertheless, previous work has suggested that signaling information is indicative of genres (e.g.Taboada and Lavid [28]; Zeldes [34]).Therefore, this project aims to anchor signaling devices on a more diverse corpus to demonstrate the inadequacy of signaling by DMs only, the abundance of open-class signals, and more importantly, the distribution of signaling devices across genres.
This paper suggests annotation guidelines to build a Universal Dependencies (UD) treebank for Korean. We discuss the part-of-speech annotation of Korean specific-categories such as prenouns, numeral classifiers, and (pre)final endings, and propose how to implement UD scheme in Korean regarding selecting a head and assigning dependency relations to dependents. UD prioritizes content words over functional words since the former exhibits less cross-linguistic variations. In a noun phrase, for instance, a core noun is always a head of the entire noun phrase independently of a language. The rest are treated as a dependent: not only a modifier such as an adjective but also a functional category such as an article, numeral quantifier, demonstrative, and so on. However, when it comes to head-less constructions such as coordination or predicate ellipsis, UD firmly advocates the head-initial strategy. The present application of UD to Korean tries to follow UD’s principles as much as possible. Korean is a head-final language, so that headed constructions are analyzed head-finally. In contrast, head-less ones are tagged head-initially. This might disregard language-specific characteristics from a linguistic perspective, but the strategy allows us to build up a set of treebanks in a cross-linguistically consistent way (i.e., the fundamental purpose of UD).