Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
This article investigated how network growth algorithms-preferential attachment, preferential acquisition, and lure of the associates-relate to the acquisition of words in the phonological language network, where edges are placed between words that are phonologically similar to each other. Through an archival analysis of age-of-acquisition norms from English and Dutch and word learning experiments, we examined how new words were added to the phonological network. Across both approaches, we found converging evidence that an inverse variant of preferential attachment-where new nodes were instead more likely to attach to existing nodes with few connections-influenced the growth of the phonological network. We suggest that the inverse preferential attachment principle reflects the constraints of adding new phonological representations to an existing language network with already many phonologically similar representations, possibly reflecting the pressures associated with the processing costs of retrieving lexical representations that have many phonologically similar competitors. These results contribute toward our understanding of how the phonological language network grows over time and could have implications for the learning outcomes of individuals with language disorders. (PsycInfo Database Record (c) 2020 APA, all rights reserved).
Achieving changes to education practices and structures is a significant issue facing reformers internationally, and researchers have confronted how such changes, and the conditions for these, might be conceptualized. These issues resonate particularly as researchers grapple with imagining a post-COVID-19 landscape where social and educational norms may change. Tyack and Tobin, in their 1994 article 'The "Grammar" of Schooling: Why has it been so hard to change?' argued that several features of the American education system are so persistent as to warrant being understood as the 'grammar' of schooling. In this article, we reconceptualize this 'grammar' by taking seriously Tyack and Tobin's insistence that 'grammar' organises meaning. Starting here, we argue that what they took to be grammatical features are the products and not the producers of meaning. We draw on the cases of the United States and England to argue that four international discourses have performed this meaning-making work: industrialization; welfarism; neoliberalism and neoconservatism. These are the 'grammars' of schooling-and of society. Their discursive products, including age grading and sorting into subjects are, we suggest, 'lexical' features that express the grammar. We use lexical features to explain the multi-directional interplay between discourse and educational feature: the lexical may endure longer than the grammatical, changes to which may be effected and/or legitimated through appealing to a lexical feature. We conclude by outlining key implications for realizing and conceptualizing educational change, including for a post-COVID-19 landscape.
The article analyzes the technique of including outdated, unused words of the Russian language, typical of M. Tyomkina's book of verses "Unvisual Aids". Such linguistic "memes" work as triggers for the self-reflection of the lyrical heroine, who has lived in America for a long time; at the same time, feeling the inextricable connection of her inner self with her mother's tongue, its meanings, phobias, and norms. The Russian words, popping up in her memory, sometimes colloquial, sometimes bookish, or professional-philological ones, form the depth of personality that cannot be realized in everyday life, in business and everyday conversations in English. The article shows that these words, which are "deeply drowned" in her "pre-memory" do not fulfill cognitive or communicative functions, but only psychological and poetic ones, acting as metonyms or "unvisual aids" for self-reflection. Immersion into the lexical and grammatical fields of the native language meanings plays the role of "sequence of times" (Tyomkina), the sequence of linguistic, historical and existential elements. This constitutes the originality of the author's cultural multi-identity. In rejecting the classical form of the Russian verse, Tyomkina develops the lyrical function of poetry through her special "recitative".
Summary This paper compares the romanization of Gaul in the 1st century BC and the gallicization of the island of Martinique during 17th-century French colonial expansion, using criteria set out by Muf- wene's Founder Principle. The Founder Principle determines key ecological factors in the formation of creole vernaculars, such as the founding populations and their proportion to the whole, language varieties spoken, and the nature and evolution of the interactions of the founding populations (also referred to as “colonization styles”). Based on the comparison, it will be claimed that new languages arise when a language undergoes vehicularization and subsequently shifts from one speech community to another. In other words, linguistic genesis would be a complicated case of language contact, where not only one, but sev- eral dialects of both superstrate and substrate varieties are involved, in a historical context where the identity function of language, or the norm, is overriden by the need to communicate. Research also indicates that language varieties spoken at the time of the shift did not pertain to normative usage, but to popular varieties, dialects, or both, since the emerging vernaculars - in Gaul, as well as in Martinique - preserved some of their phonological and lexical particularities.
This paper is an attempt to shed light on linguistic deviation in literary style. Literary language, with its three main genres; poetry, drama and prose, is a situational variety of English that has specific features which belong to the literary and elevated language of the past. Literary language has been assigned a special status since antiquity, and is still used nowadays by some speakers and writers in certain situations and contexts. It has been considered as sublime and distinctive from all other types of language; one which is deviant from ordinary use of language in that it breaks the common norms or standards of language. A basic characteristic of literary style is linguistic deviation which occurs at different levels; lexical, semantic, syntactic, phonological, morphological, graphological, historical, dialectal and register. All these types of deviations are thoroughly investigated and stylistically analyzed in this paper so as to acquaint readers, students of English, researchers, and those interested in the field, with this type of linguistic phenomenon whose data is based on selected samples from major classical works in English literature
The article provides an overview of the lexical and grammatical features as well as the sociopolitical environment of Marollien that originated in the 18th century as a dialect on the territory of Brussels. Marollien is essentially the Dutch language in its Brabantian dialect, strongly influenced by French. There are literary works, performances, and musicals written and staged in Marollien, as well as dictionaries and journals published in it. Historically, the Marollien dialect is a sociolect: it was generally used by Belgians coming to Brussels from Wallonia in search of a job and settling in one of the districts of Brussels — Marolles. A special emphasis is placed on lexical features of the dialect: gastronomic and everyday vocabulary are looked at and the examples of French loanwords and Southern Dutch language norm deviations are provided. Standard Dutch calques in French, when translating idioms in particular, are also identified. The differences between Dutch, French, and Marollien place names are illustrated. In the field of morphology and word formation, there is a regular mixture of Germanic and Romanic stems which is indicated. Examples of Marollien phonetic features are also provided. The article acknowledges frequent code switching in Marollien speech, which by and large resembles the phenomenon of linguistic interference. Due to the fact that Marollien is rapidly disappearing, the Brussels-Capital region is trying to support the dialect: various activities are being organized in order to propagate its use and enhance its prestige. Nevertheless, Marollien is not included in the well-known citizen initiative “Marnix Plan”, aimed at developing the methodology for the sequential study of several languages for all segments of the population in Brussels. This initiative is also discussed in the article.
The present study explores the process of how Korean students develop pragmatic competence when writing request emails during metapragmatic instruction. In particular, the study focuses on how students’ usage of request strategies, lexical devices, external modifications, and request perspectives change during the metapragmatic instructional period. Descriptive analysis of the usage frequency of politeness devices was conducted to examine changes before and after the instruction. In addition, the participants’ accounts during the metapragmatic discussions and retrospective interviews were analyzed to ascertain their intentions of the requests as well as their experiences during the instructional period. The results of the study showed that the participants developed pragmatic competence as they enhanced their awareness of the difference between what they had intended and what was actually conveyed. Moreover, the metapragmatic discussions among their peers in regard to the source of the pragmatic failure as well as their developing pragmatic awareness allowed them to experience meaningful interactions. The study stresses the need for explicit metapramatic instructions along with encouragements for metapragmatic discussions to reflect on the norms of the target culture along with their own pragmatic knowledge for Korean students to enhance their pragmatic competence.
The article deals with the peculiarities of linguistic and cultural changes of language structure influenced by globalization process within the language contacts’ interaction. The analysis of various aspects in the modern society proves the dominance of the English language in the formation of the world collaboration. According to the research, English hybrid languages or new Englishes, based on the Standard English norms, are forced to adapt to the local linguistic and cultural needs. These hybrid languages perform the mixture of indigenous languages’ structure and Standard English rules, thought in many cases English dominates and replaces phonetic, lexical, syntactic elements of indigenous languages. Much attention in the work is paid to the peculiarities of such hybrid language as Nigerian English, which presents the local language variant, functioning in Nigeria. Owing to language contacts’ cooperation, Nigerian English combines the language features of Standard English rules and Nigerian local languages’ peculiarities.
The article deals with the study of Jazz poetry as a part of integral jazz art which highly involves music, literature, singing and jazz dancing and presents a complex worldview. Intertextuality of jazz poetry is claimed to characteristic feature of postmodern literature. Basic ideas of democracy, equality, tolerance and struggle for freedom and justice determine such key qualities of jazz art as polyphony, polyrythmy, improvisation and emotionality. Examples of different literary stylistic devices (repetitions (lexical and syntactic; anaphor, epistrophe, parallelism), alternation of long and short phrases and syncopes, complex sentences and constructions, lack of punctuation, thematical shifts and digressions, rhetorical questions) used to express democracy, equality, tolerance and struggle for freedom and justice in jazz poems are provided extensively. Multi-aspect nature of jazz art combining music, singing, literature and dancing determines its deep symbolism. Music which since its very uprise has been Jazz foundation, its secret code, its central means of struggle for freedom, has produced the majority of symbol-dominants in Jazz poetry (saxophone, horn, drums, banjo and piano) which possess special connotations of ancient traditions, authenticity, cultural uniqueness and sensuousness of Afro-Americans. Dance and singing imagery are also of great importance in jazz verses. Blues and Broadway jazz represent both deep voiceless grief of former slaves and unrestrained merriment inherent in Afro-American culture. Jazz dance as well as jazz music, is a symbol of resistance, disobedience, rebellion against former white masters. Jazz dance is a rebellion against the traditions of classical dance reflecting the struggle against established social norms.
The article analyzes the concept of comic as linguistic category.The article focuses attention on such aspects of the category of comic as classification, parameters of comic and the difficulties of translation of comic.It is singled out that the notion of concept comic goes beyond the linguistic aspect.That is why this language phenomenon is also studied from the point of view of socio-cultural aspect and the norms of human behavior in society.The category of comic is complex and ambiguous.In general, within this notion we understand the influence of jokes, for example, causing laughter.There are several methods that help to achieve laughter.Among them we distinguish irony, sarcasm, pun, allusion, periphrases, oxymoron, metaphor etc.Because of the multiaspect character of comic, this term is often used simultaneously with its similar synonymous humor, laughter, comic, funny, meaningless, cute, witty, joke, absurdity, irony, sarcasm, satire, etc.As for the categories of comic, scholars single out: analysis of style of a certain comic text, allocation and description of special features of comic text style of the specific authors; defining and studying the language and methods of realization of the category of comic on the example of a particular language; characteristics of the language parameters of specific subspecies of the category of comic.The category of comic is also subdivided into three groups.The first group is represented by understandable and easily recognizable humor based on the comic situation.The second type is presented by jokes based on the cultural base of the source language.And to the third type we refer linguistic humor.This type of humor in its structure is the most difficult to decode into a different language, because it is based on a game of words (pun), which, unfortunately, is usually considered to be not subject to translation.And, as for the translation aspect, it is necessary for translators to find a situational equivalent, an equivalent with an increased level of emotionality; to use transcoding or the combination of transcription and transliteration with the addition of word-forming morpheme; and to apply lexical and semantic and phonetically-imitating transformations.
Lexical normalization is the task of translating non-standard social media data to a standard form. Previous work has shown that this is beneficial for many downstream tasks in multiple languages. However, for Italian, there is no benchmark available for lexical normalization, despite the presence of many benchmarks for other tasks involving social media data. In this paper, we discuss the creation of a lexical normalization dataset for Italian. After two rounds of annotation, a Cohen’s kappa score of 78.64 is obtained. During this process, we also analyze the inter-annotator agreement for this task, which is only rarely done on datasets for lexical normalization,and when it is reported, the analysis usually remains shallow. Furthermore, we utilize this dataset to train a lexical normalization model and show that it can be used to improve dependency parsing of social media data. All annotated data and the code to reproduce the results are available at: http://bitbucket.org/robvanderg/normit.
The problem of language interference being a process which retards the mastering of a second language, having appeared as a result of transference of speech skills from one contact language into another (from the native language into the foreign language, from the first foreign language into the second one), has concerned researchers for decades. This phenomenon has a direct influence on the success of an individual’s mastery of a foreign language and its use—involving both receptive and productive types of speech activities. Interference resulting from the negative impact of one language on another covers all linguistic levels of the language being studied, including lexical, which leads to deviations from the language norm and numerous lexical errors of students. Linguists and methodologists are trying to find ways to reduce the interference of the language being studied at the lexical level in order to optimize the process of mastering a foreign language and minimize lexical errors of students. The purpose of the current study is to investigate ways to overcome intra-language and inter-language lexical interference in junior courses of the Azerbaijan University of Languages and to verify the validity of these methods in the course of a practical experiment.
Focus of the CONcreTEXT task is conceptual concreteness: systems were solicited to compute a value expressing to what extent target concepts are concrete (i.e., more or less perceptually salient) within a given context of occurrence. To these ends, we have developed a new dataset which was annotated with concreteness ratings. Interestingly, these works extend information on conceptual concreteness available in existing (non contextual) norms derived from human judgments with new knowledge from recently developed neural architectures, in much the same multidisciplinary spirit whereby the CONcreTEXT task was organized.
English has become ‘the world’s default mode’ (McArthur, 2002: 13) for communication. As a de facto lingua franca, English and its associated cultures are increasingly pluralistic. According to Kachru (1996: 135), ‘the term “Englishes” is indicative of distinct identities of the language and literature. “Englishes” symbolizes variation in form and function, use in linguistically and culturally distinct contexts, and a range of variety in literary creativity.’ As far as Chinese English is concerned, Kirkpatrick & Xu (2002: 278) suggest that since ‘the great majority of the estimated 350 million Chinese’ who have been learning English are far more likely to use it with other speakers of world Englishes, the development of Chinese English ‘with Chinese characteristics’ will be ‘an inevitable result’. Kirkpatrick & Xu also predict that such a variety of English will be characterized by linguistic and cultural norms derived from Chinese. This chapter will review the definitions of Chinese English, and then identify a selection of lexical, syntactic, discourse and pragmatic features of Chinese English based on an analysis of a variety of data including interviews, newspaper articles, and literary works. The chapter will conclude by considering the likelihood of Chinese English becoming a powerful variety of English.
In neural machine translation (NMT), sequence distillation (SD) through creation of distilled corpora leads to efficient (compact and fast) models.However, its effectiveness in extremely low-resource (ELR) settings has not been well-studied.On the other hand, transfer learning (TL) by leveraging larger helping corpora greatly improves translation quality in general.This paper investigates a combination of SD and TL for training efficient NMT models for ELR settings, where we utilize TL with helping corpora twice: once for distilling the ELR corpora and then during compact model training.We experimented with two ELR settings: Vietnamese-English and Hindi-English from the Asian Language Treebank dataset with 18k training sentence pairs.Using the compact models with 40% smaller parameters trained on the distilled ELR corpora, greedy search achieved 3.6 BLEU points improvement in average while reducing 40% of decoding time.We also confirmed that using both the distilled ELR and helping corpora in the second round of TL further improves translation quality.Our work highlights the importance of stage-wise application of SD and TL for efficient NMT modeling for ELR settings.
The position of the word final –s, after a weakening in archaic Latin, seems to be fixed in the spoken language in the classical period. Then, it partially disappeared in the Romance languages: in modern languages, it is conserved only north and west of the Massa–Senigallia line, while we cannot find it neither in the eastern regions nor in South Italy. Based on this fact, linguists generally claim that the weakening of the final –s started only after the intensive dialectal diversification of Latin, simultaneously with the evolution of the Romance languages. However, the data of the Computerized Historical Linguistic Database of Latin Inscriptions of the Imperial Age (LLDB) do not verify this generally accepted opinion. We can find almost as many examples of the lack of word final –s as that of –m also from the earlier centuries of the Imperial age. The aim of this paper is to explore the reasons behind the inconsistencies between the scholarly consensus and the epigraphical data.
We present a method for conducting morphological disambiguation for South Sámi, which is an endangered language. Our method uses an FST-based morphological analyzer to produce an ambiguous set of morphological readings for each word in a sentence. These readings are disambiguated with a Bi-RNN model trained on the related North Sámi UD Treebank and some synthetically generated South Sámi data. The disambiguation is done on the level of morphological tags ignoring word forms and lemmas; this makes it possible to use North Sámi training data for South Sámi without the need for a bilingual dictionary or aligned word embeddings. Our approach requires only minimal resources for South Sámi, which makes it usable and applicable in the contexts of any other endangered language as well.
Background: Virtual reality (VR) allows people to embody avatars that are different from themselves in appearance and ability. These experiences provide opportunities to challenge bodily perceptions. We devised a novel VR Body Image Training (VR-BIT) approach to target self-perceptions and pain in people with persistent pain. Methods: A 45-year old male with a 5-year history of disabling chronic low back pain participated in a four week VR-BIT intervention. Pain began following a fall from a first-floor deck. Pain was central and on the right side of his lower back, radiating to his right buttock and thigh. Pain was constant and varying at a 5/10 average intensity. The 4-week intervention consistent of three face-to-face sessions one week apart, followed by 1-week of in-home VR-BIT. During the first face-to-face session, the participant embodied three athletic avatars: a superhero (Incredible Hulk), a boxer, and a rock climber. Since the participant strongly identified with the boxer, only boxing experiences were subsequently used. Primary outcomes relating to body image (self-perceived strength, vulnerability, agility and confidence with activity) and pain intensity were assessed using numerical rating scales (0 to 10 NRS). Disabilility, kinesiophobia, overall change, and self-efficacy were assessed as secondary outcomes. Outcomes were assessed during each face-to-face session, and at 1-week and 3-month follow-up. Results: The participant reported a high degree of engagement. Positive changes were noted during and after VR for all body image and pain assessments. Improvements were retained at 3-months for body image ratings (mean change: 4.5/10 NRS) and average pain intensity (change: 2/10 NRS)). Improvements in disability (45% improvement); self-efficacy (pre: 2/12; post: 10/12); and overall change (‘Very much improved’) were noted at 3-month follow-up. No change in kinesiophobia was detected. No adverse advents were recorded. Conclusion: The participant engaged strongly with the intervention and showed clinically meaningful changes in body image, pain, disability and self-efficacy. Despite his long history of pain and rapid improvements, reported changes may be due to non-treatment effects. Nonetheless, VR-BIT clearly warrants further investigation as a potential addition to usual care.
The author reflects upon the place of the Russian language in modern Russia, its distribution in the world, the importance of basic Russian studies for the development of science and Russian society, and the activity of the Russian Academy of Sciences in maintaining the stability of linguistic norms and the culture of Russian speech. On the one hand, scientific research in the field of the Russian language is oriented at obtaining basic theoretical knowledge, which favors comprehensive study of man and society. On the other hand, there is a social order, which is formulated by society proceeding from the need to document language resources and adapt them to topical communicative requirements. The Russian Academy of Sciences carries out expert assessment of speech innovations and codification of the norms of the literary language in normative dictionaries, grammars, and reference books on the culture of speech. The current state of research on the Russian language is analyzed, special attention being paid to problems of the codification of the norms of Russian speech and related tasks.
The concept of Ontologies has been used in a wide range of application domains, due to the fact that ontologies provide a useful mean for establishing a formal, shared and collective understanding of the concepts and their underlying relations at a certain domain of interest, which allows for interoperability and information exchange in a formal an understandable way for both humans and machines. In Cultural Heritage (CH) domain, ontologies serve as a fundamental building block for the traceability of the cultural heritage objects, especially with the increasing demand of providing digital formats for cultural objects and make them available for public. In this paper we implement OntoM; an Ontology model that incorporates the relevant concepts of the Cultural Heritage (CH) domain in Qatar. Then, we will use such an ontology to perform inferences about cultural object classifications via two approaches: string matching, that allows for direct matching between the object and the ontology concepts, and semantic matching, in which we use WordNet lexical database to find all possible synonyms for properties of a given anonymous object.
This paper describes the automatic construction of FinnMWE: a lexicon of Finnish Multi-Word Expressions (MWEs). In focus here are syntactic frames: verbal constructions with arguments in a particular morphological form. The verbal frames are automatically extracted from FinnWordNet and English Wiktionary. The resulting lexicon interoperates with dependency tree searching software so that instances can be quickly found within dependency treebanks. The extraction and enrichment process is explained in detail. The resulting resource is evaluated in terms of its coverage of different types of MWEs. It is also compared with and evaluated against Finnish PropBank.
Many government schemes were unsuccessful because lack of proper feedback on the ongoing schemes, where billion dollars investment is going to be in vain. Sentiment analysis is one of best approach to analyse opinions of the peoples on various government schemes. Sentiment analysis and machine learning techniques emerged to analyse huge social media corpora to track people's views on government policies, products and services. Sentiment analysis process consists of various phases which include data discovery, data collection, data pre-processing, and data analysis. Stemming is a process to generate the morphemes in natural language sentences for various applications such as sentiment analysis, information retrieval, and domain analysis. The stemming process involved two major errors, which are over-stemming and under-stemming errors. Most of sentiment analysis natural languages processing applications used Lancaster and Porter stemming algorithms where more than one word inflected into same morpheme, which causes the etymology behaviour of the stemming word and prone to classify the tweets false positives and false negative. The proposed un-prejudice light stemming algorithm prevent etymology behaviour of morpheme and sustain its meaning during stemming process by selecting a word which has maximum number of synonyms in lexical database.
The present paper seeks to explore the phenomenon of linguistic creativity. Over the years there have been numerous theoretical and experimental studies on the topic of language and creativity. However, among the research papers few discuss linguistic creativity in cognitive and communicative aspects, which are of particular relevance to the current study. The findings are discussed in the light of cognitive-discursive approach. In today’s world violations of norms are manifested at all levels of a language and in almost all types of discourse. The existing standards determine the use of language tools in accordance with the rules of a language, its laws of register, genre, code, function, rules regarding the appropriateness of language units, their collocability, derivation, etc., as well as, in a broader sense, with the objectives of communication. Taking into account the latter statement, the question arises if a linguistic personality should prioritize the choice of preserving the linguistic norm or violate it in their lingua-creative activity to achieve a particular goal of communication. A lingua-creative personality while searching for a name to some innovative mental formations, those that have not yet been verbalized by linguistic means, either produces novel linguistic units and categories by further exploiting the productive potential of a language; or rethinks the existing models, bending the rules and norms of a language; or violates those rules and norms.
OBJECTIVE: To compare "virtual" unenhanced (VUE) computed tomography (CT) images, reconstructed from rapid kVp-switching dual-energy computed tomography (DECT), to "true" unenhanced CT images (TUE), in clinical abdominal imaging. The ability to replace TUE with VUE images would have many clinical and operational advantages. METHODS: VUE and TUE images of 60 DECT datasets acquired for standard-of-care CT of pancreatic cancer were retrospectively reviewed and compared, both quantitatively and qualitatively. Comparisons included quantitative evaluation of CT numbers (Hounsfield Units, HU) measured in 8 different tissues, and 6 qualitative image characteristics relevant to abdominal imaging, rated by 3 experienced radiologists. The observed quantitative and qualitative VUE and TUE differences were compared against boundaries of clinically relevant equivalent thresholds to assess their equivalency, using modified paired t-tests and Bayesian hierarchical modeling. RESULTS: Quantitatively, in tissues containing high concentrations of calcium or iodine, CT numbers measured in VUE images were significantly different from those in TUE images. CT numbers in VUE images were significantly lower than TUE images when calcium was present (e.g. in the spine, 73.1 HU lower, p < 0.0001); and significantly higher when iodine was present (e.g. in renal cortex, 12.9 HU higher, p < 0.0001). Qualitatively, VUE image ratings showed significantly inferior depiction of liver parenchyma compared to TUE images, and significantly more cortico-medullary differentiation in the kidney. CONCLUSIONS: Significant differences in VUE images compared to TUE images may limit their application and ability to replace TUE images in diagnostic abdominal CT imaging.
<h3>Introduction</h3><br> BOLT Egyptian Arabic Treebank -- Discussion Forum was developed by the Linguistic Data Consortium (LDC) and consists of Egyptian Arabic web discussion forum data with part-of-speech annotation, morphology, gloss and syntactic tree annotation. <br> The DARPA <a href="https://www.ldc.upenn.edu/collaborations/current-projects/bolt">BOLT</a> (Broad Operational Language Translation) program developed machine translation and information retrieval for less formal genres, focusing particularly on user-generated content. LDC supported the BOLT program by collecting informal data sources -- discussion forums, text messaging and chat -- in Chinese, Egyptian Arabic and English. The collected data was translated and annotated for various tasks including word alignment, treebanking, propbanking and co-reference. <br> The unannotated Egyptian Arabic source data is released as BOLT Arabic Discussion Forums (<a href="../../../LDC2018T10">LDC2018T10</a>). <br> The annotations in this release follow Penn Arabic Treebank (PATB) annotation guidelines. The PATB project consists of two distinct phases: (a) part-of-speech tagging which divides the text into lexical tokens and gives relevant information about each token such as lexical category, inflectional features and a gloss; and (b) Arabic treebanking, which characterizes the constituent structures of word sequences, provides categories for each non-terminal node and identifies null elements, co-reference, traces and so on. <br> There are two kinds of morphological analysis synchronized in the corpus. LDC Standard Morphological Analyzer (SAMA) Version 3.1 (<a href="../../../LDC2010L01">LDC2010L01</a>) was used for Modern Standard Arabic tokens, and CALIMA (Columbia Arabic Language and dIalect Morphological Analyzer) was used for Egyptian-Arabic tokens. <br> <h3>Data</h3><br> This release contains 440,448 tokens before clitics were split and 508,548 tree tokens after clitics were split for treebank annotation. The source material is web discussion forums collected by LDC from various sources. <br> Data is presented in a a variety of UTF-8 encoded text formats, specifically plain text, XML, tdf and Penn Treebank. See the included documentation for more information about the specific formats. <br> <h3>Acknowledgement</h3><br> This material is based upon work supported by the Defense Advanced Research Projects Agency (DARPA) under Contract No. HR0011-11-C-0145. The content does not necessarily reflect the position or the policy of the Government, and no official endorsement should be inferred. <br> <h3>Samples</h3><br> Please view the following samples: <br> <ul><br> <li><a href="desc/addenda/LDC2018T23-int.txt">Integrated</a></li><br> <li><a href="desc/addenda/LDC2018T23.tree">Penn Treebank</a></li><br> <li><a href="desc/addenda/LDC2018T23-pos.txt">POS</a></li><br> <li><a href="desc/addenda/LDC2018T23-su_xml.xml">SU Annotation</a></li><br> <li><a href="desc/addenda/LDC2018T23.tdf">SU tdf</a></li><br> <li><a href="desc/addenda/LDC2018T23.xml">Annotation Graph</a></li><br> </ul><br> <h3>Updates</h3><br> None at this time. </br> Portions © 2011-2018 Trustees of the University of Pennsylvania
With the tremendous success of deep learning models on computer vision tasks, there are various emerging works on the Natural Language Processing (NLP) task of Text Classification using parametric models. However, it constrains the expressability limit of the function and demands enormous empirical efforts to come up with a robust model architecture. Also, the huge parameters involved in the model causes over-fitting when dealing with small datasets. Deep Gaussian Processes (DGP) offer a Bayesian non-parametric modelling framework with strong function compositionality, and helps in overcoming these limitations. In this paper, we propose DGP models for the task of Text Classification and an empirical comparison of the performance of shallow and Deep Gaussian Process models is made. Extensive experimentation is performed on the benchmark Text Classification datasets such as TREC (Text REtrieval Conference), SST (Stanford Sentiment Treebank), MR (Movie Reviews), R8 (Reuters-8), which demonstrate the effectiveness of DGP models. © European Language Resources Association (ELRA), licensed under CC-BY-NC
BACKGROUND: How dental education influences students' dental and dentofacial esthetic perception has been studied for some time, given the importance of esthetics in dentistry. However, no study before has studied this question in a large sample of students from all grades of dental school. This study sought to fill that gap. The aim was to assess if students' dentofacial esthetic autoperception and heteroperception are associated with their actual stage of studies (grade) and if autoperception has any effect on heteroperception. METHODS: Between October 2018 and August 2019, a questionnaire was distributed to 919 dental students of all 5 grades of dental school at all four dental schools in Hungary. The questionnaire consisted of the following parts (see also the supplementary material): 1. Demographic data (3 items), Self-Esthetics I (11 multiple- choice items regarding the respondents' perception of their own dentofacial esthetics), Self-Esthetics II (6 Likert-type items regarding the respondents' perception of their own dentofacial esthetics), and Image rating (10 items, 5 images each, of which the respondents have to choose the one they find the most attractive). Both the self-esthetics and the photo rating items were aimed at the assessment of mini- and microesthetic features. RESULTS: The response rate was 93.7% (861 students). The self-perception of the respondents was highly favorable, regardless of grade or gender. Grade and heteroperception were significantly associated regarding maxillary midline shift (p < 0.01) and the relative visibility of the arches behind the lips (p < 0.01). Detailed analysis showed a characteristic pattern of preference changes across grades for both esthetic aspects. The third year of studies appeared to be a dividing line in both cases, after which a real preference order was established. Association between autoperception and heteroperception could not be verified for statistical reasons. CONCLUSION: Our findings corroborate the results of most previous studies regarding the effect of dental education on the dentofacial esthetic perception of students. We have shown that the effect can be demonstrated on the grade level, which we attribute to the specific curricular contents. We found no gender effect, which, in the light of the literature, suggests that the gender effect in dentofacial esthetic perception is highly culture dependent. The results allow no conclusion regarding the relation between autoperception and heteroperception.
The article analyzes differences in the description of discourse relations in corpus research, in particular with the reference to the use of discourse markers – expressions that tie together subsequent fragments of the text and provide information about the nature of these relations. The text presents three concepts of the description of explicitness and implicitness of the content: Rhetorical Structure Theory, Penn Discourse Treebank and the author’s original proposal and indicates the consequences of each solution. The analysis of relations with particles as metatexual expressions defined in accordance with The Nest Dictionary of Polish reveals the possibility of expressing explicitness as a representation of elements of informational structure shaped by the use of a given particle, and implicitness as a lack of representation of certain elements of this type.
Neural machine translation (NMT) models are typically trained using a softmax cross-entropy loss where the softmax distribution is compared against smoothed gold labels. In low-resource scenarios, NMT models tend to over-fit because the softmax distribution quickly approaches the gold label distribution. To address this issue, we propose to divide the logits by a temperature coefficient, prior to applying softmax, during training. In our experiments on 11 language pairs in the Asian Language Treebank dataset and the WMT 2019 English-to-German translation task, we observed significant improvements in translation quality by up to 3.9 BLEU points. Furthermore, softmax tempering makes the greedy search to be as good as beam search decoding in terms of translation quality, enabling 1.5 to 3.5 times speed-up. We also study the impact of softmax tempering on multilingual NMT and recurrently stacked NMT, both of which aim to reduce the NMT model size by parameter sharing thereby verifying the utility of temperature in developing compact NMT models. Finally, an analysis of softmax entropies and gradients reveal the impact of our method on the internal behavior of NMT models.
This journal article follows the research line opened on the search for semantic primes’ exponents in Old English within the frame of the Natural Semantic Metalanguage theory (Goddard 1997, 2012; Goddard and Wierzbicka 2002). The aim of this study is to complete the line of research on prime identification opened on the category Actions, events, movement, contact by establishing the Old English exponent of the prime DO. With this purpose, this paper discusses the adequacy of different OE verbs as possible prime exponent on the basis of textual frequency, morphology, semantics and syntactic complementation. Relevant data of analysis have been retrieved mainly from the lexical database of Old English Nerthus, the Dictionary of Old English (Healey et al. 2018) and the Dictionary of Old English Corpus (Healey et al. 2009).
Background/Context Inclusion of African immigrant youth voices in educational and research discourses remains rare despite the steady growth of this population in the United States over the past four decades. Consequently, the multilingual abilities of these youth remain typically unnoticed or ignored in the classroom, and little is specifically known about their histories, cultures, expectations, and achievements. Purpose Using the narrative inquiry approach and the Natural, Institutional, Discursive, Affinity, Learner, and Solidarity (NIDALS) theoretical lens, we explore the lived experiences of one African immigrant high school student in the midwestern United States. Research Design Using narrative inquiry, we qualitatively explored the lived cultural, racial, and ethnic identities and self-images experienced by a Ghanaian-born female high school student, Akosua (pseudonym), as she navigated and resisted identities ascribed to her in the midwestern U.S. Findings The student's narratives speak to issues of culture, identity, and self-image, as well as her literate life in multiple languages and literacy contexts in and out of school. The findings reveal narratives of ascribed identities, racialization, and perceived language hierarchies in the participant's daily life and indicate a need to challenge such narratives about African immigrant students and disrupt the reproduction of linguistic and racial inequality in the school system. Recommendations While school systems do follow state-sanctioned linguistic norms and ideologies, when educators draw on students’ experiences and funds of knowledge as resources already in the room in order to find ways of negotiating and disrupting language hierarchies and the ascribed identities they support, it allows all students, including multilinguals, to have their identity affirmed, even in school systems that have historically marginalized them. This, in turn, supports educational achievement, broadly realized, not only psychologically for all students but also economically and nationally for the country—a critical accomplishment in an era when educational quality in the U.S. is losing ground to foreign achievements.
The aim of this paper is to retrieve the most relevant expansion words for expanding the initial query of the user in order to enhance the outcomes of web search results. Query expansion plays a major role in reformulating a user’s initial query to a one more pertinent to the user’s intended meaning. The reformulated query is then used to obtain more appropriate outcomes from a large amount of information on the web. The proposed semantic query expansion technique uses Wikipedia and WordNet as data sources. Wikipedia is taken as a base for all query expansions because it is one of the most diversified and relevant databases available on the web. To further improve the proposed query expansion technique,WordNet—a lexical database—is used as the as another data source because the synonyms (synsets) of the query term provided by it can be quite useful for query expansion. The proposed expansion technique successfully combines the two data sources to retrieve the most relevant expansion terms from the data sources in response to the user’s original query. The proposed work has been divided into four phases: (1) extraction of relevant words from Wikipedia (2) extraction of relevant words from WordNet (3) merging of the expansion terms obtained from Wikipedia and WordNet, and (4) query formulation by combining the expansion terms using Boolean operators. This reformulated query is then fired on the web to find the desired result. The Experimental result shows a significant improvement in information retrieval using query expansion.
Using incorrect worked examples during mathematics instruction can improve student learning. However, teachers worry that students may confuse correct and incorrect examples over time, and memory research supports this fear. To examine if this forgetting occurs, we had undergraduates rate the correctness of correct and incorrect worked examples immediately and one week later (Experiment 1). Previously studied incorrect examples were rated as slightly more correct after the delay, but this did not affect ratings of unstudied examples or problem-solving accuracy. In Experiment 2, we more closely mimicked how incorrect worked examples are used in classroom settings. Again, we found only small changes in students’ memory for studied worked examples after the delay, and no changes for unstudied examples or problem-solving accuracy. Our findings suggest the costs of teaching with incorrect worked examples are limited to the specific studied problems, and do not affect learning of the underlying mathematical rule.
BACKGROUND: Instrumental activities of daily living (IADL) impairment can begin in mild cognitive impairment (MCI), and is the core criteria for diagnosing dementia in both Alzheimer's (AD) and Parkinson's (PD) diseases. The Functional Activities Questionnaire (FAQ) has high discriminative power for dementia and MCI in older age populations, but is influenced by demographic factors. It is currently unclear whether the FAQ is suitable for assessing cognitive-associated IADL in non-demented PD patients, as motor disorders may affect ratings. OBJECTIVE: To compare IADL profiles in MCI patients with PD (PD-MCI) and AD (AD-MCI) and to verify the discriminative ability of the FAQ for MCI in patients with (PD-MCI) and without (AD-MCI) additional motor impairment. METHODS: Data of 42 patients each of PD-MCI, AD-MCI, PD cognitively normal (PD-CN), and healthy controls (HC), matched according to age, gender, education, and global cognitive impairment were analyzed. ANCOVA and binary regressions were used to examine the relationship between the FAQ scores and groups. FAQ cut-offs for PD-MCI (versus PD-NC) and AD-MCI (versus HC) were separately identified using receiver operating characteristic analyses. RESULTS: FAQ total score did not differentiate between MCI groups. PD-MCI subjects had greater difficulties with tax records and traveling while AD-MCI individuals were more impaired in managing finances and remembering appointments. Classification accuracy of the FAQ was good for diagnosing AD-MCI (69%, cut-off ≥1) compared to HC, and sufficient for differentiating PD-MCI (38.1%, cut-off ≥3) from PD-CN. CONCLUSION: The FAQ task profiles and classification accuracy differed between MCI related to PD and AD.
This paper analyses data to address a specific linguistic problem, i.e. the acquisition of the modification potential of the three more or less synonymous Dutch degree modifiers heel, erg and zeer, all meaning 'very', which show syntactic differences in modification potential. It continues the research reported on in The analysis makes crucial use of linguistic applications developed in the CLARIN infrastructure, in particular the treebank search applications PaQu (Parse and Query) and GrETEL Version 4.00. The analysis benefits from the use of parsed corpora (treebanks) in combination with the search and analysis options offered by PaQu and GrETEL. Earlier work showed that despite little data for zeer modifying adpositional phrases adult speakers end up with a generalised modification potential for this word. In this paper, I extend the dataset considered, and find more (but still little) data for this phenomenon. However, I also find a similar amount of data that form counterexamples to the non-generalisation of the modification potential of heel. I argue that the examples with heel concern constructions with idiosyncratic semantics and therefore are not counted as evidence for the general rule of modification. I suggest a simple statistical analysis to account for the fact that children 'learn' that heel cannot modify verbs or adpositions though there is no explicit evidence for this and they are not explicitly taught so.
Scene graph is a graph representation that explicitly represents high-level semantic knowledge of an image such as objects, attributes of objects and relationships between objects. Various tasks have been proposed for the scene graph, but the problem is that they have a limited vocabulary and biased information due to their own hypothesis. Therefore, results of each task are not generalizable and difficult to be applied to other down-stream tasks. In this paper, we propose Entity Synset Alignment(ESA), which is a method to create a general scene graph by aligning various semantic knowledge efficiently to solve this bias problem. The ESA uses a large-scale lexical database, WordNet and Intersection of Union (IoU) to align the object labels in multiple scene graphs/semantic knowledge. In experiment, the integrated scene graph is applied to the image-caption retrieval task as a downstream task. We confirm that integrating multiple scene graphs helps to get better representations of images.
This paper represents the development of the Myanmar Named Entity Recognition (NER) system using Conditional Random Fields (CRFs). In order to develop the system, a manually annotated Named Entities (NEs) corpus - collected from Myanmar news websites and Asia Language Treebank(ALT)-Parallel-Corpus has been used. We compare the performance of the system getting syllable-based input to the one getting character-based input. We observed that training data has more impact on the performance of the system. The experimental results show that the syllable-based system performs better than the character-based system. It achieves that Precision, Recall and F1-score values of 93.62%, 91.64% and 92.62% respectively.
OBJECTIVE: To design and evaluate the effectiveness of a stimulus material in eliciting the N400 event related potential (ERP). DESIGN: A set of 700 semantically congruent and incongruent sentences was developed in accordance with current linguistic norms, and validated with an electroencephalography (EEG) study, in which the influence of age and gender on the N400 ERP magnitude was analysed. STUDY SAMPLE: Forty-five normal-hearing subjects (19-57 years, 21 females) participated in the EEG study. RESULTS: The stimulus material used in the EEG study elicited a robust N400 ERP, with a morphology consistent with the literature. Results also showed no statistically significant effect of age or gender on the N400 magnitude. CONCLUSIONS: The material presented in this paper constitutes the largest complete stimulus set suitable for both auditory and text-based N400 experiments. This material may help facilitate the efficient implementation of future N400 ERP studies, as well as promote standardisation and consistency across studies.
The research work is dealt with the culture of speech. Culture of speech is identified by language of speech, social surrounding, language and psycology, language and pragmatics. As a result of culture of speech, personal culture, human quality, linguistic knowledge of a person is realized. Several types of antropolinguistic analyses are mentioned. Public speech (in auditorium, in crowd) shows the social aspect of communicative, linguistic norms and humans morality in speaking expresses wisdom and psycholinguistic aspect of speaker. Literary norm, functional grammar, cognitive pragmatics, literary language are thoroughly explained.
Increasing popularity of electronic dictionaries, ontologies, thesauri and lexical databases makes them an effective tool for language learning purposes (Dash, 2013; Fellbaum, 2010; Miller & Fellbaum, 1992; Shimodaira et al, 2006; Sun et al, 2011). The aim of this research is to study the educational potential of electronic lexical database for the English language WordNet (Miller, 1995; Fellbaum, 1998) and electronic thesauri for the Russian language RuWordNet (Loukachevitch, 2011; Loukachevitch, Lashevich, 2016) in teaching English as a foreign language. In this research the authors focus on teaching colours, in particular a colour term white, as colours represent complex linguistic and culture-specific phenomena, reflected in the “cultural memory” of people, the very concept of ‘colour’ being a function of language and culture (Wierzbicka, 2006). Thus, understanding colours helps students both study a foreign language and learn its history and culture.This is a mixed method study based on comparison of synonyms for colour term white in WordNet and RuWordNet. The main relation among words in these thesauri is synonymy. Based on their meanings, words are grouped into unordered sets of synonyms expressing one underlying concept (synsets). This allows to consider specific senses of words and semantic relations. Firstly, Russian students studying English as a foreign language were asked to analyse the meanings and synonyms for colour term white in RuWordNet. Secondly, the students compared the representation of colour term white in WordNet. Additionally, they examined set phrases with the adjective white in English dictionaries. Then a qualitative method (interviewing) was used to reveal the students’ perceptions of studying English by means of electronic dictionaries, thesauri and lexical databases.The study allowed to claim that electronic thesauri and lexical databases are of educational value, increasing students’ linguistic awareness and language proficiency.
BACKGROUND: Analyze intrarater and interrater reliability for evaluating endoscopic images of velopharyngeal (VP) physiology. METHOD: Speakers produced 9 speech stimuli representing 4 stimulus types: sustained phonemes, repetitions of "puh," single words, and short phrases. The 37-speaker participants included 16 patients with VP dysfunction and 21 control participants. Five raters independently rated the video images for degree of VP opening, location of opening, and pattern of closure. Outcome measures included intrarater and interrater measures of reliability and the effects of raters and stimulus type on ratings. RESULTS: Intrarater reliability was acceptable, and ratings were logically consistent. Fixed effects regression coefficients for the patient and the control groups showed that raters were a significant source of variability for degree of opening and pattern of closing. Stimulus type was not a significant source of variation for any metric for the controls, but stimulus type was a significant determinant for degree of opening for patients. The degree of opening was larger for sustained phonemes than for the other speech stimuli. Ratings for degree of opening were most similar for repeated "puh." CONCLUSIONS: Interrater reliability needs to be improved so that the assessment procedure produces more consistent findings among clinicians, thus strengthening our evidence base for this procedure. Interrater additional research is needed to understand how the stimulus affects ratings of VP physiology, to identify stimuli that yield the most useful clinical information, and to understand how training affects the ratings of VP physiology.
Our work on the automatic detection of English discourse connectives in the Penn Discourse Treebank (PDTB) shows that syntactic information from the Universal Dependencies (UD) framework is a viable alternative to that from the Penn Treebank (PTB) framework. In fact, we found minor increases when comparing between the use of gold standard PTB part-of-speech (POS) tag information and automatically parsed UD information. The former has traditionally been used for the task but there are now much more UD corpora and in many more languages than that available in the PTB framework. As such, this finding is promising for areas in discourse parsing such as in multilingual as well as under production settings, where gold standard PTB information may be scarce.
In order to effectively respond to the increased linguistic and cultural diversity in the U.S. schools and close the consistently documented achievement gap between culturally and linguistically diverse (CLD) students and mainstream students, teachers need to take an asset-based approach and be able to draw on CLD students’ entire funds of linguistic knowledge. However, few studies have examined CLD students’ linguistic choices in multiple discursive spaces with different linguistic norms, values and practices. This article addresses this research gap through a case study of Elif, a Turkish American student and her linguistic boundary crossing experiences within and across three discursive spaces: her home, her Turkish heritage language school, and her mainstream school. Through in-depth analysis of interviews, observations, and field notes, the study revealed that Elif experienced different linguistic environments and boundary types. She negotiated experiences that ranged from smooth to managed to insurmountable boundaries. Finally, translanguaging practices acted as a key boundary object that mediated sociocultural discontinuities in the Turkish heritage language school, and facilitated Elif’s experiences between Turkish dominant and English dominant discursive spaces.
Corporate credit ratings (CRs) are closely related to companies’ cost of debt financing. Recent research has drawn wide attention to how nonfinancial as well as financial factors may affect ratings. By manually collecting information about the profiles of chief financial officers (CFOs) of US companies, we examine the effect of CFOs’ accounting expertise on corporate CRs. The results show that firms with accounting expert CFOs are more likely to receive higher CRs and that the effect of CFOs’ accounting expertise on the ratings is more pronounced for firms with higher default risk, suggesting that the accounting expertise of CFOs may be an important factor that affects CRs. Moreover, we find a dynamic relation between accounting expert CFOs and CRs such that a downgrade in a firm’s CR in a prior year affects the subsequent selection of an accounting expert CFO.
Neural sequence model, though widely used for modeling sequential data such as the language model, has sequential recency bias (Kuncoro et al. 2018) to the local context, limiting its full potential to capture long-distance context. To address this problem, this paper proposes augmenting sequence models with a span-based neural buffer that efficiently represents long-distance context, allowing a gate policy network to make interpolated predictions from both the neural buffer and the underlying sequence model. Training this policy network to utilize long-distance context is however challenging due to the simple sentence dominance problem (Marvin and Linzen 2018). To alleviate this problem, we propose a novel training algorithm that combines an annealed maximum likelihood estimation with an intrinsic reward-driven reinforcement learning. Sequence models with the proposed span-based neural buffer significantly improve the state-of-the-art perplexities on the benchmark Penn Treebank and WikiText-2 datasets to 43.9 and 35.2 respectively. We conduct extensive analysis and confirm that the proposed architecture and the training algorithm both contribute to the improvements.
Transformer-based pre-trained language models (PLMs) have dramatically improved the state of the art in NLP across many tasks. This has led to substantial interest in analyzing the syntactic knowledge PLMs learn. Previous approaches to this question have been limited, mostly using test suites or probes. Here, we propose a novel fully unsupervised parsing approach that extracts constituency trees from PLM attention heads. We rank transformer attention heads based on their inherent properties, and create an ensemble of high-ranking heads to produce the final tree. Our method is adaptable to low-resource languages, as it does not rely on development sets, which can be expensive to annotate. Our experiments show that the proposed method often outperform existing approaches if there is no development set present. Our unsupervised parser can also be used as a tool to analyze the grammars PLMs learn implicitly. For this, we use the parse trees induced by our method to train a neural PCFG and compare it to a grammar derived from a human-annotated treebank.
Abstract Syntactic parsing is an important topic in the field of Mongolian language information processing. Compared with English and Chinese dependency parsing, Mongolian dependency parsing is still at the beginning stage. Mongolian syntactic parsing lack of Treebank resources seriously, under such conditions, a high quality syntactic parser cannot be developed by statistical methods simply. Aiming at the characteristics that Mongolian language has rich morphological features, this paper presented a rule and statistics-based dependency parsing model using Mongolian Dependency Treebank as training and evaluation data. The morphological and syntactic rules are represented using complex features and unification operations. The statistical model is represented using lexical dependent probability. This model has now achieved accuracies of 77.18%, 69.42% and 95.44% for the unlabelled annotation score, the labeled annotation score and the head word annotation score respectively.