Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
The article attempts to clarify the role of linguistic games as a means of forming communicative competence of primary school students, which helps maintain interest in learning the Ukrainian language and aims to provide students with knowledge through their own efforts, revealing the structure and functioning of language. Linguistic play promotes the development of primary school students' idea of the basic grammatical categories and grammatical meaning of the word, which as abstract concepts are very difficult to perceive at this age. It also helps to understand the process of forming new words and master the basic word-formation patterns. Playful learning of the word promotes the development of the general speech culture of the individual, forms a subconscious understanding of the language norm and develops linguistic thinking, which is the basis of communicative competence. Systematization of teachers' experience presented on personal sites, in journal articles and methodical publications allowed to identify the most popular linguistic games used by teachers in Ukrainian language lessons in primary school, and divide them into the following groups: phonetic, lexical-phraseological, morphological, orthographic, syntactic. Given the fact that younger students are characterized by emotionality, cognitive activity, desire to fantasize, the desire to show their intelligence and skill, the role of using linguistic games in Ukrainian lessons as a means of forming communicative competence. Modern linguistic games involve competition, voluntariness, which allows students to meet their needs for mobility, get positive emotions of joy, admiration, satisfaction, realize their interests, master communication skills and more. During linguistic games, primary school students not only learn language norms and rules, but also get involved in a system of various subject-subject relations, form the ability to interact in a team, make decisions, prove their point, which will contribute to the formation of communicative competence. Key words: linguistic game, communicative competence of junior schoolchildren, phonetic games, lexical-phraseological games, morphological games, spelling games, syntactic games.
The new stage of Ukrainian literary language development began in the late 20th century, because state-making processes had legitimated the status of the Ukrainian language due to the law, following the main goal – to facilitate self-expression of the national genotype, raising prestige and full-fledged functioning of the state language in independent Ukraine. Since the reality is formed on the basis of personal activity, one of the leading tasks is upbringing the need of every speaker to use the state language fluently as the means for communication, as the means of forming intellectual culture, national self-awareness, which thus has direct impact on mentality and moral qualities of a person. The situation concerning the languages in the communicative environment of Ukraine obliges all speakers to accept the issues of the language and the speaking in a systematic and complex way. Command of a normative native language is the assignment for every aware citizen, who is obliged to know how to use the entire lexical heritage. A language is a weighty part of professional competency, a sensitive indicator of general culture, that is why every speaker should care about the high culture of their language. Perfect usage of a language becomes an important component of training experts in any field, particularly in the field of state governing, since the use of a language promotes their self-expression. Official activity definitely requires not only professionalism, but also thorough language competency. Our research is actual because we constantly need to work on the problems of language culture in the field of state governing, since, as a social phenomenon, the Ukrainian language reflects precise historical peculiarities, typical to a certain social and historical period: the development of new word constructions, the emergence of new words, a number of borrowings from other languages, etc. The purpose of our investigation is to find out and analyze the violations of lexical and semantic norms in official and business communication of state officials, to justify the ways and means to correct the violations of the lexical and semantic norm. Reaching the set purpose meant carrying out the following tasks: to analyze the official and business language of state officials; to single out the most spread lexical and semantic mistakes and downsides in speaking, to give recommendations on eliminating mistakes in order to improve the speaking culture of state officials in the field of their professional activity. In the investigation, there is applied a wide range of contemporary methods and approaches of research: language facts are considered from the position of the functional approach; by means of the methods of the generalization and classification analysis the types of mistakes met in the language of state officials were singled out; the comparing and contrasting analysis has allowed to find out the facts of interfering influence of the Russian language on Ukrainian, to single out the types of the interference consequences at the lexical and semantic level. In the proposed research, there is the analysis of the official and business language of state officials, different consequences of interferencial interaction of closely native languages (Ukrainian and Russian) are realised and typified by comparing the language of state officials to the current norms of the modern Ukrainian literary language. The most spread lexical mistakes and downsides are found to refer to: not motivated usage of words which results in “surgick”; usage of so-called words-parasites with no need; usage of the words which are inappropriate from the point of view of the literary norm and the etiquette rules; abuse of the words derived from foreign languages, especially from English; irrelevant tautology in oral speech, redundancy of words ( pleonasm ); confusing paronyms. Due to the analyzed language material there are linguistic explanations and recommendations on the ways and means of preventing, correcting and eliminating realised mistakes in the language of state officials. It is proved that state governors are obliged to obey communicative features of language culture, namely: its correctness, accuracy, logic, purity, pithiness and relevance. R
I describe several new efficient algorithms for querying large annotated corpora. The search algorithms as they are implemented in several popular corpus search engines are less than optimal in two respects: regular expression string matching in the lexicon is done in linear time, and regular expressions over corpus positions are evaluated starting in those corpus positions that match the constraints of the initial edges of the corresponding network. To address these shortcomings, I have developed an algorithm for regular expression matching on suffix arrays that allows fast lexicon lookup, and a technique for running finite state automata from edges with lowest corpus counts. The implementation of the lexicon as suffix array also lends itself to an elegant and efficient treatment of multi-valued and set-valued attributes. The described techniques have been implemented in a fully functional corpus management system and are also used in a treebank query system.
The preservation and expansion of the functioning sphere of the national languages of Russia is a requirement of modern times. Translation plays an important role in this, while, due to systemic discrepancies, literary and customary norms of the language occupy a dominant role. Translation is highly developed in the Republic of Sakha (Yakutia), works of oral folk art, fiction, some texts of the official business style are being translated, and the theoretical foundations of translation are being developed. Nevertheless, microtoponyms as an object of study of the private theory of Russian-Yakut translation have not been sufficiently studied, which determined the aim of the current study — the analysis of the phonetic, lexical and grammatical features of the models for creating nomenclature clichés in the modern Yakut language based on the materials of the translation of bus stops into the Yakut language. The work is an attempt, for the first time, to consider models for translating microtoponyms, in particular, the names of bus stops. The subject of the study was deliberately chosen: when sound accompaniment in the Yakut language was introduced in the capital’s transport routes, the public was outraged at the quality of the translation of the names of bus stops, which determined its editing. The material is the second translation, made by staff at the Institute of Humanities and Social Sciences of the Siberian Branch of the Russian Academy of Sciences, edited by teachers of the department of stylistics of the Yakut language and Russian-Yakut translation of the NEFU named after M. K. Ammosov. The methods of continuous sampling, comparative and descriptive methods were used. The names of bus stops reflect important information about the history and culture of the locality. The stops in Yakutsk are named according to the objects in close proximity to the stops: territorial zones, streets, less often historical events and dates. The analysis has shown that the translation of bus stops names is primarily aimed at an easy perception by the recipient. Phonetic methods of translation are present: transliteration and transcription in combination with equivalent substitution, addition, amplification; lexical ones are represented by equivalent and adequate substitutions, tracing, concretization, generalization, addition, amplification and explication; grammatical — by permutation and replacement.
<h3>Introduction</h3><br> BOLT Egyptian Arabic Treebank - Conversational Telephone Speech was developed by LDC and consists of Egyptian Arabic conversational telephone speech data with part-of-speech annotation, morphology, gloss and syntactic tree annotation. <br> The DARPA <a href="https://www.ldc.upenn.edu/collaborations/current-projects/bolt">BOLT</a> (Broad Operational Language Translation) program developed machine translation and information retrieval for less formal genres, focusing particularly on user-generated content. LDC supported the BOLT program by collecting informal data sources -- discussion forums, text messaging and chat -- in Chinese, Egyptian Arabic and English. The collected data was translated and annotated for various tasks including word alignment, treebanking, propbanking and co-reference. <br> The annotations in this release follow Penn Arabic Treebank (PATB) annotation guidelines. The PATB project consists of two distinct phases: (a) part-of-speech tagging which divides the text into lexical tokens and gives relevant information about each token such as lexical category, inflectional features and a gloss; and (b) Arabic treebanking, which characterizes the constituent structures of word sequences, provides categories for each non-terminal node and identifies null elements, co-reference, traces and so on. <br> There are two types of morphological analysis synchronized in the corpus. LDC Standard Morphological Analyzer (SAMA) Version 3.1 (<a href="../../../LDC2010L01">LDC2010L01</a>) was used for Modern Standard Arabic tokens, and <a href="http://www.aclweb.org/anthology/W12-2301">CALIMA</a> (Columbia Arabic Language and dIalect Morphological Analyzer) was used for Egyptian-Arabic tokens. <br> <h3>Data</h3><br> This release contains 153,171 tokens before clitics were split and 182,965 tree tokens after clitics were split for treebank annotation. The source data was selected from conversational telephone speech collected by LDC for the CALLHOME project that was transcribed and segmented into sentence units. <br> Data is presented in a variety of UTF-8 encoded text formats, specifically plain text, XML, tdf and Penn Treebank. See the included documentation for more information about the specific formats. <br> <h3>Sponsorship</h3><br> This material is based upon work supported by the Defense Advanced Research Projects Agency (DARPA) under Contract No. HR0011-11-C-0145. The content does not necessarily reflect the position or the policy of the Government, and no official endorsement should be inferred. <br> <h3>Samples</h3><br> Please view the following samples. <br> <ul><br> <li><a href="desc/addenda/LDC2021T12.integrated.xml.txt">Integrated XML</a></li><br> <li><a href="desc/addenda/LDC2021T12.xml">SU XML</a></li><br> <li><a href="desc/addenda/LDC2021T12.vowel.tree">Penn Treebank</a></li><br> <li><a href="desc/addenda/LDC2021T12.novowel.tree">Penn Treebank without vowels</a></li><br> <li><a href="desc/addenda/LDC2021T12.tdf">SU TDF</a></li><br> <li><a href="desc/addenda/LDC2021T12.before.txt">POS Before (TXT)</a></li><br> <li><a href="desc/addenda/LDC2021T12.after.txt">POS After (TXT)</a></li><br> </ul><br> <h3>Updates</h3><br> None at this time. </br> Portions © 1997, 2002, 2011-2021 Trustees of the University of Pennsylvania
This study aimed to characterize factors that influence early dialect development in a language environment with multiple dialects. Children were evaluated for these dialect effects compared with normal hearing referenced measures of speech and language development that are commonly implemented in hearing-impaired children. Dialect exposure and use were assessed longitudinally in Chinese children (2-6 years old) that were raised in a community where Putonghua (PTH) and Sichuanhua (SCH) Mandarin dialects were used. Lexical tones in these dialects are different. A total of 20 boys and 20 girls (2 years old at the beginning of the study) that attended the same nursery school were included in this study. SCH was used by the majority of subjects <4 years old. The majority of subjects >4 years old used either dialect, with a few users of both dialects at this age. PTH tone perception did not differ significantly as a function of dialect use. Tone recognition and discrimination were >90% accurate by 6 years old, in contrast to previous results for children with minimal exposure and use of PTH. Children with approximately ⩾50% PTH exposure might be accurately assessed with norm-referenced speech materials spoken in PTH, regardless of their preferred dialect. However, the current norm-referenced assessments of children with minimal PTH exposure and nonusers of the dialect might be inaccurate.
The aim of this paper is to present the utility of the Gorazd: An Old Church Digital Hub for scholars working with Old Romanian and Slavonic texts written on the territory of today’ s Romania. The Gorazd Project was realized during the years 2016–2020 and it includes an Old Church Slavonic Card Index and three Old Church Slavonic lexical databases, among which the largest one is represented by the digitized and updated version of the monumental Lexicon linguæ palæoslovenicæ (vol. I–IV, 1958–1997) composed by the Institute of Slavonic Studies of the Czech Academy of Sciences. As the Gorazd Project uses English as meta-language, its application is not limited to narrowly specialized Slavic philologists, but it is also open for scholars of neighbouring fields. The dictionaries within the Gorazd Digital Hub can serve as a reference tool not just for the oldest attested Slavonic vocabulary and its semantics, but also for the biblical concordance of the Slavonic oldest Bible redaction and the oldest attested Old Church Slavonic morphological forms.
In this article, we present our methodologies for SemEval-2021 Task-4: Reading Comprehension of Abstract Meaning. Given a fill-inthe-blank-type question and a corresponding context, the task is to predict the most suitable word from a list of 5 options. There are three sub-tasks within this task: Imperceptibility (subtask-I), Non-Specificity (subtask-II), and Intersection (subtask-III). We use encoders of transformers-based models pre-trained on the masked language modelling (MLM) task to build our Fill-in-the-blank (FitB) models. Moreover, to model imperceptibility, we define certain linguistic features, and to model non-specificity, we leverage information from hypernyms and hyponyms provided by a lexical database. Specifically, for non-specificity, we try out augmentation techniques, and other statistical techniques. We also propose variants, namely Chunk Voting and Max Context, to take care of input length restrictions for BERT, etc. Additionally, we perform a thorough ablation study, and use Integrated Gradients to explain our predictions on a few samples. Our best submissions achieve accuracies of 75.31% and 77.84%, on the test sets for subtask-I and subtask-II, respectively. For subtask-III, we achieve accuracies of 65.64% and 62.27%. The code is available here.
The problem of divergence and conflicts between literary norms is considered a particularly significant problem that arose during the formation of the modern Belarusian literary language. Practically, there are two types of literary Belarusian used in today’s Belarusian society. One is the official literary norm formed during the Soviet era (after 1933, to be precise) and used for many years in administration, education, and publication inside Belarus. The other is the literary norm Taraškievica, which was supported by some Belarusians, including intellectuals, from the Perestroika period. Taraškievica is a literary norm of modern Belarusian that was originally popularized in the 1920s and is named for the linguist Branislaŭ Taraškievič, the author of “Belarusian Grammar for Schools” (Vilnius, 1918), which was the foundation for the norm. These two literary norms are currently at odds in Belarusian society over which is the “authentic” Belarusian literary language. This article explores how divergence and conflicts between two Belarusian literary norms have emerged and how the two linguistic norms are regarded as “authentic” by their supporters. First, I overview the history of modern literary Belarusian from the perspective of corpus planning in language policy. Second, I analyze linguistic ideologies among the proponents of official norms and Taraškievica by focusing on the meta-linguistic discourse of the supporters of two standard language norms. Although Belarusian was exclusively used as a spoken language by peasants or poor szlachtas at the end of the eighteenth century, it began to be used for literary works starting in the nineteenth century. At the same time, linguistic research on the Belarusian language (to be precise, Belarusian dialects of Russian) was also promoted in the Russian Empire. In 1905 the tsarist government for the first time officially permitted publishing activity in Belarusian. After the collapse of the tsarist regime, with “Belarusian Grammar for Schools” published by B. Taraškievič in 1918, the orthographic and grammatical norms presented in this book were widely endorsed by the Belarusian intellectuals of that time and quickly brought into everyday linguistic practice. After the establishment of the Byelorussian Soviet Socialist Republic (BSSR) in 1919, it was the Soviet government, rather than individuals such as writers, linguists, and social activists, that guided the formation of the Belarusian literary language. Within the BSSR, the Institute of Belarusian Culture (later the Belarusian Academy of Sciences) played a leading role in the compilation of scientific terminology and Belarusian dictionaries, development of orthography, and other projects. When the Stalinist regime began in the 1930s, however, the linguists of the Academy of Sciences, who had led the standardization of Belarusian, were successively purged. In 1933, the BSSR’s Council of People’s Commissars implemented the orthographic reform that rendered the Belarusian orthographic norms, and some grammatical norms, closer to the Russian literary language. While the official orthography based on this reform was adopted in Soviet Belarus in 1934, the reform was accepted neither by the Belarusians in Western Belarus (then under the rule of Poland) nor by Belarusian communities abroad. They continued to use the pre-reform literary norm, i.e., Taraškievica. In the Perestroika period, with political and social liberalization in Soviet Belarus, the use of the Belarusian language based on Taraškievica gradually began to flow into the country from the expatriate Belarusian community. It rapidly gained popularity among the intellectuals as a more “authentic” Belarusian literary norm. Nowadays, while official literary norms of Belarusian are used in the field of administration, education, and mass media, Taraškievica is also widely used in the more personal sphere of language use, such as blogs and social networking sites. Analysis of the meta-linguistic discourse among the proponents of each norm shows that the proponents of Taraškievica have a linguistic ideology emphasizing the “authenticity” of Taraškievica in terms of its legitimacy and significance as a “symbol of the community.” Simultaneously, they try to highlight the authenticity of Taraškievica by repeatedly denying the legitimacy of the official norm as a national symbol. Supporters of the official norm, meanwhile, embrace a linguistic ideology placing a particular emphasis on authenticity for Belarusians within Belarus. In proving the “authenticity” of the Belarusian literary language, they are inclined to argue that the legitimacy of the official norm has persistently thrived within Belarus, denouncing Taraškievica as having developed outside Belarus.
AbstractData were checked for univariate outliers using standardised scores (<i>z</i> > ± 3.29) and for multivariate outliers using the Mahalanobis distance test (<i>p</i> <.001; Tabachnick & Fidell, 2019). Data were also examined for the parametric assumptions that underlie within-subjects ANOVA (Tabachnick & Fidell, 2019), other than the behavioural data from video analysis, which were derived from frequency counts. Where the assumption of sphericity was violated, Greenhouse–Geisser-adjusted <i>F</i> tests were used. Initial analyses employed repeated-measures (RM) 2 (Load) × 3 (Tempo) (M)ANOVAs for the three psychological measures (i.e., RSME, NASA-TLX, and Affect Grid) and cardiac measures (HR and HRV indices). Additionally, exploratory analyses were conducted using a mixed-model approach, adopting the between-subject factors of personality (introvert vs. extrovert), sex (women vs. men), and age group (young adults vs. middle-aged adults). Significant <i>F</i> tests were followed up with pairwise/multiple comparisons, or in the case of interaction effects, examination of 95% confidence intervals (95% CIs) to identify where differences lay. Behavioural data were collated for the urban environment (high load) simulation under the following categories: (a) video data pertaining to four triggers (pedestrian, garbage truck, traffic lights, and vehicle cutting); (b) and simulator-derived data from the accelerator and brake pedal positions (i.e., 0 = no pressure applied, 1 = maximum braking); (c) mean speed (mph), and (d) course completion time (min). For the highway environment (low load), simulator data were collated for: (a) accelerator and brake pedal positions; (b) mean speed, and (c) completion time. From among these data, where parametric assumptions were not met and transformations would not serve to normalise the distribution, nonparametric analyses was adopted using rank-based, nonparametric tests. Specifically, the Wald-type statistic (WTS) and the ANOVA-type statistic (ATS) was computed within the nparLD package (Noguchi et al., 2012) of data analysis software R. In the absence of the load factor for the trigger and pedal data, a within-subjects, one-way ANOVA for the effect of music tempo was computed. In the exploratory analyses, a factorial approach was used, with a series of mixed-model ANOVAs 3 ([Tempo] × 2 [Personality], 3 [Tempo] × 2 [Sex], and 3 [Tempo] × 2 [Age Group]). Note that the main effect of tempo for trigger and pedal data is relevant to the main analysis but in the interest of parsimony is incorporated within exploratory factorial analyses.<br>Detailed Description of Data FileThis SPSS data file contains the demographic data (i.e. sex, age, age group [1 = young adult, 2 = middle-aged adult], personality [1 = introvert, 2 = extrovert]) for each of the 46 participants (presented with one participant per row). Behavioural measures relating to the driving simulation are included. These include the elapsed time (mins) for each trial. Also, mean speed (mph), brake pedal use (i.e. 0 = no pressure applied, 1 = maximal braking), accelerator pedal use (i.e., 0 = no pressure applied, 1 = maximal acceleration) and risk ratings (on a scale from 1 [<i>safe driving</i>] to 4 [<i>reckless driving</i>]). Note that these performance-related measures appear 12 times in total; that is for each simulator trigger (i.e. a pedestrian who walked at 5 km/h across a zebra crossing, a garbage truck that moved slowly in the left-hand lane and prompted an overtaking manoeuvre, traffic lights that changed to red, a slow vehicle on a stretch of road on which overtaking was prohibited and a vehicle that cut across unexpectedly at a four-way intersection) across all three high-load (urban) conditions. Additionally, the scores across all conditions for the measures of the NASA Task Load Index (NASA-TLX), Affect Grid (affective valence and affective arousal), Rating Scale Mental Effort (RSME) and wordsearch task are included. The psychophysiological measures of heart rate variability (HRV) and mean heart rate (HR) are also included. For HRV and HR, specifically, we present mean HR, minimum HR, maximum HR, standard deviation of normal RR intervals (SDNN), HR standard deviation and root mean square of successive differences (RMSSD). <i>z</i>-scores (i.e. standardised scores) for each variable are also included. Note that each participant was exposed to six experimental conditions (high load/fast tempo, high load/slow tempo, high load/no music, low load/fast music, low load/slow music and low load/no music). Accordingly, the measures that pertain to each trial (i.e. NASA-TLX, RSME, Affect Grid, wordsearch task, HRV indices, risk ratings, mean speed, brake pedal use and accelerator pedal use) appear six times in the data file.
Despite the increasing number of U.S. born Latinos, placing heritage and native speakers in the Spanish curriculum is still a challenge (MacGregor-Mendoza & Moreno, 2020). The present article (a) addresses the unique needs of heritage speakers in the Spanish curriculum; (b) problematizes traditional grammar-based placement exams; and (c) describes a multiple-choice placement exam (free upon request) designed and used at Georgia State University (GSU), a major urban university in the Southeastern U.S. Taking a sociolinguistic approach to the dialectical nature of Spanish, the GSU Spanish Language Program Coordinator developed a placement test based on what students—heritage, native, and non-native—do when asked to perform language tasks. The placement test design is outlined using distinctions of linguistic norms, both local/ regional and general. Reference is made to the ways in which diverse types of Spanish speakers align linguistically with general Spanish. This essay responds to the call for language standardization studies that recognize diglossia within a single named language by examining the role of heteroglossia to challenge monolingual language standardization ideologies (McLelland, 2021). Pedagogical implications for identifying and placing K-16 learners in a meaningful Spanish for Heritage Speakers classroom are discussed.
Abstract This article contains an edition of a 17th-century Low German letter addressed to the German congregation in Stockholm. Seen against the shift of writing language from Low to High German, this letter is analyzed in respect to code-mixing, which is shown to fulfill a communicative function. Furthermore the author here suggests that the code-mixing observed in the letter can be described as congruent lexicalization, where Low German syntactic structures are filled with both Low and High German lexical material.
The mirror effect is the finding that in recognition tests, a manipulation that increases the hit rate also decreases the false alarm rate. For example, low frequency words have a higher hit rate and a lower false alarm rate than high frequency words. Because the mirror effect is held to be a regularity of memory, it has had a pronounced influence on theories of recognition. We took advantage of the recent increase in the number of linguistic databases to create sets of stimuli that differed on one dimension (contextual diversity, frequency, or concreteness) but were more fully equated on other dimensions known to affect memory. Experiment 1 (contextual diversity), Experiment 3 (frequency), and Experiment 5 (concreteness) found no evidence of a mirror effect. We also conducted parallel experiments which used previously published stimuli that could not avail of the new databases and which therefore contained confounds. Experiment 2 (contextual diversity), Experiment 4 (frequency), and Experiment 6 (concreteness) all resulted in mirror effects. If this pattern of results is replicable, it has broad implications for theories of recognition, which typically view the mirror effect as a benchmark finding. Unfortunately, few articles on the mirror effect include the stimuli, rendering the past literature of little use in testing this hypothesis. We encourage researchers to create and assess other pools of highly controlled stimuli to establish whether the stimulus-based mirror effect obtains when confounds are eliminated or whether it is due to the presence of these confounds. (PsycInfo Database Record (c) 2023 APA, all rights reserved).
WordNet is a large lexical database of English. Nouns, verbs, adjectives, and suffixes are grouped into a set of cognitive synonyms (syncs), each of which represents a separate concept. Synsets are interconnected using conceptual-semantic and lexical relationships. The resulting network of meaningful words and concepts is managed using a browser. WordNet is also open to the public for download. The WordNet structure is a useful tool fo r the process of computer linguistics and natural languages.
RST-style discourse parsing plays a vital role in many NLP tasks, revealing\nthe underlying semantic/pragmatic structure of potentially complex and diverse\ndocuments. Despite its importance, one of the most prevailing limitations in\nmodern day discourse parsing is the lack of large-scale datasets. To overcome\nthe data sparsity issue, distantly supervised approaches from tasks like\nsentiment analysis and summarization have been recently proposed. Here, we\nextend this line of research by exploiting distant supervision from topic\nsegmentation, which can arguably provide a strong and oftentimes complementary\nsignal for high-level discourse structures. Experiments on two human-annotated\ndiscourse treebanks confirm that our proposal generates accurate tree\nstructures on sentence and paragraph level, consistently outperforming previous\ndistantly supervised models on the sentence-to-document task and occasionally\nreaching even higher scores on the sentence-to-paragraph level.\n
It is proposed to create an information and reference system containing complete and comprehensive information about the scientific and scholarly results achieved by the RAS institutions, scientific departments and researchers in the field of linguistics. The reference system should provide information support for high-quality scientific and methodological guidance from the relevant branch of the Russian Academy of Sciences. The classification of information objects describing scientific and scholarly results includes both traditional forms of writing (publications, reports, dissertations) and the new digital ones (linguistic databases, websites, corpora, accounts, etc.). The preliminary results of creating such a system are described. The functions of the help system are listed. The updated version of the section “Linguistics” of the State rubricator of scientific and technical information is offered.
Binary Stanford Sentiment Treebank (SST2) is a binary version of SST and Movie Review dataset (the neutral class was removed), that is, the data was classified only into positive and negative classes. The files:<br> texts.txt: Document set (text). One per line.<br> score.txt: Document class whose index is associated with texts.txt<br> split_<k>.pkl: pandas DataFrame with k-cross validation partition
This paper explores the difficulties of annotating transcribed spoken Dutch-Frisian code-switch utterances into Universal Dependencies. We make use of data from the FAME! corpus, which consists of transcriptions and audio data. Besides the usual annotation difficulties, this dataset is extra challenging because of Frisian being low-resource, the informal nature of the data, code-switching and non-standard sentence segmentation. As a starting point, two annotators annotated 150 random utterances in three stages of 50 utterances. After each stage, disagreements where discussed and resolved. An increase of 7.8 UAS and 10.5 LAS points was achieved between the first and third round. This paper will focus on the issues that arise when annotating a transcribed speech corpus. To resolve these issues several solutions are proposed.
Widespread access to social media ensures that new and emergent coinages are noticed by population masses, attain domains of usage, and strengthen their place within a language. Of such new linguistic constructs, metaphorical neologisms are usually most adopted and frequently used. The present study aims to examine the x + head phrase structure widely used in Turkish and to evaluate its semantic properties. Examples semantically identical to this structure are included in the linguistic database through the Turkish National Corpus. The appearance of such phrases was observed on different social media platforms. Additionally, the opinions of people were sought on the use of this structure in social interactions. Further, the differences between the usage of the mentioned structure on its own and its use within a context were measured via a survey.
Machine learning training methods depend plentifully and intricately on\nhyperparameters, motivating automated strategies for their optimisation. Many\nexisting algorithms restart training for each new hyperparameter choice, at\nconsiderable computational cost. Some hypergradient-based one-pass methods\nexist, but these either cannot be applied to arbitrary optimiser\nhyperparameters (such as learning rates and momenta) or take several times\nlonger to train than their base models. We extend these existing methods to\ndevelop an approximate hypergradient-based hyperparameter optimiser which is\napplicable to any continuous hyperparameter appearing in a differentiable model\nweight update, yet requires only one training episode, with no restarts. We\nalso provide a motivating argument for convergence to the true hypergradient,\nand perform tractable gradient-based optimisation of independent learning rates\nfor each model parameter. Our method performs competitively from varied random\nhyperparameter initialisations on several UCI datasets and Fashion-MNIST (using\na one-layer MLP), Penn Treebank (using an LSTM) and CIFAR-10 (using a\nResNet-18), in time only 2-3x greater than vanilla training.\n
Among the various challenges regarding distance education is the necessity of reducing the student dropout rate. In this sense, the present research aimed to contribute to the design of a lexical database focused on emotions and opinions that can be incorporated into a predictive evasion software. For the database design, we used the Scup tool to collect 150 tweets containing distance education students’ opinions and analyzed them in the light of Martin and White’s Appraisal Framework, along with five resources related to the sentiment Analysis field, which were taken from Liu’s work. In addition, we used the Aulete dictionary to describe the lexical units found in our corpus to better fit them into the analysis categories. Results showed 220 opinion tokens, which were identified and labeled according to their polarity. Moreover, these tokens were included in the domains attitude (judgment and appreciation) and graduation (sharp and strong) from the linguistic framework used. The results also indicated the necessity of another resource to help identify the use of figurative language, slangs, and extralinguistic elements, such as GIFS and emojis.
While some heritage languages enjoy large numbers of speakers and vibrant communities, centuries-old and ongoing sociohistorical and sociolinguistic oppression has resulted in the extreme endangerment of many Indigenous languages. To counter this linguistic and cultural loss, a growing number of communities have engaged in language revitalization efforts that are tied to broader objectives of ethnic reclamation and cultural resistance, aiming not only to maintain but also to strengthen what has been lost. Heritage language revitalization is a long-term project that demands change and engagement across many aspects of community life, work that is ripe with tensions and contradictions. This chapter considers three recurrent questions in heritage language revitalization: what efforts should be prioritized in language revitalization, who should take responsibility in revitalizing a language, and how should revitalization efforts navigate the perceived need to establish linguistic norms and standards while concomitantly supporting linguistic diversity. To date, these questions have been described as tensions or problems that reveal conflicting priorities, often the result of historical inequalities, and that frequently hinder language revitalization efforts. Rather than framing these questions as problems, the present chapter considers how communities have responded to these challenges to create new opportunities for collaboration and new approaches that embrace ambiguity and pluralism.
The article considers the system of principles for creating a modern textbook for adults as a basic linguodidactic tool in teaching Russian to foreigners at an advanced stage of training, with the focus of the educational process on the formation of intercultural competence. Russian language textbook for foreigners contains topics and situations of reality that are relevant to Russian linguoculture in their verbal embodiment in texts of different genres and styles of the modern Russian language. The author demonstrates that with a cognitive-communicative approach, the discursive base of the textbook on the Russian language for foreigners contains relevant topics and situations of reality in their verbal embodiment in texts of different genres and styles of the modern Russian language.; the linguistic concept of the textbook is based on the achievements of theoretical Russian studies; the methodological interpretation of the educational material reflects the understanding of the experience of teaching Russian to foreigners. The textbook serves as a basis for creative use both by the teacher in the educational process and by the students in their independent work, which plays a crucial role in learning. The author emphasizes the importance of the modern textbook as a means of improving the professional competence of the teacher. The author notes that the textbook takes into account the specifics of the forms and norms of modern communication, including the use of information and communication technologies. The textbook is based on clearly defined goals based on the internal motivation of students, focuses on the development and improvement of their communication skills with the active cooperation of the teacher and the student.
Parsing, i.e., identifying the underlying hierarchical structure of natural language expressions is important for several natural language processing applications. In recent times Machine Learning (ML) approaches have been developed for this study for many languages. Most of the effective techniques require an annotated corpus of the language for training and validation. For the Manipuri language of the Tibeto-Burman family, neither such a corpus nor a grammar framework to automatically analyse and represent the structure of sentences exists yet. This study proposes a Context-Free Grammar (CFG) that provides the framework to represent the structure of Manipuri sentences. This paves the way for parsing Manipuri sentences using CFG-based parsers for various applications and to conveniently build a Treebank for developing ML-based parsers for Manipuri. The rules of the proposed CFG are handcrafted after extensive analysis of the structure of Manipuri sentences. The grammar covers simple, compound, complex and compound-complex sentences. For evaluation, we induce an Earley's parser with the proposed CFG and test it over a collection of sentences that covers the possible varieties of structure. A recognition rate of 83.20% achieved in these experiments indicates the effectiveness of the proposed grammar.
Liking and pleasantness are common concepts in psychological emotion theories, and everyday language related to emotions. Despite obvious similarities between the terms, several empirical and theoretical notions support the idea that pleasantness and liking are cognitively different phenomena, becoming most evident in the context of emotion regulation and art enjoyment. In this study it was investigated whether liking and pleasantness indicate behaviourally measurable differences, not only in the long timespan of emotion regulation, but already within the initial affective responses to visual and auditory stimuli. A cross-modal affective priming protocol was used to assess whether there is a behavioural difference in the response time when providing an affective rating to a liking or pleasantness task. It was hypothesized that the pleasantness task would be faster as it is known to rely on rapid feature detection. Furthermore, an affective priming effect was expected to take place across the sensory modalities and the presentative and non-presentative stimuli. A linear mixed effect analysis indicated a significant priming effect, as well as an interaction effect between the auditory and visual sensory modalities and the affective rating tasks of liking and pleasantness: While liking was rated fastest across modalities, it was significantly faster in vision compared to audition. No significant modality dependent differences between the pleasantness ratings were detected. The results demonstrate that liking and pleasantness rating scales refer to separate processes already within the short time scale of a one to two seconds. Furthermore, the affective priming effect indicates that an affective information transfer takes place across modalities and the types of stimuli applied. Unlike hypothesized, liking rating took place faster across the modalities. This is interpreted to support emotion theoretical notions where liking and disking are crucial properties of emotions perception and homeostatic self-referential information, possibly overriding pleasantness-related feature analysis. Conclusively, the findings provide empirical evidence for a conceptual delineation of common affective processes.
Literary language and poetic aspects arise from deviating from the linguistic norms for artistic purposes. In the formalism theory, this is called "defamiliarization". One of the most important aspects of deviating from the norm of automatic language and defamiliarization is the semantic norm-breaking obtained by using imaginary forms and creating novel semantic relationships. The Indian style poetry pays special attention to defamiliarization and breaking the norms of automatic language in semantic domains; therefore, most of the poets of this period have focused on creativity in the field of new themes and images. Rashid Tabrizi, known as Rashid Abbasi, is one of the poets of the 11th century AH and the Indian style period. In his poems, he has used various methods of imaginary elements. Therefore, the study of semantic norm-breaking in Rashid Tabrizi's poetry shows an important dimension and stylistic aspect of his poetry. Among Rashid Tabrizi's works, the two Masnavis, Hosne Glousoz and Saqinameh, concerning the unity in form, subject, and content, as well as the quantity, which differ by only about ten percent, are suitable for analysis of semantic norm-breaking. The quantity of these two works provides sufficient data for frequency studies. In this article, using the descriptive-analytical method,, semantic norm-breakings in the two mentioned Masnavis have been examined. In addition to a comparative study, this article aimed at identifying the outstanding features of Rashid Tabrizi's literary style in the field of semantic relations and imageries. Different types of semantic norm-breaking found in Rashid Tabrizi's poetry in this article are as follows: simile (compound simile, predicative simile, and allegorical simile), metaphor, personification, implicit metaphor, metaphor in the verb, synesthesia, paradoxical (paradoxical compounds, predicative paradoxes), exaggeration, kenning (paradoxical kenning, predicative kenning), and poetic scales. Also, in this article, the images and similarities of Rashid Tabrizi's poetry are examined and his most frequent semantic networks are investigated Based on frequency studies, the simile is the most frequent method in semantic norm-breaking of Rashid Tabrizi's poetry seen in most of the verses of the studied Masnavi. Also, most similes are additional combinations, which are also mentioned in Hosne Glousoz and Saqiynameh. Many of these combinations are novel, and some are paradoxical. The new combinations in Rashid Tabrizi's semantic norm-breaking, except for simile, include metaphorical additions, paradoxical combinations, and ironic combinations. Also, in some cases, synesthesia is observed. Hence, after many uses of simile and metaphorical additions, the construction of additional or descriptive novel compositions is the most important feature of the poet's stylistics in the field of semantic norm-breaking. In the studied works, the explicit metaphor has no significant use; on the contrary, the implicit metaphor, personification and metaphor in the verb as a whole are frequent. Therefore, after the simile and new compositions, the most commonly used semantic norm-breaking constructs are types of virtual predicates (prediction of the verb to a virtual subject or virtual verb to a real subject). The use of these constructs has created movement and vitality in Rashid Abbasi's poetry. Also, the two dominant rhetorical features of the studied Masnavis related to their mystical themes are paradoxical and exaggeration. Paradoxical (contradiction) in the Masnavis in question is predominant in two ways: The mystical nature of these works: Because paradox is one of the characteristics of mystical texts and due to the paradoxical experiences of mystics, it is considered one of the main levels of mystical language. The creative and innovative aspect of these images: This is so because these images, like most synesthesia examples, are not fit into literary clichés. There are significantly more paradoxical images in Hosne Glousoz; but in Saqinameh a large part of the paradoxical combinations goes back to the paradox of mystical truths and unseen affairs, and it has phenomenological aspects. Although synesthesia and paradox are highly used in Hosne Glousoz, exaggeration has been used more in Saqiynameh. The reason for such high frequency is the mystical aspect of Saqiynameh and the abundance of images in it. Exaggeration is a feature of epic texts, and mystical texts have epic infrastructures. In the last section, we examined the high-frequency networks of Masnavi’s images. Images of plants, fiery and lighted things, pub images, water and sea images, and animals' images form frequent video networks in the text. Overall, the variety of networks in Rashid Tabrizi’s poetry highly shows his creative style in imagery. This study concludes that Rashid Tabrizi has used a significant variety of expressive and innovative constructions in a semantic norm-breaking way.
Emotional reactions to movies are typically similar between people. However, depressive symptoms decrease synchrony in brain responses. Less is known about the effect of depressive symptoms on intersubject synchrony in conscious stimulus-related processing. In this study, we presented amusing, sad and fearful movie clips to dysphoric individuals (those with elevated depressive symptoms) and control participants to dynamically rate the clips' valences (positive vs. negative). We analysed both the valence ratings' mean values and intersubject correlation (ISC). We used electrodermal activity (EDA) to complement the measurement in a separate session. There were no group differences in either the EDA or mean valence rating values for each movie type. As expected, the valence ratings' ISC was lower in the dysphoric than the control group, specifically for the sad movie clips. In addition, there was a negative relationship between the valence ratings' ISC and depressive symptoms for sad movie clips in the full sample. The results are discussed in the context of the negative attentional bias in depression. The findings extend previous brain activity results of ISC by showing that depressive symptoms also increase variance in conscious ratings of valence of stimuli in a mood-congruent manner.
Background: This manuscript evaluates patient and provider perspectives on the core components of a Behavioral Health Home (BHH) implemented in an urban, safety-net health system. The BHH integrated primary care and wellness services (e.g., on-site Nurse Practitioner and Care Manager, wellness groups and tools, population health management) into an existing outpatient clinic for people with serious mental illness (SMI). Methods: As the qualitative component of a Hybrid Type I effectiveness-implementation study, semi-structured interviews were conducted with providers and patients 6 months after program implementation, and responses were analyzed using thematic analysis. Valence coding (i.e., positive vs. negative acceptability) was also used to rate interviewees' transcriptions with respect to their feedback of the appropriateness, acceptability, and feasibility/sustainability of 9 well-described and desirable Integrated Behavioral Health Core components (seven from prior literature and two additional components developed for this intervention). Themes from the thematic analysis were then mapped and organized by each of the 9 components and the degree to which these themes explain valence ratings by component. Results: Responses about the team-based approach and universal screening for health conditions had the most positive valence across appropriateness, acceptability, and feasibility/sustainability by both providers and patients. Areas of especially high mismatch between perceived provider appropriateness and measures of acceptability and feasibility/sustainability included population health management and use of evidence-based clinical models to improve physical wellness where patient engagement in specific activities and tools varied. Social and peer support was highly valued by patients while incorporating patient voice was also found to be challenging. Conclusions: Findings reveal component-specific challenges regarding the acceptability, feasibility, and sustainability of specific components. These findings may partly explain mixed results from BHH models studied thus far in the peer-reviewed literature and may help provide concrete data for providers to improve BHH program implementation in clinical settings. Plain language abstract: Many people with serious mental illness also have medical problems, which are made worse by lack of access to primary care. The Behavioral Health Home (BHH) model seeks to address this by adding primary care access into existing interdisciplinary mental health clinics. As these models are implemented with increasing frequency nationwide and a growing body of research continues to assess their health impacts, it is crucial to examine patient and provider experiences of BHH implementation to understand how implementation factors may contribute to clinical effectiveness. This study examines provider and patient perspectives of acceptability, appropriateness, and feasibility/sustainability of BHH model components at 6-7 months after program implementation at an urban, safety-net health system. The team-based approach of the BHH was perceived to be highly acceptable and appropriate. Although providers found certain BHH components to be highly appropriate in theory (e.g., population-level health management), their acceptability of these approaches as implemented in practice was not as high, and their feedback provides suggestions for model improvements at this and other health systems. Similarly, social and peer support was found to be highly appropriate by both providers and patients, but in practice, at months 6-7, the BHH studied had not yet developed a process of engaging patients in ongoing program operations that was highly acceptable by providers and patients alike. We provide these data on each specific BHH model component, which will be useful to improving implementation in clinical settings of BHH programs that share some or all of these program components.
<strong>ACCEPTED ABSTRACT:</strong> <strong>Introduction:</strong> Studies consistently report that patients with schizophrenia exhibit qualitative abnormalities on language production tasks. These abnormalities are possibly associated with the severity of psychotic symptoms. Despite this, some studies have conflictingly suggested that patients with schizophrenia exhibit similar word frequency (WF) effects on lexical tasks compared to healthy subjects. Given that previous studies calculated WFs from language corpora, we aimed to investigate the relationship between WF and psychotic symptoms using a novel, simple method for calculating WF. <strong>Methods:</strong> Thirty-six patients with schizophrenia were included in the study. The severity of positive symptoms was measured using the Scale for the Assessment of Positive Symptoms (SAPS). One semantic and one letter fluency task were administered with the patients instructed to produce as many animal anmes and words beginning with the letter p in 60 s, respectively. Every response in the output was assigned (1) a corpus-based WF, extracted from the German-language lexical database dlexDB, and (2) a within-sample WF. The within-sample WF was calculated as the raw number of participants who produced the word. Spearman’s correlations were computed between the WF variables and symptoms. <strong>Results:</strong> Corpus-based WF exhibited skewed, kurtic, and/or non-normal distribution. Contrastingly, within-sample WF displayed normal, non-skewed, and non-kurtic distribution. There were no significant correlations between corpus-based WF and symptoms on both tasks. Conversely, within-sample WF on semantic fluency was significantly negatively and weakly correlated with the global SAPS score, as well as subscales measuring delusions and bizarre behavior. Further, within-sample WF on letter fluency was significantly positively and weakly correlated with the subscale measuring bizarre behavior of the SAPS scale. <strong>Conclusion:</strong> The differences in the data distribution patterns between corpus-based WF and within-sample WF indicate that different methodological frameworks may have better use of one or the other variable type. Further, significant correlations with positive symptoms were observed only for within-sample WF. It can be concluded that within-sample WF may be more appropriate for analyzing verbal fluency output in psychiatric research compared to corpus-based WF.
Image beauty assessment is an important subject of computer vision. Therefore, building a model to mimic the image beauty assessment becomes an important task. To better imitate the behaviours of the human visual system (HVS), a complete survey about images of different categories should be implemented. This work focuses on image beauty assessment. In this study, the pairwise evaluation method was used, which is based on the Bradley-Terry model. We believe that this method is more accurate than other image rating methods within an image group. Additionally, Convolution neural network (CNN), which is fit for image quality assessment, is used in this work. The first part of this study is a survey about the image beauty comparison of different images. The Bradley-Terry model is used for the calculated scores, which are the target of CNN model. The second part of this work focuses on the results of the image beauty prediction, including landscape images, architecture images and portrait images. The models are pretrained by the AVA dataset to improve the performance later. Then, the CNN model is trained with the surveyed images and corresponding scores. Furthermore, this work compares the results of four CNN base networks, i.e., Alex net, VGG net, Squeeze net and LSiM net, as discussed in literature. In the end, the model is evaluated by the accuracy in pairs, correlation coefficient and relative error calculated by survey results. Satisfactory results are achieved by our proposed methods with about 70 percent accuracy in pairs. Our work sheds more light on the novel image beauty assessment method. While more studies should be conducted, this method is a promising step.
The research is focused on definitions of discourse relations, a topic that is currently little-studied. The paper gives a brief overview of existing solutions for discourse relations definitions: Rhetorical Structure Theory (RST), Segmented Discourse Representation Theory (SDRT), Penn Discourse Treebank (PDTB), and Cognitive approach to Coherence Relations. The author shows criteria used to define a discourse relation, or, in case of a narrower definition, a logical-semantic relation, in these approaches and outlines the shortcomings of the described definitions. The author also describes the principles used to build the classification and the definitions of logical-semantic relations (LSR) in the Supracorpora Database of connectives (SDB). The classification is based on four basic semantic operations upon which rests every LSR's definition: implication, location on the chronological scale, comparison, correlation between specific and general or an element and a set. The classification consistently distinguishes the levels at which the LSR can be established: propositional, illocutionary, and metalinguistic. Each LSR is defined on the basis of these two criteria. Thus, for example, for the LSR of alternative based on the comparison operation, one has the choice between the LSR of propositional, illocutionary and metalinguistic alternative (We will go to the mountains or to the sea vs. Put the gun away, or are you scared? vs. The symbol of the year or, simply speaking, cutie-pie). In case of LSRs based on implication or comparison, the polarity criterion is added, distinguishing whether the LSR is established between p and q or their negative correlates p and q are also to be taken into account in order to obtain a correct interpretation (cf. well-known descriptions of how the Russian conjunction no 'but' functions). In addition, semantic and pragmatic characteristics of the context are also considered in the classification. For example, in the case of the LSR of specification and generalization, the semantic correlation between p and q (together with their intensional and extensional interpretations) is taken heed of. Several definitions of LSR and corresponding examples are provided. Thus, the LSR of extensional specification is defined as follows: based on the operation of correlation between the general and the particular; established at the propositional level; X contains a generalized notion or state of things p; Y contains a more particular q-notion, limiting p-extensional. And the LSR of intensional specification is defined as follows: based on the operation of correlation between the general and the particular; established at the metalinguistic level; X contains a generalized concept or state of things p; Y contains a more particular q-notion, limiting p-intensional. The definitions used in the SDB definitions make it possible to evaluate, on the basis of the proposed criteria, the semantic closeness of relations and increase the level of consistency in the work of experts and annotators. That in turn increases the value of the annotated material, and therefore its reliability.
Home clutter can adversely affect work performance, health and well-being. Clinical-level hoarding disorders usually manifest during early adolescence, so early detection and prevention of subclinical hoarding tendencies are essential. This study aimed to evaluate a community-based programme for individuals with poor organising and decluttering skills who volunteered to receive education on how to organise their homes. We conducted an open-label randomised controlled trial beginning in January 2016 in Tokyo. We enrolled 61 volunteers aged 12-55 years with problems with organising and decluttering. A workshop and home visit group (n = 30) attended four workshop sessions on organising skills and received a visit from a home organiser. The home visit only group (n = 31) only received the home organiser visit. The primary outcome was Saving Inventory-Revised (SI-R; Japanese version) scores. The secondary outcomes were Clutter Image Rating Scale and Rosenberg Self-Esteem Scale (Japanese version) scores. Between-group changes from baseline to 7 months were analysed using a general linear model. At follow-up, the SI-R scores of both groups had improved. The mean change from baseline in SI-R scores was -20.8 (standard deviation = 9.8) and -13.1 (standard deviation = 14.3) in the workshop and home visit and home visit only groups, respectively. The estimated between-group difference in SI-R score changes from baseline (adjusted for baseline SI-R score) was non-significant at -5.7 (95% confidence interval, -12.4 to 0.9; p =.089). However, the difference was significant in the univariate model: -7.2 (95% confidence interval, -13.7 to -0.8; p =.029). Although both groups improved, after adjusting for baseline values and participant characteristics, there was no significant difference between the groups. Our results suggest that a workshop-style educational intervention and assistance and advice from professional organisers may help to improve the living conditions of people with hoarding tendencies.
This article presents a method for automatic assignment of syntactic dependency relations to the corpus of American Norwegian speech (CANS). Different machine learning techniques and corpora are used. Finally, an accuracy measure is computed and compared with a relatively new treebank for spoken Norwegian.
Abstract Depression is characterized by disturbed emotional processing, with increased amygdala reactivity during perception of negative emotional stimuli. Here, we assessed whether intermittent theta-burst stimulation (iTBS) modulates amygdala activity during an emotional picture anticipation paradigm using functional magnetic resonance imaging (fMRI) in a patient sample with depression. Patients were randomized to either active (n=21) or sham (n=21) iTBS over the dorsomedial prefrontal cortex (DMPFC), delivered twice daily for ten days at target intensity. Depression symptom assessment and fMRI scanning took place just before treatment start and once again four weeks later. During fMRI scanning, picture stimuli of negative and positive valence were presented, indicated by a prior red or green screen, respectively. Behavioral valence ratings of the picture stimuli were conducted outside of the scanner. Amygdala activation during perception of negative picture stimuli was reduced after active, but not sham, iTBS (left amygdala: F(1,25)= 7.22, p=.013, right amygdala: F(1,25)=8.65, p=.007). Baseline amygdala reactivity was not correlated with depressive symptom levels. Behavioral valence ratings remained unchanged after active iTBS. The findings suggest a treatment effect of active iTBS in an emotional processing network in depression, spanning the DMPFC and amygdala. Keywords: dorsomedial prefrontal cortex, emotion, iTBS, transcranial magnetic stimulation
Constituency parsing is generally evaluated superficially, particularly in a multiple language setting, with only F-scores being re-ported. As new state-of-the-art chart-based parsers have resulted in a transition from traditional PCFG-based grammars to span-based approaches (Stern et al., 2017; Gaddy et al.,2018), we do not have a good understanding of how such fundamentally different approaches interact with various treebanks as results show improvements across treebanks (Kitaev and Klein, 2018), but it is unclear what influence annotation schemes have on various treebank performance (Kitaev et al., 2019). In particular, a span-based parser’s capability of creating novel rules is an unknown factor. We perform an analysis of how span-based parsing performs across 11 treebanks in order to examine the overall behavior of this parsing approach and the effect of the treebanks’ specific annotations on results. We find that the parser tends to prefer flatter trees, but the approach works well because it is robust enough to adapt to differences in annotation schemes across treebanks and languages.
Subject of the work: to find out how S. Maugham was able to use stylistic means in this story. Purpose of the work: to find out what stylistic means were used in this story. Relevance: Stylistics is the science that studies styles of speech and the use of linguistic means in them. This helps to make speech stylistically correct. And the correctness of speech is the basis of speech culture, that is, the ability to assimilate linguistic norms and use the expressive means of language. Stylistics also helps with the formation of skills in the coherent exposition of thoughts in oral and written form. Stylistics introduces the patterns of language use in different spheres of communication, their stylistic originality, and thereby enriches knowledge about the functional aspect of the language. Stylistics as a branch of linguistics is of great importance for the development and theory of language. Conclusion: The purpose of the work has been achieved. We found out what stylistic means were used. Предмет работы: выяснить каким образом С.Моэм смог использовать стилистические средства в этом рассказе. Цель работы: выяснить какие стилистические средства были использованы в этом рассказе. Актуальность: Стилистика - это наука, изучающая стили речи и использование в них языковых средств. Это помогает сделать речь стилистически правильной. А правильность речи - основа речевой культуры, то есть умение усваивать языковые нормы и пользоваться выразительными средствами языка. Стилистика также помогает в формировании навыков связного изложения мыслей в устной и письменной форме. Стилистика знакомит с закономерностями использования языка в разных сферах общения, их стилистической оригинальностью и тем самым обогащает знания о функциональной стороне языка. Стилистика как раздел языкознания имеет большое значение для развития и теории языка. Вывод: Цель работы была достигнута. Мы выяснили какие стилистические средства были использованы. Жұмыс тақырыбы: бұл әңгімеде С.Моэм стилистикалық құралдарды қалай қолдана білгенін білу. Жұмыстың мақсаты: бұл әңгімеде қандай стилистикалық құралдар қолданылғанын білу. Өзектілігі: Стилистика - сөйлеу мәнерлерін және оларда тілдік құралдарды қолдануды зерттейтін ғылым. Бұл сөйлеуді стилистикалық тұрғыдан дұрыс жасауға көмектеседі. Ал сөйлеудің дұрыстығы - сөйлеу мәдениетінің негізі, яғни тілдік нормаларды сіңіріп, тілдің экспрессивті құралдарын қолдана білу. Стилистика сонымен қатар ойды ауызша және жазбаша түрде үйлестіру дағдысын қалыптастыруға көмектеседі. Стилистика қарым-қатынастың әр түрлі салаларында тілдің қолданылу заңдылықтарын, олардың стилистикалық ерекшелігін енгізеді және сол арқылы тілдің функционалдық аспектісі туралы білімді байытады. Стилистика тіл білімінің бір саласы ретінде тілдің дамуы мен теориясы үшін үлкен маңызға ие. Қорытынды: Жұмыстың мақсаты орындалды. Біз қандай стилистикалық құралдар қолданылғанын білдік.
The present article considers linguistic norms and pronunciation standard of the Spanish language. It is shown that from the theoretical point of view the normative standards of dialectic variation of the Spanish language are considered. The problem of geographically diversified pronunciation standard of the Spanish consonant sounds is established. It is being noted that the unified standard language norm for all Spanish- speaking countries to be used in terms of teaching standards of the Spanish language as a foreign. The conclusion is made that knowledge of functioning particularities of the language system at cultural and social level is highly needed. In addition to that dialectic and accent variations of the Spanish language are to be taken into consideration
Controlling the presented forms (or structures) of generated text are as important as controlling the generated contents during neural text generation. It helps to reduce the uncertainty and improve the interpretability of generated text. However, the structures and contents are entangled together and realized simultaneously during text generation, which is challenging for the structure controlling. In this paper, we propose an efficient, straightforward generation framework to control the structure of generated text. A structure-aware transformer (SAT) is proposed to explicitly incorporate multiple types of multi-granularity structure information to guide the text generation with corresponding structure. The structure information is extracted from given sequence template by auxiliary model, and the type of structure for the given template can be learned, represented and imitated. Extensive experiments have been conducted on both Chinese lyrics corpus and English Penn Treebank dataset. Both automatic evaluation metrics and human judgement demonstrate the superior capability of our model in controlling the structure of generated text, and the quality ( like Fluency and Meaningfulness) of the generated text is even better than the state-of-the-arts model.
In this paper, we describe the use of recurrent neural networks to capture sequential information from the self-attention representations to improve the Transformers. Although self-attention mechanism provides a means to exploit long context, the sequential information, i.e. the arrangement of tokens, is not explicitly captured. We propose to cascade the recurrent neural networks to the Transformers, which referred to as the TransfoRNN model, to capture the sequential information. We found that the TransfoRNN models which consists of only shallow Transformers stack is suffice to give comparable, if not better, performance than a deeper Transformer model. Evaluated on the Penn Treebank and WikiText-2 corpora, the proposed TransfoRNN model has shown lower model perplexities with fewer number of model parameters. On the Penn Treebank corpus, the model perplexities were reduced up to 5.5% with the model size reduced up to 10.5%. On the WikiText-2 corpus, the model perplexity was reduced up to 2.2% with a 27.7% smaller model. Also, the TransfoRNN model was applied on the LibriSpeech speech recognition task and has shown comparable results with the Transformer models.
<h3>Introduction</h3><br> BOLT Egyptian Arabic Treebank - SMS/Chat, Linguistic Data Consortium (LDC) catalog number LDC2021T17 and ISBN 1-58563-976-1, was developed by LDC and consists of Egyptian Arabic SMS/Chat data with part-of-speech annotation, morphology, and syntactic tree annotation. <br> The DARPA <a href="https://www.ldc.upenn.edu/collaborations/current-projects/bolt">BOLT</a> (Broad Operational Language Translation) program developed machine translation and information retrieval for less formal genres, focusing particularly on user-generated content. LDC supported the BOLT program by collecting informal data sources -- discussion forums, text messaging and chat -- in Chinese, Egyptian Arabic and English. The collected data was translated and annotated for various tasks including word alignment, treebanking, propbanking and co-reference. <br> The unannotated Egyptian Arabic source data is released as <a href="../../../LDC2017T07">BOLT Egyptian Arabic SMS/Chat and Transliteration (LDC2017T07)</a>. <br> The annotations in this release follow Penn Arabic Treebank (PATB) annotation guidelines. The PATB project consists of two distinct phases: (a) part-of-speech tagging which divides the text into lexical tokens and gives relevant information about each token such as lexical category, inflectional features and a gloss; and (b) Arabic treebanking, which characterizes the constituent structures of word sequences, provides categories for each non-terminal node and identifies null elements, co-reference, traces and so on. <br> There are two types of morphological analysis synchronized in the corpus. LDC Standard Morphological Analyzer (SAMA) Version 3.1 (<a href="../../../LDC2010L01"> LDC2010L01</a>) was used for Modern Standard Arabic tokens, and <a href="http://www.aclweb.org/anthology/W12-2301">CALIMA</a> (Columbia Arabic Language and dIalect Morphological Analyzer) was used for Egyptian-Arabic tokens. <br> <h3>Data</h3><br> This release contains 349,414 tokens before clitics were split and 435,677 tree tokens after clitics were split for treebank annotation. The source data was collected by LDC from its collection platform or by donation and was manually reviewed to exclude material not in the target language or with sensitive content. Originally written in Arabizi (or Romanized/Latin characters) script, the source data was transliterated to Arabic script and manually corrected prior to treebank annotation. <br> Data is presented in a variety of UTF-8 encoded text formats, specifically plain text, XML, tdf and Penn Treebank. See the included documentation for more information about the specific formats. <br> <h3>Sponsorship</h3><br> This material is based upon work supported by the Defense Advanced Research Projects Agency (DARPA) under Contract No. HR0011-11-C-0145. The content does not necessarily reflect the position or the policy of the Government, and no official endorsement should be inferred. <br> <h3>Samples</h3><br> Please view the following samples: <br> <ul><br> <li><a href="desc/addenda/LDC2021T17.tdf">SU TDF (TXT)</a></li><br> <li><a href="desc/addenda/LDC2021T17.integrated.txt">Integrated (TXT)</a></li><br> <li><a href="desc/addenda/LDC2021T17.w-vowel.tree">Penn Treebank (TXT)</a></li><br> <li><a href="desc/addenda/LDC2021T17.su.xml">SU XML</a></li><br> <li><a href="desc/addenda/LDC2021T17.xml">Annotation Graph (XML)</a></li><br> </ul><br> <h3>Updates</h3><br> None at this time. </br> Portions © 2013-2021 Trustees of the University of Pennsylvania
[Context:] Causal relations (e.g., If A, then B) are prevalent in functional requirements. For various applications of AI4RE, e.g., the automatic derivation of suitable test cases from requirements, automatically extracting such causal statements are a basic necessity. [Problem:] We lack an approach that is able to extract causal relations from natural language requirements in fine-grained form. Specifically, existing approaches do not consider the combinatorics between causes and effects. They also do not allow to split causes and effects into more granular text fragments (e.g., variable and condition), making the extracted relations unsuitable for automatic test case derivation. [Objective & Contributions:] We address this research gap and make the following contributions: First, we present the Causality Treebank, which is the first corpus of fully labeled binary parse trees representing the composition of 1,571 causal requirements. Second, we propose a fine-grained causality extractor based on Recursive Neural Tensor Networks. Our approach is capable of recovering the composition of causal statements written in natural language and achieves a F1 score of 74% in the evaluation on the Causality Treebank. Third, we disclose our open data sets as well as our code to foster the discourse on the automatic extraction of causality in the RE community.
Affective experiences occur across the wake-sleep cycle-from active wakefulness to resting wakefulness (i.e., mind-wandering) to sleep (i.e., dreaming). Yet, we know little about the dynamics of affect across these states. We compared the affective ratings of waking, mind-wandering, and dream episodes. Results showed that mind-wandering was more positively valenced than dreaming, and that both mind-wandering and dreaming were more negatively valenced than active wakefulness. We also compared participants' self-ratings of affect with external ratings of affect (i.e., analysis of affect in verbal reports). With self-ratings all episodes were predominated by positive affect. However, the affective valence of reports changed from positively valenced waking reports to affectively balanced mind-wandering reports to negatively valenced dream reports. These findings show that (1) the positivity bias characteristic to waking experiences decreases across the wake-sleep continuum, and (2) conclusions regarding affective experiences depend on whether self-ratings or verbal reports describing these experiences are analysed.
Morphological tagging of code-switching (CS) data becomes more challenging especially when language pairs composing the CS data have different morphological representations. In this paper, we explore a number of ways of implementing a language-aware morphological tagging method and present our approach for integrating language IDs into a transformerbased framework for CS morphological tagging. We perform our set of experiments on the Turkish-German SAGT Treebank. Experimental results show that including language IDs to the learning model significantly improves accuracy over other approaches.