Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
Linguistic bias is the differential use of linguistic abstraction (as defined by the Linguistic Category Model) to describe the same behaviour for members of different groups. Essentially, it is the tendency to use concrete language for belief-inconsistent behaviours and abstract language for belief-consistent behaviours. Having found that linguistic bias is produced without intention or awareness in many contexts, researchers argue that linguistic bias reflects, reinforces, and transmits pre-existing beliefs, thus playing a role in belief maintenance. Based on the Linguistic Category Model, this assumes that concrete descriptions reduce the impact of belief-inconsistent behaviours while abstract descriptions maximize the impact of belief-consistent behaviours. However, a key study by Geschke, Sassenberg, Ruhrmann, and Sommer [2007] found that concrete descriptions of belief-inconsistent behaviours actually had a greater impact than abstract descriptions, a finding that does not fit e)
Abstract In the context of the current heated debate surrounding the pervasive influence of the English language and Anglo-American culture on other languages, as well as the widespread purist attitude towards some contact-induced language change phenomena, both abroad and in Romania, our article discusses the situation of English lexical borrowings in present-day Romanian, focusing on the perception and processing of the so-called luxury Anglicisms ( Sections 2 and 3 ) by young Romanian native speakers, in an attempt to see whether such an analysis can help clarify their acceptability and diffusion across our target population. We propose an alternative cognitive, psycholinguistic approach to the study of contact-induced lexical borrowings, aiming to show that there is no difference in the young Romanian native speakers’ processing of sentences containing luxury Anglicisms and their established Romanian counterparts. Such findings may support our claim that the acceptability and diffusion of such Anglicisms are pervasive across our target population, even if the official position generally condemns such uses, considering them gratuitous and a burden in communication, even making it unintelligible sometimes. Our analysis starts from the observation that most (but not all) Romanian academics, whether linguists or not, tend to embrace a purist attitude, while on the other hand young Romanians accept such Anglicisms and tend to use them extensively. In fact, such uses are not limited to young people, who have been the subjects of our research, but are the ‘norm’ in daily conversations and elsewhere across the general population ( Stoichițoiu Ichim 2006 ). Thus, there seems to be a gap between the actual acceptability and diffusion of luxury Anglicisms among Romanians and the ‘official’ recommendations. Based on the results of a sensicality task, meant to show how 188 Romanians, aged 18–22, process and perceive sentences with or without luxury Anglicisms (see Section 6 ), we will try to show that luxury Anglicisms are accepted and, by recurrent use, diffused among the Romanian community. For a more accurate picture of their diffusion, the findings will be further correlated with data from CoRoLa, the only official corpus of present-day Romanian (beginning 1989) made available under the auspices of the Romanian Academy, as well as a corpus currently in the making, and the Internet (see Section 7 ). Besides showing that luxury Anglicisms cannot really be blamed for burdening or impairing processing, and thus communication, and explaining why such uses should not be censured or disapproved, we hope that our study of acceptability and diffusion will demonstrate that we are dealing with a complex, multi-layered phenomenon that can be better understood by going beyond a diachronic and synchronic analysis of particular words and a frequency count, and should incorporate more experimental data. Last but not least, we suggest that, on the practical side, such experimental studies as the one described here could be used as an additional criterion for the lexicographic inclusion of lexical borrowings.
Purpose Our goal was to evaluate an updated version of the "Cookie Theft" picture by obtaining norms based on picture descriptions by healthy controls for total content units (CUs), syllables per CU, and the ratio of left-right CUs. In addition, we aimed to compare these measures from healthy controls to picture descriptions obtained from individuals with poststroke aphasia and primary progressive aphasia (PPA) to assess whether these measures can capture impairments in content and efficiency of communication. Method Using an updated version of this picture, we analyzed descriptions from 50 healthy controls to develop norms for numbers of syllables, total CUs, syllables per CU, and left-right CU. We provide preliminary data from 44 individuals with aphasia (19 with poststroke aphasia and 25 with PPA). Results A total of 96 CUs were established based on the written transcriptions of spoken picture descriptions of the 50 control participants. There was a significant effect of group on total CUs, syllables, syllables per CU, and left-right CUs. The poststroke participants produced significantly fewer total CU and syllables than those with PPA. Each aphasic group produced significantly fewer total CUs, fewer syllables, more syllables per CU, and lower left-right CUs (indicating a right-sided bias) compared to controls. Conclusions Results show that the measures of numbers of syllables, total CUs, syllables per CU, and left-right CUs can distinguish language output of individuals with aphasia from controls and capture impairments in content and efficiency of communication. A limitation of this study is that we evaluated only 44 individuals with aphasia. In the future, we will evaluate other measures, such as CUs per minute, lexical variability, grammaticality, and ratio of nouns to verbs. Supplemental Material https://doi.org/10.23641/asha.7015223.
We explore two solutions to the problem of mistranslating rare words in neural machine translation. First, we argue that the standard output layer, which computes the inner product of a vector representing the context with all possible output word embeddings, rewards frequent words disproportionately, and we propose to fix the norms of both vectors to a constant value. Second, we integrate a simple lexical module which is jointly trained with the rest of the model. We evaluate our approaches on eight language pairs with data sizes ranging from 100k to 8M words, and achieve improvements of up to +4.3 BLEU, surpassing phrasebased translation in nearly all settings. 1
Instructional language programs in German childcare centers have shown limited effectiveness. Two reasons may be that (a) the training is unconnected with everyday situations in which children typically acquire language and (b) the programs adopt a cultural model of psychological autonomy, a model that may be inconsistent with some children’s background. In the present study, we implemented an everyday-based language intervention in four German childcare centers. In a prepost design, teachers ( N = 37, M = 32.97 years) were first trained to adopt an elaborative, socially oriented style. Their language behavior, videotaped and analyzed during daily routines over 1 year, demonstrated significant changes (e.g., asking more open-ended questions, referring to social content and decontextualized content more often). Independent of their families’ cultural orientation. children’s ( N = 85, M = 3.42 years) language competencies significantly increased beyond age-related development norms. In comparison with a control group of children who visited childcare centers implementing instructional language programs, children of the intervention group performed significantly better in nonword repetition (an indicator of lexical knowledge) after 1 year. The results demonstrate that, in a brief intervention, teachers’ conversational style could be effectively changed toward promoting language development in a culture-sensitive way. Although the direct link to children’s language development remains to be proven, results indicate that children with different cultural backgrounds could profit from this everyday-based approach without using extra settings, materials, or instructions.
Dans le cadre général d’une sémiotique des cultures, cette recherche a utilisé les propositions épistémologiques et méthodo-logiques de la sémantique interprétative pour renouveler l’analyse de textes irlandais médiévaux. Le but était d’apporter une contribution à une problématique générale intéressant les sciences du langage, mais aussi les sciences historiques: comment fonder la pertinence scientifique d’une interprétation de textes et signes anciens appartenant à une culture différente? Pour cela il a été choisi de viser les faits sémantiques qui interviennent dans les processus de transfert de sens que la tradition rhétorique nomme comparaison, métaphore ou symbole. L’approche méthodologique a nécessité l’édition d’un corpus interlinéaire offrant un accès direct aux données de l’Electronic Dictionary of the Irish Language. L’étude du lexique a confronté les possibilités définitoires aux afférences contextuelles observées par un relevé systématique des isotopies ciblées.Sur le plan sémantique, l’analyse des processus différentiels qui structurent les molécules sémiques a permis d’observer la circulation des sèmes marquant les analogies intentionnelles. Sur le plan diachronique, la description du système de valeur, pris dans sa globalité, fonde la pertinence de l’interprétation en ce qu’il intègre les normes sociales du contexte historique du signe. Sur le plan des études celtiques, l’analyse des correspondances entre les domaines de l’orientation spatiale, des cycles temporels et des fonctions sociales donne un nouvel accès à la complexité du système de pensée de cette culture. Les formes sémantiques décrites fournissent de nouveaux modèles pour des comparaisons. Sur cette base, les expressions de l’association arbre-savoir ont été décrites pour apporter une solution aux problèmes de l’étymologie de la lexie druid- et proposer le dépassement des approches lexicales monographiques par l’approche intertextuelle.
Modern Ukrainian language is characterized by interrelated tendencies of synthetical character and analyticity, which are motivated by: 1) the folk-colloquial element of the literary norm; 2) book tradition; 3) the law of language economy, etc. One of the brightest analysts expresses the dynamic development of the prepositional system, which has recently been actively replenished by semantically specialized two-component, three-component, and other entities. The ultimate manifestation of such an analyticism is the presence of dissected prepositions of the sample від... до, з... до.The textual function of the prepositions within the texts appears in the inter-phrased / intra-phrased, intersentenced manifestations; at the same time, the expansion of the paradigmatic plane of prepositions, their participation in the creation of images, are apparent. Within the context the preposition appears as a relatively independent element with its own inventory of distributions, phonetic variants, and others. The solution of the question of the status of the text of the prepositions will enable understanding of the mechanisms of interaction between the components and components of the language system and the establishment of ways for the creation and functioning of syntaxemes, the extension of the formation of their own semantic potential of the latter.Lexical and grammatical meanings of prepositions coincide, which does not mean their identity. The grammatical meaning of the preposition is the realization of the form of syntactic relation between the words, and the lexical one should consider the designation of a certain relation between the objects, the action and the object, etc., which makes it possible to enumerate the prepositions to the words-relatives. The ability of prepositions to determine its lexical meaning in the noun (more broadly – in the name) is related not to the absence of this value, but to its corresponding specifics. Interpretation of prepositions should be based on the lexical environment, since they indicate the relation between the objects. The ratio implies the presence of not less than two quantities, therefore, prepositions can not exist without these values, that is, they can not be used independently. The meaning of the prepositions implies their functioning within the phrase. And by its nature the primary prepositions are many-valued, they are characterized by homonymy. In this case, the context is a diagnostic indicator of a certain value as a virtual-system.In analyzing the semantics of relations, it is necessary to take into account, as much as possible, the particularities of the lexical meaning of prepositions and the semantic features of words that form the left and right-side distribution of the prepositional-case design. Due to this, the following classification of intratextual semantic relations, expressed by the Ukrainian primary prepositions in artistic-fiction and journalistic language-linguistic discourse practices, is real: 1) spatial relationships that are the most researched in modern linguistics. Among them differentiate the local (place) and additive (direction). Locative relations are differentiated into: suppressive (finding above the surface) and invasive (finding inside).
The traditional understanding of data from Likert scales is that the quantifications involved result from measures of attitude strength. Applying a recently proposed semantic theory of survey response, we claim that survey responses tap two different sources: a mixture of attitudes plus the semantic structure of the survey. Exploring the degree to which individual responses are influenced by semantics, we hypothesized that in many cases, information about attitude strength is actually filtered out as noise in the commonly used correlation matrix. We developed a procedure to separate the semantic influence from attitude strength in individual response patterns, and compared these results to, respectively, the observed sample correlation matrices and the semantic similarity structures arising from text analysis algorithms. This was done with four datasets, comprising a total of 7,787 subjects and 27,461,502 observed item pair responses. As we argued, attitude strength seemed to account for much information about the individual respondents. However, this information did not seem to carry over into the observed sample correlation matrices, which instead converged around the semantic structures offered by the survey items. This is potentially disturbing for the traditional understanding of what survey data represent. We argue that this approach contributes to a better understanding of the cognitive processes involved in survey responses. In turn, this could help us make better use of the data that such methods provide.
The thesis is devoted to the comparative study of phraseological units indicating the action and state intensity in languages with different structures: Germanic (English, German) and Slavic (Ukrainian, Russian). The specificity of the category of intensity is manifested in its close interaction and dependence on such categories as quantity, quality, graduality, expressiveness, emotionality, imagery, assessment. In the paper, intensity is defined as a certain degree of the actions and states properties expression, the quantitative change of which occurs within a certain quality with a deviation from the norm (reference point) in the direction of their increase or decrease. It is established that the majority of the units under study (98%) are phraseological units which denote a reinforced manifestation of the intensive action and state. To distinguish phraseological units which denote the action intensity the following general characteristics are involved: agency, activity, and dynamism; for phraseological units which denote the state intensity those are nonagency, passivity, and static. These differential features, which are used for the taxonomy of the predicate vocabulary, are transferred onto the phraseological material analysed in the paper, because the phraseological units are means of secondary nomination and in functional-grammatical terms they correlate with certain parts of the speech (nouns, adjectives, verbs) and perform certain syntactic functions. Among the phraseological units with semantics of the action intensity 6 phrasesemantic groups indicating the intensity are distinguished: 1) intellectual activity; 2) physical actions; 3) activities (social, labor); 4) speech actions and sounds; 5) physiological actions; 6) motion. Phraseological units with the semantics of the state intensity are divided into 4 phrase-semantic groups indicating the intensity: 1) positive psycho-emotional state; 2) negative psycho-emotional state; 3) physiological state; 4) the course (manifestation) of natural phenomena. The paper proposes two-stage semantic analysis: 1) the analysis of the dictionary definition of phraseological unit; 2) the analysis of the formal structure of the phraseological unit. At the first stage, in order to identify the seme of the action and state intensity in the phraseological unit, the methodology of analysing the dictionary definitions and the component analysis method are used, on the basis of which components-intensifiers are distinguished, i.e. word-markers that signal at the high (extreme) degree of an action or a state. At the second stage of the analysis, a set of formal explicit (word-building, lexical-grammatical, syntactic) means of expressing the action and state intensity is established. This consideration of phraseological units allows to reveal the mechanisms of systemic and complex interaction of multilevel linguistic means which take direct part in the formation of the action and state intensity. In phraseological units under analysis, intensity can be expressed explicitly (by different linguistic means in the formal structure of the phraseological units) or implicitly (encoded in dictionary definitions of the phraseological units, their internal form, etc.). The implicit way of expressing the intensity in the analysed phraseological units is represented by: a) adverbs with the intensifying meaning (“intensifiers”) (the most frequent in the comparable languages are: Eng. very, extremely, greatly; Ger. sehr, heftig, auserst; Rus. очень, сильно; Ukr. дуже, сильно ); b) components with hidden semantics of intensity (“intensitives”), distinguished through additional lexicographic definition of the component part in the interpretation of the phraseological unit; c) the internal form of the phraseological unit. The explicit expression of the action and state intensity is represented in the structure of the phraseological units by the following linguistic means: a) word-building (prefixes, semi-prefixes); b) lexical-grammatical (adjectives in the attributive function, words-antonyms, reflexive verbs, verbs of destruction, prepositional-noun sets, pronouns, particles with the intensifying meaning); c) syntactic (comparative constructions, lexical repetition). To represent the semantics of the phraseological units, which denote the action and state intensity, the method of their semantic structure modelling is used, aimed at identifying the meaningful connections between the participants in the phraseological situation. For this purpose, standard formulas of interpretation are created, which includ the following semantic components (the participants of the phraseological situation): the subject of the action or state (X), the object of the action or state (Y), the state of the subject / object (Cond.), which in its nature can be psycho-emotional (negative (Cond neg ) or positive (Cond pos )), physiological (Cond phys ), the action of the subject / object (V), and the location (Loc.) of the subject / object. The formula of interpretation reflects a certain standard extralingual situation, which can correspond to numerous specific phraseological units in the language.
To begin with, stylistics in the one hand is a discipline of applied linguistics that utilizes linguistic theories, perspectives and methods in analyzing all the literary narratives. More tellingly, stylistics in the one side is a field of study that stands between literary criticism and linguistics, i.e. it involves both literature and linguistics. Foregrounding in the other side refers to the use of literary devices (poetic language, parallelism and deviation for instance) for the sake of challenging the common and/or traditional literary norms and achieving deautomatization or literary –aesthetic functions. In addition, e. e. Cummings American poet;is considered as one of the most celebrated poets in the modern period. He is worldly known for his Avant-grade typography, nonconformist construction and eccentric capitalization. Therefore, the main contention of this paper is to analyze literarily and stylistically Cummings’ poem Buffalo Bill in all the stylistic levels (graphological, phonological, lexical, morphological, syntactic and semantic levels). Accordingly, this paper presents e. e. Cummings as a unique modernist poet. It explains extensively what do the concepts foregrounding and stylistics mean. Besides, this paper argues how Cummings resolved to accomplish ‘literariness’ in his poem through using eccentric typography, rebellious structures and uncommon capitalization.
Automatically recognized terminology is widely used for various domain-specific texts processing tasks, such as machine translation, information retrieval or ontology construction. However, there is still no agreement on which methods are best suited for particular settings and, moreover, there is no reliable comparison of already developed methods. We believe that one of the main reasons is the lack of state-of-the-art method implementations, which are usually non-trivial to recreate—mostly, in terms of software engineering efforts. In order to address these issues, we present ATR4S, an open-source software written in Scala that comprises 13 state-of-the-art methods for automatic terminology recognition (ATR) and implements the whole pipeline from text document preprocessing, to term candidates collection, term candidate scoring, and finally, term candidate ranking. It is highly scalable, modular and configurable tool with support of automatic caching. We also compare 13 state-of-the-art methods on 7 open datasets by average precision and processing time. Experimental comparison reveals that no single method demonstrates best average precision for all datasets and that other available tools for ATR do not contain the best methods.
The article presents translation analysis of the texts within tourism discourse. According to the authors, the Internet is the most popular source of information and thus tourist websites are aimed at forming tourism attractiveness of a certain region as well as promoting regional branding. As illustrated by examples of multilingual hotel websites, the language component of website content is an essential factor for translation. As a result, the analysis of data shows that in many translations various errors are made, which are characterized by a violation of stylistic, lexical, grammatical, spelling and punctuation norms or rules, consequently, translated texts do not correspond to their original communicative and pragmatic function. Having studied the original examples, the authors prove that the translated text in the tourism discourse performs its main function, i.e. attracts a large number of potential customers only when a professional translator while translating generates a new text, taking into account grammatical and linguistic norms of the language of translation, as well as maintaining stylistic imagery and colour in accordance with a specific lingua-culture of a foreign recipient.
The aim of this paper is to describe the phrasemes in doping discourse. The corpus created comes from a magazine specialized in sports, paying particular attention to specific headlines, found in these magazines, devoted to doping. These phrasemes appear in a high frequency rate mainly in the titles of magazine articles, mainly in the titles of magazine articles. The Speaker uses numerous linguistic mechanisms (the most important one being unfrozeness) to transgress the norms of the fixation of the phrases. The research is carried out within a Meaning Text Theory (MTT) theoretical framework. Two main families of phrasemes (= non-free phrases) are distinguished: lexical phrasemes and semantic-lexical phrasemes. Three major classes of phrasemes are presented: non- compositional idioms, compositional collocations and clichés.
Recognising the identity of conspecifics is an important yet highly variable skill. Approximately 2 % of the population suffers from a socially debilitating deficit in face recognition. More recently the existence of a similar deficit in voice perception has emerged (phonagnosia). Face perception tests have been readily available for years, advancing our understanding of underlying mechanisms in face perception. In contrast, voice perception has received less attention, and the construction of standardized voice perception tests has been neglected. Here we report the construction of the first standardized test for voice perception ability. Participants make a same/different identity decision after hearing two voice samples. Item Response Theory guided item selection to ensure the test discriminates between a range of abilities. The test provides a starting point for the systematic exploration of the cognitive and neural mechanisms underlying voice perception. With a high test-retest reliability (r=.86) and short assessment duration (~10 min) this test examines individual abilities reliably and quickly and therefore also has potential for use in developmental and neuropsychological populations.
One of the relatively recent trends in learner corpora research is building and exploiting learner translator corpora. Within corpus-based translation studies (CTS) translations are approached as a special variety of the target language. They are usually represented by texts produced by professional translators and are studied as manifestations of the current translational norm. Learner translations can be seen as a more specific variant of the said variety, which is likely to deviate from the accepted translational norm. As of now, typical linguistic features of learner translations as opposed to professional ones are only tentatively described. We hypothesize that these texts should demonstrate heavier translationese features due to the lack of professional translational skills, comparatively poor source language processing competence and target language production skills. The aim of this research is to compare learner and professional Russian translations of English mass-media texts with the reference Russian corpus of non-translations to reveal lexical differences between the three. We found that learner translations consistently showed more distance from non-translations than their professional counterparts, while both learner and professional translations undoubtedly had discursive features which made them linguistically different from naturally occurring language. These findings might help define (non)professionalism in translation and shed light on correlation between the linguistic features of a given text and translation quality, as well as contribute to pedagogical approaches to translator education.
In a variety of research fields, including linguistics, human–computer interaction research, psychology, sociology and behavioral studies, there is a growing interest in the role of gestural behavior related to speech and other modalities. The analysis of multimodal communication requires high-quality video data and detailed annotation of the different semiotic resources under scrutiny. In the majority of cases, the annotation of hand position, hand motion, gesture type, etc. is done manually, which is a time-consuming enterprise requiring multiple annotators and substantial resources. In this paper we present a semi-automatic alternative, in which the focus lies on minimizing the manual workload while guaranteeing highly accurate annotations. First, we discuss our approach, which consists of several processing steps such as identifying the hands in images, calculating motion of the hands, segmenting the recording in gesture and non-gesture events, etc. Second, we validate our approach against existing corpora in terms of accuracy and usefulness. The proposed approach is designed to provide annotations according to the McNeill (Hand and mind: what gestures reveal about thought, University of Chicago Press, Chicago, 1992) gesture space and the output is compatible with annotation tools such as ELAN or ANVIL.
Large-scale semantic norms have become both prevalent and influential in recent psycholinguistic research. However, little attention has been directed towards understanding the methodological best practices of such norm collection efforts. We compared the quality of semantic norms obtained through rating scales, numeric estimation, and a less commonly used judgment format called best-worst scaling. We found that best-worst scaling usually produces norms with higher predictive validities than other response formats, and does so requiring less data to be collected overall. We also found evidence that the various response formats may be producing qualitatively, rather than just quantitatively, different data. This raises the issue of potential response format bias, which has not been addressed by previous efforts to collect semantic norms, likely because of previous reliance on a single type of response format for a single type of semantic judgment. We have made available software for creating best-worst stimuli and scoring best-worst data. We also made available new norms for age of acquisition, valence, arousal, and concreteness collected using best-worst scaling. These norms include entries for 1,040 words, of which 1,034 are also contained in the ANEW norms (Bradley & Lang, Affective norms for English words (ANEW): Instruction manual and affective ratings (pp. 1-45). Technical report C-1, the center for research in psychophysiology, University of Florida, 1999).
The article is devoted to the study of lexical and structural characteristics of medical terminology in the English language instructions of medicines certified in Ukraine, and the means of its reproduction in the Ukrainian language translation. The analysis of the texts of English-language medical instructions shows that the medical terminology of the instructions relates to the medical condition. The problem of adequate and equivalent translation and instruction relates to the reproduction of a terminological pharmaceutical dictionary, cliche, formulas, abbreviations, and the like. The problem is illustrated by examples of grammatical, lexical, syntactic transformations in translations into Ukrainian (on the material of the preparation Panadol), where the calcula- tion makes up a third of the means. Prefixes have certain semantic functions; Productive for medical terminology of medical instructions are also suffixes. The translation of English-language instructions requires sufficient knowledge of translators in the relevant field and strict adherence to the norms of the Ukrainian language. English-language medical terminology in the text of the instructions for medical products has lexical, structural and other features that are reproduced in the relevant Ukrainian translations of the instructions of the Ministry of Health of Ukraine. Traces of the formation of medical terminology of pharmaceutical texts, means of reproduction for an adequate equivalent translation are traced. The process of development of medical products, their discussion at international conferences and further implementation takes place, mainly in English. Thus, among the less well-researched and actual ones, there is the problem of adequate translation of the medical terminology vocabulary of the English language instructions of medical devices certified by Ukraine.
Evaluation is crucial in the research and development of automatic summarization applications, in order to determine the appropriateness of a summary based on different criteria, such as the content it contains, and the way it is presented. To perform an adequate evaluation is of great relevance to ensure that automatic summaries can be useful for the context and/or application they are generated for. To this end, researchers must be aware of the evaluation metrics, approaches, and datasets that are available, in order to decide which of them would be the most suitable to use, or to be able to propose new ones, overcoming the possible limitations that existing methods may present. In this article, a critical and historical analysis of evaluation metrics, methods, and datasets for automatic summarization systems is presented, where the strengths and weaknesses of evaluation efforts are discussed and the major challenges to solve are identified. Therefore, a clear up-to-date overview of the evolution and progress of summarization evaluation is provided, giving the reader useful insights into the past, present and latest trends in the automatic evaluation of summaries.
This paper describes a support vector machine-based approach to different tasks related to sentiment analysis in Twitter for Spanish. We focus on parameter optimization of the models and the combination of several models by means of voting techniques. We evaluate the proposed approach in all the tasks that were defined in the five editions of the TASS workshop, between 2012 and 2016. TASS has become a framework for sentiment analysis tasks that are focused on the Spanish language. We describe our participation in this competition and the results achieved, and then we provide an analysis of and comparison with the best approaches of the teams who participated in all the tasks defined in the TASS workshops. To our knowledge, our results exceed those published to date in the sentiment analysis tasks of the TASS workshops.
Up to now, the potential of eye tracking in science as well as in everyday life has not been fully realized because of the high acquisition cost of trackers. Recently, manufacturers have introduced low-cost devices, preparing the way for wider use of this underutilized technology. As soon as scientists show independently of the manufacturers that low-cost devices are accurate enough for application and research, the real advent of eye trackers will have arrived. To facilitate this development, we propose a simple approach for comparing two eye trackers by adopting a method that psychologists have been practicing in diagnostics for decades: correlating constructs to show reliability and validity. In a laboratory study, we ran the newer, low-cost EyeTribe eye tracker and an established SensoMotoric Instruments eye tracker at the same time, positioning one above the other. This design allowed us to directly correlate the eye-tracking metrics of the two devices over time. The experiment was embedded in a research project on memory where 26 participants viewed pictures or words and had to make cognitive judgments afterwards. The outputs of both trackers, that is, the pupil size and point of regard, were highly correlated, as estimated in a mixed effects model. Furthermore, calibration quality explained a substantial amount of individual differences for gaze, but not pupil size. Since data quality is not compromised, we conclude that low-cost eye trackers, in many cases, may be reliable alternatives to established devices.
In this descriptive linguistic study, the lexico-grammatical complexity of placement and exit English for Academic Purposes (EAP) student writing samples was analyzed using corpus linguistic methods to explore language development as a result of student enrollment in the EAP program. Writing samples were typed, matched, and tagged. A concordance software was used to produce lexical realizations of grammatical features. A comparison was made of normed frequency counts for nine phrasal and clausal features as well as raw frequencies for type to token ratio (TTR), average word length, and word count. In addition, the contribution of variables such as advanced grammar and writing course grades, LOEP scores, and the number of semesters in the EAP program to the English Learner's (EL) lexico-grammatical complexity found in exit essays was also examined. Twelve paired parametric and non-parametric analyses of lexico-grammatical variables were performed. Dependent t test results showed that normed frequency counts for such features as pre-modifying nouns, attributive adjectives, adverbial conjunctions, coordinating conjunctions, TTR, average word length, and word count changed significantly, and students produced more of those features in their exit writing than in their placement essay. Non-parametric Wilcoxon test indicated that such a change was also observable with noun + that clauses. The frequencies of verb + that clauses and subordinating conjunction because, though non-significant, actually decreased. A split plot ANOVA allowed to see whether a change in above mentioned statistically significant lexico-grammatical features could be attributed to grammar instruction in EAP 1560. The results showed that there was no statistically significant difference between those who took EAP 1560 class and those who did not on pre-modifying nouns, coordinating conjunctions, TTR, average word length, and word count. On the other hand, those students who did not take EAP 1560 class had higher counts of attributive adjectives but lower of adverbial conjunctions, both statistically significant results, than those students who took the class. Lastly, five multiple linear regression analyses were conducted to predict frequencies of exit pre-modifying nouns, attributive adjectives, noun + that clauses, adverbial conjunctions, and TTR from EAP 1560 and EAP 1640 grades, LOEP scores, and the number of semesters students spent in the EAP program at SSC. The only significant regression analysis was with TTR, and 28% of its variance could be explained by the independent variables. LOEP Language Usage score was the only significant individual contributor to the model. Even though exit adverbial conjunctions were not predictable from the chosen IVs, LOEP Sentence Meaning score proved the only significant contributor to that model. The results indicate that compressed phrasal features are indicative of higher complexity and EL proficiency, while clausal features are acquired earlier and signal elaboration, as previously described in the literature.
This paper tests the new-dialect formation model of Peter Trudgill (1986 et seq) by examining several phonological features of Tibetan as spoken in the diaspora community of Kathmandu, Nepal. Established by an influx of migrants from many dialect regions beginning in 1959, this presents a unique opportunity to study koinéization, new dialect formation, in progress. Trudgill’s model predicts that a new dialect should largely emerge in the second generation born in the new region, exhibiting both simplification, the failure of marked variants to transmit across generations, and focusing, the selection of particular variants as a new norm for the community's new variety.Data from seventy-three sociolinguistic interviews was coded for phonological and lexical variables known to differ across Tibetan-speaking regions, and NeighborNets were constructed in SplitsTree. Results indicate that regionally marked variables were not transmitted into the first or second generation of Diaspora-raised speakers, but Diaspora speakers exhibited a high degree of variation comparable to that of speakers from the numerically- and socially-dominant U-Tsang region. That younger speakers have not yet converged on a single new variety suggests a role for additional factors to affect the rate of koinéization.
Ambivalence is a common experience that permeates a broad range of research. Unfortunately, quantifying ambivalence has proven a daunting task, with researchers limited to studying vacillating ambivalence, VA (i.e., temporal oscillations between favor/disfavor evaluations of an attitude object). Here, we demonstrate the use of the density matrix to measure both VA and what we term “simultaneous ambivalence” (SA): ambivalence that manifests itself as “in the moment” concurrent favor/disfavor evaluations. In a methodological study we gave participants the option of either single-responding or double-responding to questionnaire items regarding a controversial topic (i.e., affirmative action). Since standard statistical procedures provide no means for analyzing double responses, such data are routinely treated as “bad.” As demonstrated here, the density matrix provides an unambiguous and relatively easy means of accounting for double responses, which is our indicator of SA. Our data are well explained by a mixture model, with participants divided into two nearly equal groups of SA and non-SA participants, and provide evidence that the general phenomenon of SA transcends differences of gender and ethnicity. Further, the density matrix data are consistent with viewing SA and VA as distinct ambivalence constructs.
The author attempts to identify the place for the language norms in the modern library verbal communications. The relevancy of this study is determined by the crisis of verbal communications in the society (overuse of slang, frivolous borrowings from foreign languages, decay of the language culture, using gadgets in interpersonal and business communications). The author concludes that verbal communication skills are needed to harmonize all communicative components, including language norms and speech culture. The findings of the monitoring survey of library verbal communications are presented. The level of command of the modern Russian language norms is identified for library and information specialists. It was found that the largest part of respondents had insufficient knowledge of language norms which affects their professional activity. The author analyzes orthoepical (pronouncing and articulatory) errors, violation of lexical, grammar (word-formation, morphological and syntactic) and stylistic norms of the Russian language, observance of orthographic and punctuation rules in the written speech. She argues that revealing the gaps helps to define the problems of librarians’ speech culture and to increase the efficiency of verbal communications in the professional library environment, in particular within the system of advanced training.
Sound-symbolic word classes are found in different cultures and languages worldwide. These words are continuously produced to code complex information about events. Here we explore the capacity of creative language to transport complex multisensory information in a controlled experiment, where our participants improvised onomatopoeias from noisy moving objects in audio, visual and audiovisual formats. We found that consonants communicate movement types (slide, hit or ring) mainly through the manner of articulation in the vocal tract. Vowels communicate shapes in visual stimuli (spiky or rounded) and sound frequencies in auditory stimuli through the configuration of the lips and tongue. A machine learning model was trained to classify movement types and used to validate generalizations of our results across formats. We implemented the classifier with a list of cross-linguistic onomatopoeias simple actions were correctly classified, while different aspects were selected to build onomato)
We estimate lexical Concreteness for millions of wordsacross 77 languages. Using a simple regression framework,we combine vector-based models of lexical semantics withexperimental norms of Concreteness in English and Dutch.By applying techniques to align vector-based semantics acrossdistinct languages, we compute and release Concreteness esti-mates at scale in numerous languages for which experimentalnorms are not currently available. This paper lays out thetechnique and its efficacy. Although this is a difficult datasetto evaluate immediately, Concreteness estimates computedfrom English correlate with Dutch experimental norms at ρ=.75 in the vocabulary at large, increasing to ρ =.8 amongNouns. Our predictions also recapitulate attested relationshipswith word frequency. The approach we describe can be readilyapplied to numerous lexical measures beyond Concreteness.
We present an open-source software platform that transforms emotional cues expressed by speech signals using audio effects like pitch shifting, inflection, vibrato, and filtering. The emotional transformations can be applied to any audio file, but can also run in real time, using live input from a microphone, with less than 20-ms latency. We anticipate that this tool will be useful for the study of emotions in psychology and neuroscience, because it enables a high level of control over the acoustical and emotional content of experimental stimuli in a variety of laboratory situations, including real-time social situations. We present here results of a series of validation experiments aiming to position the tool against several methodological requirements: that transformed emotions be recognized at above-chance levels, valid in several languages (French, English, Swedish, and Japanese) and with a naturalness comparable to natural speech.
Word embedding, has been a great success story for natural language processing in recent years. The main purpose of this approach is providing a vector representation of words based on neural network language modeling. Using a large training corpus, the model most learns from co-occurrences of words, namely Skip-gram model, and capture semantic features of words. Moreover, adding the recently introduced character embedding model to the objective function, the model can also focus on morphological features of words. In this paper, we study the impact of training corpus on the results of word embedding and show how the genre of training data affects the type of information captured by word embedding models. We perform our experiments on the Persian language. In line of our experiments, providing two well-known evaluation datasets for Persian, namely Google semantic/syntactic analogy and Wordsim353, is also part of the contribution of this paper. The experiments include computation of word embedding from various public Persian corpora with different genres and sizes while considering comprehensive lexical and semantic comparison between them. We identify words whose usages differ between these datasets resulted totally different vector representation which ends to significant impact on different domains in which the results vary up to 9% on Google analogy and up to 6% on Wordsim353. The resulted word embedding for each of the individual corpora as well as their combinations will be publicly available for any further research based on word embedding for Persian.
We present LEAR (Lexical Entailment Attract-Repel), a novel post-processing method that transforms any input word vector space to emphasise the asymmetric relation of lexical entailment (LE), also known as the IS-A or hyponymy-hypernymy relation. By injecting external linguistic constraints (e.g., WordNet links) into the initial vector space, the LE specialisation procedure brings true hyponymyhypernymy pairs closer together in the transformed Euclidean space. The proposed asymmetric distance measure adjusts the norms of word vectors to reflect the actual WordNetstyle hierarchy of concepts. Simultaneously, a joint objective enforces semantic similarity using the symmetric cosine distance, yielding a vector space specialised for both lexical relations at once. LEAR specialisation achieves state-of-the-art performance in the tasks of hypernymy directionality, hypernymy detection, and graded lexical entailment, demonstrating the effectiveness and robustness of the proposed asymmetric specialisation model.
This study explores the use of kin terms in a corpus of Vietnamese–English bilingual spontaneous conversation. While the corpus features a range of single Vietnamese lexical items in otherwise English discourse, kin terms, as in the example below, are overwhelmingly the most frequent (accounting for 84%, 164/196, of single Vietnamese words in an English context). Borrowing or Code-switching? Traces of community norms in Vietnamese-English speechAll authorsLi Nguyen http://orcid.org/0000-0001-8632-7909https://doi.org/10.1080/07268602.2018.1510727Published online:09 October 2018Table Download CSVDisplay Table The study puts forward an empirical attempt at determining whether such items should be considered code-switches or borrowings, and the role that pragmatic norms play in shaping this linguistic behaviour. Discourse distribution of the kin terms in terms of person reference and syntactic role are used as cross-language ‘conflict sites’, to determine the level of integration of such items as a test of their status as code-switches or borrowings. This reveals that the distribution of Vietnamese kin terms in an otherwise English context mirrors that of Vietnamese kin terms in monolingual Vietnamese, and is distinct from that of English kin terms. This measure of integration suggests that these may be single-word code-switches. Nonetheless, the high frequency of use and their diffusion across the community are suggestive of borrowings. Follow-up interviews with the participants reveal specific community norms that underlie the use of these terms, namely as a linguistic resource to retain, promote and conform to community cultural practice. While the paper acknowledges the difficulty in determining the exact status of these forms based on existing criteria, it demonstrates how judicious application of empirical methodology enables us to pinpoint such strategies in studying language in contact.
Rumour is an old social phenomenon used in politics and other public spaces. It has been studied for only hundred years by sociologists and psychologists by qualitative means. Social media platforms open new opportunities to improve quantitative analyses. We scanned all scientific literature to find relevant features. We made a quantitative screening of some specific rumours (in French and in English). Firstly, we identified some sources of information to find them. Secondly, we compiled different reference, rumouring and event datasets. Thirdly, we considered two facets of a rumour: the way it can spread to other users, and the syntagmatic content that may or may not be specific for a rumour. We found 53 features, clustered into six categories, which are able to describe a rumour message. The spread of a rumour is multi-harmonic having different frequencies and spikes, and can survive several years. Combinations of words (n-grams and skip-grams) are not typical of expressivity betwee)
The dual route model predicts that idiomatic phrases show a processing advantage over matched novel phrases. This model postulates that familiar phrases are processed by a faster direct route, and novel phrases are processed by an indirect route. This thesis investigated the role of familiar form and concept in direct route activation. Study 1 provided norming evidence for experimental stimuli selection. Study 2 examined whether direct route can be activated for translated Chinese idioms in Chinese-English bilinguals. Bilinguals listened to the idiom up until the last word (e.g., draw a snake and add), then saw either the idiom ending (e.g., feet) or the matched control ending (e.g., hair); to which they made lexical decision and reaction times were recorded. Results showed evidence for dual route model and provided preliminary support for both familiar concept and lexical association as drivers of direct route activation.
In this article, in the context of the character of the relationship between the orthographic norm and the lexical and phonetic levels of the language, the peculiarities of spelling words and morphemes as one of the aspects of the orthographic norm are compared. An attempt is made to ascertain specific characterological features inherent in the orthographic norm of the Middle English and Old Slavonic languages by the material of the XII-century manuscripts corresponding to these languages, by comparing the involved principles of orthography and considering the nature of their relationship with intravariance and extravariance of spelling.
<p><em>Abstrak</em><em> - </em><strong>Penelitian ini berjudul Metonimia dan Metafora dalam Norma dan Eksploitasi Tipe Semantis Adjektiva <em>Value</em> Frasa Nomina <em>Eye</em> Pada COCA ‘Penelitian ini mengkaji kolokasi terdekat dengan kata <em>eye</em> untuk mendapatkan makna prototipe dalam norma dan makna eksploitasi norma. Analisis kajian bertumpu pada <em>The Theory of Norms and Exploitations,</em> TNE karya Hanks (2013), sebuah teori bahasa yang berfokus pada kajian leksikal, berbasis kelola korpus dan teori bawah atas. Metodologi yang digunakan adalah metode pendekatan gabungan antara kualitatif sebagai pendekatan yang utama dan kuantitatif berdasarkan frekuensi kata dalam korpus. Lima puluh frasa nomina tertinggi dan lima puluh frasa nomin terendah dari 500 frekuensi di seleksi dan dipilah berdasarkan kategori tipe semantis ajektiva dengan fokus pada tipe semantis <em>value</em>. Jenis makna dalam norma dan eksploitasi bervariasi dengan inti perluasan makna literal terhadap metonimia dan metafora. Metonimia konseptual dan metafora konseptual di tingkat dasar yang diterapkan untuk frasa nomina <em>eye</em> adalah organ perseptual bersanding sebagai persepsi dan metafora konseptual melihat adalah menyentuh. Pada tingkat abstrak metafora konseptual menjadi berpikir, mengetahui atau mengerti adalah melihat.</strong></p><p> </p><p><strong><em>Kata Kunci</em></strong><em> – Norma dan Eksploitasi, Jenis dari Nilai Semantik, metonymy, metaphor, Frase kata benda “ eye” </em></p><p> </p><p><em>Abstract</em> - <strong>This reseach entitled ‘Metonymy and Metaphor in Norm and Exploitation Semantic Types Adjective Value of Noun Phrase Eye in COCA’. This research analyse adjacent collocation the noun eye in oder to identify the prototype meaning of norms and extention meaning of the exploitations. The research is based on The Theory of Norms and Exploitations, TNE by Hanks (2013), a lexical and bottom-up theory, based on corpus data. The methodology used is a mixed-method of qualitative and quantitative of frequency of word in corpus. 50 highest frequency of noun phrase eye and 50 lowest frequency noun phrase from 500 frequncy are selected and sorted out within the semantic type of the adjectives and focus on the semantic types of value. Type of meaning in norms and exploitations are varied with the core literal meaning extension towards metonymy and metaphor. The basic conceptual Metonymy and the conceptual of metaphor for eye is perceptual organ stands for perception and for metaphor seeing is touching.In the abstract level of conceptual metaphor is describes as thimking, knowing aand understanding is seeing.</strong></p><p><em> </em></p><p><strong><em>Keywords</em></strong><strong><em> </em></strong><em>-</em><strong><em> </em></strong><em>N</em><em>orms and </em><em>E</em><em>xploitations, </em><em>S</em><em>emantic </em><em>T</em><em>ype of </em><em>V</em><em>alue, </em><em>M</em><em>etonymy, </em><em>M</em><em>etaphor, </em><em>E</em><em>ye noun phrase.</em><em></em></p>
The peculiar phenomenon in the field of mass communication is religious periodicals. On the one hand, it has common peculiarities by which periodicals are characterized, on the other hand it has a special communicative purpose and a range of topics covered. The aim of the article is to carry out a general analysis of the lexical composition of religious Christian periodical texts. The classification of the periodical religious publication lexical composition was carried out in the article on the publication of the newspaper «Volyn diocesan reports» (2004–2018). The specificity of the lexical composition is determined by topics of publications. In religious periodicals, first of all, problems of faith, spirituality, Christian ethics, norms of moral behavior of a Christian, church and religious life are violated, as a result of which the vocabulary on named realities designation becomes a stylish one. Among them, the vocabulary on the designation of the highest God’s people of the Christian religion, names of religious holidays, posts, memorable days is represented most quantitatively, as well as the vocabulary reflecting the organizational life of the church as an establishment and public institution (names of clergy posts, holy dignitaries, items of church use). Onomastic vocabulary and vocabulary on the designation of religious holidays are usually used in canonical forms, thus folk forms and newly created words are witnessed. Medical vocabulary is widely presented.A specific feature of the researched journal is that popular science materials about volyn shrines – icons are being constantly published in it, using a special art terminology.In religious periodical issues of political, social, economic, cultural life are actively discussed which determines the use of vocabulary denoting these realities.Religious periodicals, as well as secular media editions, demonstrate processes of a spoken vocabulary active usage.
This article describes the development of a free/open-source rule-based machine translation system for Catalan to Sardinian based on the Apertium platform. Special attention is given to the components of the system related with transfer (structural and lexical) and lexical selection, drawing attention to issues stemming from the current state of the Sardinian written norm. The system has a word-error rate (WER) of 20.5% and a position-independent word-error rate (PER) of 13.9%. We analyse the remaining errors by doing a qualitative analysis of the translation of four articles from the encyclopaedic domain.
Setting the problem. Formulation of the problem. One of the important factors in the formation of Ukrainian statehood is the use of the state language in all spheres. In this regard, before the high school, the question arises of improving the language skills and language skills of students. After all, language competences are the key to the success of future professionals both in professional and in the social environment. Analysis of publications. Questions of language training were studied by Voloshchak M., Kulbabskaya Yu.V., Matsko L., Yarmolenko S., Pentiluk M. However, in our opinion, a more detailed study requires the practical aspect of lexical, morphological and syntactic components in the professional activity of students- builders. The aim of the article is to analyze the peculiarities of teaching Ukrainian professional language for students of construction professions. Conclusions. Consequently, it is safe to assert that the basis of language training of specialists in the construction industry is vocabulary, because communicative competence depends primarily on the proper use of certain established expressions. Ability to correctly choose the forms of generic case singularities of nouns of the masculine second-order abandonment, adhere to the norms of use of designs with prepositions in the professional language of student-builders testifies to the qualitative level of their training.
The article analyzes, compares and summarizes the definitions of the polysemantic noun “terra” fixed in explanatory dictionaries. Summarizing lexicographical definitions helps to discover important information on the word semantics, its place in the lexical norm of the modern Portuguese language. The study aims to compile a new list of the word definitions, revised and supplemented on the basis of the conducted analysis. Such an approach allows identifying the total amount, composition and structure of the meanings and can serve as a basis for more comprehensive semantic analysis.
Spelling correction is a fundamental task in text mining. In this study, we assess the real-word error correction model proposed by Mays, Damerau and Mercer and describe several drawbacks of the model. We propose a new variation which focuses on detecting and correcting multiple real-word errors in a sentence, by manipulating a probabilistic context-free grammar to discriminate between items in the search space. We test our approach on the Wall Street Journal corpus and show that it outperforms Hirst and Budanitsky’s WordNet-based method and Wilcox-O’Hearn, Hirst, and Budanitsky’s fixed windows size method.
The paper addresses the problem of representation of grammatical information in dictionaries and reference-books by considering one particular title focused on the oikonymic system of the Volgograd region. A new concept of grammatical description is introduced, which is the degree of uniqueness of a geographical name within the region. Developing the rating scale for this parameter required consideration of such linguistic criteria as the word-formation pattern, the identity of the root morpheme or the entire lexical unit, the morphemic affi nity of the names with the same root. The new parameter proves particularly relevant in the view of occasional changes in the administrative-territorial division of the region, especially when deciding whether it is necessary/unnecessary to rename a particular locality. The study also takes to specify the grammatical norm regulating the choice of the case form of the toponym at its use with the generic term. Given the signifi cant number of differential features that are complementary, rather than hierarchical, it is proposed that each dictionary article has to include specific recommendations on the choice of the form of the proper name. Different patterns in the names of residents of particular settlements (katoikonyms) are described, related to the frequency of their use in regional, city and district newspapers. The authors also reveal the factors these variants of katoikonyms may be caused by, some pertinent to a certain historical period or else resulting from the peculiarities of the local toponymic system. The paper presents the structure of an article adopted for the forthcoming reference book the authors are working on, along with the samples of dictionary articles.
Considering the importance of rational, effective professional language (for thinking and communication), the subject of the study, the results of which we presented in this article, is the analysis of the consistency of logic terms with the norms of Ukrainian terminological standards. This article is about the observance of linguistic norms, in particular, giving preference to Ukrainian-speaking terms against foreign-language ones (a very large percentage of foreign-language terms in the field of logic is evident) as well as the delimitation of the names of action, events and consequences by the form of the word. Having achieved these objectives would, at the same time, lead to the adoption of terms of logic and unification of formally-linguistic means during the creation of terms. The research was to do the following: 1) the discovery of those terms of logic that do not meet the requirements of the DSTU on terminology; 2) the analysis of the possible ways to achieve the correspondence between terms and terminological standards of Ukraine. As a result of the research, we have constructed the series of interrelated process logic terms: name of the action by the verb – name of the action by verbal noun – name of the completed action (event) by the verb – name of the event by the verbal noun – name of the consequence of action by the verbal noun. We constructed such rows for terms that are the names of operations for obtaining new knowledge: generalization, restriction, derivation, proof, refutation, making a conclusion, deduction, making an assumption, making of a hypothesis, implication. We used the following requirements during the construction of these series of terms: 1) all the terms of each series must be created on the same lexical basis; 2) As the terms we should use such words, in which the meaning of the word is consistent with its form, that is, if the main purpose in accordance with a particular form of the word is an action, an event, or a consequence of an event, then the meaning of a word must be the action, the event or the consequence of the event, respectively. If, for example, the main purpose of the verbal noun with suffixes ‑annia, -ennia is to denote an unfinished or completed action, then it is incorrect to use it to denote the consequence of an event. When creating a system of terms, we have established the following relationships between them: proof – is a deduction in the case, when as a conclusion we try to confirm the truthfulness of a given thesis; refutation – is a deduction in the case, when as a conclusion we try to confirm the falseness of a predefined thesis; derivation – is a deduction in the case, when the desired conclusion is not predefined; making an assumption – is the creation of allegedly true affirmations; making of a hypothesis – is the creation of probably true affirmations in the field of science. As a result of the study of the coherence of widely used logic terms with the requirements of terminological standards, we have found out that in the Ukrainian terminology system of logic there is a large number of terms that are not consistent with the terminological standards and, therefore, we proposed new linguistically correct terms (in a number of cases we proposed linguistically correct termsduplicates in order to be able to choose the most appropriate one).
Design and implementation of automatic evaluation methods is an integral part of any scientific research in accelerating the development cycle of the output. This is no less true for automatic machine translation (MT) systems. However, no such global and systematic scheme exists for evaluation of performance of an MT system. The existing evaluation metrics, such as BLEU, METEOR, TER, although used extensively in literature have faced a lot of criticism from users. Moreover, performance of these metrics often varies with the pair of languages under consideration. The above observation is no less pertinent with respect to translations involving languages of the Indian subcontinent. This study aims at developing an evaluation metric for English to Hindi MT outputs. As a part of this process, a set of probable errors have been identified manually as well as automatically. Linear regression has been used for computing weight/penalty for each error, while taking human evaluations into consideration. A sentence score is computed as the weighted sum of the errors. A set of 126 models has been built using different single classifiers and ensemble of classifiers in order to find the most suitable model for allocating appropriate weight/penalty for each error. The outputs of the models have been compared with the state-of-the-art evaluation metrics. The models developed for manually identified errors correlate well with manual evaluation scores, whereas the models for the automatically identified errors have low correlation with the manual scores. This indicates the need for further improvement and development of sophisticated linguistic tools for automatic identification and extraction of errors. Although many automatic machine translation tools are being developed for many different language pairs, there is no such generalized scheme that would lead to designing meaningful metrics for their evaluation. The proposed scheme should help in developing such metrics for different language pairs in the coming days.
The Trail Making Test (TMT) is used in neuropsychological clinical practice to assess aspects of attention and executive function. The test consists of two parts (A and B) and requires drawing a trail between elements. Many patients are assessed with their non-dominant hand because of motor dysfunction that prevents them from using their dominant hand. Since drawing with the non-dominant hand is not an automatic task for many people, we explored the effect of hand use on TMT performance. The TMT was administered digitally in order to analyze new outcome measures in addition to total completion time. In a sample of 82 healthy participants, we found that non-dominant hand use increased completion times on the TMT B but not on the TMT A. The average completion time increased by almost 5 seconds, which may be clinically relevant. A substantial number of participants who performed the TMT with their non-dominant hand had a B/A ratio score of 2.5 or higher. In clinical practice, an abnormally high B/A ratio score may be falsely attributed to cognitive dysfunction. With our digitized pen data, we further explored the causes of the reduced TMT B performance by using new outcome measures, including individual element completion times and interelement variability. These measures indicated selective interference between non-dominant hand use and executive functions. Both non-dominant hand use and performance of the TMT B seem to draw on the same, limited higher-order cognitive resources.
Recent years have seen an increased interest in machine learning-based predictive methods for analyzing quantitative behavioral data in experimental psychology. While these methods can achieve relatively greater sensitivity compared to conventional univariate techniques, they still lack an established and accessible implementation. The aim of current work was to build an open-source R toolbox – “PredPsych” – that could make these methods readily available to all psychologists. PredPsych is a user-friendly, R toolbox based on machine-learning predictive algorithms. In this paper, we present the framework of PredPsych via the analysis of a recently published multiple-subject motion capture dataset. In addition, we discuss examples of possible research questions that can be addressed with the machine-learning algorithms implemented in PredPsych and cannot be easily addressed with univariate statistical analysis. We anticipate that PredPsych will be of use to researchers with limited programming experience not only in the field of psychology, but also in that of clinical neuroscience, enabling computational assessment of putative bio-behavioral markers for both prognosis and diagnosis.
Language culture creation is one of the most urgent questions nowadays. This is not only philological problem, but social as well – as it is related to different communication methods.The article covers linguistic principles of language culture creation for pupils provided dialect environment. Proved that the necessary condition for high level language culture for future primary school teachers provided dialect environment is compliance principles of oral speaking: orthoepic, lexical, grammar, stylistic. The most important their properties are accuracy, cleanliness, purity etc.Also there is covered speech environment role in creating language culture of individual.
 We determine language culture for junior pupilsas possession of verbal and written forms of language on all levels, ability to use optimal language tools for current situation. Language norm is main concept of language culture. We believe that main requirement for any spoken phrase is its correctness. As a result of these factors, requirements for communication are created. We thought that during junior pupils’ speech improving the primary importance is work on language accuracy. Non-normative accents and speaking are often effect of negative impact of dialect environment on junior pupils. And this danger stores permanently.Іnformation technologies help to individualize and differentiate the studies of Ukrainian in initial classes. The uses of ICT do the lessons of Ukrainian and reading dynamic, bright, more effective.
 Improvement language culture for future primary school teacher is an integral part of the formation of his professiogram. Language environment is important factor for creating language culture. Dialect environment has both positive and negative influence. The worth-while experiment of the use of ICT at initial school we saw at Ivano-Frankivsk school №26. In spring in 2014 department of education entered in Ukraine a pedagogical experiment «Smart Kids». Within the framework of this experiment in the initial classes of school set projectors and interactive boards on that children execute educational tasks in a playing form. Games are a didactics, bright and interesting. So, regional dialects may do speech richer, but at the same time do it more complex: phonetically dialects are understandable by all speakers, however lexical are not understandable for people from another regions. Using dialects by students is natural phenomenon. This communication provides tight connection between history, way of life, customs of his native land.