Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
The languages developed by deaf communities are unique for using visual signs produced by the hand. In the present study, we explored the cognitive effects of employing the hand as articulator. We focused on the arbitrariness of the form-meaning relationship—a fundamental feature of natural languages—and asked whether sign languages change the processing of arbitrary non-linguistic stimulus-response (S-R) associations involving the hand. This was tested using the Simon effect, which specifically requires such type of associations. Differences between signers and speakers (non-signers) only appeared in the Simon task when hand stimuli were shown. Response-time analyses revealed that the distinctiveness of signers’ responses derived from an increased ability to process memory traces of arbitrary S-R pairs related to the hand. These results shed light on the interplay between language and cognition as well as on the effects of sign language acquisition. [ABSTRACT FROM AUTHOR], Copyright )
We investigated the effect of auditory noise added to speech on patterns of looking at faces in 40 toddlers. We hypothesised that noise would increase the difficulty of processing speech, making children allocate more attention to the mouth of the speaker to gain visual speech cues from mouth movements. We also hypothesised that this shift would cause a decrease in fixation time to the eyes, potentially decreasing the ability to monitor gaze. We found that adding noise increased the number of fixations to the mouth area, at the price of a decreased number of fixations to the eyes. Thus, to our knowledge, this is the first study demonstrating a mouth-eyes trade-off between attention allocated to social cues coming from the eyes and linguistic cues coming from the mouth. We also found that children with higher word recognition proficiency and higher average pupil response had an increased likelihood of fixating the mouth, compared to the eyes and the rest of the screen, indicating stron)
This paper deals with the skills related to the early reading acquisition in two countries that share language. Traditionally on reading readiness research there is a great interest to find out what factors affect early reading ability, but differ from other academic skills that affect general school learnings. Furthermore, it is also known how the influence of pre-reading variables in two countries with the same language, affect the development of the reading. On the other hand, several studies have examined what skills are related to reading readiness (phonological awareness, alphabetic awareness, naming speed, linguistic skills, metalinguistic knowledge and basic cognitive processes), but there are no studies showing whether countries can also influence the development of these skills.Our main objective in this study was to establish whether there were differences in the degree of acquisition of these skills between Spanish (119 children) and Peruvian (128 children), five years old)
Ancestral Polynesian society is the formative base for development of the Polynesian cultural template and proto-Polynesian linguistic stage. Emerging in western Polynesia ca 2700 cal BP, it is correlated in the archaeological record of Tonga with the Polynesian Plainware ceramic phase presently thought to be of approximately 800 years duration or longer. Here we re-establish the upper boundary for this phase to no more than 2350 cal BP employing a suite of 44 new and existing radiocarbon dates from 13 Polynesian Plainware site occupations across the extent of Tonga. The implications of this boundary, the abruptness of ceramic loss, and the shortening of duration to 350 years have substantive implications for archaeological interpretations in the ancestral Polynesian homeland. [ABSTRACT FROM AUTHOR], Copyright of PLoS ONE is the property of Public Library of Science and its content may not be copied or emailed to multiple sites or posted to a listserv without the copyright holder's ex)
Manual action verbs modulate the right-hand grip force in right-handed subjects. However, to our knowledge, no studies demonstrate the ability to accomplish this modulation during bimanual tasks nor describe their effect on left-hand behavior in unimanual and bimanual tasks. Using load cells and word playlists, we evaluated the occurrence of grip force modulation by manual action verbs in unimanual and symmetrical bimanual tasks across the three auditory processing phases. We found a significant grip force increase for all conditions compared to baseline, indicating the occurrence of modulation. When compared to each other, the grip force variation from baseline for the three phases of both hands in the symmetrical bimanual task was not different from the right-hand in the unimanual task. The left-hand grip force showed a lower amplitude for auditory phases 1 and 2 when compared to the other conditions. The right-hand grip force modulation became significant from baseline at 220 ms af)
In this study, we leverage human evaluations, content analysis, and computational modeling to generate a comprehensive analysis of readers’ evaluations of authors’ communication quality in social media with respect to four factors: author credibility, interpersonal attraction, communication competence, and intent to interact. We review previous research on the human evaluation process and highlight its limitations in providing sufficient information for readers to assess authors’ communication quality. From our analysis of the evaluations of 1,000 Twitter authors’ communication quality from 300 human evaluators, we provide empirical evidence of the impact of the characteristics of the reader (demographic, social media experience, and personality), author (profile and social media engagement), and content (linguistic, syntactic, similarity, and sentiment) on the evaluation of an author’s communication quality. In addition, based on the author and message characteristics, we demonstrate)
Background: Idea density (ID), a natural language processing–based index, was developed to aid in the detection of dementia through the analysis of English narratives. However, it has not been applied to non-English languages due to the difficulties in translating grammatical concepts. In this study, we defined rules to count ideas in Japanese narratives based on a previous study and proposed a novel method to estimate ID in Japanese text using machine translation. Materials: The study participants comprised 42 Japanese patients with dementia aged 69–98 years (mean: 84.95 years). We collected free narratives from the participants to build a speech corpus. The narratives of the patients were translated into English using three machine translation systems: Google Translate, Bing Translator, and Excite Translator. The ID in the translated text was then calculated using the Dependency-based Propositional ID (DEPID), an English ID scoring tool. Results: The maximum correlation coefficie)
Emotions have crucial influence on vocabulary learning and text comprehension. However, whether morphosyntactic learning is influenced by emotional conditions has remained largely unclear. In this study, we investigated how induced positive and negative emotions affect the learning of morphosyntactic rules in a foreign language. It was found that negative emotion increased the accuracy and efficiency of syntactic learning, but had no significant effect on the learning of morphological marking rules. Positive emotion was not found to be significantly associated with learning outcomes. The findings shed light on the effects of affective states on the structural aspects of foreign language learning. [ABSTRACT FROM AUTHOR], Copyright of PLoS ONE is the property of Public Library of Science and its content may not be copied or emailed to multiple sites or posted to a listserv without the copyright holder's express written permission. However, users may print, download, or email articles fo)
We analyze the model of social interactions with coevolution of the topology and states of the nodes. This model can be interpreted as a model of language change. We propose different rewiring mechanisms and perform numerical simulations for each. Obtained results are compared with the empirical data gathered from two online databases and anthropological study of Solomon Islands. We study the behavior of the number of languages for different system sizes and we find that only local rewiring, i.e. triadic closure, is capable of reproducing results for the empirical data in a qualitative manner. Furthermore, we cancel the contradiction between previous models and the Solomon Islands case. Our results demonstrate the importance of the topology of the network, and the rewiring mechanism in the process of language change. [ABSTRACT FROM AUTHOR], Copyright of PLoS ONE is the property of Public Library of Science and its content may not be copied or emailed to multiple sites or posted to a l)
Speech understanding can be thought of as inferring progressively more abstract representations from a rapidly unfolding signal. One common view of this process holds that lower-level information is discarded as soon as higher-level units have been inferred. However, there is evidence that subcategorical information about speech percepts is not immediately discarded, but is maintained past word boundaries and integrated with subsequent input. Previous evidence for such subcategorical information maintenance has come from paradigms that lack many of the demands typical to everyday language use. We ask whether information maintenance is also possible under more typical constraints, and in particular whether it can facilitate accent adaptation. In a web-based paradigm, participants listened to isolated foreign-accented words in one of three conditions: subtitles were displayed concurrently with the speech, after speech offset, or not displayed at all. The delays between speech offset and)
This paper describes a quick aphasia battery (QAB) that aims to provide a reliable and multidimensional assessment of language function in about a quarter of an hour, bridging the gap between comprehensive batteries that are time-consuming to administer, and rapid screening instruments that provide limited detail regarding individual profiles of deficits. The QAB is made up of eight subtests, each comprising sets of items that probe different language domains, vary in difficulty, and are scored with a graded system to maximize the informativeness of each item. From the eight subtests, eight summary measures are derived, which constitute a multidimensional profile of language function, quantifying strengths and weaknesses across core language domains. The QAB was administered to 28 individuals with acute stroke and aphasia, 25 individuals with acute stroke but no aphasia, 16 individuals with chronic post-stroke aphasia, and 14 healthy controls. The patients with chronic post-stroke aph)
In this paper we discuss the project of digitization of the Dictionary of the Serbo-Croatian Standard and Vernacular \nLanguage. Scanning and character recognition were a particular challenge, since various non-standard \ncharacter set encoding was used in the course of the almost 60-year long production of the dictionary. The first \naim of the project was to formalize the micro-structure of the dictionary articles in order to parse the digitized \ntext of and transform it into structured data stored in relational lexical database. This approach is compatible \nwith several standard structured forms and ontologies (TEI, LMF, Ontolex, LexInfo). A lexical database model \nwas designed in compliance with these structured forms, following mostly the lemon model. Mapping of \nthe lexical entry markers to LexInfo and TEI enabled export of the lexical data to the mentioned formats. A \nsoftware solution for the dictionary text analysis, parsing and lexical database population was developed and \ntested on the first and the last published volumes of the dictionary (which contain 27,141 articles in total). An \nevaluation of the results shows that the developed model and software solution can be successfully used for \nthe other volumes as well.
We present a method for detecting annotation errors in manually and automatically annotated dependency parse trees, based on ensemble parsing in combination with Bayesian inference, guided by active learning.We evaluate our method in different scenarios: (i) for error detection in dependency treebanks and (ii) for improving parsing accuracy on in-and out-of-domain data.
Version 1 of the Late Latin Charter Treebank (LLCT1). Early medieval Latin documentary texts with morphological and syntactic annotation. Ancient Language Dependency Treebank compatible linguistic annotation, Prague style treebank format (PML). See full description in Korkiakangas & Lassila, 2013, "Abbreviations, fragmentary words, formulaic language: treebanking medieval charter material". Will be replaced by LLCT2 in 2018 (version 2).
A simple multiple imputation-based method is proposed to deal with missing data in exploratory factor analysis. Confidence intervals are obtained for the proportion of explained variance. Simulations and real data analysis are used to investigate and illustrate the use and performance of our proposal.
The syntax and semantics of human language can illuminate many individual psychological differences and important dimensions of social interaction. Accordingly, psychological and psycholinguistic research has begun incorporating sophisticated representations of semantic content to better understand the connection between word choice and psychological processes. In this work we introduce ConversAtion level Syntax SImilarity Metric (CASSIM), a novel method for calculating conversation-level syntax similarity. CASSIM estimates the syntax similarity between conversations by automatically generating syntactical representations of the sentences in conversation, estimating the structural differences between them, and calculating an optimized estimate of the conversation-level syntax similarity. After introducing and explaining this method, we report results from two method validation experiments (Study 1) and conduct a series of analyses with CASSIM to investigate syntax accommodation in social media discourse (Study 2). We run the same experiments using two well-known existing syntactic metrics, LSM and Coh-Metrix, and compare their results to CASSIM. Overall, our results indicate that CASSIM is able to reliably measure syntax similarity and to provide robust evidence of syntax accommodation within social media discourse.
The English language has evolved dramatically throughout its lifespan, to the extent that a modern speaker of Old English would be incomprehensible without translation. One concrete indicator of this process is the movement from irregular to regular (-ed) forms for the past tense of verbs. In this study we quantify the extent of verb regularization using two vastly disparate datasets: (1) Six years of published books scanned by Google (2003–2008), and (2) A decade of social media messages posted to Twitter (2008–2017). We find that the extent of verb regularization is greater on Twitter, taken as a whole, than in English Fiction books. Regularization is also greater for tweets geotagged in the United States relative to American English books, but the opposite is true for tweets geotagged in the United Kingdom relative to British English books. We also find interesting regional variations in regularization across counties in the United States. However, once differences in population ar)
The broad use of computer-supported collaborative-learning (CSCL) environments (e.g., instant messenger–chats, forums, blogs in online communities, and massive open online courses) calls for automated tools to support tutors in the time-consuming process of analyzing collaborative conversations. In this article, the authors propose and validate the cohesion network analysis (CNA) model, housed within the ReaderBench platform. CNA, grounded in theories of cohesion, dialogism, and polyphony, is similar to social network analysis (SNA), but it also considers text content and discourse structure and, uniquely, uses automated cohesion indices to generate the underlying discourse representation. Thus, CNA enhances the power of SNA by explicitly considering semantic cohesion while modeling interactions between participants. The primary purpose of this article is to describe CNA analysis and to provide a proof of concept, by using ten chat conversations in which multiple participants debated the advantages of CSCL technologies. Each participant’s contributions were human-scored on the basis of their relevance in terms of covering the central concepts of the conversation. SNA metrics, applied to the CNA sociogram, were then used to assess the quality of each member’s degree of participation. The results revealed that the CNA indices were strongly correlated to the human evaluations of the conversations. Furthermore, a stepwise regression analysis indicated that the CNA indices collectively predicted 54% of the variance in the human ratings of participation. The results provide promising support for the use of automated computational assessments of collaborative participation and of individuals’ degrees of active involvement in CSCL environments.
The author attempts to identify the place for the language norms in the modern library verbal communications. The relevancy of this study is determined by the crisis of verbal communications in the society (overuse of slang, frivolous borrowings from foreign languages, decay of the language culture, using gadgets in interpersonal and business communications). The author concludes that verbal communication skills are needed to harmonize all communicative components, including language norms and speech culture. The findings of the monitoring survey of library verbal communications are presented. The level of command of the modern Russian language norms is identified for library and information specialists. It was found that the largest part of respondents had insufficient knowledge of language norms which affects their professional activity. The author analyzes orthoepical (pronouncing and articulatory) errors, violation of lexical, grammar (word-formation, morphological and syntactic) and stylistic norms of the Russian language, observance of orthographic and punctuation rules in the written speech. She argues that revealing the gaps helps to define the problems of librarians’ speech culture and to increase the efficiency of verbal communications in the professional library environment, in particular within the system of advanced training.
This article is devoted to analyse the stylistic characteristics of the verbal innovations taken from the German magazines. Each functional style has its own special features. Stylistic peculiarities of the modern German journalism consist in the evaluative connotation, in the metaphorical usage of the verbs and in the usage of grammar categories of the verbs. The evaluative connotation can beshown in semantics of the whole word as well as in its components. There are 27 evaluative models and only 3 of them are not active. The metaphorical sense of the innovations lies in the usage of the verbs in the fi gurative meaning, in the usage of the lexical items in the unusual communicative situation and in the obtaining of new shades of meaning. The grammatical peculiarities of the journalistic style consist in the digression of the grammar rules and norms of German in the usage of such categories as person, number, tense, voice and mood.
The article considers variants of paronyms compatibility and some mistakes of Turkmen students in the use of paronyms – words closed in pronunciation but not identical, and sometimes different in meaning. The vocabulary of foreign students entering the university does not always correspond to the needs of their language practice, often there is no knowledge of the paronyms necessary for it. The study of various paronyms combining variants with other words, semantic connections are as necessary as knowledge of grammatical rules and spelling norms. The correct using of the word in speech suggests, firstly, knowledge of the word structure and the lexical-semantic variants of the polysemy; secondly, analysis of the word-building composition with the individual morphemes meaning identification; thirdly, the ability to choose the word needed for a given context, for which it is necessary to determine its place in the lexical-semantic system, i.e. to find its connection with other words, to identify general and differential features in a given lexical group making up a certain semantic unity. As a result of observations of Turkmen students“ oral and written speech, the main reasons for the erroneous paronyms substitution are revealed, involving roots consonance. The same root paronyms are also confused because of inaccurate understanding of the prefixes, suffixes meaning and difference they bring to the word meaning. Considering the same root-words, there are many problems due to the fact that they are heterogeneous in the semantic and word-formation aspect. Paronyms like the same-root words, similar to synonyms. They have a number of common features which leads to the similarity of these two language units, and finally to the error occurrence while using paronyms in speech. The article also contains examples of training exercises for speech skills developing in the paronyms using at Russian language classes in the Turkmen audience.
In this article there are given some data about a brief observation of studying lexical system of the Uzbek language, the names of scientists, whose efforts have helped arouse opportunities of the language.This article thoroughly highlights several studies done in the field of lexical system, especially semantic features of terms relating to food, clothes and architecture. The author of this article states that ornithological terms which are considered as a part of zoonims, can be the subject of special investigation for linguists. And also author studied the some birds’ names are not included in literary language or their literary norms are not assigned clearly for not having been studied thoroughly.
We examined the effect of language proficiency on the status and dynamics of proactive inhibitory control in an occulo-motor cued go-no-go task. The first experiment was designed to demonstrate the effect of second language proficiency on proactive inhibitory cost and adjustments in control by evaluating previous trial effects. This was achieved by introducing uncertainty about the upcoming event (go or no-go stimulus). High- and low- proficiency Hindi-English bilingual adults participated in the study. Saccadic latencies and errors were taken as the measures of performance. The results demonstrate a significantly lower proactive inhibitory cost and better up-regulation of proactive control under uncertainty among high- proficiency bilinguals. An analysis based on previous trial effects suggests that high- proficiency bilinguals were found to be better at releasing inhibition and adjustments in control, in an ongoing response activity in the case of uncertainty. To further understand )
The article is devoted to the study of associative field of the word using the method of experiment. The paper presents a brief review of the definitions of concepts such as: free associative experiment, association, content analysis. The article attempts to analyze the results of the free associative experiment, associative field study of steppe. The factual material for illustration of the main provisions was the direct lexical association of the respondents. The experimental data allow us to observe mental stereotypes of society, to reveal its cultural memory, modern values verbalized in associations. The following methods were used in the research: questionnaire survey, descriptive, generalization, systematization, observation, content analysis.
Building on the literature that approaches self-disclosure as a decision-making process, we proposed a self-reported Sensitive Information Disclosure (SID) measure and tested the measure’s reliability and validity in two studies across a variety of interview modes and settings. We used theory to identify potential dimensions of sensitive information disclosures, created potential scale items, performed two separate card sorts, and validated the resulting pool of items in two separate experiments. Participants answered the SID scale items following an interview involving sensitive information, potential risk, and after-disclosure vulnerability. Study 1 was a laboratory experiment conducted with 165 university students. Exploratory factor analysis results revealed a two-factor structure, Personal Discomfort and Revealing Personal Information. Study 2 replicated these procedures using confirmatory factor analysis to confirm the factor structure and demonstrate the scale’s reliability and validity, with a sample of 77 students and 275 participants from Amazon’s M-Turk. Together, these results demonstrate that the proposed 11-item SID scale has good convergent and discriminant validity as well as good reliability. A quasi-experimental application of the measure is illustrated using the substantive findings from Study 2. This research fills a gap in the literature by developing a topic-free scale to measure SID as a dependent variable. The ability to accurately measure sensitive information disclosure is an important and necessary step toward developing a more thorough understanding of how people feel and react when asked to provide personal information in diverse interview settings.
This study examines Swedish morning TV’s framing of the phenomenon of exercising. Morgonstudion in SVT, and Nyhetsmorgon in TV4, is the morning shows that has been investigated. The aim of the study is to enlighten and enhance the understanding of how the phenomenon of exercising are framed in Swedish morning television, as well as to contribute to the theorization of media’s representation of exercise. A qualitative content analysis has been used to capture the language, to see how they present exercising and if they legitimate it. Theoretical framework applied are framing theory, representation, legitimize, healthism and public service vs commercial television. The research fields are health communication, exercise in media and morning television journalism. The result shows that Swedish morning television, through different methods and lexical choices, legitimize the phenomenon of exercise. SVT is more focused on exercise that suits everyone and that viewers can change their lifestyle by making small changes in their everyday life. For example, they have an idea of how to get the pulse up by exercises in the garden while TV4 is aiming at reaching out to those that is already exercising and specifies types of exercising during the show. They give advice on how to find out what training tools you need to complete a triathlon for example. Lexical choices reinforce that exercise is something positive and both the guests and hostess sees exercise as a norm. The elements of exercise differ between the programs, as well as the studio environment and the content. On the other hand, there are similarities such as the language that is used and they are both visited by experts.
This paper tests the new-dialect formation model of Peter Trudgill (1986 et seq) by examining several phonological features of Tibetan as spoken in the diaspora community of Kathmandu, Nepal. Established by an influx of migrants from many dialect regions beginning in 1959, this presents a unique opportunity to study koinéization, new dialect formation, in progress. Trudgill’s model predicts that a new dialect should largely emerge in the second generation born in the new region, exhibiting both simplification, the failure of marked variants to transmit across generations, and focusing, the selection of particular variants as a new norm for the community's new variety.Data from seventy-three sociolinguistic interviews was coded for phonological and lexical variables known to differ across Tibetan-speaking regions, and NeighborNets were constructed in SplitsTree. Results indicate that regionally marked variables were not transmitted into the first or second generation of Diaspora-raised speakers, but Diaspora speakers exhibited a high degree of variation comparable to that of speakers from the numerically- and socially-dominant U-Tsang region. That younger speakers have not yet converged on a single new variety suggests a role for additional factors to affect the rate of koinéization.
he object of this paper is the variant of quasi-standard language, i.e. the variant of the perceived standard language formed by the young generation of the Aukštaitian area. The aims of the study are to examine whether the assessments made by respondents representing the Aukštaitian area suggest the presence of such a quasi-standard and, if they do, to provide the characterisation of this quasi-standard variant on the basis of the collected data. The data of the study consists of eight audio texts-stimuli which represent six Lithuanian regiolects (A and D represent Southern Aukštaitian; B and E represent Žemaitian, C represents Northwestern Aukštaitian, F represents the western part of Eastern Aukštaitian, G represents Southwestern Aukštaitian, while H represents the eastern part of East Aukštaitian) and the responses to two questions in the questionnaire designed according to the principles of perceptual dialectology which ask the respondents to rate the similarity of the audio text-stimulus to the standard language. The study demonstrated that the respondents perceived texts-stimuli B and E (representative of the Žemaitian dialect) as the least similar to the standard language. Texts-stimuli H (representative of the eastern part of East Aukštaitian regiolect), A and D (representative of East Aukštaitian regiolect), on the contrary, were seen as the closest to the standard language. Since the ratings of these texts-stimuli in the respondents’ assessment were substantially higher in comparison to the rest of the texts-stimuli, the results suggest the existence of a quasi-standard. The analysis of respondents’ motives of giving high scores to audio texts-stimuli A, D, and H demonstrates that the morphological and lexical characterisation of the quasi-standard of the young generation representing the Aukštaitian area is only fragmentary. The most prominent are phonetic features, namely: more open and more closed pronunciation of vowels i and u, shortening of unstressed long vowels o, u, and i, lengthening of the stressed short vowels u and i, correct accentuation and non-reduced endings. Based on the analysis carried out, it is possible to assume that the quasi-standard variety formed by the young generation representing the Aukštaitian area consists of some tertiary phonetic features of the eastern parts of East Aukštaitian and South Aukštaitian regiolectal zones, norms of standard language pronunciation, shortened verb forms typically characteristic of dialects and standard language and mixed lexis (containing that of standard language / dialects / borrowings).
Dans le cadre général d’une sémiotique des cultures, cette recherche a utilisé les propositions épistémologiques et méthodo-logiques de la sémantique interprétative pour renouveler l’analyse de textes irlandais médiévaux. Le but était d’apporter une contribution à une problématique générale intéressant les sciences du langage, mais aussi les sciences historiques: comment fonder la pertinence scientifique d’une interprétation de textes et signes anciens appartenant à une culture différente? Pour cela il a été choisi de viser les faits sémantiques qui interviennent dans les processus de transfert de sens que la tradition rhétorique nomme comparaison, métaphore ou symbole. L’approche méthodologique a nécessité l’édition d’un corpus interlinéaire offrant un accès direct aux données de l’Electronic Dictionary of the Irish Language. L’étude du lexique a confronté les possibilités définitoires aux afférences contextuelles observées par un relevé systématique des isotopies ciblées.Sur le plan sémantique, l’analyse des processus différentiels qui structurent les molécules sémiques a permis d’observer la circulation des sèmes marquant les analogies intentionnelles. Sur le plan diachronique, la description du système de valeur, pris dans sa globalité, fonde la pertinence de l’interprétation en ce qu’il intègre les normes sociales du contexte historique du signe. Sur le plan des études celtiques, l’analyse des correspondances entre les domaines de l’orientation spatiale, des cycles temporels et des fonctions sociales donne un nouvel accès à la complexité du système de pensée de cette culture. Les formes sémantiques décrites fournissent de nouveaux modèles pour des comparaisons. Sur cette base, les expressions de l’association arbre-savoir ont été décrites pour apporter une solution aux problèmes de l’étymologie de la lexie druid- et proposer le dépassement des approches lexicales monographiques par l’approche intertextuelle.
Human free association (FA) norms are believed to reflect thestrength of links between words in the lexicon of an averagespeaker. Large-scale FA norms are commonly used as a datasource both in psycholinguistics and in computational mod-eling. However, few studies aim to analyze FA norms them-selves, and it is not known what are the most important factorsthat guide speakers’ lexical choices in the FA task. Here, wefirst provide a statistical analysis of a large-scale data set ofEnglish FA norms. Second, we argue that such analysis caninform existing computational models of semantic memory,and present a case study with the topic model to support thisclaim. Based on our analysis, we provide the topic model withdictionary-based knowledge about word synonymy/antonymy,and demonstrate that the resulting model predicts human FAresponses better than the topic model without this information.
In this study we developed and evaluated a crowdsourcing-based latent semantic analysis (LSA) approach to computerized summary scoring (CSS). LSA is a frequently used mathematical component in CSS, where LSA similarity represents the extent to which the to-be-graded target summary is similar to a model summary or a set of exemplar summaries. Researchers have proposed different formulations of the model summary in previous studies, such as pregraded summaries, expert-generated summaries, or source texts. The former two methods, however, require substantial human time, effort, and costs in order to either grade or generate summaries. Using source texts does not require human effort, but it also does not predict human summary scores well. With human summary scores as the gold standard, in this study we evaluated the crowdsourcing LSA method by comparing it with seven other LSA methods that used sets of summaries from different sources (either experts or crowdsourced) of differing quality, along with source texts. Results showed that crowdsourcing LSA predicted human summary scores as well as expert-good and crowdsourcing-good summaries, and better than the other methods. A series of analyses with different numbers of crowdsourcing summaries demonstrated that the number (from 10 to 100) did not significantly affect performance. These findings imply that crowdsourcing LSA is a promising approach to CSS, because it saves human effort in generating the model summary while still yielding comparable performance. This approach to small-scale CSS provides a practical solution for instructors in courses, and also advances research on automated assessments in which student responses are expected to semantically converge on subject matter content.
The study of translation norms is one of the areas in translation studies which identify regularities of behavior (i.e. trends of relationships and correspondences between ST and TT segments) by comparing source texts and their translations. Norms of translation are mostly done in areas other than religious texts. Therefore, it seems necessary to do a research on religious texts. Textual–linguistic norms govern the selection of TT linguistic material: lexical items, phrases, and stylistic features. To do so, translation strategies adopted by translators were identified through comparing translations and source texts. Translation strategies proposed by Chesterman (1997) are investigated in samples of texts translated by World Ahlubayt assembly, an organization in charge of religious translation in Iran. The texts included seven books from seven translators in World Ahlulbayt Assembly. The strategies investigated in corpus dealt with three linguistic levels: semantic, syntactic and pragmatic strategies and changes done at these three levels. The results showed that syntactic changes were of the highest frequency in all texts. At semantic level, synonymy was the most frequent translation strategy. At syntactic level, clause structure changes and at pragmatic level and explicitness change were the most frequent changes.
What happens when a new social convention replaces an old one? While the possible forces favoring norm change-such as institutions or committed activists-have been identified for a long time, little is known about how a population adopts a new convention, due to the difficulties of finding representative data. Here, we address this issue by looking at changes that occurred to 2,541 orthographic and lexical norms in English and Spanish through the analysis of a large corpora of books published between the years 1800 and 2008. We detect three markedly distinct patterns in the data, depending on whether the behavioral change results from the action of a formal institution, an informal authority, or a spontaneous process of unregulated evolution. We propose a simple evolutionary model able to capture all of the observed behaviors, and we show that it reproduces quantitatively the empirical data. This work identifies general mechanisms of norm change, and we anticipate that it will be of interest to researchers investigating the cultural evolution of language and, more broadly, human collective behavior.
In the second half of the nineteenth century, the Romanian literary language lexicon, illustrated by literary and scientific texts, periodicals, manuals, underwent a process of renewal through an extensive acceptance of neologisms, in the awaited correlation with the overall modernization of culture. Dicţionarul limbei române, an academic project, printed between 1871 and 1877, tries to impose an ideal norm, that aimed, consistent with a certain linguistic conception, at preserving and enriching the native Latin vocabulary, by selecting certain loans and lexical creations, concurrently with the exclusion of variants (phonetic, morphological or lexical variants) that conflicted with the „spirit” of the Romanian language.
This chapter explores the written performance of Arabic native speakers (NSs) and upper-level/advanced learners of Arabic via a presentation of descriptive statistics of a number of direct measures of written complexity, accuracy, and fluency (CAF). It discusses that upper-level learners indeed resemble NSs in many ways, although NSs are shown to be more complex, accurate, and fluent writers. Issues of Arabic literacy and literacy development are of critical importance to Arabic learners, educators, and millions of native speakers. Instances of NS production that were coded as an error involved departures from the morpho-syntactic norms of Modern Standard Arabic (MSA) rather than the use of colloquial lexical items or common spelling variants. The CAF framework is typically employed in relation to Skehan's trade-off hypothesis, in which learners are assumed to possess finite attentional resources available for devotion to either linguistic complexity, accuracy, or fluency in their L2 production.
The topic suggested for this paper is the effect of linguistic developments and synchronic standardization. I take the assumption to be tested here is the hypothesis to the effect that language change is as well as was in the past balanced between language development and standardization. I illustrate this with one aspect of the recent lexical diffusion or spreading of the innovation paradigms of verb class ‘바라->바래-’(to wish) in addition to other similar cases of verb stems, such as ‘놀라->놀래-(to be surprised), 모자라->모자래-(to be deficient), 나무라->나무래-(to reproach)’ et cetera in now-days Korean. As a result, out of this study concerned, I could draw a concluding remark that a on-going natural morphological change with regard to suffix -i in synchronic Korean can be slowed or stymied by the process of standardization and enforcement of strict written various form of norms. All in all, I d like to suggest in this paper that the tension between natural developments in folk spoken language and strict standardization has led to a linguistic rift between every day life of communication and Modern Korean.
Abstract The article deals with basic requirements to the translation for specific purposes, namely legal translation. The problem posed here is defining object and theoretical basis of legal translation. The question of the necessity of information search as an integral part of translation strategy has been raised. Detailed analysis revealed that the requirements of professional translators include knowledge of lexical and grammatical peculiarities of both languages in legal sphere; deep understanding of the concepts employed by specialists in particular field and the specialist terms used to express these concepts and their relationships in the source and target languages. It is recommended that evaluation of the translation may be done on the following principles: communicative pragmatic norms of translation; equivalent norms of translation; absence of contextual, cultural, functional, lexico-grammatical mistakes.
The article is devoted to a modern problems research of television titles editing of the media addressee by the media addressor. The objectives of work are achieved by application of methods of deductive and inductive logical analysis, descriptive method, content analysis, lexical and semantic, lexical and grammatical, and stylistic analysis, comparative analysis, deep interview and poll of informants. Titles as fragments of media texts of the television program “Time Will Show” act as material of the research. In the article, attention is paid to a problem of television titles editing of the mediaaddressee in which editorial work of the media addressor is often limited to an inscription: “The spelling and a punctuation of the author are kept”. Authors in details analyze television titles of the mass media addressee of the “Time Will Show” program broadcast on Channel 1 of the Russian television; sort a number of examples from other elements of media system subject to influence of the research object and from fiction with justified violation of language norms. Authors offer the answer to a question how the editor has to work with text elements on the screen, support the position with opinions of scientists and results of the comparative analysis of the actual material, deep interview and poll of social networks users. The received results are significant for development of psycholinguistics, pragmalinguistics, cognitive linguistics, linguosemiotics, cultural linguistics, discursive linguistics, in particular, of media discourse theory, influence theory. This is because the peculiarities of verbal self-presentation of the media addressee in television titles characterizing the language personality of the media addressee as the media addressor producing the media message as reaction to a television message of the media addressor, and making influence on the mass addressee – television audience are revealed. The article will be useful to philologists, editors, journalists not only in the theoretical plan, but also in practical work.
The article analyses the structure, semantics and word building of the early grammatical terminology as a reflection of spiritual cultural values of the ethnos. The attempts and ways of the authors of the “Grammar book” are shown in order to form native scientific tradition of the definitions of the parts of speech and their categories. The actuality of the subject is stipulated by the necessity to clarify the role of the ancient grammatical science in the development of national peculiarities of the formation of terms, comprehension of the reception of the work of Greek-Roman grammatical science in the formation of Ukrainian morphological terminology. The aim of the work is to study the structure of terms, the regularities of their creation, the study of systemic relations between the units of the terminological vocabulary and the formation of the Ukrainian grammatical terminology in the aspect of modern national terminology through the prism of the ancient reception. Research methods. Descriptive scientific method was used to characterize extralinguistic factors of the development of Ukrainian linguistics, comparative-historical in order to find out the origin of terminological lexemes and historical changes in their semantic structure, comparable was used to determine the relationship between the names of parts of the language at the time of antiquity and in the period of the formation of Ukrainian grammatical science, its tracing ways, etc. The parts of speech are relevant in the same relation to a grammatical and lexical description of the language. The theory of the parts of speech is the foundation for grammar because all grammar books are built as a description of the parts of speech by their morphological, mainly word-changing, characteristics. Ukrainian grammatical theory was considerably affected by the Greek-Latin canon of a grammatical description. According to ancient tradition, in the first Ukrainian grammar book there were eight parts of speech to distinguish and the names of the categories often arose by copying the corresponding Latin accidents. Conclusions.Constructional characteristics taken into grammatical description of the Ukrainian language modeled the tendencies of the further grammatical abstraction as to the norms of the Ukrainian language, including specific specialization of the grammatical characteristics directed to the disclosure of the ethnical essence of the native literary language.
This paper describes the development of the first syntactically-annotated corpus of Breton. The corpus is part of the Universal Dependencies project. In the paper we describe how the corpus was prepared, some Breton-specific constructions that required special treatment, and in addition we give results for parsing Breton using a number of off-the-shelf data-driven parsers.
The problem of interference is one of the most complex issues related to language interaction, so it is especially important to investigate its workings on the example of the language of Russian Germans in the Kirov region. The article realises the historical and linguocultural approaches to the study of the interrelationship between folk-colloquial speech and the traditional culture of Russian Germans, residing on the territory of the Kirov region. The authors present the results of an in-depth analysis of interference features in the Russian speech of German bilinguals under the influence of the German language and its dialects, namely, the phonetic, lexical and grammatical features that occur under the influence of interference with the German language. The Russian speech of German bilinguals is heterogeneous and varies from "virtually without an accent" to "unnatural" for Russian monolingual hearing. The interaction of the Russian and German languages in the speech of German bilinguals resulted in the increased invasion of the norms of one language system into the framework of another language. This leads to the so-called levelling of the interacting languages. In other words, we see the emergence of a third–intermediate system that does not coincide either with the German or Russian languages and performs in the bilingual consciousness an adaptive function to the environment language. This study contributes to German dialectology, enriching both the theory and typology of island dialects, which retain archaic features and the theory and practice of scientifically grounded language policy and language preservation.
The ability to track non-adjacent dependencies (the relationship between ai and bi in an aiXbi string) has been hypothesized to support detection of morpho-syntactic dependencies in natural languages (‘The princess reluctantly kiss the frog’). But tracking such dependencies in natural languages entails being able to generalize dependencies to novel contexts (‘The general angrily berat his troops’), and also tracking co-occurrence patterns between functional morphemes like and (a class of elements that often lack perceptual salience). We use the Headturn Preference Procedure to investigate (i) whether infants are capable of generalizing dependencies to novel contexts, and (ii) whether they can track dependencies between perceptually non-salient elements in an artificial grammar aXb. Results suggest that 18-month-olds extract abstract knowledge of a_b dependencies between non-salient a and b elements and use this knowledge to subsequently re-familiarize themselves with specific ai_bi co)
Since its inception the issue of absence has preoccupied both the practitioners of corpus linguistics and its detractors. To the latter it is self-evident, a truism, that a corpus can yield no information about phenomena it does not contain, a criticism which we hope to demonstrate is based on a failure to grasp the complexity of the notion of absences and an ignorance of the flexibility of corpus techniques. However the former, the exponents of CL, have also worried greatly about the significance of not finding something, say, a particular set of lexical items or a certain syntactic structure in their corpus. Is this (non) discovery telling me something about the discourse type(s) under study or about what is usually termed the ‘representativity’ of the corpus (i.e. how typical of the discourse type is the subset of it contained in the corpus)? And the CL literature is replete with warnings ‘not confuse corpus data with language itself’ (McEnery & Hardie 2012: 26), to which we would add that observations arising from corpus data can only be generalised with the utmost care. Following Kant, we must not confuse the tangible, the phenomenal (corpus) with the intangible noumenal (language). \nIn this chapter we will discuss, on the basis of a number of case studies, what can reasonably be inferred about discourses from corpus analysis, particularly with regards to absences. Along with Scott, we maintain ‘much can be inferred from what is absent’ (2004), and following Taylor (2012) we will argue that corpus tools provide an ‘armory’ for locating and verifying absence. In particular, the comparison and contrast among different corpora can firstly reveal absences, both those being searched for and others accidentally stumbled upon, and then allow the analyst to track the appearance and disappearance of linguistic elements or discoursal notions once they have come in some way to the analyst’s attention. \nFinally, since most things are absent from most places most of the time we need to decide the parameters of relevant or salient or meaningful absence/s, those which are worth either looking for if somehow suspected or instead, if stumbled upon, are worthy of further investigation. One indication could be unexpectedness or non-obviousness, that is, discovering absence when a presence is expected. This however raises the question of expected by whom and why, especially since researchers have their own unique past primings (Hoey 2005) which influence expectations in the present. And then, when an absence is discovered, how does one decide whether the absence is intentional or otherwise, especially given, as already stated, that absence is the norm? Far too often, particularly in the field of critical discourse analysis, it is taken for granted that a silence or absent message or voice must have been deliberately suppressed with little evidence of intentionality. Finally, once an absence is adjudged relevant and worthy of investigation, do we attempt to explain it? If so, what kinds of explanations are valid and interesting? Which are trivial and which non-trivial, that is, are themselves non-obvious and unexpected?
Themass media as a sphere of speech activity represent a certain style of material presentation, the purpose of which is to inform and captivate the addressees. The functioning of dialecticisms in the texts are usually studied in comparision with the stylistic norms of the literary language and considered as stylization aimed at interaction with the reader, approaching to his worldview. Intensive use of territorially specific language elements not only in modern fiction works but also in the language of the regional periodicals and radio broadcasting caused an increase of interest to the scientific studies of functional-stylistic peculiarities of the dialectal words in the language of modern mass media.
 The purpose of the research is to analyze the stylistic motivation of the use of dialectal elements in the language of the periodicals, to find out the genre specification of the penetration of regional spoken units into the language of periodicals and their stylistic function.
 The article considers the use of the dialectal elements in the language of regional periodical “Volyn” as a means of emotionally-colored attitude of the authors of publications towards certain objects. The importance is placed on the mechanism of comparison, harping different lexical units of the national language in order to achieve artistic and aesthetic effect in journalism. It is proved that the use of geographically differentiated language elements is determined by the expressive function of language in the media. It has the effect of immediacy, informality, communicative interaction with the reader in his own language. Colloquial lexical items in journalism are a manifestation of updating verbalness in the newspaper discourse. The orientation of the media to the daily/spoken language patterns allows to bring the content of the texts closer to the reader's perception.
 The analytical and journalistic materials of the newspaper «Volyn» use live linguistic units (nouns, adjectives, verbs, adverbs, rarely –numerals) in order to reproduce the regional linguistic flavour, the linguistic characteristics of the subjects of publications. Territorially differentiated units in journalism are a convincing manifestation of the implementation of verbal stylistic means in the newspaper discourse. Dialect and colloquial tokens help to identify the individual elements of the ethno-cultural picture of the world of the inhabitants of Volyn in the language of the regional printed editions.
Processing of nouns and action verbs can be differentially compromised following lesions to posterior and anterior/motor brain regions, respectively. However, little is known about how these deficits progress in the course of neurodegeneration. To address this issue, we assessed productive lexical skills in a patient with posterior cortical atrophy at two different stages of his pathology. On both occasions, he underwent a structural brain imaging protocol and completed semantic fluency tasks requiring retrieval of animals (nouns) and actions (verbs). Imaging results were compared with those of controls via voxel-based morphometry, whereas fluency performance was compared to age-matched norms through Crawford’s t-tests. In the first assessment, the patient exhibited atrophy of more posterior regions supporting multimodal semantics (medial temporal and lingual gyri), together with a selective deficit in noun fluency. Then, by the second assessment, the patient’s atrophy had progressed mainly towards fronto-motor regions (rolandic operculum, inferior and superior frontal gyri) and subcortical motor hubs (cerebellum, thalamus), and his fluency impairments had extended to action verbs. These results offer unprecedented evidence of the specificity of the pathways related to noun and action-verb impairments in the course of neurodegeneration, highlighting the latter’s critical dependence on damage to regions supporting motor functions, as opposed to multimodal semantic processes.
Theory-driven text analysis has made extensive use of psychological concept dictionaries, leading to a wide range of important results. These dictionaries have generally been applied through word count methods which have proven to be both simple and effective. In this paper, we introduce Distributed Dictionary Representations (DDR), a method that applies psychological dictionaries using semantic similarity rather than word counts. This allows for the measurement of the similarity between dictionaries and spans of text ranging from complete documents to individual words. We show how DDR enables dictionary authors to place greater emphasis on construct validity without sacrificing linguistic coverage. We further demonstrate the benefits of DDR on two real-world tasks and finally conduct an extensive study of the interaction between dictionary size and task performance. These studies allow us to examine how DDR and word count methods complement one another as tools for applying concept dictionaries and where each is best applied. Finally, we provide references to tools and resources to make this method both available and accessible to a broad psychological audience.
In this descriptive linguistic study, the lexico-grammatical complexity of placement and exit English for Academic Purposes (EAP) student writing samples was analyzed using corpus linguistic methods to explore language development as a result of student enrollment in the EAP program. Writing samples were typed, matched, and tagged. A concordance software was used to produce lexical realizations of grammatical features. A comparison was made of normed frequency counts for nine phrasal and clausal features as well as raw frequencies for type to token ratio (TTR), average word length, and word count. In addition, the contribution of variables such as advanced grammar and writing course grades, LOEP scores, and the number of semesters in the EAP program to the English Learner's (EL) lexico-grammatical complexity found in exit essays was also examined. Twelve paired parametric and non-parametric analyses of lexico-grammatical variables were performed. Dependent t test results showed that normed frequency counts for such features as pre-modifying nouns, attributive adjectives, adverbial conjunctions, coordinating conjunctions, TTR, average word length, and word count changed significantly, and students produced more of those features in their exit writing than in their placement essay. Non-parametric Wilcoxon test indicated that such a change was also observable with noun + that clauses. The frequencies of verb + that clauses and subordinating conjunction because, though non-significant, actually decreased. A split plot ANOVA allowed to see whether a change in above mentioned statistically significant lexico-grammatical features could be attributed to grammar instruction in EAP 1560. The results showed that there was no statistically significant difference between those who took EAP 1560 class and those who did not on pre-modifying nouns, coordinating conjunctions, TTR, average word length, and word count. On the other hand, those students who did not take EAP 1560 class had higher counts of attributive adjectives but lower of adverbial conjunctions, both statistically significant results, than those students who took the class. Lastly, five multiple linear regression analyses were conducted to predict frequencies of exit pre-modifying nouns, attributive adjectives, noun + that clauses, adverbial conjunctions, and TTR from EAP 1560 and EAP 1640 grades, LOEP scores, and the number of semesters students spent in the EAP program at SSC. The only significant regression analysis was with TTR, and 28% of its variance could be explained by the independent variables. LOEP Language Usage score was the only significant individual contributor to the model. Even though exit adverbial conjunctions were not predictable from the chosen IVs, LOEP Sentence Meaning score proved the only significant contributor to that model. The results indicate that compressed phrasal features are indicative of higher complexity and EL proficiency, while clausal features are acquired earlier and signal elaboration, as previously described in the literature.
<p>Gender identity, one of the most important social categories in people’s lives, is socially constructed and language is claimed to have a significant role in constructing the gender identity. This paper studies the construction of Sundanese women through five Sundanese nouns referring to women found in the corpus of <em>Manglè </em>magazine, published between 1958–2013. The research employs a mixed-method design in which quantitative analysis is combined with qualitative analysis to investigate how the nouns referring to women are used to construct Sundanese women from the periods of Guided Democracy (1958–1965) to Reformation (2004–2013). The quantitative analysis is used to examine the frequency of word occurrence diachronically. The frequency of word accurrence is subsequently interpreted qualitatively by considering social and cultural contexts, such as the norms of speech levels in Sundanese, Sundanese belief about marriage, and gender issues. The result of analysis shows that women are constructed in various identities by every noun referring to them. The lexical choices used to contruct women are greatly influenced by the social and cultural contexts. </p>
The translation of certain toponyms that had not yet been assimilated to the Romanian language at the beginning of the 19th century represented a real challenge for translators at that time. A first aspect to be considered here is the linguistic status as proper names and the possible translation options that could not be correlated to any tradition. A second aspect is the precarious stage of Romanian geographical terminology, reflected by the terminological variation for the same concept and the lack of semantic affinity, either real or related to the actual terminology. This article addresses mainly the first aspect mentioned above. The issues addressed are as follows: a concise presentation of the concept of proper names translation, the distinction between untranslatable and translatable or partially translatable proper names, the factors motivating the option of translating or not the translatable terms from a toponymic collocation. Our corpus reflects the incipient stage of the translation of translatable or partially translatable toponyms in Romanian, a stage in which the translator is free to decide upon translatability. Compared to the actual norm, the different choices from one translator to another or even those opted for by the same translator—especially the option of not translating toponyms that are nowadays translated in most languages—reveal the lack of importance of linguistic meaning (that is the lexical meaning of the etymon) of the proper name as far as its functioning was concerned, as well as the role of this non-functionality in identifying the linguistic status of a proper name.
Linguistic register reflects changes in speech that depend on the situation, especially the status of listeners and listener-speaker relationships. Following the sociolinguistic rules of register is essential in establishing and maintaining social interactions. Recent research suggests that children over 3 years of age can understand appropriate register-listener relationships as well as the fact that people change register depending on their listeners. However, given previous findings that infants under 2 years of age have already formed both social and speech categories, it may be possible that even younger children can also understand appropriate register-listener relationships. The present study used Infant-Directed Speech (IDS) and formal Adult-Directed Speech (ADS) to examine whether 20-month-old toddlers can understand register-listener relationships. In Experiment 1, we used a violation-of-expectation method to examine whether 20-month-olds understand the individual associatio)