Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
The BHSA (Biblia Hebraica Stuttgartensia Amstelodamensis) is the BHS text plus the linguistic annotations of the Eep Talstra Centre for Bible and Computer.
 The BHSA is available as a data set in Text-Fabric format. Text-Fabric is a minimalistic model to represent text: it provides addresses for all textual objects, so that it is easy to add arbitrary information at all textual levels, precisely and firmly anchored. A Text-Fabric resource resembles an IKEA ware house. The parts are nicely separated and stacked, so that they can be retrieved easily, to be combined into meaningful output later on. A consequence is that different teams with divergent purposes still can add to the same body of work, with a minimum of interference or duplication of work. Text-Fabric has helped with various types of data construction work, of which the most visible is the website SHEBANQ. We focus on two recent data combination jobs, (A) treebanks from the BHSA data and (B) a detailed comparison of the morphology in the BHSA and in the Open Scriptures effort. As the OSM is not yet finished, the comparison is repeatable.
Abstract Introduction Drug cue reactivity (DCR) is widely used in experimental settings for both assessment and intervention. There is no validated database of pictorial cues available for methamphetamine and opioids. Methods 360 images in three-groups (methamphetamine, opioid and neutral (control)) matched for their content (objects, hands, faces and actions) were selected in an initial development phase. 28 participants with a history of both methamphetamine and opioid use (37.1 ± 8.11 years old, 12 female) with over six months of abstinence were asked to rate images for craving, valence, arousal, typicality and relatedness. Results All drug images were differentiated from neutral images. Drug related images received higher arousal and lower valence ratings compared to neutral images (craving (0-100) for neutral (11.5±21.9), opioid (87.7±18.5), and methamphetamine (88±18), arousal (1-9) for neutral (2.4±1.9), opioid (4.6±2.7), and methamphetamine (4.6±2.6), and valence (1-9) for neutral (4.8±1.3), opioid (4.4±1.9), and methamphetamine (4.4±1.8)). There is no difference between methamphetamine and opioid images in craving, arousal and valence. There is a significant positive relationship between the amount of time that participants spent on drug-related images and the craving they reported for the image. Every 10 points of craving were associated with an increased response time of 383millisecond. Three image sets were automatically selected for equivalent fMRI tasks (methamphetamine and opioids) from the database (tasks are available at github). Conclusion LIBR MOCD provides a resource of validated images/tasks for future DCR studies. Additionally, researchers can select several sets of unique but equivalent images based-on their psychological/physical characteristics for multiple assessments/interventions.
Abstract Semantically ambiguous and emotional words occur frequently in language, and the different meanings of ambiguous words can sometimes have different emotional loads. For example, the Spanish word heroína (heroin/heroine) can refer to a drug or to a woman who performs a heroic act. Because both ambiguity and emotionality affect word processing, there is a need for normative databases that include data on the emotionality of the different meanings of such words. Thus far, no bases of this type are available in Spanish. With this in mind, the current study will present meaning-dependent affective (valence) ratings for 252 Spanish ambiguous words. The analyses performed show that (a) among ambiguous words, those words with meanings that have distinct affective valence are quite frequent, (b) ambiguous words rated as neutral in isolation can have meanings of opposite valence (i.e., negative-positive or positive-negative), and (c) the valence estimated for ambiguous words in isolation is better explained by the weighted average of the valence of their meanings by dominance. A database of this kind can be useful both for basic research (e.g., relationship between emotion and language and ambiguity processing) and for applied research (e.g., cognitive and emotional biases in emotional disorders and second language learning).
WordNet is a large lexical database originally conceived as a model of human semantic memory. While its design loosely follows ontological principles, it was developed by linguists and psychologists without the benefit of input from philosophers. WordNet's unexpected popularity as a tool for Natural Language Processing and Knowledge Engineering revealed the irregularities of the lexicon and the need to clearly distinguish the lexical from the conceptual level consistent with ontological research, as articulated in the work of Nicola Guarino and his colleagues. This paper reviews WordNet's development from a lexicon to a resource that incorporates some ontological principles.
Named entity recognition (NER) is widely used in natural language processing applications and downstream tasks. However, most NER tools target flat annotation from popular datasets, eschewing the semantic information available in nested entity mentions. We describe NNE-a fine-grained, nested named entity dataset over the full Wall Street Journal portion of the Penn Treebank (PTB). Our annotation comprises 279,795 mentions of 114 entity types with up to 6 layers of nesting. We hope the public release of this large dataset for English newswire will encourage development of new techniques for nested NER.
Parsers are available for only a handful of the world's languages, since they require lots of training data. How far can we get with just a small amount of training data? We systematically compare a set of simple strategies for improving low-resource parsers: data augmentation, which has not been tested before; cross-lingual training; and transliteration. Experimenting on three typologically diverse low-resource languages---North Sámi, Galician, and Kazah---We find that (1) when only the low-resource treebank is available, data augmentation is very helpful; (2) when a related high-resource treebank is available, cross-lingual training is helpful and complements data augmentation; and (3) when the high-resource treebank uses a different writing system, transliteration into a shared orthographic spaces is also very helpful.
A lot of lexical research conducted by experts use existing words in the English Thesaurus. However, the English Thesaurus only gives synonyms of the words searched for and does not provide similarities between words that can be called WordNet. So in this study, a WordNet can be made that can help research that uses databases for English. The similarity between words makes many researchers today are still looking for word relations by manual method, or still use the English Thesaurus. The making of WordNet is expected to be very useful for researchers who want a lexical database for their research, which is currently still relatively small. Therefore it is better to make an English WordNet which will later accommodate words that have the same meaning or Synonym Sets and this WordNet focuses on grouping those words. So that researchers can do lexical research more broadly and unlimitedly with the existence of words that are still unclear in their similarities. Calculation with clustering gets an F1 Score Result at 10.68%, Recall at 8.53% and Precision at 14.28%.
A lexical database of eastern Indonesia
This dataset is a PostgreSQL dump of a database generated using autocode_big from the data in the LexiRumah lexical database of eastern Indonesia and Timor-Leste. The dataset contains the forms from the parent dataset, together with automatically generated sound correspondence scorers, pairwise similarity scores, cognate classes, and alignments, all created using Lingpy's LexStat algorithm.
It is essential in emotion science to have standardized stimuli to aid in replication as well as the elicitation of basic defensive and appetitive states. There are numerous standardized stimulus sets for eliciting emotion perception including scenes, faces, videos, and audio scripts (Coan & Allen, 2007). These stimuli are typically rated subjectively on pleasantness and activation/arousal scales. There are, however, few studies that provide normative pleasantness and arousal ratings from scripts useful in predicting emotional states for future studies on emotional imagery. The goals of the current study were to add to the literature 1) standardized pleasant, neutral, and unpleasant text-driven scripts rated on pleasantness and arousal, and 2) a more substantial set of stimuli that could be useful in electrophysiological studies requiring large amounts of experimental trials. Thirty-five undergraduate students were presented 133 text-driven scripts, which were classified as pleasant (i.e., excitement, erotica, relaxation), neutral, or unpleasant (i.e., contamination, embarrassment, threatening). While reading the script, participants actively imagined themselves in the presented scenario and then provided pleasantness and arousal ratings. As expected, from pleasant to neutral to unpleasant conditions there was a positive linear trend with pleasantness ratings (p<.001). Further, from pleasant to neutral conditions arousal ratings declined, whereas from neutral to unpleasant conditions arousal ratings rose. Thus, we found a significant quadratic trend for the arousal ratings (p<.001). Taken together, these preliminary data indicate that these novel text-driven scripts are eliciting the predicted emotional responses. Keywords: imagery, emotion, text-driven scripts, arousal, hedonic valence, norming study
Abstract This paper presents the Slovene Training Corpus ssj500k 2.2, which has been annotated on the levels of tokenization, sentence segmentation, part-of-speech tagging, lemmatization, syntactic dependencies, named entities, verbal multi-word expressions, and semantic role labeling. It describes the individual layers of annotation and shows the scope of using the training corpus in the production of various lexicons, such as the lexicon of multi-word units and the valency lexicon of modern Slovene. It concludes by presenting our future work, i.e. the annotation of multi-word expressions based on the Slovene Lexical Database.
We recently introduced the EmojiGrid as an intuitive graphical self-report tool to measure food&shy;evoked valence and arousal. The EmojiGrid is a Cartesian grid, labeled with facial icons (emoji) expressing different degrees of valence and arousal. The lack of verbal labels makes it a valuable, language-independent tool for cross-cultural research. Users can efficiently report their subjective ratings of both valence and arousal with a single click on the location of the grid that best represents their affective state after perceiving a stimulus. The EmojiGrid has previously been validated for the assessment of emotions evoked by food images. In this study we validated the EmojiGrid for the affective appraisal of odors. Observers (N=56, 24 males, mean age=24.3±4.6) smelled 40 randomly presented odors (27 food and 13 non&shy;food smells), ranging from very unpleasant and arousing (e.g., feces, fish), via pleasant and calming (e.g., clove, cinnamon), to very pleasant and stimulating (e.g., peach, caramel). The odor samples consisted of felt pens, with tips that were impregnated with 4 ml of fluid odorant substance. Each pen was presented once, for about 5 seconds at 2 cm below both nostrils. The participants sniffed following a verbal command. Immediately after sniffing the pen was removed, and participants were given at least 30 s to smell fresh air. The participants reported their affective appraisal of each odor using the EmojiGrid. The resulting mean valence and arousal ratings closely agree with those from previous studies in the literature that were obtained with alternative rating tools. In addition, we find that the EmojiGrid yields the typical universal U-shaped relation between mean valence and arousal that is commonly observed for a wide range of affective sensory (visual, auditory, tactile, gustatory) stimuli. We conclude that the EmojiGrid is also a valid affective self-report tool for the assessment of odor evoked emotions.
The CHLG at a glance A parsed (= syntactically annotated) corpus of historical Low German in Penn-Treebank style. Allows for efficient, reproducible searches for a large number of morphosyntactic structures.
The electronic lexical databases WordNets, have become essential for many computer applications, especially in linguistic research. Free French WordNet is an XML lexical database for French language based on Princeton WordNet for the English language and other multilingual resources. So far, research on Free French WordNet has focused on the construction and relevance of lexico-semantic information. However, no effort is made to facilitate the exploitation of this database under the Java language. In this context, this paper proposes our approach for the development of a new Java API based on Java Architecture for XML Binding. This Java API will make it easier for developers to exploit and use Free French WordNet to create applications for natural language processing. In order to assess the usefulness of our API, The API performance has been evaluated in the context of a Browser that we developed to extract semantic and lexical relations connecting synsets contained in this database, such as: the tree of hypernymy, the tree of hyponymy, synonyms, etc. The results showed that our API perfectly meets the needs of programmatically exploitation, exploration and consultation of this database in a Java application.
Evidence from both behavioral and neuropsychological studies suggest that different types of organizational principles govern semantic representations of abstract and concrete words. The reviewed neuroimaging studies provide new evidence about the role of brain areas of the semantic network involved in the encoding of some types of information during processing of abstract and concrete concepts, better characterizing the neural underpinnings and the organizational principles of semantic representation of these types of word.
We present an ongoing project of enriching an annotation of a parallel dependency treebank, namely the Prague Czech-English Dependency Treebank, with verb-centered semantic annotation using a bilingual synonym verb class lexicon, CzEngClass. This lexicon, in turn, links the predicate occurrences in the corpus to various external lexicons, such as FrameNet, VerbNet, PropBank frame files, OntoNotes, and WordNet. We briefly describe the content of the CzEngClass synonym class lexicon and then we focus on its use for an enrichment of corpus annotation, which proceeds in two steps -automatic preprocessing and manual correction. This paper describes a first milestone of a long-term project; so far, approx. 100 CzEngClass classes, containing about 1800 different verbs each for both Czech and English, are available for such annotation. The corpus coverage at the moment is about 50%, allowing us to extract some basic statistics and discover a set of issues that appeared during the annotation process. The ultimate goal is to have a high-coverage, multilingual verbal synonym lexicon and corpora with all events annotated by such lexicon, to serve both theoretical studies in lexical semantic, translatology, corpus annotation studies etc. as well as a usable resource for training automatic semantic text processing systems for event/participant detection and linking and for general information extraction.
This article describes the first release version of a new lexicostatistical database of Northern Eurasia, which includes Europe as the most well-researched linguistic area. Unlike in other areas of the world, where databases are restricted to covering a small number of concepts as far as possible based on often sparse documentation, good lexical resources providing wide coverage of the lexicon are available even for many smaller languages in our target area. This makes it possible to attain near-completeness for a substantial number of concepts. The resulting database provides a basis for rich benchmarks that can be used to test automated methods which aim to derive new knowledge about language history in underresearched areas.
Lexical database of the lesser Sunda islands
In Receive Our Memories, José Orozco, associate professor of history at Whittier College, offers a unique set of primary documents accompanied by a family history and local history, insightfully contextualized within the first half of twentieth-century Mexican history. It examines the life of a campesino (and great-grandfather of the author) Luz Moreno (1877–1953) from the town of San Miguel el Alto, Jalisco, through the 170 letters written to his eldest daughter Pancha (1901–2002), who had migrated to Stockton, California. Luz's letters reveal, from the perspective of a poor man in his own words, aspects of his philosophical inner self—one quite different from the taciturn and stoic outward persona. Indeed, Moreno's letters—sent to his eldest daughter from 1950, after she married rather late in life and emigrated, until his death a few years later—are as complex as they were frequent. Reserved and a man of few words in person, the ailing patriarch reveals in his letters to his favorite daughter his intense love for her and an “interior monologue,” while his prose “displays a conscious and persistent literary intent” (p. 10). In the historiography of immigration, Orozco's book is unique in that most letters examined to tell stories of migration are from the migrant's perspective, not from that of a family member left behind to struggle with the physical and emotional separation from a loved one. Orozco explains that Moreno “saw emigration as a threat to his religion, to the coherence of his family, and to the viability of a way of life that he once believed was immutable” (p. 40). In Moreno's own words from 1951, “Many go illegally to that Promised Land in search of the Dollar and they give more importance to it than to tending the corn in their own country. Compared to the Dollar, everything seems to be stacked against corn... What shall we do? Shall we only eat Dollars?” (p. 40). His letters and Orozco's analysis reveal the negative effects of modernity and modernization.Orozco has organized the content thematically and translated 80 of the original letters in these chapters, providing historical and historiographical context for each one. The themes include affective bonds, religion, poverty, letter writing and newspaper reading, and old age and dying. The first chapter tells the history of San Miguel el Alto and weaves the story of the entire Moreno clan into local and national political developments, with a focus on religious resistance. Especially rich in details about Pancha's life in Jalisco, this chapter decenters the somewhat persistent triumphalist narrative of the 1910 revolution through a careful depiction of these fervent Catholics and their participation in both Cristero rebellions (from 1926 and 1929 and from the mid-1930s to early 1940s, respectively) and the Sinarquista movement (from 1937 to 1950). Indeed, Orozco states that some Alteños considered the Cristero rebellions to be the true revolution rather than the 1910 movement against Porfirio Díaz. The story of the Moreno family highlights Catholic resistance to state power and the complexities of traditional, yet at times quite fluid, gender roles and norms. For example, Pancha served as secretary of feminine action for the Sinarquista party in the late 1930s and 1940s, for which she performed work considered acceptable for her gender such as sewing the party's flags, feeding members, and coordinating children's activities. But she also actively recruited new members, traveling with her uncle to neighboring locales. In 1916, the family had refused to let Pancha marry her sweetheart Juan; she defied them when he returned in 1950 to marry her, and the male family members eventually accepted the situation.Receive Our Memories demonstrates the elderly Luz Moreno's self-awareness of his life in poverty and how avid newspaper reading and his Catholic faith made him conscious of the global interconnectedness of “el miserable pueblo.” The letters also speak with sophistication to international political events and ideologies. Of particular concern to Moreno when writing to his daughter were the godlessness, greed, and capitalistic excesses of the United States and the threats that these posed to her Catholic soul. Perhaps less predictably, Moreno weighed in on heavy political and economic issues like the nuclear arms race, the Cold War, the Korean War, and the bracero program with tremendous detail. The book sheds light on the importance of the family economy (both in Jalisco and transnationally through remittances) and the indissoluble bonds of family in the face of migration. The text is accompanied by sketches by artist José Lozano, meant to harken back to “lexical-visual collaborations undertaken in the 1930s and 1940s” between Mexican artists and American radical intellectuals (p. x). Moreno's ability to connect to global political, economic, and social developments and his self-awareness about poverty and his relation to others in that struggle worldwide in his writings make them a unique source. It will be of interest to historians and anthropologists of Mexico and scholars studying the elderly, migration, and poverty in any geographic region. The letters combined with Orozco's careful contextualization of the economic, political, and religious milieus make the book ideal for undergraduate students and scholars alike.
This article introduces a package developed for R (R Core Team, 2017) for performing an integrated analysis of multiple data blocks (i.e., linked data) coming from different sources. The methods in this package combine simultaneous component analysis (SCA) with structured selection of variables. The key feature of this package is that it allows to (1) identify joint variation that is shared across all the data sources and specific variation that is associated with one or a few of the data sources and (2) flexibly estimate component matrices with predefined structures. Linked data occur in many disciplines (e.g., biomedical research, bioinformatics, chemometrics, finance, genomics, psychology, and sociology) and especially in multidisciplinary research. Hence, we expect our package to be useful in various fields.
The volume collects articles which discuss complexity, conventionality and creativity in the English language from perspectives as diverse as specialised discourse, language teaching and learning, language varieties, lexical creativity, stylistics, knowledge dissemination through the media and audio-visual translation. It offers a multifaceted picture of the ways in which opposing forces exerted by conventionality and creativity contribute to shaping all levels of the linguistic system. The interpretive paradigm is offered by the theory of complex systems, a rich research framework attempting to describe and explain the dynamics which emerge in the many forms of situational adaptation of natural systems. Norms and conventions are, in fact, constantly exploited and manipulated through the creative behaviour of language users. This may lead to unpredictable synchronic effects and variation and, ultimately, to diachronic innovation.
Communication is an essential part of human life. For the person with hearing and speaking disability, it is inconvenient to communicate with other people. In this paper, an end-to-end system to convert English voice to Indian Sign Language (ISL) gloss (written form of sign language) is proposed which will help deaf to communicate with others and vice versa. This system accepts English voice as an input and, converts it into the text using the speech recognition. From the recognized English text, ISL gloss is generated using the lexical database called WordNet. The focus of our work is to build a robust sign language machine translation system to convert the English text to ISL gloss using the linguistic database WordNet.
espanolEn el presente trabajo se delinea la aportacion de la lengua vasca a la formacion de la norma castellana durante la Edad Media y el Siglo de Oro. Para llegar a tal fin, se perfila primero el marco historico de redes sociales y linguisticas operantes en el contacto del castellano con el euskera desde los origenes de la convivencia vasco-latino-romanica hasta el siglo XVII y se tiene en cuenta, despues, la doble direccion en el contacto vasco-castellano-romanico, sobrevenido historicamente tanto en la direccion del euskera hacia el castellano como del castellano al vascuence, sin olvidar que ha afectado tambien a territorio hoy frances. EnglishThe present article studies the Basque Language contribution to the Castilian linguistic Norm in the Middle Ages and the Golden Age. For this purpose will be first designed the historical frames operating since ancient times on the linguistic contact between Basque and Romance Languages in order to explain how the direction of the Basque-Romance contact has occurred in both directions respectively, taken into account that France domain has been also concerned.
This article presents a new method for reducing socially desirable responding in Internet self-reports of desirable and undesirable behavior. The method is based on moving the request for honest responding, often included in the introduction to surveys, to the questioning phase of the survey. Over a quarter of Internet survey participants do not read survey instructions, and therefore, instead of asking respondents to answer honestly, they were asked whether they responded honestly. Posing the honesty message in the form of questions on honest responding draws attention to the message, increases the processing of it, and puts subsequent questions in context with the questions on honest responding. In three studies (nStudy I = 475, nStudy II = 1,015, nStudy III = 899), we tested whether presenting the questions on honest responding before questions on desirable and undesirable behavior could increase the honesty of responses, under the assumption that less attribution of desirable behavior and/or admitting to more undesirable behavior could be taken to indicate more honest responses. In all studies the participants who were presented with the questions on honest responding before questions on the target behavior produced, on average, significantly less socially desirable responses, though the effect sizes were small in all cases (Cohen’s d ranging between 0.02 and 0.28 for single items, and from 0.17 to 0.34 for sum scores). The overall findings and the possible mechanisms behind the influence of the questions concerning honest responding on subsequent questions are discussed, and suggestions are made for future research.
The purpose of this work is to determine the role of the koine of the Bakhchisaray’s capital and its environs in the formation of the supra-dialect koine and the literary (standard) Crimean Tatar language in the seventeenth and eighteenth centuries. The period that we are considering in this article is quite indicative precisely in terms of the development of the Crimean Tatar literary language and its oral form – the supra-dialect koine. This time was the last stage of the fully functioning Crimean language before the Crimean Khanate lost its independence totally (1783). Phonological, lexical and grammatical norms which determined the vector of further development of the literary language of the Crimean Tatars, based mainly on the Bakhchisaray urban koine, had already crystallized in that epoch’s language. The material of this study consists of legal documents. They provide the best way to trace the processes of the formation of norms in the general Crimean supra-dialect koine, which based on the capital’s koine. Of particular value are the records of Sharia courts of the Crimean kadiys, on the one hand, and the khan’s yarliks along with letters, on the other. Both types of documents demonstrate two literary styles that were forming by different Turkic linguo-cultural traditions: that of the Golden Horde and the Crimean proper, the latter being a regional one which was influenced by the Ottoman language. The fact of lingual archaism and the mixing the phonetic, lexical and grammatical traditions of different Turkic languages in the texts of the manuscripts of official and business writing testify to the mixed character of norms in the Bakhchisaray, pre-dialect koine norms, and the norms of the literary language on the basis of the interaction of homogeneous Turkic idioms (Cumanian and Seljukian) with a small share of heterogeneous, mostly lexical, borrowings.
This study aims to explore cross-language intensification in affirmative sentences by examining the translation of standard amplifiers, words that scale upward towards an assumed norm to emphasize a quality of any entities, from Thai into English. The data comprises 602 parallel concordance lines with 17 intensifying patterns, which were drawn from a corpus of eight works of fiction in Thai and their English translations translated by qualified translators. The analysis of the data found that in the English translation, English amplifiers (e.g. very, really) were found with the highest frequency, followed by intensified lexemes and comparative and superlatives respectively. The findings suggest that the tendency to transfer standard amplifiers was through lexical (TL amplifiers, intensified lexemes, emphasizing adjectives) and syntactic means (comparatives and superlatives, exclamatory constructions, and metaphors), and that the selection was made in accordance with the context. Compared with the Thai standard amplifier maak2 ‘much-many’, the linguistic devices used in the English translations tend to reveal a stronger force of intensity. The findings can provide pedagogical implications in translations. They, for instance, can raise students’ awareness of the various linguistic forms used in transferring intensity expressed in the source text and also provide norms in translating amplifiers from Thai to English, which might be useful for students in translation programs. In addition, students may realize that if a literary work loses the expressivity of feelings or emotion, it becomes uninteresting and lacks vivacity, thus losing appeal to the TL reader.
English Abstract: The special status of the German language, its multinational character and national variability are always in the focus of attention of different sciences and give practical material for research. It is the example of the German language that clearly demonstrates how strong the nonlinguistic factors can influence the history of its development. Thus, a pronounced dialectic and regional differentiation stems from the Reformation and feudal fragmentation of the Middle Ages. As a result of the political influence of one or another dynasty, the German-speaking region is far from heterogeneity, despite the geographical unity. The consequence of these factors has become the formation of different versions of the German language: Hochdeutsch (in the German’s territory), Osterreichisches Deutsch (in the Austria’s territory), Schweizer Deutsch (in the Switzerland’s territory). There are certainly distinctive features of the German language in Luxembourg and Liechtenstein. Being the dynamic system the language is responsive to all transformations taking place in the society and changes together with it. Unfortunately, not always these changes have positive and progressive character. In the XXI century the situation worsens by practically unlimited influence of the media, first of all, due to the possibilities of the global network. If in the second half of the XX century the statement that the media is the fourth power was rather declarative, then nowadays it has become evident. Online versions of different editions provide mass audience with almost real time information, sometimes in telegraphic style. In this case, unfortunately, it is not possible to preserve the purity of language in the broadest sense. The language becomes simpler from lexical, grammatical and syntactical points of view, saturated with foreign words (der Think-Tank, point of no return, das Statement) and uncharacteristic constructions (in 2019 instead of 2019), even the norms of pronunciation are violated (China). In the pursuit of the audience’s attention, in the desire not only to inform the public about a particular event but to portray it as a sensation, the media often resorts to different creative means, particularly on the lexical level (lindnern, sodern, der Kurz-Schluss, das Leyen-Theater, das Erdogate). The analysis of such items demonstrates the clear link between usually high-profile events taking place in the society and a rise of language phenomena that differ from normative rules. It is worth noting that therefore through language and the imposition of foreign norms manipulative influence impacts the mass audience, which can be countered by the well-considered and strict language policy at the national level. Russian Abstract: Особый статус немецкого языка, его полинациональность и национальная вариативность всегда находятся в центре внимания различных наук и дают практический материал для исследований. И именно на примере немецкого языка наглядно видно, насколько сильным может быть влияние экстралингвистических факторов на историю его развития. Так, достаточно ярко выраженная диалектная и региональная дифференциация берет свое начало в Реформации и феодальной раздробленности средних веков. В результате политического влияния той или иной династии немецкоязычный регион не отличается гетерогенностью, несмотря на географическое единство. Следствием подобных факторов стало формирование различных вариантов немецкого языка: Hochdeutsch (на территории Германии), Osterreichisches Deutsch (на территории Австрии), Schweizer Deutsch (на территории Швейцарии). Безусловно, есть также отличительные особенности немецкого языка в Люксембурге или Лихтенштейне. Будучи подвижной системой, язык чутко реагирует на все происходящие в обществе преобразования и изменяется вместе с ним. К сожалению, не всегда эти изменения носят положительный и прогрессивный характер. В XXI веке ситуация усугубляется практически безграничным влиянием СМИ, прежде всего, благодаря возможностям всемирной сети. Если во второй половине прошлого века утверждение, что СМИ являются четвертой властью, было скорее декларативным, то в настоящее время это стало очевидным. Online-версии различных изданий предоставляют массовой аудитории информацию практически в режиме реального времени, иногда в телеграфном стиле. В этом случае, к сожалению, не удается сохранять чистоту языка в самом широком смысле слова. Язык упрощается, как с лексической или грамматической точек зрения, так и синтаксически, насыщается иноязычными словами (der Think-Tank, point of no return, das Statement) и нетипичными конструкциями (in 2019 вместо 2019), даже нарушаются нормы произношения (China). В погоне за вниманием аудитории, желанием не просто проинформировать о том или ином событии, а представить его сенсационно, в СМИ зачастую прибегают к различным креативным способам, прежде всего, на лексическом уровне (lindnern, sodern, der Kurz-Schluss, das Leyen-Theater, das Erdogate). При анализе подобных наименований прослеживается четкая взаимосвязь между происходящими в обществе событиями, как правило, резонансными, и всплеском языковых явлений, отличных от нормативных правил. Стоит отметить, что таким образом, через язык и навязывание «чужих» норм оказывается манипулятивное воздействие на массовую аудиторию, противостоять которому возможно посредством хорошо продуманной и жесткой языковой политики на государственном уровне.
Most research groups studying human navigational behavior with virtual environment (VE) technology develop their own tasks and protocols. This makes it difficult to compare results between groups and to create normative data sets for any specific navigational task. Such norms, however, are prerequisites for the use of navigation assessments as diagnostic tools—for example, to support the early and differential diagnosis of atypical aging. Here we start addressing these problems by presenting and evaluating a new navigation test suite that we make freely available to other researchers (https://osf.io/mx52y/). Specifically, we designed three navigational tasks, which are adaptations of earlier published tasks used to study the effects of typical and atypical aging on navigation: a route-repetition task that can be solved using egocentric navigation strategies, and route-retracing and directional-approach tasks that both require allocentric spatial processing. Despite introducing a number of changes to the original tasks to make them look more realistic and ecologically valid, and therefore easy to explain to people unfamiliar with a VE or who have cognitive impairments, we replicated the findings from the original studies. Specifically, we found general age-related declines in navigation performance and additional specific difficulties in tasks that required allocentric processes. These findings demonstrate that our new tasks have task demands similar to those of the original tasks, and are thus suited to be used more widely.
Differences in nomenclature, regarding the legal reasoning of judex facti and judex juris decisions, occur in determining the act of taking part in a criminal act. This is due to different reasoning methods. The legal consideration approach to judex facti decisions, in verifying facts as norms, is performed lexically. The way the judge's logic works is by using deductive logic and verifying the facts of the defendant's actions to normalise elements that are merely restrictive. The judex juris decisions of judges and the judex facti legal judgments understand the act of participation in corruption case by using an inductive reasoning method. Judex juris decisions examine judex facti legal considerations by determining the major premise more extensively. Judges search for the legal principles underlying norms to verify the facts of the defendant's condition. The results of the verification and the conclusion of judex juris state that the defendant's actions are proven but there are no faults. Thus, judex facti decisions are cancelled and it is decided that the defendant is free from all legal charges.
36 Feminist Studies 45, no. 1. © 2019 by Feminist Studies, Inc. Sonny Nordmarken Queering Gendering: Trans Epistemologies and the Disruption and Production of Gender Accomplishment Practices Those who are deemed “unreal” nevertheless lay hold of the real, a laying hold that happens in concert, and a vital instability is produced by that performative surprise. —Judith Butler, Gender Trouble Beginning in the 1960s, scholars began to theorize gender as a contextually specific process rather than a universal category reflecting an essential pre-discursive sex. Two interrelated traditions developed: a discursive approach, which theorized gender as performative, and an interactionist approach, which investigated the interactional achievement of gender. For Judith Butler, “what we take to be an internal essence of gender is manufactured through a sustained set of acts, posited through the gendered stylization of the body.”1 Gender is therefore performative: it is a series of effects produced through the repetition and citation of stylized acts, which are named via and thus produced through discourse; discourse also produces the defining limits of subjects.2 Candace West and Don Zimmerman theorized gender as a “routine, methodical, 1. Judith Butler, Gender Trouble, 10th anniversary ed. (New York: Routledge, 1999), xv. 2. Butler, Gender Trouble; Judith Butler, Bodies that Matter: On the Discursive Limits of Sex (New York: Routledge, 1993); Judith Butler, Undoing Gender (New York: Routledge, 2004). Sonny Nordmarken 37 and recurring accomplishment” produced in social interaction.3 They observed how, in the relational process of “doing gender,” social actors display gender, presenting an appearance to others, who attribute gender by interpreting this appearance. In this article, I investigate how actors interactionally challenge and construct discursive structures in order to contribute to scholarship that analyzes the role language plays in such interactions.4 Following Sandy Stone, who suggests that transsexuals are not a class, nor a third gender, but a genre, “a set of embodied texts,” who, through their interpretation, might potentially disrupt dichotomous sexuality and gender categories, I examine the spaces in which the discursive and the interactional merge to investigate how gender minorities, as simultaneous subjects, texts, social actors, and cultural workers, queer hegemonic gender practices.5 I argue that members of trans linguistic communities and gender nonconforming individuals queer the normative gender process in two ways: by productively linguistically communicating third-person gender pronouns and by disruptively inhibiting gender’s hegemonic attribution. Lal Zimman has made parallel observations, explaining linguistic gender self-determination practices as a new cultural phenomenon, focusing on terminology for types of gendered persons, grammatical gender forms (i.e., pronouns), and lexical items that relate to embodied sex.6 My analysis builds on Zimman’s observations using sociological, performance 3. Candace West and Don H. Zimmerman, “Doing Gender,” Gender & Society 1, no. 2 (1987): 126. 4. For example, see Mimi Schippers, “Recovering the Feminine Other: Masculinity, Femininity, and Gender Hegemony,” Theory and Society 36, no. 1 (2007): 85–102. 5. Sandy Stone, “The Empire Strikes Back: A Posttranssexual Manifesto,” in Body Guards: The Cultural Politics of Gender Ambiguity, ed. Julia Epstein and Kristina Straub (New York: Routledge, 1991), 296. The term “gender minorities ” encompasses individuals identifying with gender identity terms other than those they were assigned and individuals with gender nonconforming appearance. 6. See Lal Zimman, “Transgender Language Reform: Some Challenges and Strategies for Promoting Trans-affirming, Gender-Inclusive Language,” Journal of Language and Discrimination 1, no. 1 (2017): 84–105; Lal Zimman, “Trans People’s Linguistic Self-Determination and the Dialogic Nature of Identity,” in Representing Trans: Linguistic, Legal, and Everyday Perspectives, ed. Evan Hazenberg and Miriam Meyerhoff (Wellington, New Zealand: Victoria University Press, 2017). 38 Sonny Nordmarken studies and performativity frameworks to theorize these practices. By approximating new gender pronoun-attribution norms that bring a trans queer paradigm to life in interactions and by disregarding their perceptions of each other’s bodies, actors accomplish gender pronouns linguistically, override hegemonic gender attribution norms, and reorganize gender accountability. These practices institutionalize a new interpretive frame and accountability structure through which social actors create and recognize a variety of gender expressions, identities, and pronouns, reworking performativity to produce gender minorities as subjects. I argue that this gendering-queering is a form of disidentificatory gender accomplishment. In addition, I...
The article is devoted to the study of the causes of communicative barriers in the process of dialogue at manufacturing. Special attention is given to parcelled statements as they function in an oral text, having a double semantic load and expressing in addition to external substantive contents the author's implied meaning, which is often a priority in the information perspective, eludes the addressee. The relevance of the proposed work is due to the need for a detailed description of this area of business communication (alongside with the expressive syntax) and identify the causes of misunderstanding of the very nature of the designated language phenomenon by recipients entering into speech cooperation.The expediency of the consideration of the parcelling in the designated key is justified by the fact that it, representing a multilevel semantic structure in structural terms, in addition to the visual shell endowed with a deep subtext interpreting the linguistic picture of the social environment, acts in the business style as an auxiliary element serving as a certain universal code for the allocation of the information segment in the speech flow and the strengthening of suggestion. On the other hand, the incorrect division of phrases in the speech flow not only undermines the literary norms, but also can cause communicative difficulties for the negotiators. Consequently, the strategy of conducting business conversations should be carried out taking into account the complex knowledge of psycholinguistics, culture of speech and office work in order to both recognise the parcelled units in the lexical and grammatical complex and take into consideration their structural and semantic features in the production of the text. This, in our opinion, forms a successful personality in terms of communication.
The article is devoted to the assessment of the network community as a collective subject, as a group of interconnected and interdependent persons performing joint activities. According to the main research hypothesis, various forms of group subjectness, which determine its readiness for joint activities, are manifested in the discourse of the network community. Discourse constitutes a network community, mediates the interaction of its participants, represents ideas about the world, values, relationships, attitudes, sets patterns of behavior. A procedure is proposed for identifying discernible traces of the subjectness of a network community at various levels (lexical, semantic, content-analytical scales, etc.). The subjective structure of the network community is described based on experts' implicit representations. The revealed components of the subjectness of network communities are compared with the characteristics of the subjectness of offline social groups. It is shown that the structure of the subjectness of network communities for some components is similar to the structure of the characteristics of the subjectness of offline social groups: the discourse of the network community represents a discussion of joint activities, group norms, and values, problems of civic identity. The specificity of network communities' subjectness is revealed, which is manifested in the positive support of communication within the community, the identification and support of distinction between "us" and "them". Two models of the relationship between discursive features and the construct "subjectness" are compared: additive-cumulative and additive. The equivalence of models is established based on the discriminativeness and the level of consistency with expert evaluation by external criteria.
The article presents the results of an original research into a hypothetical dependence of translation techniques selection on the term structure in the target text. The data was obtained following the analysis of translation techniques applied to render into Ukrainian 932 English terminological word combinations related to Teaching Foreign Languages and Applied Linguistics. It was established that word combinations constitute the most numerous category in the English terminology corpus selected for the analysis. It was also found that the share of two-component terminological word combinations considerably supersedes the share of lexical units with a larger amount of words. Adjective-Noun and Noun-Noun models turned out to be the most frequent models the two-component word combinations are based on in the said sphere. In the category of the two-component word combinations, the share of adjective-headed lexical units amounts to 52%, while the Adjective-Noun model accounts for half (49%) of them. The share of the noun-headed word combinations, where the Noun-Noun model predominates, is 37%. The rest of the models have a tendency to follow the adjective-headed word combinations in their behaviour. The analysis of the correlation of translation techniques selected to render into Ukrainian the Language Pedagogy English terminological word combinations allowed to assume that the choice of the techniques is dependent on their structure. The two-component adjective-headed word combinations tend to be translated by means of calque. However, in rendering the two-component noun-headed word combinations the share of calque diminishes by half. The increase of the amount of components in a word combination is accompanied by a sharp fall in the share of calque and the simultaneous rise in the proportion of transformations, dominated by transposition, often in combination with word addition or deletion. The research results do not give any ground to assume that the selection of translation techniques in rendering English terminological word combinations related to Teaching Foreign Languages and Applied Linguistics into Ukrainian has any specific features as compared to other specialized areas, because the said results are quite similar to those observed in other spheres of human activity. Like in those spheres, calquing is used if the principles of the word combination structure in English and Ukrainian coincide, while transformations occur in case of their discrepancies. Words are added into the target text to ensure a greater degree of rendering the meaning of the term from the source text, while the reason for the word deletion is the redundancy of the word combination in the source text, i.e. the possibility to render its meaning in the target text with fewer words. Transposition is applied to meet the target language norms, and the simultaneous use of several types of transformation is explained by the desire to comply with several requirements mentioned above.
The paper deals with studying language deviations of different types in James Joyce’s Ulysses and Finnegans Wake. Deviations in general are known to be a departure from a norm or accepted standard; in linguistics deviations are viewed as an artistic device that can be applied in different forms and at various textual levels. The author’s language deformation is analyzed as a form of deviation used for expressing the writer’s language knowledge. It is concluded that in Ulysses the destruction of the language is thoroughly thought out and multi-aimed. For instance, occasional compound units that dominate the novel imitate the style of Homer, reviving the ancient manner in contemporary language. Despite the use of conventional word-building patterns, rich semantic abundance being the basic principle of Joyce's poetics seriously complicates interpretation of the new words in the source language. The attempt is also made to systematize deviation techniques in Finnegans Wake. In particular, multilinguality is found to be the base of the lexical units created by J. Joyce. Such hybrid nonce words produce the polyphony effect and trigger the mechanism of polysemantism together with unlimited associativity of the textual material, broadening the boundaries of linguistic knowledge as a whole. Additionally, certain results of a deeper comparative analysis of the ways to translate the author’s deviations into Russian are given. The analysis of three Russian versions of Ulysses and the experimental fragmentary translation of Finnegans Wake show that there exists some regularity in the choice of translation method, particularly its dependence on the structural similarities/ differences of the source and the target languages, as well as the language levels affected by J. Joyce in the process of lingual destruction. The impossibility of complete conveyance of the semantic depth of the text and stylistic features in the target language is noted.
Cross-disciplinary communication is often impeded by terminological ambiguity. Hence, cross-disciplinary teams would greatly benefit from using a language technology-based tool that allows for the (at least semi-) automated resolution of ambiguous terms. Although no such tool is readily available, an interesting theoretical outline of one does exist. The main obstacle for the concrete realization of this tool is the current lack of an effective method for the automatic detection of the different meanings of ambiguous terms across different disciplinary jargons. In this paper, we set up a pilot study to experimentally assess whether the word sense induction technique of ‘context clustering’, as implemented in the software package ‘SenseClusters’, might be a solution. More specifically, given several sets of sentences coming from a cross-disciplinary corpus containing a specific ambiguous term, we verify whether this technique can classify each sentence in accordance to the meaning of the ambiguous term in that sentence. For the experiments, we first compile a corpus that represents the disciplinary jargons involved in a project on Bone Tissue Engineering. Next, we conduct two series of experiments. The first series focuses on determining appropriate SenseClusters parameter settings using manually selected test data for the ambiguous target terms ‘matrix’ and ‘model’. The second series evaluates the actual performance of SenseClusters using randomly selected test data for an extended set of target terms. We observe that SenseClusters can successfully classify sentences from a cross-disciplinary corpus according to the meaning of the ambiguous term they contain. Hence, we argue that this implementation of context clustering shows potential as a method for the automatic detection of the meanings of ambiguous terms in cross-disciplinary communication.
In this paper we present a pipeline for the detection of spelling variants, i.e., different spellings that represent the same word, in non-standard texts. For example, in Middle Low German texts in and ihn (among others) are potential spellings of a single word, the personal pronoun ‘him’. Spelling variation is usually addressed by normalization, in which non-standard variants are mapped to a corresponding standard variant, e.g. the Modern German word ihn in the case of in. However, the approach to spelling variant detection presented here does not need such a reference to a standard variant and can therefore be applied to data for which a standard variant is missing. The pipeline we present first generates spelling variants for a given word using rewrite rules and surface similarity. Afterwards, the generated types are filtered. We present a new filter that works on the token level, i.e., taking the context of a word into account. Through this mechanism ambiguities on the type level can be resolved. For instance, the Middle Low German word in can not only be the personal pronoun ‘him’, but also the preposition ‘in’, and each of these has different variants. The detected spelling variants can be used in two settings for Digital Humanities research: On the one hand, they can be used to facilitate searching in non-standard texts. On the other hand, they can be used to improve the performance of natural language processing tools on the data by reducing the number of unknown words. To evaluate the utility of the pipeline in both applications, we present two evaluation settings and evaluate the pipeline on Middle Low German texts. We were able to improve the F1 score compared with previous work from \(0.39\) to \(0.52\) for the search setting and from \(0.23\) to \(0.30\) when detecting spelling variants of unknown words.
Previous research findings supporting the advantages of the go/no-go choice over the yes/no choice in lexical decision task (LDT) have suggested that the go/no-go choice might require less cognitive resources in the non-decisional processes. This study aims to test such an idea using the event-related potential method. In this study, the tasks (yes/no LDT and go/no-go LDT) and word frequency (high and low) were manipulated, and the difference between the go/no-go choice and yes/no choice were examined with BP, pN, pN1, P200, N400, and P3 components that were assumed to be closely related with the various parameters in the diffusion model. The results showed that BP, pN and pN1 amplitudes reflecting the preparation stage were not differently affected by word frequency and the task type. However, ERPs after stimulus onset showed differences. The P200 amplitudes were smaller in the go/no-go task than in the yes/no task only for low-frequency words. N400 and P3 amplitudes were only affect)
Through advances in neural language modeling, it has become possible to generate artificial texts in a variety of genres and styles. While the semantic coherence of such texts should not be over-estimated, the grammatical correctness and stylistic qualities of these artificial texts are at times remarkably convincing. In this paper, we report a study into crowd-sourced authenticity judgments for such artificially generated texts. As a case study, we have turned to rap lyrics, an established sub-genre of present-day popular music, known for its explicit content and unique rhythmical delivery of lyrics. The empirical basis of our study is an experiment carried out in the context of a large, mainstream contemporary music festival in the Netherlands. Apart from more generic factors, we model a diverse set of linguistic characteristics of the input that might have functioned as authenticity cues. It is shown that participants are only marginally capable of distinguishing between authentic )
It is easier to indicate the ink color of a color-neutral noun when it is presented in the color in which it has frequently been shown before, relative to print colors in which it has been shown less often. This phenomenon is known as color-word contingency learning. It remains unclear whether participants actually learn semantic (word-color) associations and/or response (word-button) associations. We present a novel variant of the paradigm that can disentangle semantic and response learning, because word-color and word-button associations are manipulated independently. In four experiments, each involving four daily sessions, pseudowords—such as enas, fatu or imot—were probabilistically associated with either a particular color, a particular response-button position, or both. Neutral trials without color-pseudoword association were also included, and participants’ awareness of the contingencies was manipulated. The data showed no influence of explicit contingency awareness, but clear )
Time Perspective (TP) is an important area of research within the ‘psychological time’ paradigm. TP, or the manner in which individuals conduct themselves as a reflection of their cogitation of the past, the present, and the future, is considered as a basic facet of human functioning. These perceptions of time have an influence on our actions, perceptions, and emotions. Assessment of TP based on human language on Twitter opens up a new avenue for research on subjective view of time at a large scale. In order to assess TP of users’ from their tweets, the foremost task is to resolve grammatical tense into the underlying temporal orientation of tweets as for many tweets the tense information, and their temporal orientations are not the same. In this article, we first resolve grammatical tense of users’ tweets to identify their underlying temporal orientation: past, present, or future. We develop a minimally supervised classification framework for temporal orientation task that enables in)
Many Canadians experience unequal access to primary care services, despite living in a country with a universal health care system. Health inequalities affect all Canadians but have a much stronger impact on the health of vulnerable populations. Health inequalities are preventable differences in the health status or distribution of health resources as experienced by vulnerable populations. A geospatial approach was applied to examine how closely the distribution of primary care providers (PCPs) in London, Ontario meet the needs of vulnerable populations, including people with low income status, seniors, lone parents, and linguistic minorities. Using enhanced two step floating catchment area (E2SFCA) method, an index of geographic access scores for all PCPs and PCPs speaking French, Arabic, and Spanish were separately developed at the dissemination area (DA) level. To analyze how PCPs are distributed, comparative analyses were performed in association with specific vulnerable groups. G)
The authors of the article focus on the difficulties experienced by younger students in mastering the spelling norms of the Russian language. This is the inability to immediately distinguish in the consciousness of the signified and signifying, the inability to correctly determine the word stress and a number of others. The teacher should know the methods of formation of students ' concept of "phoneme" and the ability to recognize other phonetic units of the language. It is emphasized that the phonetic work should precede the graphic one, based on the development of the speech-motor apparatus. The authors present a description of some methods of formation of the phonetic competence, such as: exercises on the distinction between words as lexical units and as a "phonetic word", the correct syllabification, accent, modelling, awareness similarsocial functions.
Lately, discourse structure has received considerable attention due to the benefits its application offers in several NLP tasks such as opinion mining, summarization, question answering, text simplification, among others. When automatically analyzing texts, discourse parsers typically perform two different tasks: i) identification of basic discourse units (text segmentation) ii) linking discourse units by means of discourse relations, building structures such as trees or graphs. The resulting discourse structures are, in general terms, accurate at intra-sentence discourse-level relations, however they fail to capture the correct inter-sentence relations. Detecting the main discourse unit (the Central Unit) is helpful for discourse analyzers (and also for manual annotation) in improving their results in rhetorical labeling. Bearing this in mind, we set out to build the first two steps of a discourse parser following a top-down strategy: i) to find discourse units, ii) to detect the Cen)