Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
The recent rise in digitized historical text has made it possible to quantitatively study our psychological past. This involves understanding changes in what words meant, how words were used, and how these changes may have responded to changes in the environment, such as in healthcare, wealth disparity, and war. Here we make available a tool, the Macroscope, for studying historical changes in language over the last two centuries. The Macroscope uses over 155 billion words of historical text, which will grow as we include new historical corpora, and derives word properties from frequency-of-usage and co-occurrence patterns over time. Using co-occurrence patterns, the Macroscope can track changes in semantics, allowing researchers to identify semantically stable and unstable words in historical text and providing quantitative information about changes in a word’s valence, arousal, and concreteness, as well as information about new properties, such as semantic drift. The Macroscope provides information about both the local and global properties of words, as well as information about how these properties change over time, allowing researchers to visualize and download data in order to make inferences about historical psychology. Although quantitative historical psychology represents a largely new field of study, we see this work as complementing a wealth of other historical investigations, offering new insights and new approaches to understanding existing theory. The Macroscope is available online at http://www.macroscope.tech.
One of the strategies that researchers have used to investigate the role of sensorimotor information in lexical–semantic processing is to examine the effects of words’ rated body–object interaction (BOI; i.e., the ease with which the human body can interact with a word’s referent). Processing tends to be facilitated for words with high as compared with low BOI, across a wide variety of tasks. Such effects have been referenced in debates over the nature of semantic representations, but their theoretical import has been limited by the fact that BOI is a fairly coarse measure of sensorimotor experience with words’ referents. In the present study, we collected ratings for 621 words on seven semantic dimensions (graspability, ease of pantomime, number of actions, animacy, size, danger, and usefulness), in order to investigate which attributes are most strongly related to BOI ratings and to lexical–semantic processing. BOI ratings were obtained from previous norming studies (Bennett, Burnett, Siakaluk, & Pexman in Behavior Research Methods, 43, 1100–1109, 2011; Tillotson, Siakaluk, & Pexman in Behavior Research Methods, 40, 1075–1078, 2008), and measures of lexical–semantic processing were obtained from previous behavioral megastudies involving either the semantic categorization task (concrete/abstract decision; Pexman, Heard, Lloyd, & Yap in Behavior Research Methods, 49, 407–417, 2017) or the lexical decision task (Balota et al., Behavior Research Methods, 39, 445–459, 2007). The results showed that the motor dimensions of graspability, ease of pantomime, and number of actions were all related to BOI, and that these dimensions together explained more variance in semantic processing than did the BOI ratings alone. These ratings will be useful for researchers who wish to study how different kinds of bodily interactions influence lexical–semantic processing and cognition.
Reverse correlation is an influential psychophysical paradigm that uses a participant’s responses to randomly varying images to build a classification image (CI), which is commonly interpreted as a visualization of the participant’s mental representation. It is unclear, however, how to statistically quantify the amount of signal present in CIs, which limits the interpretability of these images. In this article, we propose a novel metric, infoVal, which assesses informational value relative to a resampled random distribution and can be interpreted like a z score. In the first part, we define the infoVal metric and show, through simulations, that it adheres to typical Type I error rates under various task conditions (internal validity). In the second part, we show that the metric correlates with markers of data quality in empirical reverse-correlation data, such as the subjective recognizability, objective discriminability, and test–retest reliability of the CIs (convergent validity). In the final part, we demonstrate how the infoVal metric can be used to compare the informational value of reverse-correlation datasets, by comparing data acquired online with data acquired in a controlled lab environment. We recommend a new standard of good practice in which researchers assess the infoVal scores of reverse-correlation data in order to ensure that they do not read signal in CIs where no signal is present. The infoVal metric is implemented in the open-source rcicr R package, to facilitate its adoption.
Psychological researchers have traditionally focused on lab-based experiments to test their theories and hypotheses. Although the lab provides excellent facilities for controlled testing, some questions are best explored by collecting information that is difficult to obtain in the lab. The vast amounts of data now available to researchers can be a valuable resource in this respect. By incorporating this new realm of data and translating it into traditional laboratory methods, we can expand the reach of the lab into the wilderness of human society. This study demonstrates how the troves of linguistic data generated by humans can be used to test theories about cognition and representation. It also suggests how similar interpretations can be made of other research in cognition. The first case tests a long-standing prediction of Gentner’s natural partition hypothesis: that verb meaning is more subject to change due to the textual context in which it appears than is the meaning of nouns. Within a diachronic corpus, verbs and other relational words indeed showed more evidence of semantic change than did concrete nouns. In the second case, corpus statistics were employed to empirically support the existence of phonesthemes—nonmorphemic units of sound that are associated with aspects of meaning. A third study also supported this measure, by demonstrating that it corresponds with performance in a lab experiment. Neither of these questions can be adequately explored without the use of big data in the form of linguistic corpora.
This paper introduces a new experimental protocol for studying mental representations of urban soundscapes through a simulation process. Subjects are asked to create a full soundscape by means of a dedicated software tool, coupled with a structured sound data set. This paradigm is used to characterize urban sound environment representations by analyzing the sound classes that were used to simulate the auditory scenes. A rating experiment of the soundscape pleasantness using a seven-point bipolar semantic scale is conducted to further refine the analysis of the simulated urban acoustic scenes. Results show that (1) a semantic characterization in terms of presence/absence of sound sources is an effective way to characterize urban soundscape pleasantness, and (2) acoustic pressure levels computed for specific sound sources better characterize the appraisal than the acoustic pressure level computed over the overall soundscape.
This article investigates the role(s) of the katakana syllabary in Japanese discourse, with a focus on the discourse producer’s underlying motivation for using katakana for native Japanese terms and how it influences word perception. Specifically, this study analyses 1) a corpus of texts to identify patterns in use and 2) a survey of professional writers (e.g., journalists, column writers) to triangulate results obtained from text analysis. Results show that the katakana syllabary is used to indicate the word in question is somehow ‘different’ from the norm, making a visual and mental distinction in the commonly shared word or concept. While each writer’s motivations may widely vary, the findings of this study suggest that in any written discourse, katakana may be employed to conceptualise a dichotomous view in otherwise common concepts. This also suggests that over time the original role of the katakana syllabary has been extended to becoming the linguistic choice of convenience, with roles ranging from filling lexical gaps to creating meaning gaps in Japanese native words.
The article explores the problems of lexical and grammatical and specially legal interpretation of certain criminal procedural norms, in particular, those that regulate the victim's right to procedural communication in criminal proceedings. It is noted that the guiding principle of lawful interpretation is the principle of dialogical communication, where dialogue is understood as a dynamic and constructive way of thinking, creating, interpreting, leading from the analysis of legal and technical errors of legal norms to effective law-making and enforcement and developing in a spiral, because every dialogue must continue the previous ones and prepare the next ones.The notion of a lawful interpretation of a sectoral (criminal-procedural) legal norm is defined asa special independent form of official and informal interpretation, an actual need in which it exists before, during and after the application of the criminal-procedural norm, and is intended to ensure its interpretative evolution for the sake of effective legalization.It is emphasized that the importance and necessity of lexical and grammatical and special-legal interpretation is dictated by: the presence of gaps in sectoral (criminal-procedural) legislation; the existence of conflicts in sectoral (criminal procedural) legislation; availability of valuation concepts in sectoral (criminal procedural) legislation; the presence of issues related to legal and technical errors in sectoral (criminal) law; the presence of problems of the degree of legal regulation of the compositional construction of criminal procedural norms; the presence of issues related to the appeal to the linguistic structural elements of the criminal procedural norm.
The article explores the techniques of creating a humorous effect in the Italian film «The Taming of the Scoundrel», taking into account the audiovisual nature of the material under study. Modern approaches to defining the concept of humour in linguistics are examined and the main linguistic means of realization of humour are discussed, including contextual changes of lexical meanings of words, expletive constructions, rhetorical questions, accumulation of homogeneous terms, sentence, use of incompatible concepts, repetitions, contrasting comparisons, metaphors, black humour, hyperbole, grotesque synecdoche, antithesis etc. It was revealed that the humorous effect of the movie «The Taming of the Scoundrel» is based on a combination of two incompatible stereotyped images – residents of the Italian village and the city. The main method of humour creation was a combination of incompatible concepts and depictions of deviations from the norm (disrespect for a sick person, arrogant behaviour with a beautiful woman, rudeness towards a woman, secular nature of religion, when a monk trains the volleyball team, love and understanding towards animals and disrespect towards people). Special attention in the realization of humour is drawn to the visual elements, built on the principle of a combination of the inseparable, the deviation from the norm and hyperbolization. The dialogues show patterns of conversational style and simplicity of vocabulary and syntax, so they do not require the use of a variety of translation transformations. The article describes ways of translating Italian humour into Ukrainian and mechanisms for conservation the humorous effect. The vast majority of sentences are translated by literal translation. In some cases, lexical substitutions are used. For the most part, the humorous effect persists, except the case of translation of wordplay, realities, artistic comparisons. In some cases, the humorous effect could be conserved by choosing other translation strategies. The movie «The Taming of the Scoundrel» gained high popularity because of its simplicity of perception. Although Western humour differs from Eastern humour because of national cultural, socio-psychological differences, the humour in the movie «The Taming of the Scoundrel» is clear and understandable. The humorous effect is lost only in cases of linguistic or cultural identity of the phrases. However, the problem of creating a humorous effect in audiovisual contexts and its conservation in different cultural areas remains always urgent.
Temperament and Psychological Types can be defined as innate psychological characteristics associated with how we relate with the world, and often influence our study and career choices. Furthermore, understanding these features help us manage conflicts, develop leadership, improve teaching and many other skills. Assigning temperament and psychological types is usually made by filling specific questionnaires. However, it is possible to identify temperamental characteristics from a linguistic and behavioral analysis of social media data from a user. Thus, machine-learning algorithms can be used to learn from a user’s social media data and infer his/her behavioral type. This paper initially provides a brief historical review of theories on temperament and then brings a survey of research aimed at predicting temperament and psychological types from social media data. It follows with the proposal of a framework to predict temperament and psychological types from a linguistic and behaviora)
In opera theatres, a speech consultant assists singers in pronouncing foreign languages, as well as helping them with the pronunciation of Slovenian. A sung opera text needs to be intelligible and linguistically standardised. The articulation of the sounds of Slovenian literary language in opera singing depends on the linguistic norm and on musical laws. The opera theatre speech consultant is involved in the process of creating a performance from the very beginning. Detailed knowledge of the text, as well as of music and directing concepts, also enables the consultant to proofread the texts in the theatre programme in a quality way and to ensure that the surtitles correspond to the stage action.
This paper describes a set of Prolog modules with predicates to access the information stored in the lexical database WordNet. The aim is to use the defined predicates for empowering (fuzzy) logic programming systems for approximate reasoning tasks. To achieve this goal it is necessary to have means to calculate the degree of relationship between the words. Because WordNet relates words but does not give graded information of the relation between those words, it is necessary to implement standard methods to compute that gradation.
The article argues that social transformations in the last decades have actively influenced the language issues. The vocabulary and semantic system have expanded the resources of the language emotional and evaluative means, mainly at the expense of slang vocabulary. The study focuses on the functioning features of the argot vocabulary in the novel “Fox” by Myroslav Dochynets since it is the basic lexical stratum used in communication by the criminal world in the novel. The author considers important issues of the argot vocabulary life, its entry into the canvas of a work of art, its compliance with stylistic and aesthetic norms. The key findings of the study argue that creative originality is determined by the author's focus on the unusual language of his characters, his desire to convey their world outlook and philosophy by means of argot lexicon, and thus to draw attention of the reader. On the one hand, negative lexical nominations prevail (musora, slidak, kent, shushval, chuika). On the other hand, the author appeals to national language experience of a reader and achieves desired semantic and stylistic effect using units of lexical, phraseological, syntactic language levels. The author analyzes contextual synonyms which emphasize the local coloring of the individual narration: koroidka; vypravna colonia; terytoria, de zakon vidpochyvav; nevilnyi svit; dorosla zona. The study sheds light on the processes that take place in the literary language and assumes the need to evaluate all the components of the speech stream, including the dynamics of the Ukrainian argot system development. The conducted research proves that the language of the criminal world is constantly changing, developing, expanding the sphere of influence and use in fiction, that is, the language develops not in its traditional hypostasis, but in oral, spontaneous, socially uncontrolled speech. Thus, jargon is an element of the linguistic picture of the world, a powerful semiosphere of a particular time layer of culture, which opens a meaningful perspective in the word as a quintessence of the socio-cultural, spiritual, and psychological climate of the epoch.
Whether states have Article III standing is a question that has in recent years induced a puzzling and nonstandard patterning of votes amongst the Justices of the Supreme Court. It is, of course, not uncommon for that bench to be characterized by sharp ideological divides. What is unusual and symptomatic in the state standing litigation context is rather this: Specific Justices seem to adopt divergent, seemingly inconsistent, positions on the same basic question of constitutional law when it is presented in different litigation matters.1 When it comes to state standing, the Court’s ideological divide is not merely acute but also inconstant and seemingly unstable.\nConsider two recent cases in an evenly divided eight-member Court has been unable to reach decision on this Article III issue. The Court as a result demurred from a decision in both cases, albeit to divergent effect in the two matters.2 Although we do not know the break down in votes in either case, I think it is reasonable to assign the “liberal” and “conservative”3 Justices to the opposite sides of the state standing issue in these two cases based on the questions and preferences evinced in the oral arguments and other indicia of judicial preferences. 4 That is, the liberals (conservatives) sometimes embraced state standing, and sometimes uniformly rejected it. This suggests that the question of state standing does not have an obvious and unidirectional ideological valence. It rather implies that its ideological valence is unstable for individual Justices, even holding constant the bench’s composition.\nA rather dismayingly plausible interpretation of this dynamic would begin with the basic unpredictability of Article III standing doctrine and its consequent vulnerability to partisan polarization effects among the Justices in high-profile public law litigation. Where a state presses a left-leaning position, the logic goes, Justices and commentators take predictable positions pro and contra—and vice versa. This happens because the doctrine either cannot or more contingently does not impose a frictional constraint on the expression of their normative priors. The ensuing constellation of votes and hence majority or dissenting opinions can be predicted with some confidence if one knows which president appointed a Justice and how they would vote on the merits of a case.\nSuch a view would not break new ground. The law reviews resound with complaints about standing doctrine’s mutability5 as well as its mismatch with attractive normative accounts of Article III ends.6 But complaints about its unique incoherence are somewhat overstated. Some degree of instability is probably inevitable in multimember bodies such as the Supreme Court given social choice dynamics.7 That this instability would take on familiar partisan form in cases concerning policy questions with obvious and strong partisan coloring—such as immigration law,8 environmental law,9 and national healthcare policy10 —is by no stretch surprising given the larger pattern of partisan polarization among the Justices.\nStill, it is not very satisfying to end the analysis with this stark “legal realist” conclusion.11 Nor did I think it is enough to simply assume it is possible to assert some “principled” account of state standing without thinking about why the doctrine has generated these concurrent but diametrically opposed votes on similar cases. Brute resort to partisanship as explanans is insufficient not because it lacks predictive power, or because it is somehow false. Rather, in the United States of the early twenty-first century, national partisan divides tend to track deep and consequential normative divides.12 Resiling to partisanship to account for doctrinal difference may be accurate,13 but it obscures far more interesting questions about how and why recondite matters of federal jurisdiction take on more readily cognized political colors. It fails to illuminate why a division of votes happens. To the extent that legal scholarship aims to map, and then plot potential pathways across, normative contestation, partisanship-based explanans can be both powerful and simultaneously unavailing for the task at hand. They beg the question of how we are to interpret ideological divisions on the Court by ousting an analysis of ideas with a brute act of taxonomy.\nNor do I think it is plausible to stipulate by fiat a single normative key to the state standing problem by appealing to text, original public meaning, or the like. There is already some air between the lexical anchors of Article III standing—the terms “case” and “controversy” in disconnected elements of the Constitution’s text—and the normative motors of current standing doctrine. That doctrine has further developed largely in terms of cases lodged by private litigants; its translation to state actors is not necessarily a neat or obvious one.14 There are hence a large number of disarticulated joints in the doctrinal armature tying constitutional meaning to its application in specific circumstances.\nAs a result of these gaps, theoretical ipse dixits are decidedly underwhelming. The litigated world is just too fluid to be nailed down by formalist or originalist certainties. This problem undermines perhaps the most cogent alternative analytic method to the approach I take here. This approach would turn to a historical consideration of states’ ability to lodge certain kinds of suits in federal courts.15 The leading historical approach in this vein, however, implicitly assumes that the background relationships of states vertically with the federal government and its own citizens, and also horizontally to other states, have been constant and stable enough to enable meaningful transhistorical comparisons. I am not sure that is right (in fact, I am pretty sure it is pretty clearly wrong). The need for some translation of historical doctrine to a contemporary context creates a need for normative criteria to evaluate whether the linkages between anterior doctrinal forms and constitutional norms persist or have evaporated.16 History, in short, entails normative exegesis as much as any other modality of constitutional inquiry.\nIn what follows, I offer a quite modest contribution to debates on state standing.17 I do not offer “right answers.” Rather, I posit that it is useful to understand the “stakes” of state standing. By “stakes,” I mean the practical consequences of resolving, one way or another, the unsettled doctrinal choices respecting to the ability of states to initiate a matter in federal courts. Why, that is, does state standing mater? An inquiry into stakes can usefully proceed step-wise. A first task is to identify the subset of state standing cases that presently elicit division among the Justices. A second task is to articulate the interesting normative consequences of narrowing or widening the Article III gauge in this contested class. Parts I and II attend respectively to these tasks.\nIn particular, I aim here to flesh out the multifarious character of downstream consequences plausibly related to state standing doctrine. For example, it is already a familiar claim in litigation over this Article III question that a denial of state standing will lead to an issue’s nonjusticiability. My analysis suggests we should be a bit skeptical of that notion. This skepticism, in turn, helps decenter what has become a modal concern in state standing debates. Instead, it suggests the value of attending to other, less familiar institutional-design implications, such as effects on the structural constitution and the incentives of state officials. In the end, I suggest that the latter may well be more important than any other concern.\nMy conclusion then draws back from the specifics of state standing to develop some more general reflections on the contents and aims of federal-court scholarship in an era of obvious and powerful partisan and ideological polarization. Put crudely, the animating worry there is whether the deepingly polarizing of American society, which the Court cannot escape, alters the way that scholars—putatively above the partisan fray—should talk about and think about the law of federal jurisdiction.
Neurodegenerative diseases causing dementia are known to affect a person’s speech and language. Part of the expert assessment in memory clinics therefore routinely focuses on detecting such features. The current outpatient procedures examining patients’ verbal and interactional abilities mainly focus on verbal recall, word fluency, and comprehension. By capturing neurodegeneration-associated characteristics in a person’s voice, the incorporation of novel methods based on the automatic analysis of speech signals may give us more information about a person’s ability to interact which could contribute to the diagnostic process. In this proof-of-principle study, we demonstrate that purely acoustic features, extracted from recordings of patients’ answers to a neurologist’s questions in a specialist memory clinic can support the initial distinction between patients presenting with cognitive concerns attributable to progressive neurodegenerative disorders (ND) or Functional Memory Disorder ()
The article presents the results of the study of phraseological meaning, in particular the relation between cognitive (denotative-significative) and connotative components, as well as the completeness of the latter. The suitability of distinguishing a special, distinct from the lexical, meaning in idioms, which has long been one of the central problems of linguistic studies, tells them from words, variables phrases and sentences. Based on the achievements of modern linguistics, which postulate the “block organization” of semantics of linguistic units, the author combines denotation and phraseological meaning in one macrocomponent – cognitive, or denotative-significative, which will cover the formal and semantic parts of the meaning. Concerning the connotative component, in the semantics of idioms, it merges with the subject-logical information, since the process of nomination and perception of the reality with these units has an emotional and intellectual nature. In the structure of the connotative component of idioms, the author distinguishes between four interrelated and mutually dependent components: emotional, expressive, appraisal, and cultural. Due to this structure, the idioms appear as polylexical nominative-characterizing signs, the internal form of which retains the collective experience of sensory perception of objects of reality with varying degrees of expressiveness, a system of ideological positions, aesthetic assessments, norms of speech and extra-language behavior of a language team.
The subsistence of Neolithic populations is based on agriculture, whereas that of previous populations was based on hunting and gathering. Neolithic spreads due to dispersal of populations are called demic, and those due to the incorporation of hunter-gatherers are called cultural. It is well-known that, after agriculture appeared in West Africa, it spread across most of subequatorial Africa. It has been proposed that this spread took place alongside with that of Bantu languages. In eastern and southeastern Africa, it is also linked to the Early Iron Age. From the beginning of the last millennium BC, cereal agriculture spread rapidly from the Great Lakes area eastwards to the East African coast, and southwards to northeastern South Africa. Here we show that the southwards spread took place substantially more rapidly (1.50–2.27 km/y) than the eastwards spread (0.59–1.27 km/y). Such a faster southwards spread could be the result of a stronger cultural effect. To assess this possibility,)
Public communication in the contemporary world constitutes a particularly multifaceted phenomenon. The Internet offers unlimited possibilities of contact and public expression, locally and globally, yet exerts its power too, inducing the use of the Internet lingo, loosening language norms, and often encourages the use of a lingua franca, English in particular. This leads to linguistic choices that are liberating for some and difficult for others on ideological grounds, due to the norms of the discourse community, or simply because of insufficient language skills and linguistic means available. Such choices appear to particularly characterise post-colonial states, in which the co-existence of multiple local tongues with the language once imperially imposed and now owned by local users makes the web of repertoires especially complex. Such a case is no doubt India, where the use of English alongside the nationally encouraged Hindi and state languages stems not only from its historical past, but especially its present position enhanced not only by its local prestige, but also by its global status too, and also as the primary language of Online communication. The Internet, however, has also been recognised as a medium that encourages, and even revitalises, the use of local tongues, and which may manifest itself through the choice of a given language as the main medium of communication, or only a symbolic one, indicated by certain lexical or grammatical features as identity markers. It is therefore of particular interest to investigate how members of such a multilingual community, represented here by Hindi users, convey their cultural identity when interacting with friends and the general public Online, on social media sites. This study is motivated by Kachru’s (1983) classical study, and, among others, a recent discussion concerning the use of Hinglish (Kothari and Snell, eds., 2011). This research analyses posts generated by Hindi users on Facebook (private profiles and fanpages) and Twitter, where personalities of users are largely known, and on YouTube, where they are often hidden, in order to identify how the users mark their Indian identity. Investigated will be Hindi lexical items, grammatical aspects and word order, cases of code-switching, and locally coloured uses of English words and spelling conventions, with an aim to establish, also from the point of view of gender preferences, the most dominating linguistic patterns found Online.
This paper presents a case study of the use of the NINJAL Parsed Corpus of Modern Japanese (NPCMJ) for syntactic research. NPCMJ is the first phrase structure-based treebank for Japanese that is specifically designed for application in linguistic (in addition to NLP) research. After discussing some basic methodological issues pertaining to the use of treebanks for theoretical linguistics research, we introduce our case study on the status of the Coordinate Structure Constraint (CSC) in Japanese, showing that NPCMJ enables us to easily retrieve examples that support one of the key claims of Kubota and Lee (2015): that the CSC should be viewed as a pragmatic, rather than a syntactic constraint. The corpus-based study we conducted moreover revealed a previously unnoticed tendency that was highly relevant for further clarifying the principles governing the empirical data in question. We conclude the paper by briefly discussing some further methodological issues brought up by our case study pertaining to the relationship between linguistic research and corpus development.
Abstract The article focuses on a comparison of English-Estonian code-copying in blogs and in vlogs. This paper applies a usage-based approach, combining a cognitive angle with the code-copying framework. The aim is to provide a holistic view on contact-induced language change by applying bottom-up analysis of naturalistic multilingual language use. In contact linguistic literature it has been observed that there is a preference for insertions vs. alternations depending on sociolinguistic setting (community type, generation, proficiency) and structural properties of the languages involved. Research on English-Estonian language contacts has shown that in blogs there is no clear preference for either. From a different point of view, it has been observed that lexical impact (global copying in the terms of code-copying framework) precedes semantic and structural impact (selective copying) but provide no explanation. Comparisons between blogs (750 entries and 275,263 words from 45 bloggers) and vlogs (5,5 hours, approximately 36,854 words from 5 vloggers) shows that global copies and alternations heavily prevail over other types of copying, yet the number of selective copies is somewhat higher in vlogs and of mixed copies in blogs. Selective copies are loan translations rather than structural changes. We assume that (1) prevalence of global copies and alternations depends on genre norms (blogs and vlogs are constructed as non-monolingual, highly individualized genres); (2) as selective copies are mostly loan translations, it implies the role of meaning and cognitive aspects: idioms and fixed expressions are figurative and cognitively prominent; combinational properties and grammatical meanings are abstract; so, the more abstract the meaning is, the more cognitive effort and time is required for entrenchment and conventionalization; (3) copying and alternation is denser in vlogs because the genre is oral and spontaneous vs. written, edited genre of blog.
This article analyzes two translations and one adaptation of Gulliver’s travels, by Jonathan Swift, a novel originally published in the eighteenth century. The main goal of this study is to observe if translators and adaptors have applied the standard norm of Brazilian Portuguese as a tool to carry out a temporal distancing effect derived from both lexical selection and adoption of particular linguistic structures that attain a semantic, syntactic and stylistic representation of such effect. The results show that, in fact, translators has resorted to strategic use of certain lexical items and grammatical structures that bring out a temporal distancing effect, but some editions have displayed hybrid linguistic structures in line with some informal register, which includes one of the adaptations analyzed, particularly destined for young readers.
How various types of focus differ with respect to exhaustivity has been a topic of enduring interest in language studies. However, most of the theoretical work explicating such associations has done so cross-linguistically, and little research has been done on how people process and respond to them during language comprehension. This study therefore investigates the associations between the concept of exhaustivity and three focus types in Chinese (wh, cleft, and only foci) using a trichotomous-response design in two experiments: a forced-choice judgment and a self-paced reading experiment, both with adult native speakers. Its results show that, whether engaged in conscious decision-making or an implicit comprehension process, the participants distinguished only-focus and cleft-focus from wh-focus clearly, and also that there are specific differences between only-focus and cleft-focus in conscious decision-making. This implies that, in terms of the relationship between exhaustivity and)
Throughout the 20th century, there have been many different forms of abstract painting. While works by some artists, e.g., Piet Mondrian, are usually described as static, others are described as dynamic, such as Jackson Pollock’s ‘action paintings’. Art historians have assumed that beholders not only conceptualise such differences in depicted dynamics but also mirror these in their viewing behaviour. In an interdisciplinary eye-tracking study, we tested this concept through investigating both the localisation of fixations (polyfocal viewing) and the average duration of fixations as well as saccade velocity, duration and path curvature. We showed 30 different abstract paintings to 40 participants — 20 laypeople and 20 experts (art students) — and used self-reporting to investigate the perceived dynamism of each painting and its relationship with (a) the average number and duration of fixations, (b) the average number, duration and velocity of saccades as well as the amplitude and curvature area of saccade paths, and (c) pleasantness and familiarity ratings. We found that the average number of fixations and saccades, saccade velocity, and pleasantness ratings increase with an increase in perceived dynamism ratings. Meanwhile the saccade duration decreased with an increase in perceived dynamism. Additionally, the analysis showed that experts gave higher dynamic ratings compared to laypeople and were more familiar with the artworks. These results indicate that there is a correlation between perceived dynamism in abstract painting and viewing behaviour — something that has long been assumed by art historians but had never been empirically supported.
We developed a method to automatically assess texts for features that help readers produce gist inferences. Following fuzzy-trace theory, we used a procedure in which participants recalled events under gist or verbatim instructions. Applying Coh-Metrix, we analyzed written responses in order to create gist inference scores (GISs), or seven variables converted to Z scores and averaged, which assess the potential for readers to form gist inferences from observable text characteristics. Coh-Metrix measures reflect referential cohesion and deep cohesion, which increase GIS because they facilitate coherent mental representations. Conversely, word concreteness, hypernymy for nouns and verbs (specificity), and imageability decrease GIS, because they promote verbatim representations. Also, the difference between abstract verb overlap among sentences (using latent semantic analysis) and more concrete verb overlap (using WordNet) should enhance coherent gist inferences, rather than verbatim memory for specific verbs. In the first study, gist condition responses scored nearly two standard deviations higher on GIS than did the verbatim condition responses. Predictions based on GIS were confirmed in two text analysis studies of 50 scientific journal article texts and 50 news articles and editorials. Texts from the Discussion sections of psychology journal articles scored significantly higher on GIS than did texts from the Method sections of the same journal articles. News reports also scored significantly lower than editorials on the same topics from the same news outlets. GIS proved better at discriminating among texts than did alternative formulae. In a behavioral experiment with closely matched text pairs, people randomly assigned to high-GIS versions scored significantly higher on knowledge and comprehension.
Verb bias facilitates parsing of temporarily ambiguous sentences, but it is unclear when and how comprehenders use probabilistic knowledge about the combinatorial properties of verbs in context. In a self-paced reading experiment, participants read direct object/sentential complement sentences. Reading time in the critical region was investigated as a function of three forms of bias: structural bias (the frequency with which a verb appears in direct object/sentential complement sentences), lexical bias (the simple co-occurrence of verbs and other lexical items), and global bias (obtained from norming data about the use of verbs with specific noun phrases). For reading times at the critical word, structural bias was the only reliable predictor. However, global bias was superior to structural and lexical bias at the post-critical word and for offline acceptability ratings. The results suggest that structural information about verbs is available immediately, but that context-specific, semantic information becomes increasingly informative as processing proceeds.
The determination of the semantic extent of lexical units which refer to the nomination of an individual has led to a more precise lexical image of the speech of the Svrljig area, but also to the understanding of the cultural identity of a single linguistic community. In a word, the formation of these lexical units is primarily based on visual perception (color, size, shape and the like). At the same time, their final form was also influenced by the socio-cultural specificities of the speech community which rested on the traditional human life values of people living in a rural environment. The volume and diversity of the nominations used to denote physical characteristics reflect a clear cultural specificity of the speech community in which everything that disrupts the ideal of a natural image and is not in accordance with the social norms is deemed negative.
INTRODUCTION: The aim of the present study was to investigate how emotions influence pain, measured by one subjective self-rated measure, the numeric rating scale (NRS), and one objective physiological measure, the number of skin conductance responses (NSCR). METHOD: Eighteen volunteers were exposed to conditions with pictorial emotional stimuli (neutral, positive, negative), authentic ICU-sound (noise, no-noise) and electrical stimulation (pain, no-pain) individually titrated to induce moderate pain. When using all combinations of picture inducing emotions, sound, and pain, each of these conditions (12 conditions lasting for 60 seconds each) were followed by pain ratings. Ratings of arousal (low to high) and valence (pleasant to unpleasant) were used as indicators of affective state for each condition. Mean NSCR was also measured throughout the experiment for each condition. RESULTS: Even though NRS and NSCR increased during painful stimuli, they did not correlate during the trial. However, NSCR was positively correlated with the strength of the electrical stimulation, r = 0.48, P = 0.046, whereas NRS showed positive correlations with the anxiety level, assessed by affective ratings (arousal, r = 0.61, P < 0.001, and valence, r = 0.37, P < 0.001). CONCLUSIONS: The NRS was strongly influenced by affective state, with higher pain ratings during more anxiety-like states, whereas NSCR correlated to the strength of electrical pain stimulation. That reported pain is moderated by anxiety, puts forward a discussion whether reduction of the anxiety level should be considered during analgesia treatment.
Abstract Previous studies have established some differences in the use of nominal address terms in Cameroon French and in Western-based varieties of French. For instance, the use of kinship terms and social titles in Cameroon French diverges in many ways from their use in Hexagonal French. Also, Cameroon French speakers tend to use a wide range of lexical strategies (e.g., semantic shift, derivation, conversion, compounding, lexical borrowing) to generate address terms that take into account their communicative needs, interpersonal relationships involved, and sociocultural exigencies and norms. The aim of this chapter is to highlight the creativity involved in the emergence of a system of nominal forms of address original to Cameroon French and to show the pragmatic intents behind the emergence of these forms. Taking a postcolonial pragmatic perspective, I look at some innovative processes in the use of nominal address terms and discuss contexts in which they are employed. I also examine how some address terms are used in the realization of different types of speech acts by highlighting the sociocultural constraints and illocutionary intents behind the choices of Cameroon French speakers.
Abstract Meaning Representation (AMR) is a meaning representation framework in which the meaning of a full sentence is represented as a single-rooted, acyclic, directed graph. In this article, we describe an on-going project to build a Chinese AMR (CAMR) corpus, which currently includes 10,149 sentences from the newsgroup and weblog portion of the Chinese TreeBank (CTB). We describe the annotation specifications for the CAMR corpus, which follow the annotation principles of English AMR but make adaptations where needed to accommodate the linguistic facts of Chinese. The CAMR specifications also include a systematic treatment of sentence-internal discourse relations. One significant change we have made to the AMR annotation methodology is the inclusion of the alignment between word tokens in the sentence and the concepts/relations in the CAMR annotation to make it easier for automatic parsers to model the correspondence between a sentence and its meaning representation. We develop an annotation tool for CAMR, and the inter-agreement as measured by the Smatch score between the two annotators is 0.83, indicating reliable annotation. We also present some quantitative analysis of the CAMR corpus. 46.71% of the AMRs of the sentences are non-tree graphs. Moreover, the AMR of 88.95% of the sentences has concepts inferred from the context of the sentence but do not correspond to a specific word.
The same concept can mean different things or be instantiated in different forms, depending on context, suggesting a degree of flexibility within the conceptual system. We propose that a feature-based network model can be used to capture and predict this flexibility. We modeled individual concepts (e.g., banana, bottle) as graph-theoretical networks, in which properties (e.g., yellow, sweet) were represented as nodes and their associations as edges. In this framework, networks capture within-concept statistics that reflect how properties relate to one another across instances of a concept. We extracted formal measures of these networks that capture different aspects of network structure, and explored whether a concept’s network structure relates to its flexibility of use. To do so, we compared network measures to a text-based measure of semantic diversity, as well as to empirical data from a figurative-language task and an alternative-uses task. We found that network-based measures were predictive of the text-based and empirical measures of flexible concept use, highlighting the ability of this approach to formally capture relevant characteristics of conceptual structure. Conceptual flexibility is a fundamental attribute of the cognitive and semantic systems, and in this proof of concept we reveal that variations in concept representation and use can be formally understood in terms of the informational content and topology of concept networks.
When a shift in writing style is noticed in a document, doubts arise about its originality. Based on this clue to plagiarism, the intrinsic approach to plagiarism detection identifies the stolen passages by analysing the writing style of the suspicious document without comparing it to textual resources that may serve as sources for the plagiarist. Character n-grams are recognised as a successful approach to modelling text for writing style analysis. Although prior studies have investigated the best practice of using character n-grams in authorship attribution and other problems, there is still a need for such investigations in the context of intrinsic plagiarism detection. Moreover, it has been assumed in previous works that the ways of using character n-grams in authorship attribution remain the same for intrinsic plagiarism detection. In this paper, we study the effect of character n-grams frequency and length on the performance of intrinsic plagiarism detection. Our experiments utilise two state-of-the-art methods and five large document collections of PAN labs written in English and Arabic. We demonstrate empirically that the low- and the high-frequency n-grams are not equally relevant for intrinsic plagiarism detection, but their performance depends on the way they are exploited.
Combinatory Categorial Grammars provide a transparent interface between surface syntax and underlying semantic representation. Discourse Representation Theory allows the handling of meaning across sentence boundaries. Based on the foundations of these two theories along with the work of Johan Bos on the Boxer framework for English language, we propose an approach to the task of semantic parsing with Discourse Representation Structure for the French language. By giving an example of discourse analysis on French sentences and experimenting on 4,525 sentences taken from the French Treebank corpus, we demonstrate and evaluate the outcomes of our framework.
Abstract The Wedding Invitation is one of the significant text genres. Following genre analysis approach and discourse analysis (DA), the present research analysed the wedding invitation genres in Pakistan to explore generic structures, as well as the role played by the broader socio-cultural norms and values in shaping this genre. Therefore, a corpus of 50 wedding invitations in Urdu and English was randomly selected from cards received from January to June 2018. The results of this genre analysis revealed seven obligatory and one optional move in Urdu, while six obligatory and one optional move in English invitations. Through discourse analysis, it has been uncovered how religious association and cultural influence in Pakistani society shape textual selection. Little variation was displayed in the invitations of the two languages, presumably due to regional cultural reflections and recent influence of western values. A comparison of Pakistani and UK invitations showed differences not only in move selection but also in lexical choices which are shaped by the respective cultures.
<strong>🌙 This is an alpha pre-release of spaCy v2.1.0 and available on pip as <code>spacy-nightly</code>. It's not intended for production use. See here for the updated nightly docs.</strong> <pre><code class="lang-bash">pip install -U spacy-nightly </code></pre> If you want to test the new version, we recommend using a new virtual environment. Also make sure to download the new models – see below for details and benchmarks. ⚠️ Due to difficulties linking our new <code>blis</code> for faster platform-independent matrix multiplication, this nightly release currently doesn't work on Python 2.7 on Windows. We expect this to be corrected in the future. ✨ New features and improvements Tagger, Parser, NER and Text Categorizer <strong>NEW:</strong> Experimental ULMFit/BERT/Elmo-like pretraining (see #2931) via the new <code>spacy pretrain</code> command. This pre-trains the CNN using BERT's cloze task. A new trick we're calling <em>Language Modelling with Approximate Outputs</em> is used to apply the pre-training to smaller models. The pre-training outputs CNN and embedding weights that can be used in <code>spacy train</code>, using the new <code>-t2v</code> argument. <strong>NEW:</strong> Allow parser to do joint word segmentation and parsing. If you pass in data where the tokenizer over-segments, the parser now learns to merge the tokens. Make parser, tagger and NER faster, through better hyperparameters. Add simpler, GPU-friendly option to <code>TextCategorizer</code>, and allow setting <code>exclusive_classes</code> and <code>architecture</code> arguments on initialization. Add <code>EntityRecognizer.labels</code> property. Remove document length limit during training, by implementing faster Levenshtein alignment. Use Thinc v7.0, which defaults to single-thread with fast <code>blis</code> kernel for matrix multiplication. Parallelisation should be performed at the task level, e.g. by running more containers. Models & Language Data <strong>NEW:</strong> 2-3 times faster tokenization across all languages at the same accuracy! <strong>NEW:</strong> Small accuracy improvements for parsing, tagging and NER for 6+ languages. <strong>NEW:</strong> The English and German models are now available under the MIT license. <strong>NEW:</strong> Statistical models for Greek. <strong>NEW:</strong> Alpha support for Tamil, Ukrainian and Kannada, and base language classes for Afrikaans, Bulgarian, Czech, Icelandic, Lithuanian, Latvian, Slovak, Slovenian and Albanian. Improve loading time of <code>French</code> by ~30%. CLI <strong>NEW:</strong> <code>pretrain</code> command for ULMFit/BERT/Elmo-like pretraining (see #2931). <strong>NEW:</strong> New <code>ud-train</code> command, to train and evaluate using the CoNLL 2017 shared task data. Check if model is already installed before downloading it via <code>spacy download</code>. Pass additional arguments of <code>download</code> command to <code>pip</code> to customise installation. Improve <code>train</code> command by letting <code>GoldCorpus</code> stream data, instead of loading into memory. Improve <code>init-model</code> command, including support for lexical attributes and word-vectors, using a variety of formats. This replaces the <code>spacy vocab</code> command, which is now deprecated. Add support for multi-task objectives to <code>train</code> command. Add support for data-augmentation to <code>train</code> command. Other <strong>NEW:</strong> Enhanced pattern API for rule-based <code>Matcher</code> (see #1971). <strong>NEW:</strong> <code>Doc.retokenize</code> context manager for merging and splitting tokens more efficiently. <strong>NEW:</strong> Add support for custom pipeline component factories via entry points (#2348). <strong>NEW:</strong> Implement fastText vectors with subword features. <strong>NEW:</strong> Built-in rule-based NER component to add entities based on match patterns (see #2513). <strong>NEW:</strong> Allow <code>PhraseMatcher</code> to match on token attributes other than <code>ORTH</code>, e.g. <code>LOWER</code> (for case-insensitive matching) or even <code>POS</code> or <code>TAG</code>. <strong>NEW:</strong> Replace <code>ujson</code>, <code>msgpack</code>, <code>msgpack-numpy</code>, <code>pickle</code>, <code>cloudpickle</code> and <code>dill</code> with our own package <code>srsly</code> to centralise dependencies and allow binary wheels. <strong>NEW:</strong> <code>Doc.to_json()</code> method which outputs data in spaCy's training format. This will be the only place where the format is hard-coded (see #2932). <strong>NEW:</strong> Built-in <code>EntityRuler</code> component to make it easier to build rule-based NER and combinations of statistical and rule-based systems. <strong>NEW:</strong> <code>gold.spans_from_biluo_tags</code> helper that returns <code>Span</code> objects, e.g. to overwrite the <code>doc.ents</code>. Add warnings if <code>.similarity</code> method is called with empty vectors or without word vectors. Improve rule-based <code>Matcher</code> and add <code>return_matches</code> keyword argument to <code>Matcher.pipe</code> to yield <code>(doc, matches)</code> tuples instead of only <code>Doc</code> objects, and <code>as_tuples</code> to add context to the <code>Doc</code> objects. Make stop words via <code>Token.is_stop</code> and <code>Lexeme.is_stop</code> case-insensitive. Accept <code>"TEXT"</code> as an alternative to <code>"ORTH"</code> in <code>Matcher</code> patterns. Use <code>black</code> for auto-formatting <code>.py</code> source and optimse codebase using <code>flake8</code>. You can now run <code>flake8 spacy</code> and it should return no errors or warnings. See <code>CONTRIBUTING.md</code> for details. 🔴 Bug fixes Fix issue #1487: Add <code>Doc.retokenize()</code> context manager. Fix issue #1537: Make <code>Span.as_doc</code> return a copy, not a view. Fix issue #1574: Make sure stop words are available in medium and large English models. Fix issue #1585: Prevent parser from predicting unseen classes. Fix issue #1642: Replace <code>regex</code> with <code>re</code> and speed up tokenization. Fix issue #1665: Correct typos in symbol <code>Animacy_inan</code> and add <code>Animacy_nhum</code>. Fix issue #1748, #1798, #2756, #2934: Add simpler GPU-friendly option to <code>TextCategorizer</code>. Fix issue #1773: Prevent tokenizer exceptions from setting <code>POS</code> but not <code>TAG</code>. Fix issue #1782, #2343: Fix training on GPU. Fix issue #1816: Allow custom <code>Language</code> subclasses via entry points. Fix issue #1865: Correct licensing of <code>it_core_news_sm</code> model. Fix issue #1889: Make stop words case-insensitive. Fix issue #1903: Add <code>relcl</code> dependency label to symbols. Fix issue #1963: Resize <code>Doc.tensor</code> when merging spans. Fix issue #1971: Update <code>Matcher</code> engine to support regex, extension attributes and rich comparison. Fix issue #2014: Make <code>Token.pos_</code> writeable. Fix issue #2329: Correct <code>TextCategorizer</code> and <code>GoldParse</code> API docs. Fix issue #2369: Respect pre-defined warning filters. Fix issue #2390: Support setting lexical attributes during retokenization. Fix issue #2396: Fix <code>Doc.get_lca_matrix</code>. Fix issue #2464, #3009: Fix behaviour of <code>Matcher</code>'s <code>?</code> quantifier. Fix issue #2482: Fix serialization when parser model is empty. Fix issue #2644: Add table explaining training metrics to docs. Fix issue #2648: Fix <code>KeyError</code> in <code>Vectors.most_similar</code>. Fix issue #2671, #2675: Fix incorrect match ID on some patterns. Fix issue #2693: Only use <code>'sentencizer'</code> as built-in sentence boundary component name. Fix issue #2728: Fix HTML escaping in <code>displacy</code> NER visualization and correct API docs. Fix issue #2754, #3028: Make <code>NORM</code> a <code>Token</code> attribute instead of a <code>Lexeme</code> attribute to allow setting context-specific norms in tokenizer exceptions. Fix issue #2769: Fix issue that'd cause segmentation fault when calling <code>EntityRecognizer.add_label</code>. Fix issue #2772: Fix bug in sentence starts for non-projective parses. Fix issue #2779: Fix handling of pre-set entities. Fix issue #2782: Make <code>like_num</code> work with prefixed numbers. Fix issue #2833: Raise better error if <code>Token</code> or <code>Span</code> are pickled. Fix issue #2838: Add <code>Retokenizer.split</code> method to split one token into several. Fix issue #2870: Make it illegal for the entity recognizer to predict whitespace tokens as <code>B</code>, <code>L</code> or <code>U</code>. Fix issue #2871: Fix vectors for reserved words. Fix issue #2901: Fix issue with first call of <code>nlp</code> in Japanese (MeCab). Fix issue #2924: Make IDs of displaCy arcs more unique to avoid clashes. Fix issue #3012: Fix clobber of <code>Doc.is_tagged</code> in <code>Doc.from_array</code>. Fix issue #3027: Allow <code>Span</code> to take unicode value for <code>label</code> argument. Fix issue #3048: Raise better errors for uninitialized pipeline components. Fix issue #3064: Allow single string attributes in <code>Doc.to_array</code>. Fix issue #3093, #3067: Set <code>vectors.name</code> correctly when exporting model via CLI. Fix issue #3112: Make sure entity types are added correctly on GPU. Fix issue #3122: Correct docs of <code>Token.subtree</code> and <code>Span.subtree</code>. Fix issue #3128: Improve error handling in converters. Fix issue #3248: Fix <code>PhraseMatcher</code> pickling and make <code>__len__</code> consistent. Fix issue #3277: Add en/em dash to tokenizer prefixes and suffixes. Fix serialization of custom tokenizer if not all functions are defined. Fix bugs in beam-search training objective. Fix problems with model pickling. ⚠️ Backwards incompatibilities This version of spaCy requires downloading <strong>new models</strong>. You can use the <code>spacy validate</code> command to find out which models need upda
This paper investigates the historical (1850s–2000s) evolution of semantics in the English language using contemporaneous, decade-specific computational estimates of word concreteness. Study 1 describes the computational method of generating time-locked estimates of concreteness based on the Corpus of Historic American English, and makes available the computed scores for 25,000 English words over 15 decades. We also report several tests of reliability and validity, demonstrating that our historical concreteness scores have high levels of both. Study 2 uses concreteness scores to revisit findings of studies that use a static set of contemporary human concreteness norms to examine historical trends of semantic change. Specifically, we observed (contra Hills & Adelman, (Cognition, 143, 87–92 2015)) that distinct word types of the English language become increasingly more concrete over time and (in line with Hills & Adelman, (Cognition, 143, 87–92 2015) & Hills, Adelman & Noguchi, (The Quarterly Journal of Experimental Psychology, 70(8), 1603–1619 2016)) that relatively concrete words tend to be used more often than abstract ones. We discuss both contrastive and corroborative claims in light of recent work on semantic evolution and argue for the use of time-locked computed estimates over static human norms when examining diachronic linguistic phenomena.
Test publishers usually provide confidence intervals (CIs) for normed test scores that reflect the uncertainty due to the unreliability of the tests. The uncertainty due to sampling variability in the norming phase is ignored. To express uncertainty due to norming, we propose a flexible method that is applicable in continuous norming and allows for a variety of score distributions, using Generalized Additive Models for Location, Scale, and Shape (GAMLSS; Rigby & Stasinopoulos, 2005). We assessed the performance of this method in a simulation study, by examining the quality of the resulting CIs. We varied the population model, procedure of estimating the CI, confidence level, sample size, value of the predictor, extremity of the test score, and type of variance-covariance matrix. The results showed that good quality of the CIs could be achieved in most conditions. The method is illustrated using normative data of the SON-R 6-40 test. We recommend test developers to use this approach to arrive at CIs, and thus properly express the uncertainty due to norm sampling fluctuations, in the context of continuous norming. Adopting this approach will help (e.g., clinical) practitioners to obtain a fair picture of the person assessed.
In this article, we discuss and give examples of how word-form frequency information derived from existing corpora statistics can be used to improve dictionary content. The frequency information is used in combination with rule-based morphological data based on derivational and inflectional information from the Swedish Morphological Database compiled at the University of Gothenburg, and the lexical database owned by the Swedish Academy. The method currently used in the ongoing project for updating the monolingual Contemporary Dictionary of the Swedish Academy is described, and some examples of dictionary entries identified as candidates for update based on frequency measures are given. Different aspects of morphological dictionary content are discussed and highlighted by comparison between the above-mentioned definition dictionary and a learner’s dictionary. The role of headword or lemma as well as cross-referencing methods in a digital dictionary as compared to a printed dictionary is also discussed. Finally, a few examples of suggested modifications and enhancements are given.
Grammar anomalies are frequent in texts produced by second language learners. When describing these anomalies, two main issues arise: What is anomalous use of grammar? And how are grammar anomalies distinct from orthographic and lexical anomalies? We review earlier error-definitions and suggest defining grammar anomalies according to an explicit norm instead of L1 usage. We propose a broader definition of grammar than in previous studies, based on Boye & Harder (2012). The distinction between grammar, lexicon and orthography is illustrated with data from 28 adult L1 English learners of L2 Danish. In the corpus, 55.9 % of the anomalies were related to grammar. Finally, we discuss how definitions and procedures can be used in future studies of naturally occurring grammar anomalies.
Semantic alignment is a key process underlying interpersonal and team communication. However, semantic similarity is difficult to quantify, and statistical approaches designed to measure it often rely on methods that make the identification of the relative importance of key words difficult. This study outlines how conceptual recurrence analysis (CRA) can address these issues and can be used to detect conceptual structure in interpersonal communication. We developed several novel CRA metrics to analyze communication data reported previously by Mancuso, Finomore, Rahill, Blair, and Funke (Proceedings of the Human Factors and Ergonomics Society Annual Meeting, 58, 405–409, 2014), gathered from teams who worked cooperatively on a logic puzzle under different cognitive biasing contexts. CRA, like other measures of semantic coordination, relies on parameters whose values affect estimates of semantic alignment. We evaluated how the dimensionality of semantic spaces affects metrics quantifying the conceptual similarity of communicative exchanges, and whether metrics calculated from top-down, a priori semantic spaces or bottom-up semantic spaces empirically derived from each data set were more sensitive to biasing context. We found that the novel CRA measures were sensitive to manipulations of cognitive bias, and that higher-dimensional, bottom-up semantic spaces generally yielded more sensitivity to the experimental manipulations, though when the communication was evaluated with respect to specific key concepts, lower-dimensional, top-down spaces performed nearly as well. We conclude that CRA is sensitive to experimental manipulations in ways consistent with prior findings and that it presents a customizable framework for testing predictions about interpersonal communication patterns and other linguistic exchanges.
The synthesis of standardized regression coefficients is still a controversial issue in the field of meta-analysis. The difficulty lies in the fact that the standardized regression coefficients belonging to regression models that include different sets of covariates do not represent the same parameter, and thus their direct combination is meaningless. In the present study, a new approach called concealed correlations meta-analysis is proposed that allows for using the common information that standardized regression coefficients from different regression models contain to improve the precision of a combined focal standardized regression coefficient estimate. The performance of this new approach was compared with that of two other approaches: (1) carrying out separate meta-analyses for standardized regression coefficients from studies that used the same regression model, and (2) performing a meta-regression on the focal standardized regression coefficients while including an indicator variable as a moderator indicating the regression model to which each standardized regression coefficient belongs. The comparison was done through a simulation study. The results showed that, as expected, the proposed approach led to more accurate estimates of the combined standardized regression coefficients under both random- and fixed-effect models.
Discourse Relations, also known as coherence or rhetorical relations, characterize the semantic or pragmatic relationships between clauses or sentences in discourse.Such relations are established in order to facilitate effective communication.In addition to the inventory of relations, previous research has also investigated how discourse relations are established or signaled.Discourse markers (DMs) are considered to be the most typical signals in discourse; however, focusing merely on DMs is inadequate as they can only account for a small number of relations in discourse.Thus, researchers have been exploring textual signals beyond DMs such as the Penn Discourse Treebank 2.0 (PDTB, Prasad et al. [22]) and the Rhetorical Structure Theory Signalling Corpus (RST-SC, Das and Taboada [5]).Despite their different theoretical groundings and approaches to relation signaling, both corpora annotated the Wall Street Journal (WSJ) section of the Penn Treebank (PTB, Marcus et al. [19]), i.e. the news articles.Nevertheless, previous work has suggested that signaling information is indicative of genres (e.g.Taboada and Lavid [28]; Zeldes [34]).Therefore, this project aims to anchor signaling devices on a more diverse corpus to demonstrate the inadequacy of signaling by DMs only, the abundance of open-class signals, and more importantly, the distribution of signaling devices across genres.
The article deals with the semantics of the term “populism” and its re-lexicalization via changing the meaning from neutral to negative. The author gives examples of the third-wave American populist leaders’ speeches, namely those made by Ugo Chaves, Nicolas Maduro, Evo Morales and Daniel Ortega, and reveals linguistic means of populism, as well as discursive strategies: proliferation (lexical loading) of key concepts, such as “pueblo”, “nación”, “patria” (“people”, “nation”, “motherland”), constructing discursive “I” in a pattern “I am my people”, the use of discursive strategies of polarization that helps political leaders delegitimize the opponent and construct the para-reality to legalize and symbolize self-power. The populist discourse is marked semantically by key words pueblo, nación, patria, special contextual coloring is designed with set expressions like la voz del pueblo (voice of the people), by forming occasional derivatives that break language norms, by actualizing oppositions “friend - enemy”, “acquaintance – stranger”, “people – outsider”, by applying to slogans and precedent phenomena. The participation of precedent phenomena in discursive strategies of populism is highlighted in the choice of precedent names (i.e. Simón Bolívar, Augusto César Sandino, Rubén Darío), it is directed at affiliating modern political leaders with prominent leader of the past, thus pointing to national values and providing legitemisy of the politicains in power.
In the present research, we investigated whether people’s everyday language contains sufficient signal to predict the future occurrence of mental illness. Language samples were collected from the social media website Reddit, drawing on posts to discussion groups focusing on different kinds of mental illness (clinical subreddits), as well as on posts to discussion groups focusing on nonmental health topics (nonclinical subreddits). As expected, words drawn from the clinical subreddits could be used to distinguish several kinds of mental illness (ADHD, anxiety, bipolar disorder, and depression). Interestingly, words drawn from the nonclinical subreddits (e.g., travel, cooking, cars) could also be used to distinguish different categories of mental illness, implying that the impact of mental illness spills over into topics unrelated to mental illness. Most importantly, words derived from the nonclinical subreddits predicted future postings to clinical subreddits, implying that everyday language contains signal about the likelihood of future mental illness, possibly before people are aware of their mental health condition. Finally, whereas models trained on clinical subreddits learned to focus on words indicating disorder-specific symptoms, models trained to predict future mental illness learned to focus on words indicating life stress, suggesting that kinds of features that are predictive of mental illness may change over time. Implications for the underlying causes of mental illness are discussed.
The article discusses the influence of speech acts on the formation of personality. The rules of speech etiquette are traced in the stream of speech and are perceived by the interlocutor as part of the education and culture of people. In this case, it becomes necessary to know some factors of the linguistic culture of the people. Speech etiquette for each native speaker has its own language form. In the speech etiquette of all the peoples of the world, it is possible to identify common features in all languages there are stable formulas of greeting and farewell, forms of respectful treatment of elders. Analyzing speech acts and highlighting the features of the communicative behavior of native speakers, the authors pay attention to speech etiquette as a set of norms and stereotypes of communication that have developed in society due to historical traditions and social structure. If native speakers do not follow the rules of verbal communication, then it is impossible to speak about a high level of language proficiency. Since speech etiquette emphasizes the culture and traditions of the people, it is necessary to take into account discrepancies in the speech actions of different nationalities, since each language has its own, formed system of addresses. Knowledge of the norms of speech etiquette of the country with a different sociocultural reality is the key to success in interpersonal and professional communication. Intercultural communication is a very complex and multidimensional process, therefore, only knowledge of grammatical structures and lexical content of the language is not enough to communicate between two native speakers who have an excellent linguistic culture. It is the linguisticcultural competence and knowledge of the norms of speech acts that can serve as a basis for communication between representatives of different countries. В статье рассматривается влияние речевых поступков на формирование личности. Правила речевого этикета прослеживаются в потоке речи и воспринимаются собеседником как часть воспитания и культуры людей. При этом возникает необходимость знания некоторых факторов языковой культуры народа. Речевой этикет для каждого носителя языка имеет свою языковую форму. В речевом этикете всех народов мира можно выделить общие черты, во всех языках существуют устойчивые формулы приветствия и прощания, формы уважительного обращения к старшим. Анализируя речевые поступки и выделяя особенности коммуникативного поведения носителей языка, авторы уделяют внимание именно речевому этикету как совокупности норм и стереотипов общения, сложившихся в социуме в силу исторических традиций, социального устройства. Если носители родного языка не соблюдают правила речевого общения, то говорить о высоком уровне владения языком невозможно. Поскольку речевой этикет подчёркивает культуру и традиции народа, необходимо учитывать расхождения в речевых поступках разных национальностей, так как в каждом языке существует своя сформировавшаяся система обращений. Знание норм речевого этикета страны с иной социально культурной действительностью является залогом успеха в межличностном и профессиональном общении. Межкультурная коммуникация очень сложный и многоаспектный процесс, поэтому только знания грамматических структур и лексического наполнения языка недостаточно для общения двух носителей языка, обладающих отличной языковой культурой. Именно лингвострановедческая компетенция и знание норм речевых поступков может послужить основой для общения представителей разных стран.
The article deals with the relevant linguistic issue of correlation between word spelling and the distinction of units belonging to different grammatical classes. The concepts of word and part of speech are contrasted. The author has revealed the peculiarities of lexical units functioning in written speech, which enables their part-of-speech status identification. The analysis of criteria, suggested by the linguists for differentiation of homonymous adverbs, preposition-and-case-form combinations, and derivative prepositions resulted in proposing a procedure of consecutive operations, accomplished to identify the part-of-speech status of grammatically homonymous words. The results of the linguistic experiment show how in speech practice native speakers solve the problems of part of speech determination and establishing the spelling of such words. It was revealed that the methods for distinguishing grammatical homonyms used by recipients, in many cases, do not lead to the correct solution. The existing codified guidelines are not applied while writing, as spelling of the major part of grammatically homonymous words does not meet the requirements of the norm. To solve the problem under consideration, it is necessary to adjust the content of spelling rules and change the spelling of a number of words, where traditional spelling principles are reflected.
Introduction. The linguistic norm is a defining feature of literary language at all stages of its development. This linguistic phenomenon is characterized by complexity and multidimensionality predetermined by internal and external conditionality of language development. It explainsinsufficient research of the language norm. Every language realizes the implicit and explicit norm. A codified component of the latter is referred to as a prescriptive norm and non-codified one is labeled as descriptive.Purpose. The article focuses on determining the status of prescriptive and descriptive norms in linguistics.Methods. The paper applies linguistic methods and techniques, including descriptive analysis, as well as methods of comparison, classification, and generalization.Results. The research shows that the descriptive norm creates selection of already existing speech facts, based on the usage and the prescriptive norm. The descriptive standard (Ist-Norm) in German is related to the rules of a written language, rules of pronunciation and oral language.The spelling norm of the German language is always prescriptive, and the orthoepic norm is, on the contrary, descriptive, except for the media and public style. The process of German orthoepic norm codification is fostered by two trends. The representatives of the first movement considerthe norm of pronunciation as an ideal and understand it as a prescriptive norm. Their opponents believe that the process of codification is descriptive. The prescriptive norm is ideal. Generally, the prescriptive approach covers the issues on standards in pronunciation, syntax, correctstylistic use of lexical means. The prescriptive norm is labelled as a “regulatory norm” because it has passed its own way, and is considered as fixed one in a form for a certain period of time and regulates the use of linguistic means in speech. The frame construction is part of the prescriptive norm (Soll-Norm) of the German language, is characterized by a high degree of representation in texts of different styles and performs structural-syntactic and communicative-pragmatic functions in the sentence.Conclusion. The descriptive norm represents various possibilities of a particular language system, accepted and implemented by the linguistic society of the world. Characteristic features of the descriptive norm are as follows: its determinism, development in the process of change and simultaneously with the change in the language system, variability. The properties of the prescriptive norm are codification, awareness and obligation for all speakers, taking into account social status. In German, prescriptive (prдskriptive) and descriptive (deskriptive) rules are called Regeln rules.
<strong>Abstract </strong> <br />Using new methods to analyze poems with particular language is a fundamental requirement. One of these methods is theory of "literary reading" in the process of interpretation of text, especially poetry, which was introduced by Michel Riffater, a French literary critic and professor of poetry in 20th century, in his book "Poem Semiotics". In his opinion, semiotic method helps reader in interpretation and analysis of metaphorical basis of texts, especially in poetry, and reveals hidden semantic layers for him or her. Although Riffater′s method isn′t true in all kinds of poetry especially narrative, it would be helpful in certain types of poetry, especially lyric poems, in our own literature. <br />Here reader travels to mental background of writer or poet through considering level of manifestation of metaphors, conventional poetry′s interpretations and beliefs, hi or her own poetry metaphors and beliefs, exploring semantic infrastructure of text and the magic of vocabulary alignments. He also finds out underlying layers of text, thereby analyzing complex and obscure text. <br />Because Bidel's poem is always associated with consecutive deviation from norms and systematic rules, and through accumulation of lexical and cluster combinations, surprises readers with new metaphor, interpretation of his poems can be a model for explaining poet's plan and mental coherence. This method forces reader to reflect about Bidel's poetry and find new implications of his poetry. This process advances reader mathematically and systematically, and through hypograms, he or she understands and receives final matrix of poetry. Also, by this kind of reading of Bidel sonnet, we can reach the predominant structure in most of his sonnets. <br /><strong>Hypogram</strong> <br />Hypogram is a picture connoisseur of poetry in minds of reader and presents central concept of poetry in colorful, widespread, strange and unfamiliar clothes. In fact, hypograms are the key issues of a poem, and poetry matrix is compiled through collecting hypograms. <br /><strong>Matrix</strong> <br />Matrix is the dominant meaning of a piece of poetry and the fundamental proposition of poetry and the roots of the hypograms, which gradually develops in the reading of poetry in the mind of audience, but practically does not exist in poetry and produces poetry hypograms. Matrix exists as an inanimate soul in semantic context of poetry and shows itself behind all the metaphors and poetic images, leading to unity and integrity of poetry. <br /> The process of Riffater′s interpretation can be summarized as follows: 1. reading text for regular reception; 2. highlighting elements that look unusual and preventing a conventional interpretation; 3. discovering hypograms or unclear parts in texts; and, 4. finding a matrix of hypograms that is a sentence or word and can produce hypograms and text. <br />By reflecting on matrix of sonnet and finding hypograms, one can conclude that all verses serve a single basis and inner meaning. With discovery of various types of hypograms, one can find a correlation between words of sonnet and unique message of poet, and it helps in interpretation of other similar poetries. In this study, for the first time, we have applied theory of Riffater in understanding Bidel Dehlavi poetry. In the process of interpreting a sonnet of Bidel, we reached hypograms such as distress, tears in the time of dawn, retreat, apologies, love and moving and understanding inner of human, all of which had a mystical meaning and thus we arrived at the matrix of "excellence and reaching God." This matrix is seen in the background of most of Bidel's sonnets. <br />Investigating resurrection of vocabularies in Bidel's poem would remove stain of failure from Persian language. Combining parts of a sonnet showed perfect power of Persian language and ability of Bidel in combining words. In the end, we must say that in interpretation of Persian poetry, which is an unfamiliar space, this method can be used. But, still remains a big space for such studies concerning adapting theory of Riffater to other Persian poems. It is hoped that we have taken a step in this direction.
Part of speech (POS) tagging, the assignment of syntactic categories for words in running text, is significant to natural language processing as a preliminary task in applications such as speech processing, information extraction, and others. Urdu language processing presents a challenge due to the dual behaviour of various Urdu POS tags in differing situations (morphosyntactic ambiguity). This paper addresses this challenge by developing a novel tagging approach using linear-chain conditional random fields (CRF). Our work is the first instance of a CRF approach for Urdu POS tagging. The proposed model employs a strong, stable and balanced language-independent as well as language dependent feature set. The language-dependent feature considered includes part-of-speech tag of the previous word and suffix of the current word while the language-independent features includes the ‘context words window’. Our approach was evaluated against support vector machine techniques for Urdu POS—considered as state of the art—on two benchmark datasets. The results show our CRF approach to improve upon the F-measure of prior attempts by 8.3–8.5%.
Machine translation (MT) is directly linked to its evaluation in order to both compare different MT system outputs and analyse system errors so that they can be addressed and corrected. As a consequence, MT evaluation has become increasingly important and popular in the last decade, leading to the development of MT evaluation metrics aiming at automatically assessing MT output. Most of these metrics use reference translations in order to compare system output, and the most well-known and widely spread work at lexical level. In this study we describe and present a linguistically-motivated metric, VERTa, which aims at using and combining a wide variety of linguistic features at lexical, morphological, syntactic and semantic level. Before designing and developing VERTa a qualitative linguistic analysis of data was performed so as to identify the linguistic phenomena that an MT metric must consider (Comelles et al. 2017). In the present study we introduce VERTa’s design and architecture and we report the experiments performed in order to develop the metric and to check the suitability and interaction of the linguistic information used. The experiments carried out go beyond traditional correlation scores and step towards a more qualitative approach based on linguistic analysis. Finally, in order to check the validity of the metric, an evaluation has been conducted comparing the metric’s performance to that of other well-known state-of-the-art MT metrics.
Translating proper names in earlier Romanian versions of the Bible raised different challenges. Some of them were solved in the main text, some other in marginal notes. Such notes are to be found in the second complete translation of the Old Testament into Romanian, kept in the manuscript no. 4389 from the Romanian Academy Library and dated in the second half of the 17th century. The marginal notes from this old Romanian translation refer to the relation of the text with its Slavonic source, in terms of correcting the translation errors, with the secondary sources (in Latin, Romanian, and Greek), pointing to some denomination models different from the main source, and with the linguistic norm of the translated text, in terms of grammatical and lexical adaptations to the system and vocabulary of Romanian. This article explores the strategies related to the translation into Romanian of biblical names based on their treatment in the marginal notes of the mentioned text; it also aims at clarifying, as far as possible, the sources and how the translator relates to them.