Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Sleep enhances memories, particularly emotional memories. As such, it has been suggested that sleep deprivation may reduce posttraumatic stress disorder. This presumes that emotional memory consolidation is paralleled by a reduction in emotional reactivity, an association that has not yet been examined. In the present experiment, we used an incidental memory task in humans and obtained valence and arousal ratings during two sessions separated either by 12 h of daytime wake or 12 h including overnight sleep. Recognition accuracy was greater following sleep relative to wake for both negative and neutral pictures. While emotional reactivity to negative pictures was greatly reduced over wake, the negative emotional response was relatively preserved over sleep. Moreover, protection of emotional reactivity was associated with greater time in REM sleep. Recognition accuracy, however, was not associated with REM. Thus, we provide the first evidence that sleep enhances emotional memory while preserving emotional reactivity.
Background: Electronic health records are invaluable for medical research, but much of the information is recorded as unstructured free text which is time-consuming to review manually. Aim: To develop an algorithm to identify relevant free texts automatically based on labelled examples. Methods: We developed a novel machine learning algorithm, the 'Semi-supervised Set Covering Machine' (S3CM), and tested its ability to detect the presence of coronary angiogram results and ovarian cancer diagnoses in free text in the General Practice Research Database. For training the algorithm, we used texts classified as positive and negative according to their associated Read diagnostic codes, rather than by manual annotation. We evaluated the precision (positive predictive value) and recall (sensitivity) of S3CM in classifying unlabelled texts against the gold standard of manual review. We compared the performance of S3CM with the Transductive Vector Support Machine (TVSM), the original fully-supe)
Language change takes place primarily via diffusion of linguistic variants in a population of individuals. Identifying selective pressures on this process is important not only to construe and predict changes, but also to inform theories of evolutionary dynamics of socio-cultural factors. In this paper, we advocate the Price equation from evolutionary biology and the Pó lyaurn dynamics from contagion studies as efficient ways to discover selective pressures. Using the Price equation to process the simulation results of a computer model that follows the Pó lya-urn dynamics, we analyze theoretically a variety of factors that could affect language change, including variant prestige, transmission error, individual influence and preference, and social structure. Among these factors, variant prestige is identified as the sole selective pressure, whereas others help modulate the degree of diffusion only if variant prestige is involved. This multidisciplinary study discerns the primary and co)
The amino acid sequences of proteins determine their three-dimensional structures and functions. However, how sequence information is related to structures and functions is still enigmatic. In this study, we show that at least a part of the sequence information can be extracted by treating amino acid sequences of proteins as a collection of English words, based on a working hypothesis that amino acid sequences of proteins are composed of short constituent amino acid sequences (SCSs) or "words". We first confirmed that the English language highly likely follows Zipf's law, a special case of power law. We found that the rank-frequency plot of SCSs in proteins exhibits a similar distribution when low-rank tails are excluded. In comparison with natural English and "compressed" English without spaces between words, amino acid sequences of proteins show larger linear ranges and smaller exponents with heavier low-rank tails, demonstrating that the SCS distribution in proteins is largely scal)
Reviewed by: Nations of Nothing But Poetry: Modernism, Transnationalism, and Synthetic Vernacular Writing Patrick Redding Matthew Hart. 2010. Nations of Nothing But Poetry: Modernism, Transnationalism, and Synthetic Vernacular Writing. New York: Oxford University Press. $55.00 hc. 256 pp. Matthew Hart's Nations of Nothing But Poetry is the first volume in the new Oxford series "Modernist Literature and Culture" to focus primarily on poetry and to engage with issues of transnationalism. Hart's book, and the Oxford series more generally, are representative of what Douglas Mao and Rebecca Walkowitz have called the "the New Modernist Studies." Two of the most distinctive features of this methodological shift in the study of modernism are an emphasis on approaching literary objects and languages from [End Page 143] beyond the confines of the nation state and a new openness to literature's engagement with mass media. Much of the intellectual energy of Hart's book derives from the ambitious and elegant way it conceives of modernism as a transnational formation. Hart's specific intervention in the field rests on his claim that modernist vernacular poetry—often associated with the regional or the local—is in fact deeply aware of its position within multinational and global networks of language, culture, and political power. In terms of its desire to sustain a rigorously comparative angle of vision, Hart's book succeeds at almost every turn. Hart builds on the work of Anita Patterson, Brent Hayes Edwards, and Charles Pollard to call our attention to a unique constellation of vernacular poets from Scotland, the United States, Africa, and the Caribbean who are nonetheless engaged in a common intellectual project. The central figures discussed in Nations of Nothing But Poetry are Hugh MacDiarmid, Basil Bunting, Kamau Brathwaite, Melvin B. Tolson, Harryette Mullen, and Mina Loy, though W. H. Auden, Ezra Pound, and T. S. Eliot also make significant cameos. Such an eclectic list of poets indicates the synoptic nature of Hart's interests and the surprising nature of his claims about relationships across national and temporal boundaries. His ambition is nothing less than an account of twentieth-century vernacular poetry as a social force across the globe. Fittingly enough for a book concerned with acts of synthesis, Nations of Nothing But Poetry fuses theoretical models from across several disciplines. Hart draws upon major currents in modernist and postcolonial studies: theories of the vernacular and "minor" literature, "Afro-modernity," and arguments about diasporic and world literature. The writings of philosophical critics of political power—Giorgio Agamben, Etienne Balibar, Michel Foucault, Michael Hardt and Antonio Negri, David Harvey, Fredric Jameson—as well as leading theorists of postcolonial and diasporic thought—K. Anthony Appiah, Homi Bhabha, Dipesh Chakrabarty, Frantz Fanon, Dilip Gaonkar, Simon Gikandi, Paul Gilroy, Edouard Glissant, Ngugi wa Thiong'o, Edward Said, Gayatri Spivak—are cited frequently and with great precision in this book. Unlike some academic studies with a prominent theoretical armature, Hart also has a good ear for the vernacular. He is a careful close reader, offering nuanced and pointed remarks about the lexical and formal elements of individual lines of poetry. His discussion of Pound's translation of Women of Trachis and Mullen's Muse & Drudge were particularly insightful to this reader, casting those works in an entirely new light. While attentive to details of poetic form, Nations of Nothing But Poetry persistently focuses its attention on the relation between different scales and temporalities of social analysis. Hart conceives this book not as a series of commentaries on diverse poets across the globe, but rather as a justification [End Page 144] for fresh conceptual frameworks that can attend to the experience of 'alternative modernities.' He argues that the uneven economic, spatial, and temporal developments of modern life must be understood without reference to a single norm (historically derived from the West). For this reason, the category of 'nation' receives extended critical treatment in Hart's analysis. Modernist vernacular writing, he explains, exposes the seams between "local identities, the redoubtable nation-state form of government, and the increasingly globalized nature of twentieth-century culture" (5). "Synthetic vernacular writing" is the central concept of the book. The phrase is derived from what MacDiarmid called "Synthetic Scots," a...
In the Diaries of Marin Sanudo, there are two strange reports sent to Venetian authorities from Hvar, in August and September 1512. The documents are strange because they are in Latin, and not, as is the norm for Cinquecento reports from Dalmatia by Venetian officials, in the Veneto dialect of Italian. The author of these reports is Sebastiano Giustinian, the provveditore generale of Dalmatia in 1512, on a mission to supress several local revolts. Sebastiano Giustinian 1459-1543 was a successful Venetian diplomat, serving in Hungary around 1500, on the Ferrarese court of Alfonso I d'Este in 1506, in Brescia at the time of Venetian defeat at Agnadello 1509. After his Istrian and Dalmatian engagement in 1510-1512, Giustinian will go on to England, to the court of Henry VIII 1514-1519; during this period Giustinian exchanged letters with Erasmus and Thomas More, to Crete 1520-1523 and to France 1526-1531, ending his career as the procurator of St Mark's in Venice. Giustinian was obviously well suited to courts and diplomacy; however, a peace-keeping mission in Dalmatia required abilities of a different type. At first, Giustinian successfully suppressed uprisals in Zadar, Sibenik, and Split, with a simple demonstration of Venetian military power. The rebellious citizens and peasants of Hvar, however, were by this time well organised guerillas. Giustinian employed paramilitaries from nearby Poljica, Brac and Trogir to attack Vrboska, a rebel village on Hvar, but this action ended in uncontrolled looting, which was not well received in Venice. Giustinian then tried something else, an almost theatrical public performance in Stari Grad, where he offered the inhabitants a choice between war and peace, celebrating the peace they have chosen in the cathedral of the City of Hvar. But soon afterwards Giustinian suffered a defeat by guerillas in Jelsa, with rebels later taking political action against him in the Venetian Senate. The Latin reports were written from Hvar, on August 3, 1512 this is a letter Giustinian sent to his son Marino, intending it for public circulation, and on September 2, 1512. The letter to Marino reports Giustinian's successes in Zadar and Sibenik; the report to the Senate is an apology, where Giustinian tries to balance the looting of Vrboska with good news from Split and Stari Grad. The performance in Stari Grad, obviously inspired by a scene from Livy Liv. 21, 18-19, when Q. Fabius Maximus in 218 B. C., holding two ends of his toga, theatrically offered the Carthaginians a choice between war and peace, is itself reported in a high humanist style, with lexical echoes from Curtius Rufus and Paulinus Petricordiae, with a Ciceronian antithesis between mansuetudo and severitas cf. Cic. off. 1,88, and Ambrosius, Epistles 9, 64, 10. In a similar way, the report from Zadar contrasts teachings from Scripture, from Aristotle and Cicero, presented by Giustinian in a public speech, with laughable cowardice of fearful rebels, who try to escape disguised as females, or hide in holes barely fit for mice. The rhetoric of Giustinian's reports is a characteristical Renaissance humanist strategy, based on words and wisdom of the Ancients, on a strong belief that the Antiquity can explain the present day and offer solutions for current problems. Seen in this light, Giustinian's reports from Hvar testify to a breakdown of humanist rhetoric; they are written in Latin and styled as humanist texts as long as the provveditore believed that the situation can conform to ancient models. When events get out of hand, Giustinian drops the Latin; neither the language nor its literary models are fit for reporting one's own defeats and unclean, tangled issues of impasse. Moreover, such defeats frustrate the very essence of Renaissance humanism, its idea that, if we can control words, we can control reality as well.
The usefulness of parallel corpora in translation studies and machine translation is strictly related to the availability of aligned data. In this paper we discuss the issues related to the design of a tool for the alignment of data from a parallel treebank, which takes into account morphological, syntactic and semantic knowledge as annotated in this kind of resource. A preliminary analysis is presented which is based on a case study, a parallel treebank for Italian, English and French, i.e. ParTUT. The paper will focus, in particular, on the study of translational divergences and their implications for the development of an alignment tool of parallel parse trees that, benefitting from the linguistic information provided in ParTUT, could properly deal with such divergences.
Logical Forms are an exceptionally important linguistic representation for highly demanding semantically related tasks like Question / Answering and Text Understanding, but their automatic production at runtime is higly error-prone. The use of a tool like XWNet and other similar resources would be beneficial for all the NLP community, but not only. The problem is: Logical Forms are useful as long as they are consistent, otherwise they would be useless if not harmful. Like any other resource that aims at providing a meaning representation, LFs require a big effort in manual checking order to reduce the number of errors to the minimum acceptable – less than 1 %- from any digital resource. As will be shown in detail in the paper, the available resources – XWNet, WN30-lfs, ILF-suffer from lack of a careful manual checking phase, and the number of errors is too high to make the resource usable as is. We classified mistakes by their syntactic or semantic type in order to facilitate a revision of the resource that we intend to do using regular expressions. We also commented extensively on semantic issues and on the best way to represent them in Logical Forms. 1.
International audience
Scots: Studies in its Literature and Language John M. Kirk and Iseabail Macleod (eds). Rodopi, 2013 ISBN 9789042037397, 65 [euro], 309pp. Scholarly Festschrifts dedicated to a specific scholar usually mark the coronation of the achievements of a lifelong career, and express the esteem and consideration in which the honorand is held by their peers; in the present case, the unquestionable importance and the extremely high quality of J. Derrick McClure's committed involvement in most aspects of Scots research makes it only too appropriate that this collection should gather some of the best names in Scottish Studies to celebrate him with diversified articles focussing on his field of expertise. Jeremy Smith's 'Textual Afterlives: Barbour's Bruce and Hary's Wallace' skilfully presents the application of historical pragmatics to five editions of John Barbour's The Bruce and Blind Hary's The Wallace, namely John Ramsay's manuscript (1489), Robert Leprevik's print (1571), Andro Hart's edition of 1620, Robert Freebaim's of 1758 and John Pinkerton's of 1790 for the first, and John Ramsay's manuscript (1488), Robert Leprevik's edition of 1570, the Glasgow editions of 1685 and 1713, Robert Freebaim's of 1758, and Robert Morison's of 1790 for the second. Details such as layout, punctuation, capitalisation, fonts and the individual treatment of distinctively Scottish lexemes are carefully sifted in order to infer the effect which they presumably exerted on their contemporary Scottish readership; special attention is devoted to the medieval and early modern understanding of a text as a conglomerate of concepts rather than grammatical units, as well as to any indicator of the shift from an oral to a visual approach which the introduction of the printing press is known to have entailed. Variations in editorial choices during the centuries suggest a growing antiquarian interest in correctness, as well as the first signs of a romantic 'mythological historicity' which drew heavily on the epic's purported authenticity inspired by the authority of the manuscript originals; prefaces and textual interpolations, on the other hand, are noted as reflecting the evolution of society's political attitude towards its southern neighbour. It is this constant redefinition of Scottish society's identity which, according to the article's premise, is reflected in the textual minutiae, and which warrants the exploration of each edition as a culturally-embedded product of its period, rather than a mere reproduction of the original. Robert McColl Millar's To bring my language near to the language of men? Dialect and Dialect Use in the Eighteenth and Early Nineteenth Centuries: Some Observations' explores how the social, economic and political changes which marked the second half of the eighteenth century influenced the increasingly self-conscious recourse to dialect for literary purposes against the opposing tendency of widespread diffidence towards anything diverging from the accepted norm. The case study concentrates on two emigrants' letters to their families, one from a Scottish indentured servant in Maryland in the early eighteenth century, the second from an English political prisoner in New South Wales in the early 1800s: the relevant dates are posited as the two approximate temporal extremes between which the standard language is presumed to have imposed itself. The theoretical premise is then tested against the entries found in the early nineteenth-century Original Statistical Account of Scotland, which are subdivided into the two categories of overt attitudes towards language as opposed to covered ones. Both concepts have been previously developed by McColl Millar, and in this article identify the self-conscious, often ambivalent comments passed by the informants on the linguistic landscape of their district on the one hand, and the incursions of Scots lexical items, idioms and proverbs into an otherwise wholly English text on the other. …
A major computational burden, while performing document clustering, is the calculation of similarity measure between a pair of documents. Similarity measure is a function that assign a real number between 0 and 1 to a pair of documents, depending upon the degree of similarity between them. A value of zero means that the documents are completely dissimilar whereas a value of one indicates that the documents are practically identical. Traditionally, vector-based models have been used for computing the document similarity. The vector-based models represent several features present in documents. These approaches to similarity measures, in general, cannot account for the semantics of the document. Documents written in human languages contain contexts and the words used to describe these contexts are generally semantically related. Motivated by this fact, many researchers have proposed semantic-based similarity measures by utilizing text annotation through external thesauruses like WordNet (a lexical database). In this paper, we define a semantic similarity measure based on documents represented in topic maps. Topic maps are rapidly becoming an industrial standard for knowledge representation with a focus for later search and extraction. The documents are transformed into a topic map based coded knowledge and the similarity between a pair of documents is represented as a correlation between the common patterns. The experimental studies on the text mining datasets reveal that this new similarity measure is more effective as compared to commonly used similarity measures in text clustering.
Language processing is commonly characterized by an event-related increase in theta power (4-7 Hz) in scalp EEG. Oscillatory brain dynamics underlying alcohol's effects on language are poorly understood despite impairments on verbal tasks. To investigate how moderate alcohol intoxication modulates event-related theta activity during visual word processing, healthy social drinkers (N = 22, 11 females) participated in both alcohol (0.6 g/kg ethanol for men, 0.55 g/kg for women) and placebo conditions in a counterbalanced design. They performed a double-duty lexical decision task as they detected real words among non-words. An additional requirement to respond to all real words that also referred to animals induced response conflict. High density whole-head MEG signals and midline scalp EEG data were decomposed for each trial with Morlet wavelets. Each person's reconstructed cortical surface was used to constrain noise-normalized distributed minimum norm inverse solutions for theta frequencies. Alcohol intoxication increased reaction time and marginally affected accuracy. The overall spatio-temporal pattern is consistent with the left-lateralized fronto-temporal activation observed in language studies applying time-domain analysis. Event-related theta power was sensitive to the two functions manipulated by the task. First, theta estimated to the left-lateralized fronto-temporal areas reflected lexical-semantic retrieval, indicating that this measure is well suited for investigating the neural basis of language functions. While alcohol attenuated theta power overall, it was particularly deleterious to semantic retrieval since it reduced theta to real words but not pseudowords. Second, a highly overlapping prefrontal network comprising lateral prefrontal and anterior cingulate cortex was sensitive to decision conflict and was also affected by intoxication, in agreement with previous studies indicating that executive functions are especially vulnerable to alcohol intoxication.
The program part of the lexical database was developed as the Mozilla Firefox application (2005-2006); the processing software is designed as a modern lexicographic workstation. Since 2007, we have gradually complemented the database with the lexicographicallly relevant data (research project Creation of a Lexical Database of the Czech Language of the Beginning of the 21st Century, 2005-2011). We have focused predominantly on detailed treatment of the example part, which demonstrates the breadth of the collocability of the lemmas.
Contemporary society is going through a series of changes as a result of globalization. Transformations in the cultural aspect of society are reflected in the change of norms, which include "the negative prescriptions (categorical prohibitions) concerning people's behavior, the violation of which results in the appropriate sanctions" (Sotsiologia: Entsiclopedia 2003), - the taboo. Publications on the topic of taboo in mass media, productions of films with the titles containing the word "taboo", Internet blogs and journals written by representatives of different ethnic cultures in English and Russian have served as the material for the present analysis aimed at understanding how globalization affects the transformations in content and structure of proscriptive norms in different cultures. The investigation shows that the changes in the taboo systems are realized both in the violation of the prohibitive norms - detabooisation, and in the emergence of new taboos. Both aspects are expressed in verbal and non-verbal forms. One of the brightest manifestations of non-verbal violation of the traditional prohibitions is relationships between sexes: sexual relationships in younger age and before marriage, group- and homosexual love have become a norm in many world cultures. Another example of non-verbal break of taboo is global spread of obscene gestures. In verbal communication taboo is realized in the avoidance of particular language and themes in certain communicative situations. (Sternin 2000) Among the themes which are no more considered to be forbidden in different cultures are sexual relationships and connected with it naming certain parts of the human body; diseases connected with the intimate parts of the body. Parents not loving their children is no longer considered to be a shameful theme which has developed into a popular topic of a 'child-free' lifestyle. The result of the verbal detabooisation is the widely-discussed reduction of restrictions on the use of obscene language. A new kind of breaking the taboo is the way of using obscene language in the internet with the help of 'liturgatives' (term by G. Gusseinov 2008) - crossing out the text which might be perceived as indecent. The other tendency in the way taboo systems are changing is emergence of new taboos. Internet is one of the generators of new taboos, e.g., the notion of a password has gone beyond the military or detective discourse and has become part of everyday use. Analysis of the material shows among the new non-verbal taboos are the natural smell of the body, public display of a child's naked body. The globalization of young age as a value of society is leading to the rejection of old age. New taboos are also caused by the global spread of the concept of tolerance which is manifested in the use of euphemisms for the ethnic names. There is also evidence that the concept of tolerance underlies avoidance of naming unpleasant things on the whole. Taboos as well as norms are situational, that is the system of permitting and prohibiting norms regulates communication in different situations. The investigation shows that the notion of situational appropriacy is being transformed. Changes are observed in institutional discourse, e.g. in the Russian academic discourse the refusal from impersonality is caused by the influence of the English language academic culture. Changes in the fundamental values of society caused by globalization are also reflected in subcultures, e.g. the use of obscene language is considered to be the norm in rap texts irrespective of the language they are written in. "Rebellion against the norms" (Bauman 2000) is fed by mass media. Thus, numerous reality shows have lifted the prohibition to discuss and display personal life in public. The shift of the focus from the socially important aspects of life to the personal ones is the illustration of the spreading individualism as a globalizing value. Mass media have also broken the conventional restrictions on the display of violence which has become commonplace on TV and in cinema. Due to the fundamental power of taboo, descriptions of breaks of taboo are often accompanied by the epithet "shocking" and its synonyms. However, due to the nature of detabooisation regular and repetitive violation of prohibition leads to the loss of its power and to the conversion of the taboo into the norm. Detabooisation also results from the possibility of breaking norms as prohibition can only have power when the breach of it is followed by punishment. In modern society, due to the spread of multi-culturalism and pluralism of values the issue of sanctions has become questionable. Besides, due to the spread of tolerance as one of the key concepts of modern society communicative response is being limited. Discursively, the urge to break taboos is strengthened by the rational argumentation which plays off against the irrational nature of taboo. Analysis has shown that the changes in the taboo systems in different cultures are connected with the global process of emancipation. Modern world has placed freedom in the first place in the list of its values (Z. Bauman 2001) which can be illustrated by the frequency of appeal to the concept of freedom as an argument in favour of breaking taboos. The concept of freedom is often verbalized with the help of lexical units with the meaning of 'being beyond the boundaries' and interpreted with the help of the metaphor "open - close". Presumably, striving for the rejection of internal moral restrictions and the blurring of the boundaries are the interconnected aspects of one trend of globalization. Many values which underlie the process of detabooisation reflect the predominant influence of western cultures on the other world cultures as one of the dimensions of globalization. Analyzed discourse about transformations in the prohibitive norms not infrequently refers to the influence of American culture. Having its roots in the core of linguo-culture taboo is the reflection of fundamental values of society. Naturally, radical global changes in the norms provoke resistance of cultures affected. Like many other aspects of resistance to globalization, the opposition to the transformation of cultural prohibiting norms to a significant extent appeals to the threat of ethnic or national disintegration. In sum, the current tendencies of detabooisation and the emergence of new taboos are caused by the process of globalization and reflect the transformations in the value systems of the world cultures.
We examined the influence of female fertility on the likelihood of male participants aligning their choice of syntactic construction with those of female confederates. Men interacted with women throughout their menstrual cycle. On critical trials during the interaction, the confederate described a picture to the participant using particular syntactic constructions. Immediately thereafter, the participant described to the confederate a picture that could be described using either the same construction that was used by the confederate or an alternative form of the construction. Our data show that the likelihood of men choosing the same syntactic structure as the women was inversely related to the women's level of fertility: higher levels of fertility were associated with lower levels of linguistic matching. A follow-up study revealed that female participants do not show this same change in linguistic behavior as a function of changes in their conversation partner's fertility. We interpr)
In contrast with animal communication systems, diversity is characteristic of almost every aspect of human language. Languages variously employ tones, clicks, or manual signs to signal differences in meaning; some languages lack the noun-verb distinction (e.g., Straits Salish), whereas others have a proliferation of fine-grained syntactic categories (e.g., Tzeltal); and some languages do without morphology (e.g., Mandarin), while others pack a whole sentence into a single word (e.g., Cayuga). A challenge for evolutionary biology is to reconcile the diversity of languages with the high degree of biological uniformity of their speakers. Here, we model processes of language change and geographical dispersion and find a consistent pressure for flexible learning, irrespective of the language being spoken. This pressure arises because flexible learners can best cope with the observed high rates of linguistic change associated with divergent cultural evolution following human migration. Thus)
Autori istražuju otvorena pitanja leksickoga normiranja u hrvatskoj maritimoloskoj leksikografiji. S jedne strane, hrvatska leksicka norma u 19. st. u velikoj mjeri određena je kriterijima hrvatskih vukovaca, koji su leksicku “cistocu” određivali po mjeri pripadnosti leksika ruralnim novostokavskim organskim idiomima, a s druge strane golem i izuzetno bogat maritimni leksik bio je prokazan kao tuđ, tj. kao talijanski. Na taj nacin nastala je praznina u hrvatskoj normativnoj maritimoloskoj leksikografiji koja nije ni do danas popunjena, unatoc tome sto je Anicev Rjecnik hrvatskoga jezika prvi put u znatnijoj mjeri uveo brojne hrvatske maritimizme i dao im leksicki normativni legitimitet.
Using several languages has become a norm for those who want to learn and work in the European Union. However, teaching for plurilingualism is also a challenge. The present paper first clarifies the notions of plurilingualism and multilingualism, then discusses the role of crosslinguistic similarity in language learning in the case of European languages. It also shows how lexical crosslinguistic similarity can be used in teaching typologically related and unrelated languages, and discusses the key factors in noticing such similarity. The research presented reports on examining and raising language awareness of Polish‑‑ English cognate vocabulary in the case of a group of Polish teenage learners of English. It presents the results of a small‑‑ scale study in quasi‑‑ experimental design, as well as qualitative research on the learners’ opinions and attitudes. Finally, the paper presents implications for language pedagogy and focuses on the fact that awareness raising may affect the learners’ plurilingual competence.
Surveys on sensitive issues provide distorted prevalence estimates when participants fail to respond truthfully. The randomized-response technique (RRT) encourages more honest responding by adding random noise to responses, thereby removing any direct link between a participant’s response and his or her true status with regard to a sensitive attribute. However, in spite of the increased confidentiality, some respondents still refuse to disclose sensitive attitudes or behaviors. To remedy this problem, we propose an extension of Mangat’s (Journal of the Royal Statistical Society: Series B, 56, 93–95, 1994) variant of the RRT that allows for determining whether participants respond truthfully. This method offers the genuine advantage of providing undistorted prevalence estimates for sensitive attributes even if respondents fail to respond truthfully. We show how to implement the method using both closed-form equations and easily accessible free software for multinomial processing tree models. Moreover, we report the results of two survey experiments that provide evidence for the validity of our extension of Mangat’s RRT approach.
This functional magnetic resonance imaging (fMRI) study investigates a crucial parameter in spatial description, namely variants in the frame of reference chosen. Two frames of reference are available in European languages for the description of small-scale assemblages, namely the intrinsic (or object-oriented) frame and the relative (or egocentric) frame. We showed participants a sentence such as "the ball is in front of the man", ambiguous between the two frames, and then a picture of a scene with a ball and a man - participants had to respond by indicating whether the picture did or did not match the sentence. There were two blocks, in which we induced each frame of reference by feedback. Thus for the crucial test items, participants saw exactly the same sentence and the same picture but now from one perspective, now the other. Using this method, we were able to precisely pinpoint the pattern of neural activation associated with each linguistic interpretation of the ambiguity, whil)
LOLspeak is a complex and systematic reimagining of the English language. It is most often associated with the popular, productive and long-lasting Internet meme ‘LOLcats’. This style of English is characterised by the simultaneous playful manipulation of multiple levels of language. Using community-generated web content as a corpus, we analyse some of the common language play strategies (Sherzer 2002) used in LOLspeak, which include morphological reanalysis, atypical sentence structure and lexical playfulness. The linguistic variety that emerges from these manipulations displays collaboratively constructed norms and tendencies providing a standard which may be meaningfully adhered to or subverted by users. We conclude with a discussion of why people may choose to participate in such language play, and suggest that the language play strategies used by participants allow for the construction of complex identity.
The variety described here is representative of colloquial Assamese spoken in the eastern districts of Assam. Assam is a North-Eastern state of India, therefore Assamese and creoles of Assamese like Nagamese are spoken in the different North-Eastern states of Nagaland, Arunachal Pradesh, Meghalaya, and also the neighbouring country of Bhutan. Approximately 15 million people speak Assamese in India (see Ethnologue, Gordon 2005, which lists 15,374,000 speakers including those in Bhutan and Bangladesh). In the pre-British era (until 1826), the kingdom of Assam was ruled by Ahom kings and the then capital was based in the Eastern district of Sibsagar and later in Jorhat. American missionaries established the first printing press in Sibsagar and in the year 1846 published a monthly periodical Arunodoi using the variety spoken in and around Sibsagar as the point of departure. This is the immediate reason which led to the acceptance of the formal variety spoken in eastern Assam (which roughly comprises of all the districts of Upper Assam). Having said that, the language spoken in these regions of Assam also show a certain degree of variation from the written form of the ‘standard’ language. As against the relative homogeneity of the variety spoken in eastern Assam, variation is considerable in certain other districts which would constitute the western part of Assam, comprising of the district of Kamrup up to Goalpara and Dhubri (see also Kakati 1962 and Grierson 1968). In contemporary Assam, for the purposes of mass media and communication, a certain neutral blend of eastern Assamese, without too many distinctive eastern features, like /ɹ/ deletion, which is a robust phenomenon in the eastern varieties, is still considered to be the norm. The lexis of Assamese is mainly Indo-Aryan, but it also has a sizeable amount of lexical items related to Bodo among other Tibeto-Burman languages (Kakati 1962), and there are a substantial number of items borrowed from Hindi, English and Bengali in recent times.
This article focuses on the relationship between a lexical database and a derivational map, a hierarchical representation of the lexicon that specifies lexical and morphological inheritance by means of graph theory. In order to take steps towards construing a three-dimensional lexicon, this article also puts forward the concept of semantic pole. A semantic pole is a pivot of lexical organization defined as the area of lexical space comprised of the intersection of the lexical areas of one or more derivational paradigms and the major exponents of a semantic prime. This proposal is applied to the semantic pole sōð-trēowe in Old English and two main conclusions are reached. Firstly, a semantic pole constitutes a panchronic representation of lexical relations that contributes to the development of the third-generation Internet, which aims, among other things, at compiling databases and representing contents in 3D. Secondly, the concept of semantic pole constitutes an explanatory principle of derivational morphology and lexical semantics because it explains the degree of convergence between morphological and lexical inheritance, accounts for the clustering of lexical items around certain semantic poles and predicts the rise of polysemy.
Proponents of Neuro-Linguistic Programming (NLP) claim that certain eye-movements are reliable indicators of lying. According to this notion, a person looking up to their right suggests a lie whereas looking up to their left is indicative of truth telling. Despite widespread belief in this claim, no previous research has examined its validity. In Study 1 the eye movements of participants who were lying or telling the truth were coded, but did not match the NLP patterning. In Study 2 one group of participants were told about the NLP eye-movement hypothesis whilst a second control group were not. Both groups then undertook a lie detection test. No significant differences emerged between the two groups. Study 3 involved coding the eye movements of both liars and truth tellers taking part in high profile press conferences. Once again, no significant differences were discovered. Taken together the results of the three studies fail to support the claims of NLP. The theoretical and practical)
State of the art parsers are currently trained on converted versions of Penn Treebank into dependency representations which however don’t include null elements. This is done to facilitate structural learning and prevent the probabilistic engine to postulate the existence of deprecated null elements everywhere (see [15]). However it is a fact that in this way, the semantics of the representation used and produced on runtime is inconsistent and will reduce dramatically its usefulness in real life applications like Information Extraction, Q/A and other semantically driven fields by hampering the mapping of a complete logical form. What systems have come up with are “Quasi”-logical forms or partial logical forms mapped directly from the surface representation in dependency structure. We show the most common problems derived from the conversion and then describe an algorithm that we have implemented to apply to our converted Italian Treebank, that can be used on any CONLL-style treebank or representation to produce an “almost complete” semantically consistent dependency treebank.
Temporal predictability is thought to affect stimulus processing by facilitating the allocation of attentional resources. Recent studies have shown that periodicity of a tonal sequence results in a decreased peak latency and a larger amplitude of the P3b compared with temporally random, i.e., aperiodic sequences. We investigated whether this applies also to sequences of linguistic stimuli (syllables), although speech is usually aperiodic. We compared aperiodic syllable sequences with two temporally regular conditions. In one condition, the interval between syllable onset was fixed, whereas in a second condition the interval between the syllables' perceptual center (p-center) was kept constant. Event-related potentials were assessed in 30 adults who were instructed to detect irregularities in the stimulus sequences. We found larger P3b amplitudes for both temporally predictable conditions as compared to the aperiodic condition and a shorter P3b latency in the p-center condition than in)
Both facial expression and tone of voice represent key signals of emotional communication but their brain processing correlates remain unclear. Accordingly, we constructed a novel implicit emotion recognition task consisting of simultaneously presented human faces and voices with neutral, happy, and angry valence, within the context of recognizing monkey faces and voices task. To investigate the temporal unfolding of the processing of affective information from human face-voice pairings, we recorded event-related potentials (ERPs) to these audiovisual test stimuli in 18 normal healthy subjects; N100, P200, N250, P300 components were observed at electrodes in the frontal-central region, while P100, N170, P270 were observed at electrodes in the parietal-occipital region. Results indicated a significant audiovisual stimulus effect on the amplitudes and latencies of components in frontal-central (P200, P300, and N250) but not the parietal occipital region (P100, N170 and P270). Specifical)
Unsupervised dependency parsing is one of the most challenging tasks in natural languages processing. The task involves finding the best possible dependency trees from raw sentences without getting any aid from annotated data. In this paper, we illustrate that by applying a supervised incremental parsing model to unsupervised parsing; parsing with a linear time complexity will be faster than the other methods. With only 15 training iterations with linear time complexity, we gain results comparable to those of other state of the art methods. By employing two simple universal linguistic rules inspired from the classical dependency grammar, we improve the results in some languages and get the state of the art results. We also test our model on a part of the ongoing Persian dependency treebank. This work is the first work done on the Persian language. 1
The paper studies Hausa film language through the analysis of three communication strategies, namely proverbs, imperatives and forms of address. It shows that Hausa film creates a new discourse by reflecting modern and traditional Hausa society. The films preserve some accepted cultural norms of behavior and norms of communication in order to please the more conservative public. On the other hand, combination of traditional and modern Hausa lifestyle evokes changes in the discourse. The paper shows that proverbs are commonly used as communication strategy for indirectness, rather than a specialized language. It also discovers that imperatives are used as communication strategy in close relations between interlocutors (no matter what their social status is) to express the direct message. As for forms of address, traditional and borrowed terms reflect the changing style of life. The examples extracted from the Hausa films are to show how the regular grammatical and lexical means change their discourse function in new social context.
The paper presents and evaluates an efficient algorithm for measuring semantic similarity of texts. Calculating the level of semantic similarity of texts is a very difficult task and the proposed up to now methods suffer from computational complexity. This substantially limits their application area. The proposed algorithm tries to reduce the problem by merging a computationally efficient statistical approach to text analysis with a semantic component. The semantic properties of text words are extracted from the WordNet lexical database. The approach was tested using WordNets for two languages: English and Polish. The basic properties of this approach are also studied. The paper concludes with an analysis of the performance of the proposed method on a sample database and suggests some possible application areas.
This stylebook is an updated version of Telljohann et al. (2006). It describes the design principles and the annotation scheme for the German treebank TüBa-D/Z developed by the Division of Computational Linguistics (Lehrstuhl Prof. Hinrichs) at the Department of Linguistics (Seminar für Sprachwis-
Decoding pain in others is of high individual and social benefit in terms of harm avoidance and demands for accurate care and protection. The processing of facial expressions includes both specific neural activation and automatic congruent facial muscle reactions. While a considerable number of studies investigated the processing of emotional faces, few studies specifically focused on facial expressions of pain. Analyses of brain activity and facial responses elicited by the perception of facial pain expressions in contrast to other emotional expressions may unravel the processing specificities of pain-related information in healthy individuals and may contribute to explaining attentional biases in chronic pain patients. In the present study, 23 participants viewed short video clips of neutral, emotional (joy, fear), and painful facial expressions while affective ratings, event-related brain responses, and facial electromyography (Musculus corrugator supercilii, M. orbicularis oculi, M. zygomaticus major, M. levator labii) were recorded. An emotion recognition task indicated that participants accurately decoded all presented facial expressions. Electromyography analysis suggests a distinct pattern of facial response detected in response to happy faces only. However, emotion-modulated late positive potentials revealed a differential processing of pain expressions compared to the other facial expressions, including fear. Moreover, pain faces were rated as most negative and highly arousing. Results suggest a general processing bias in favor of pain expressions. Findings are discussed in light of attentional demands of pain-related information and communicative aspects of pain expressions.
We have developed EyeMap, a freely available software system for visualizing and analyzing eye movement data specifically in the area of reading research. As compared with similar systems, including commercial ones, EyeMap has more advanced features for text stimulus presentation, interest area extraction, eye movement data visualization, and experimental variable calculation. It is unique in supporting binocular data analysis for unicode, proportional, and nonproportional fonts and spaced and unspaced scripts. Consequently, it is well suited for research on a wide range of writing systems. To date, it has been used with English, German, Thai, Korean, and Chinese. EyeMap is platform independent and can also work on mobile devices. An important contribution of the EyeMap project is a device-independent XML data format for describing data from a wide range of reading experiments. An online version of EyeMap allows researchers to analyze and visualize reading data through a standard Web browser. This facility could, for example, serve as a front-end for online eye movement data corpora.
We describe a transformation-based learning method for learning a sequence of mono-lingual tree transformations that improve the agreement between constituent trees and word alignments in bilingual corpora. Using the manually annotated English Chinese Transla-tion Treebank, we show how our method au-tomatically discovers transformations that ac-commodate differences in English and Chi-nese syntax. Furthermore, when transforma-tions are learned on automatically generated trees and alignments from the same domain as the training data for a syntactic MT system, the transformed trees achieve a 0.9 BLEU im-provement over baseline trees. 1
This paper has two related purposes. First, our goal is to explain the results of recent research on twentieth century British (as well as American) English, using equivalent corpora of general written (published) English known as the ‘Brown Family’ of corpora. Limiting our attention to British corpora, the ‘Brown Family’ contains three matching corpora of a million words each, the BLOB, LOB and F-LOB corpora, sampled at roughly thirty-year intervals (1931±31 years, 1961 and 1991). (A fourth corpus from 1901±3 is under development, and one-third of it will be used in the latter part of this paper.) These enable us to trace the changing history of written (published) British English over a sixty-year period. Through changes in frequency in grammatical categories and constructions across a variety of genres, we observe largely consistent patterns of change which lend themselves to explanations in terms of what may be called general stylistic trends. To these trends we give such names as colloquialization (movement towards spoken norms of usage), densification (movement towards denser or more compact expression of meaning) and democratization (the trend towards avoidance of discrimination or inequality in the linguistic treatment of individuals). Only the first two of these trends will be explored in this paper. In the second part of the paper, we show how general stylistic norms, such as are provided by the ‘Brown Family’ corpora, can be used as a reference norm against which statistical deviations identify some of the characteristic features of style of an individual author or an individual text. For this we make use of Rayson’s Wmatrix software (http://ucrel.lancs.ac.uk/wmatrix/) for comparing (groups of) texts in terms of lexical, grammatical and semantic characteristics. Although the comparison is in some respects lacking in accuracy, it identifies typical style markers of an individual text, ordering them in terms of their differentness from the reference norm. It remains to be seen how far this computational technique can place the elusive notion of authorial style on an objective footing, but results so far are promising.
The present study explored different approaches for automatically scoring student essays that were written on the basis of multiple texts. Specifically, these approaches were developed to classify whether or not important elements of the texts were present in the essays. The first was a simple pattern-matching approach called “multi-word” that allowed for flexible matching of words and phrases in the sentences. The second technique was latent semantic analysis (LSA), which was used to compare student sentences to original source sentences using its high-dimensional vector-based representation. Finally, the third was a machine-learning technique, support vector machines, which learned a classification scheme from the corpus. The results of the study suggested that the LSA-based system was superior for detecting the presence of explicit content from the texts, but the multi-word pattern-matching approach was better for detecting inferences outside or across texts. These results suggest that the best approach for analyzing essays of this nature should draw upon multiple natural language processing approaches.
In this article, we investigate ambiguity in syntactic annotation. The ambiguity in question is inherent in a way that even human annotators interpret the meaning differently. In our experiment, we detect potential structurally ambiguous sentences with Constraint Grammar rules. In the linguistic phenomena we investigate, structural ambiguity is primarily caused by word order. The potentially ambiguous particle or adverbial is located between the main verb and the (participial) NP. After detecting the structures, we analyze how many of the potentially ambiguous cases are actually ambiguous using the double-blind method. We rank the sentences captured by the rules on a 1 to 5 scale to indicate which reading the annotator regards as the primary one. The results indicate that 67% of the sentences are ambiguous. Introducing ambiguity in the treebank/parsebank increases the informativeness of the representation since both correct analyses are presented.
After a period when the focus was essentially on mental architecture, the cognitive sciences are increasingly integrating the social dimension. The rise of a cognitive sociolinguistics is part of this trend. The article argues that this process requires a re-evaluation of some entrenched positions in linguistics: those that see linguistic norms as antithetical to a descriptive and variational linguistics. Once such a re-evaluation has taken place, however, the social recontextualization of cognition will enable linguistics (including sociolinguistics as an integral part), to eliminate the cracks in the foundations that were the result of suppressing the sociocultural underpinnings of linguistic facts. Structuralism, cognitivism and social constructionism introduced new and necessary distinctions, but in their strong forms they all turned into unnecessary divides. The article tries to show that an evolutionary account can reintegrate the opposed fragments into a whole picture that puts each of them in their ‘ecological position’ with respect to each other. Empirical usage facts should be seen in the context of operational norms in relation to which actual linguistic choices represent adaptations. Variational patterns should be seen in the context of structural categories without which there would be only ‘differences’ rather than variation. And emergence, individual choice, and flux should be seen in the context of the individual’s dependence on lineages of community practice sustained by collective norms.
In many areas of the behavioral sciences, different groups of objects are measured on the same set of binary variables, resulting in coupled binary object × variable data blocks. Take, as an example, success/failure scores for different samples of testees, with each sample belonging to a different country, regarding a set of test items. When dealing with such data, a key challenge consists of uncovering the differences and similarities between the structural mechanisms that underlie the different blocks. To tackle this challenge for the case of a single data block, one may rely on HICLAS, in which the variables are reduced to a limited set of binary bundles that represent the underlying structural mechanisms, and the objects are given scores for these bundles. In the case of multiple binary data blocks, one may perform HICLAS on each data block separately. However, such an analysis strategy obscures the similarities and, in the case of many data blocks, also the differences between the blocks. To resolve this problem, we proposed the new Clusterwise HICLAS generic modeling strategy. In this strategy, the different data blocks are assumed to form a set of mutually exclusive clusters. For each cluster, different bundles are derived. As such, blocks belonging to the same cluster have the same bundles, whereas blocks of different clusters are modeled with different bundles. Furthermore, we evaluated the performance of Clusterwise HICLAS by means of an extensive simulation study and by applying the strategy to coupled binary data regarding emotion differentiation and regulation.
Dependency parsing has attracted considerable interest from researchers and developers in natural language processing. However, to obtain a high‐accuracy dependency parser, supervised techniques require a large volume of hand‐annotated data, which are extremely expensive. This paper presents a simple and effective approach for improving dependency parsing with subtrees derived from unannotated data, which are easy to obtain. First, we use a baseline parser to parse large‐scale unannotated data. Then, we extract subtrees from dependency parse trees in the auto‐parsed data. Next, the extracted subtrees are classified into several sets according to their frequency. Finally, we design new features based on the subtree sets for parsing algorithms. To demonstrate the effectiveness of our proposed approach, we conduct experiments on the English Penn Treebank and Chinese Penn Treebank. The results show that our approach significantly outperforms baseline systems. It also achieves the best accuracy for the Chinese data and an accuracy competitive with the best known systems for the English data.
We present a detailed error analysis of a transition-based dependency parser trained on a Hindi dependency treebank. Parser error analysis has not been systematically examined from the point of view of treebanking before and this work intends to contribute in this area. We address two main questions in this paper: Can the parsing of certain structures be made easier by using alternative analyses for these structures? Are there certain linguistic cues implicit (or missing) in the current treebank that can be made explicit (or added) in order to make the parsing of complex constructions easier? These questions will guide us in examining the potential benefits of parser error analysis during treebanking. Through our experiments and analysis we were able to shed light on the causes of errors and subsequently have been able to improve the performance of the parser.
Context enables readers to quickly recognize a related word but disturbs recognition of unrelated words. The relatedness of a final word to a sentence context has been estimated as the probability (cloze probability) that a participant will complete a sentence with a word. In four studies, I show that it is possible to estimate local context–word relatedness based on common language usage. Conditional probabilities were calculated for sentences with published cloze probabilities. Four-word contexts produced conditional probabilities significantly correlated with cloze probabilities, but usage statistics were unavailable for some sentence contexts. The present studies demonstrate that a composite context measure based on conditional probabilities for one- to four-word contexts and the presence of a final period represents all of the sentences and maintains significant correlations (.25, .52, .53) with cloze probabilities. Finally, the article provides evidence for the effectiveness of this measure by showing that local context varies in ways that are similar to the N400 effect and that are consistent with a role for local context in reading. The Supplemental materials include local context measures for three cloze probability data sets.
Among various neural network language models (NNLMs), recurrent neural network-based language models (RNNLMs) are very competitive in many cases. Most current RNNLMs only use one single feature stream, i.e., surface words. However, previous studies proved that language models with additional linguistic information achieve better performance. In this study, we extend RNNLM by explicitly integrating additional linguistic information, including morphological, syntactic, or semantic factors. Our proposed RNNLM is called a factored RNNLM that is expected to enhance RNNLMs. A number of experiments are carried out that show the factored RNNLM improves the performance for all considered tasks: consistent perplexity and word error rate (WER) reductions. In the Penn Treebank corpus, the relative improvements over n-gram LM and RNNLM are 29.0 % and 13.0%, respectively. In the IWSLT-2011 TED ASR test set, absolute WER reductions over RNNLM and n-gram LM reach 0.63 and 0.73 points. Title and Abstract in another language (Chinese) ddddddddddddddd ddddddddddddNNLMddddddddddddddRNNLMddd ddddddddddddddddddddRNNLMddddddddddddddd ddddddddddddddddddddddddddddddddddddddd ddddddddddddddddddRNNLMddddddddddddddddd dddddddddfRNNLMddddddddddddddddfRNNLMdddddd dddddddRNNLMdddddddddddddddddddddWERddddd ddddfRNNLMdddddddddnddddddRNNLMddddd29.0%d13.0%d dIWSLT-2011 TEDdddddddddfRNNLMddddddddddd0.63ddd dnddddddd0.73ddddRNNLMdd
This study introduced a novel algorithm to compute similarities between natural languages. Using syntactical relationships derived from natural languages, the algorithm proposed a semantic structural model and quantified natural languages using the word similarity method based on WordNet and lexical databases. The experimental results indicated that the algorithm could yield optimal results in semantic recognition when applied to sentences or short texts that are grammatically complex or relatively long (longer than 12 words). The contribution of this study is in its conversion of the grammar of different natural languages into a unified semantic structure, through which the semantic similarity of two sentences or short texts can be obtained by comparison. This study aimed to enhance the capability of computers for fuzzy concept processing, which can be applied to the fields of search engines and artificial intelligence. For instance, in search engines, sentences or short text-based concepts may be semantically structured to replace key-word based queries when executing search tasks. In the field of artificial intelligence, this capability may be applied to intelligent agents to smooth the process of interaction between humans and machines.
We present two dependency parsers for Persian, MaltParser and MSTParser, trained on theUppsala PErsian Dependency Treebank. The treebank consists of 1,000 sentences today. Itsannotation scheme is based on Stanford Typed Dependencies (STD) extended for Persianwith regard to object marking and light verb contructions. The parsers and the treebank aredeveloped simultanously in a bootstrapping scenario. We evaluate the parsers by experimentingwith different feature settings. Parser accuracy is also evaluated on automatically generated andgold standard morphological features. Best parser performance is obtained when MaltParseris trained and optimized on 18,000 tokens, achieving 68.68% labeled and 74.81% unlabeledattachment scores, compared to 63.60% and 71.08% for labeled and unlabeled attachmentscore respectively by optimizing MSTParser.
Recently, Kuhlmann (2007, Dependency Structures and Lexicalized Grammars. PhD Thesis, Saarland University) and collaborators have shown how the derivations of generative grammars can be recast as dependency structures. This connection between the generative and dependency traditions opens the door to a fresh perspective on how to formally characterize natural language and what minimal machinery can cover such data. This article draws on both reported properties of structures in dependency treebanks and properties of informant data to determine the complexity of natural language along two dependency measures, gap degree (a measure of discontinuity) and well- versus ill-nestedness (whether interleaving substructures are permitted). We show that natural language includes constructions that require dependency analyses that are ill-nested and/or gap degree > 1, and argue that a grammar formalism on the right track for characterizing natural language should be able to generate such structures. We investigate the adequacy of tree-local multi-component tree adjoining grammar (TL-MCTAG) to cover existent data, examining the relationship between TL-MCTAG derivations and dependency representations. Though focused on TL-MCTAG, this work also advances the larger enterprise of discovering mathematically defined formal systems and testing their adequacies by using both linguistic judgments as well as large bodies of data from annotated corpora.