Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Sentiment ambiguous adjectives, which have been neglected by most previous researches, pose a challenging task in sentiment analysis. We present an evaluation task at SemEval-2010, designed to provide a framework for comparing different approaches on this problem. The task focuses on 14 Chinese sentiment ambiguous adjectives, and provides manually labeled test data. There are 8 teams submitting 16 systems in this task. In this paper, we define the task, describe the data creation, list the participating systems, and discuss different approaches.
A Poetics Sacralized:Luis de Góngora's Soledades as Religious Rhetoric in Luis de Tejeda's "Romance Sobre su vida" R. John McCaw During the 1610s, the Spanish poet Luis de Góngora (1561-1627) composed and circulated at court his most ambitious and experimental poems: the Fábula de Polifemo y Galatea and the Soledades.1 These poems received some acclaim at the time, but they generated much more hostility: characterized by intense erudition, densely packed tropes, and convoluted, latinized syntax, Góngora's innovative style sparked a literary firestorm the likes of which had never been seen before in Castilian culture.2 Góngora's detractors universally condemned the use of violent hyperbaton and wordy, metaphoric overlay, as such techniques created too much textual and interpretive difficulty, and thus too strongly contested poetry's traditional role in clearly communicating aesthetic, social, and moral objectives to readers and listeners. Though the anti-gongoristas appreciated a certain degree of textual challenge, they believed that Góngora's work sacrificed a necessary level of conceptual coherence for the sake of sensational techniques and tropes.3 Many of Góngora's most ardent detractors targeted his unconventional use of heroic and lofty poetic genres for the expression of mundane themes. Specifically, critics assailed Góngora's use of the revered octava real for mythological storytelling in the Fábula de Polifemo y Galatea, and also lambasted him for using the elegant silva in order to express the simple, terrestrial travels of the protagonist peregrino in the Soledades. Indeed, Góngora's poems do not merely attire worldly themes with the cadences and rhymes of elevated [End Page 3] genres: the Fábula and the Soledades convey a materialistic, earth-centered worldview in which conventional markers of Christian symbolism and morality are not immediately evident. In an age when the Spanish literary establishment strongly upheld the Horatian principle of harnessing moral instruction to aesthetic objectives (especially in lengthy texts written in serious verse forms), Góngora's poetry proved not only challenging and experimental, but downright cryptic and contrarian. Góngora's signature style developed over the course of decades (from the early 1580s through the 1610s and 1620s) and generally mirrored lexical and stylistic trends in Castilian poetry, but Góngora's detractors nevertheless saw the poetic experiments as a declaration of war against accepted literary conventions, linguistic standards, thematic expectations, and didactic norms. Góngora's most trenchant critics pounced on this last feature, and suggested that the absence of a Christian perspective in Góngora's poems was equivalent to an anti-Christian one.4 As John Beverley notes, the poet Francisco de Quevedo accused Góngora of being a converso, and the commentator Francisco Cascales referred to Góngora as the "Mahoma de la poesía española" ("Sobre" 35). Ultimately, in the 1610s and 1620s and even beyond, for many powerful members of Spain's literary elite, the formal and structural characteristics of Góngora's unique style instantly evoked a non-Christian, and at times heterodox, worldview: "Por haber elevado el juego lingüístico ingenioso como centro del gusto poético, no directamente relacionado con la moral o la doctrina, se pensaba que Góngora había producido un formalismo funcionalmente ateo, que su poesía era babélica" (Beverley, "Sobre" 35). In effect, despite the support of a small constellation of poets and apologists, gongorism was widely seen as a pernicious literary idiom, a diabolic babel, and a "foreign" discourse associated with themes that undermined Christianity and, consequently, Spanish identity. As the Soledades circulated in manuscript form during the 1610s and 1620s, some peninsular writers took an active but limited interest in imitating and reworking the poem's language and themes.5 In the late 1620s and 1630s, after Góngora's death, linguistic and thematic features of gongorism entered mainstream Spanish writing: Quevedo, Lope de Vega, and other writers first cultivated gongorism in order to mock it, but later wound up dabbling in it for many other literary uses. Other poets, such as the playwrights Tirso de Molina and Calderón de la Barca, co-opted gongorism and helped to make it popular and...
Question Time is a distinctive daily parliamentary routine. Its aim is to hold Ministers of the State accountable for the actions and decisions of the Government. However, in many Parliaments, including the New Zealand and Australian Federal Houses of Representatives, it is more of a theatrical performance where parties try their best to score political points. As any performance, Question Time is governed by certain rules and regulations outlined in an official document Standing Orders. As there is not much action, Standing Orders mainly describe language norms and specify „unparliamentary language‟. This research looks at and analyses the use of formulaic vocabulary used by MPs in the year preceding general elections in New Zealand and Australia. The formulaic language includes phrasal lexical items and formulae for asking / answering questions, for raising points of order and the Speakers‟ idiolectal phrasal vocabulary for quelling disorder in the Chambers and regulating the work of the House. The framework developed for this research consisted of the following steps: an ethnographic study of Question Time as a communicative performance which included the development of a database containing all the empirical material; a xii linguistic study of Question Time including genrelect study, parliamentary formulae study and disorder analysis before the elections. As a result this research has shown that Question Time is a communicative performance event in New Zealand and Australia with significant cultural, historic and linguistic differences in spite of the common origins of the two Parliaments. It has identified 60 Question Time genre-specific phrasal lexical items that MPs use in the two Parliaments, studied their structure and meaning (where necessary). It has also looked at the strategies the MPs employ for creating disorder in the House, and the ways of quelling disorder by the Speakers of the two Parliaments.
Opinion mining on conversational telephone speech tackles two challenges: the robustness of speech transcriptions and the relevance of opinion models. The two challenges are critical in an industrial context such as marketing. The paper addresses jointly these two issues by analyzing the influence of speech transcription errors on the detection of opinions and business concepts. We present both modules: the speech transcription system, which consists in a successful adaptation of a conversational speech transcription system to call-centre data and the information extraction module, which is based on a semantic modeling of business concepts, opinions and sentiments with complex linguistic rules. Three models of opinions are implemented based on the discourse theory, the appraisal theory and the marketers’ expertise, respectively. The influence of speech recognition errors on the information extraction module is evaluated by comparing its outputs on manual versus automatic transcripts. The F-scores obtained are 0.79 for business concepts detection, 0.74 for opinion detection and 0.67 for the extraction of relations between opinions and their target. This result and the in-depth analysis of the errors show the feasibility of opinion detection based on complex rules on call-centre transcripts.
This paper presents GATE Teamware—an open-source, web-based, collaborative text annotation framework. It enables users to carry out complex corpus annotation projects, involving distributed annotator teams. Different user roles are provided (annotator, manager, administrator) with customisable user interface functionalities, in order to support the complex workflows and user interactions that occur in corpus annotation projects. Documents may be pre-processed automatically, so that human annotators can begin with text that has already been pre-annotated and thus making them more efficient. The user interface is simple to learn, aimed at non-experts, and runs in an ordinary web browser, without need of additional software installation. GATE Teamware has been evaluated through the creation of several gold standard corpora and internal projects, as well as through external evaluation in commercial and EU text annotation projects. It is available as on-demand service on GateCloud.net, as well as open-source for self-installation.
The present study extended existing research on alexithymia in men, investigating whether the deficit in processing emotions occurs early in the process, as a result of dissociation or repression, or later, as a result of suppression. We also examined the assumption in Levant’s (2011) normative male alexithymia hypothesis that men with alexithymia would show the greatest deficits in identifying words for emotions discouraged by masculine norms that expressed vulnerability and attachment. Study 1, with 258 college men, showed that scores on measures of alexithymia and normative male alexithymia were more strongly and uniquely predicted by suppression than repression and dissociation, while controlling for positive and negative affect and depression. Study 2 used semantic priming with 85 college men, and revealed that men with alexithymia showed more errors in lexical decision performance using target emotion words discouraged by masculine norms as compared to men without alexithymia. In addition, men with and without alexithymia did not differ in their accuracy using target emotion words that are encouraged by masculine norms. We also found that the disruption in emotional processing among men with alexithymia occurred at 500 ms stimulus onset asynchrony, which is slow enough for conscious processing, supporting an explanation of suppression as the mechanism for the inhibition.
Patrick Hanks’ Theory of Norms and Exploitations (henceforth TNE) is a corpus-driven lexicocentric theory of language which helps us understand how words go together and how people use words to make meanings. It focuses on the identification of normal, central and stereotypical usage and introduces criteria for distinguishing between normal patterns of collocations, these typical phraseological patterns being seen as the main carriers of meaning, and creative uses of those patterns (when people start exploiting the rule-governed norms). ‘Lexical Analysis: Norms and Exploitations’ is a fascinating detailed account of that lexicocentric model, in which Hanks brilliantly revisits the entire field of lexical semantics to come up with a lexically-based, corpus-driven, bottom-up theory of language. Since meanings are associated with words in context, TNE uses the notion of lexical set and semantic types. For instance, a group of words such as gun, pistol, revolver, rifle constitutes a lexical set in relation to the verb fire, which is united by a common semantic type (firearms). Such a lexical set may be used to perform word sense disambiguation and activate a contrast with other senses of the verb fire (as in [[Human]fire[Human]], meaning ‘to dismiss someone from employment’).
OBJECTIVE: Lexical fluency tests are frequently used to assess language and executive function in clinical practice. We investigated the influences of age, gender, and education on lexical verbal fluency in an educationally-diverse, elderly Korean population and provided its' normative information. METHODS: We administered the lexical verbal fluency test (LVFT) to 1676 community-dwelling, cognitively normal subjects aged 60 years or over. RESULTS: In a stepwise linear regression analysis, education (B=0.40, SE=0.02, standardized B=0.506) and age (B=-0.10, SE=0.01, standardized B=-0.15) had significant effects on LVFT scores (p<0.001), but gender did not (B=0.40, SE=0.02, standardized B=0.506, p>0.05). Education explained 28.5% of the total variance in LVFT scores, which was much larger than the variance explained by age (5.42%). Accordingly, we presented normative data of the LVFT stratified by age (60-69, 70-74, 75-79, and ≥80 years) and education (0-3, 4-6, 7-9, 10-12, and ≥13 years). CONCLUSION: The LVFT norms should provide clinically useful data for evaluating elderly people and help improve the interpretation of verbal fluency tasks and allow for greater diagnostic accuracy.
Unlike some varieties of English in Southeast Asia, the notion that there is a ‘Thai English’ is debatable. This paper examines distinctive non-native features of a lexicon found in contemporary Thai writing in English to ascertain if English in this Expanding Circle country is developing its own linguistic norms. An analysis of features of lexical creativity in five short stories and novels is carried out to determine whether the characteristics found indicate that a Thai English vocabulary exists. An ‘integrated framework’ which combines concepts in World Englishes by Braj B, Kachru, Peter Strevens, and Edgar W. Schneider is adopted in this study. It appears that certain categories of lexical creativity in the fiction examined represent five indicators of Thai English - contextualization, innovation, nativization, transcultural creativity, and localization - and reveals a developing non-native variety of English.
It is a truism that meaning depends on context. Corpus evidence now shows us that normal contexts can be summarised and indeed quantified, while the creative exploitations of normal contexts by ordinary language users far exceed anything dreamed up in speculative linguistic theory. Human linguistic behaviour is indeed rule-governed, but in recent years, corpus analysis (e.g. Hanks 2013) has shown that there is not just a single monolithic system of rules: instead, language use is governed by two interlinked systems: one set of rules governing normal, idiomatic uses of words and another set of rules governing how we exploit those norms creatively. Types of creative exploitation include (among others): • using anomalous arguments to make novel meanings • ellipsis for verbal economy in discourse • metaphors, metonymy, and other figurative uses for stylistic effect and other purposes Traditional dictionaries do a good job of listing the many possible meanings of words. But they do a poor job of reporting phraseology and an even worse job of associating different meanings with phraseological patterns. Moreover, all too often, they list a creative use that happens to have been noticed by a lexicographer as if it were a conventional norm, with resultant confusion, for example: • A riddle does not mean a hole made by a bullet (but OED says it does). • To newspaper does not mean to work as a journalist (but Merriam Webster says it does). The idiom principle formulated by the late John Sinclair (1991, 1998) argues that many meanings depend for their realization on the presence of more than one word. The Pattern Dictionary of English Verbs (PDEV;
Abstract Since Prince (1981) and Givón (1983), studies on discourse reference have explained the grammatical realization of referents in terms of general concepts such as “assumed familiarity” or “discourse coherence.” In this paper, we develop a complementary approach based on a detailed statistical tracking of subjects in Emirati Arabic, from which two major categories of subject expression emerge. On the one hand, null subjects are opposed to overt ones; on the other, subject-verb (SV) is opposed to verb-subject (VS). Although null subjects strongly correlate with coreferentiality with the subject of the previous clause, they can also index more distant referents within a single episode. With respect to SV vs. VS, morpholexical classes are found to be biased toward one or the other: nouns are typically VS, pronouns SV. We conclude that the null subject variant is the norm in Emirati Arabic, and when an overt subject is appropriate, lexical identity biases the subject into SV or VS order, generating word order as a discourse-relevant parameter. Overall, our approach attempts to understand Arabic discourse from a microlevel perspective.
=7) control groups. Following nine sessions combining computerized rapid accelerated-reading program (RAP), which individually tailors rate of written text presentation to comprehension criterion (80%), and self-regulated strategies for attending and engaging, the treated group significantly outperformed the wait-listed group before treatment on (a) a grade-normed, silent sentence reading rate task requiring lexical- and syntactic level processing to decide which of three sentences makes sense; and (b) RAP presentation rates yoked to comprehension accuracy level. Each group improved significantly on these same outcomes from before to after instruction. Attention ratings and working memory for written words predicted post-treatment accuracy, which correlated significantly with the silent sentence reading rate score. Implications are discussed for (a) preventing silent reading disabilities during the transition to increasing emphasis on silent reading, (b) evidence-based approaches for making accommodation of extra time on timed tests requiring silent reading, and
Do task demands change the way we extract information from a stimulus, or only how we use this information for decision making? In order to answer this question for visual word recognition, we used EEG/MEG as well as fMRI to determine the latency ranges and spatial areas in which brain activation to words is modulated by task demands. We presented letter strings in three tasks (lexical decision, semantic decision, silent reading), and measured combined EEG/MEG as well as fMRI responses in two separate experiments. EEG/MEG sensor statistics revealed the earliest reliable task effects at around 150 ms, which were localized, using minimum norm estimates (MNE), to left inferior temporal, right anterior temporal and left precentral gyri. Later task effects (250 and 480 ms) occurred in left middle and inferior temporal gyri. Our fMRI data showed task effects in left inferior frontal, posterior superior temporal and precentral cortices. Although there was some correspondence between fMRI and EEG/MEG localizations, discrepancies predominated. We suggest that fMRI may be less sensitive to the early short-lived processes revealed in our EEG/MEG data. Our results indicate that task-specific processes start to penetrate word recognition already at 150 ms, suggesting that early word processing is flexible and intertwined with decision making.
The importance of vocabulary in reading comprehension emphasizes the need to accurately assess an individual's familiarity with words. The present article highlights problems with using occurrence counts in corpora as an index of word familiarity, especially when studying individuals varying in reading experience. We demonstrate via computational simulations and norming studies that corpus-based word frequencies systematically overestimate strengths of word representations, especially in the low-frequency range and in smaller-size vocabularies. Experience-driven differences in word familiarity prove to be faithfully captured by the subjective frequency ratings collected from responders at different experience levels. When matched on those levels, this lexical measure explains more variance than corpus-based frequencies in eye-movement and lexical decision latencies to English words, attested in populations with varied reading experience and skill. Furthermore, the use of subjective frequencies removes the widely reported (corpus) Frequency × Skill interaction, showing that more skilled readers are equally faster in processing any word than the less skilled readers, not disproportionally faster in processing lower frequency words. This finding challenges the view that the more skilled an individual is in generic mechanisms of word processing, the less reliant he or she will be on the actual lexical characteristics of that word.
Despite the assumption in early studies that children are monostylistic until sometime around adolescence, a number of studies since then have demonstrated that adult-like patterns of variation may be acquired much earlier. How \nmuch earlier, however, is still subject to some debate. In this paper we contribute to this research through an analysis of a number of lexical, phonological and \nmorphosyntactic variables across 29 caregiver/child pairs aged 2;10 to 4;2 in interaction with their primary caregivers. We first establish the patterns of use – both \nlinguistic and social – in caregiver speech and then investigate whether these patterns of use are evident in the child speech. Our findings show that the acquisition \nof variation is highly variable dependent: some show age differentiation, others do not; some show acquisition of style shifting, others do not; some show correlations between caregiver input and child output, others do not. We interpret these findings in the light of community norms, social recognition and sociolinguistic value in the acquisition of variation at these early stages.
This chapter addresses the role played by language and schools in the history of Spain’s nineteenth-century liberal nation-building project. Both the Spanish language and the public school system were strategic sites where national consensus could be built and, consequently, the achievement of linguistic homogeneity through education became a central goal for the state. In particular, I examine the conditions that favored the linguistic norms developed by the Royal Spanish Academy and the debates that surrounded their officialization and imposition in the emerging national school system. While the historiography of Spanish has traditionally described the selection and implementation of the RAE’s norms as if they were undisputed and ideologically neutral, this study will emphasize the political complexity of the standardization process by approaching the archive with an ethnographic and historical-materialist perspective.
This study addresses the social and linguistic constraints on relativizer omission in restrictive relative clauses in a mainstream urban variety of Canadian English. Drawing on the framework of variationist sociolinguistics, the authors test an array of factors that have been traditionally implicated in the choice of relative marker (e.g., syntactic function of the relative marker; animacy and definiteness of the antecedent head NP; length of the relative clause), in addition to investigating less widely researched factors, such as the informational content of the matrix clause and the lexical specificity of the head NP. A variable rule analysis of nonsubject relative clauses extracted from 19 speakers stratified by age, sex, and education reveals that relativizer omission is socially sensitive and that properties of the matrix clause and adjacency effects are key determinants in the selection of the zero variant. Recurrent structural configurations exhibited by zero marked relative clauses in vernacular discourse are indicative of grammaticalization. Comparison with other varieties of English reveals that relativizer omission fails to pattern uniformly, suggesting that there is no vernacular norm in this area of the grammar. This absence of uniformity calls into question recent attempts by researchers to formulate a unitary account of relativizer omission by appealing to putatively general language processing constraints.
Word Sense Induction (WSI) is the task of identifying the different uses (senses) of a target word in a given text in an unsupervised manner, i.e. without relying on any external resources such as dictionaries or sense-tagged data. This paper presents a thorough description of the SemEval-2010 WSI task and a new evaluation setting for sense induction methods. Our contributions are two-fold: firstly, we provide a detailed analysis of the Semeval-2010 WSI task evaluation results and identify the shortcomings of current evaluation measures. Secondly, we present a new evaluation setting by assessing participating systems’ performance according to the skewness of target words’ distribution of senses showing that there are methods able to perform well above the Most Frequent Sense (MFS) baseline in highly skewed distributions.
espanolPresentamos un sistema de normalizacion de tweets en espanol, que usa reglas de preproceso, un modelo de distancias de edicion adecuado al dominio y modelos de lengua para seleccionar candidatos de correccion segun el contexto. El sistema obtuvo resultados superiores a la media en la tarea Tweet-Norm de SEPLN 2013. EnglishWe present a system to normalize Spanish tweets, which uses preprocessing rules, a domain-appropriate edit-distance model, and language models to select correction candidates based on context. The system’s results at SEPLN 2013 Tweet-Norm task were above-average.
Patrick Hanks sees linguistic approaches to word meaning as divided between two unattractive extremes. Generative theories, such as were pioneered by Katz and Fodor (1963) and pursued recently e.g. by Wierzbicka (1996), attempt to capture meanings with an apparatus of quasi-mathematical rules and universal semantic primitives which is unequal to reflecting the messy realities revealed by empirical corpus studies. On the other hand, the doctrine of linguistic creativity advanced by Sampson (1980, 2001) is unduly defeatist in denying the possibility of scientific analysis. Hanks argues that theoretical linguistics and practical lexicography should both embrace an intermediate position which distinguishes between high-frequency “norms” of usage and rare “exploitations”. This allows linguists and lexicographers to produce scientific lexical description while nevertheless acknowledging messy variability.
The article deals with the functional treatment of the concept of number in the English noun system and looks into semantic features of substantive units, their lexical and grammatical collocation and ability to express the linguacultural peculiarities of communication. The authors focus on extralinguistic parametres aimed at better understanding the semantics of the unit and construing more effectively one’s own statement in accordance with the aim set and keeping within the varieties of the Modern English norm.
The paper presents our work on the annotation of intra-chunk dependencies on an English treebank that was previously annotated with Inter-chunk dependencies, and for which there exists a fully expanded parallel Hindi dependency treebank. This provides fully parsed dependency trees for the English treebank. We also report an analysis of the inter-annotator agreement for this chunk expansion task. Further, these fully expanded parallel Hindi and English treebanks were word aligned and an analysis for the task has been given. Issues related to intra-chunk expansion and alignment for the language pair HindiEnglish are discussed and guidelines for these tasks have been prepared and released.
Human non-verbal vocal bursts are evolutionary \nconservative emotional expressions. Humans can \neasily assess inner states of conspecifics based on \nthese calls. Moreover, they can attribute emotions \nto non-human animal vocalizations too. However, \nwhether the same acoustic cues are used to assess \nemotional content in conspecific and nonconspecific \nvocalizations is not clarified yet. \nTo test this, we compiled a pool of 100-100 various \ndog and human non-verbal vocalizations from \ndiverse social contexts, and designed an online \nsurvey, in which every sample could be rated along \nemotional valence and intensity. We also measured \nwithin each sample the average length of calls, the \nfundamental frequency and the harmonics-to-noise \nratio. \nWhile valence ratings did not differ across species, \nhuman vocalizations were less intense. Linear \nregressions revealed that both shorter dog and \nhuman calls were rated as more positive. In \ncontrast, subjects scored higher pitched human and \ndog sounds to be more intense. We also found dog \nvocalizations with shorter call length or with higher \nHNR were rated less intense. \nIn conclusion, acoustical parameters affected \nhumans’ emotional ratings independently from the \nsource species of these vocalizations. These findings \nsuggest that humans utilize the same mental \nmechanisms for recognizing conspecific and \nheterospecific vocal emotions.
This study investigates whether age and/or hearing loss influence the perception of the emotion dimensions arousal (calm vs. aroused) and valence (positive vs. negative attitude) in conversational speech fragments. Specifically, this study focuses on the relationship between participants' ratings of affective speech and acoustic parameters known to be associated with arousal and valence (mean F0, intensity, and articulation rate). Ten normal-hearing younger and ten older adults with varying hearing loss were tested on two rating tasks. Stimuli consisted of short sentences taken from a corpus of conversational affective speech. In both rating tasks, participants estimated the value of the emotion dimension at hand using a 5-point scale. For arousal, higher intensity was generally associated with higher arousal in both age groups. Compared to younger participants, older participants rated the utterances as less aroused, and showed a smaller effect of intensity on their arousal ratings. For valence, higher mean F0 was associated with more negative ratings in both age groups. Generally, age group differences in rating affective utterances may not relate to age group differences in hearing loss, but rather to other differences between the age groups, as older participants' rating patterns were not associated with their individual hearing loss.
Please note: This article is in Greek. Compiling a dialectal dictionary: the “Syntychies” lexical database: This paper introduces\nthe reader to the issues of making an online dialectal dictionary, presenting some\nof the matters that have arisen while producing a lexical database of the Cypriot Greek\ndialect. Most problems related to the selection of data and compilation of lemmas were\ncaused by the great variation in orthographic and/or morphological representation of\nCypriot word forms. The database which has been created as part of the “Syntychies”\nresearch program is available on the website http://lexcy.library.ucy.ac.cy. The choices\nthat have been adopted in this database after a lexical analysis of a large amount of data\noutline a framework for compiling other dialectical dictionaries of Greek.
Methods are proposed for measuring affective valence and arousal in speech. The methods apply support vector regression to prosodic and text features to predict human valence and arousal ratings of three stimulus types: speech, delexicalized speech, and text transcripts. Text features are extracted from transcripts via a lookup table listing per-word valence and arousal values and computing per-utterance statistics from the per-word values. Prediction of arousal ratings of delexicalized speech and of speech from prosodic features was successful, with accuracy levels not far from limits set by the reliability of the human ratings. Prediction of valence for these stimulus types as well as prediction of both dimensions for text stimuli proved more difficult, even though the corresponding human ratings were as reliable. Text based features did add, however, to the accuracy of prediction of valence for speech stimuli. We conclude that arousal of speech can be measured reliably, but not valence, and that improving the latter requires better lexical features.
This paper presents unpublished materials of Ivan Pankevitch’s dictionary of Southern Carpathian dialects that are now handled in the form of an electronic lexical database in the Slavonic Institute of the Academy of Sciences of the Czech Republic. Analysis of the selected materials showed representation of individual sources in the lexical database, allowed a preliminary determination of the literary sources and in the case of direct field records provided an opportunity to specify their geographical distribution.
The fundamental goal of this dissertation is to establish that deep, efficient, accurate parsing models can be acquired for Chinese, through parsers founded on Combinatory Categorial Grammar (CCG), a grammar formalism which has already enabled the creation of rich parsing models for English. We harness these CCG analyses of cross-linguistic syntax, harmonising them with modern accounts from Chinese generative syntax, contributing the first analysis of Chinese syntax through CCG in the literature. Supervised statistical parsing approaches rely on the availability of large annotated corpora. To avoid the cost of manual annotation, we adopt the corpus conversion methodology, in which an automatic corpus conversion algorithm projects annotations from a source corpus into the target formalism. The central contribution of this thesis is Chinese CCGbank, a corpus of 750,000 words automatically extracted from the Penn Chinese Treebank, reifying the abstract analysis through corpus conversion. We then take three state-of-the-art CCG parsers from the literature — the split-merge PCFG parser of Petrov and Klein, the transition-based CCG parser of Zhang et al., and the maximum entropy parser of Clark and Curran — and train and evaluate all three on Chinese CCGbank, achieving the first Chinese CCG parsing models in the literature. We demonstrate that while the three parsers are only separated by a small margin trained on English CCGbank, a substantial gulf of 4.8% separates the same parsers trained on Chinese CCGbank. We also confirm that the gap between the states-of-the-art in English and Chinese PSG parsing can be observed in CCG parsing. Our parsing experiments establish Chinese CCG parsing as a new and substantial challenge, a line of empirical investigation directly enabled by Chinese CCGbank.
African Languages WordNet is an ongoing project which is based on the English WordNet.WordNet is an electronic lexical database that groups words in synonym sets (synsets).In this project, words are translated from English to African Languages.As such, this paper comments on the semantic aspects of verbs in one of the African Languages, namely Northern Sotho, that are translated from English.Verbs are generally understood as expressions of action or state of being.These two languages are typologically dissimilar, with different cultural-historical backgrounds.Northern Sotho is a Bantu 1 language, agglutinating with extensive and productive use of affixes while English is not.In the first place, structural differences between these two languages pose equivalence challenges, both linguistically and computationally.Secondly, verb equivalents may not be affected by the same collocational restrictions in the source and target languages.Another issue is that a verb in one language may invoke certain connotations, which may not apply to its equivalent in another language.Finally, some concepts may be foreign and others culture-specific to one language and not the other, thus resulting in omission of some target language concepts.Attention to these equivalence challenges may enhance technological development of the target language lexicon.
Thresholded two-tone ("Mooney") images are of interest for vision science because the object hidden within the image can be hard to recognize, with recognition times in the second to minute range. However, once a subject has seen the original grayscale image from which the Mooney is generated, recognition is much accelerated. Typically, suitable "Mooney" images need to be painstakingly generated by hand. Here, we present an approach for automatically generating a two-tone image database. This is based on large number of images collected from the internet. We first selected concrete words from a linguistic database. Using these words as search words, we automatically downloaded images from an online image database (www.flickr.com). Subsequently, the images were preprocessed and thresholded using a histogram based thresholding algorithm to generate the two-tone images. We provide an image set with 330 Mooney images and psychophysical results obtained from six subjects. With a presentation time of 20 s, the average recognition time was 9.36 s ± 7.40 s. Additionally, subjective ratings (confidence, Aha, and difficulty ratings) were obtained and are presented for each subject and image. This image set is, to our knowledge, the largest two-tone image set available to the vision and cognitive science research community (https://sites.google.com/site/hayneslab/links). We provide a Matlab toolbox that makes the extension of the image database possible. Using this toolbox, the researcher can add new object names as search words and create new two-tone images easily. Furthermore, we will present possibilities to extent this toolbox using another image database called ImageNet and introduce the use of Amazon Mechanical Turk to select useful images for a particular experiment. This image database can be useful for studying the mechanisms of rapid learning of visual image recognition, and has applications in research on conscious vision, learning, priming, reward and insight. Meeting abstract presented at VSS 2013
We have introduced here a new type of corpus annotation which we call Etymological Annotation (EA). We propose this new type because although, over the years, scientists have proposed corpus annotation of various types (Atkins, Clear and Ostler 1992, Biber 1993, Leech 2005), nobody has ever suggested that words included within corpora need to be annotated at their etymological level so that one can retrieve necessary linguistic information relating to antiquity of words and terms used in corpora. The applicational relevance of etymologically annotated corpora may be visualized in language description, language planning, language education, lexicology, language technology as well as in compilation of general, historical, learner and special dictionaries. In case of those languages, where one comes across large number of words borrowed from neighbouring and foreign languages, the proper identification of source of origin of words carries tremendous referential relevance in cross-lingual lexical database generation, morphological processing, part-of-speech tagging, e-learning, digital lexical profile generation, information retrieval, machine learning, and language documentation. Thus, etymologically annotated corpora become an essential resource of applied linguistics and language technology. We propose here to define this new event with necessary direction and guidance to develop etymologically tagged language corpora for all natural languages.
This experiment investigated whether affective information from unfamiliar people can influence the affective ratings for unfamiliar foods, before and after participants have ate the foods. The participants rated a food product’s appearance in terms of how palatable it looked on a 7-point Likert scale before they had eaten it (affective expectation). They also rated how palatable the food actually was after they had eaten it (affective evaluation). Results showed that there was a significant interaction between when participants provided the ratings and whether they had been informed of other people’s affective evaluation. Exposure to affective information did not influence affective expectation, but it did increase affective evaluation. These results suggest that affective information that is presented simultaneously with a visual experience might only be influential when perceivers are able to test the information through their perceptual experience.
Much of the research exploring the relationship between taste quality and affective state suggests that sweet-tasting foods are associated with pleasant feelings, and sour- and spicy-tasting foods are associated with unpleasant feelings. The findings of arousal response as a component of overall affective state are less clear with respect to taste quality. The present study investigated the relationship between taste quality and affective state by comparing arousal and pleasantness ratings of neutral images from the International Affective Picture System (IAPS). Participants (N = 55) recorded these ratings during consumption of sprays which varied in taste quality (sweet, sour, or spicy). As hypothesized, sweet sprays elicited significantly higher ratings of pleasantness than sour or spicy sprays ( η 2 p =.14) on the neutral images. However, arousal ratings did not differ among the three taste quality conditions. Implications of the findings in a broader framework and suggestions for future research are discussed.
Prior research has suggested that configural resemblance between a current scene and a previously experienced but forgotten one may trigger déjà vu experiences. The present study examined whether there is a relationship between the frequency of actual déjé vu experiences, measured by questionnaires, and sensitivity to a configural resemblance between past and present events, measured by questionnaires, and between two scenes presented simultaneously in the laboratory. We measured familiarity ratings and remember–know judgements of several scenes. Some scenes had been previously presented, some were similar to previously presented scenes and the others were dissimilar. Déjà vu tendencies were significantly correlated with sensitivity to similarity in the measured questionnaires and in the laboratory, as well as to a feeling of familiarity for similar scenes. In this study, we found for the first time that people who more frequently experience déjé vu states were also more likely to regard themselves as sensitive to similarity and more likely to notice the similarity between two scenes in the laboratory.
AIM: A group-based multisensory activity program (Sensory Day) for residents with dementia was developed, to address the challenge of providing personalised activities within tight operational constraints in residential aged care facilities. METHOD: Fourteen participants with severe and very severe dementia were observed before, during and after participation in one of four Sensory Day sessions. The Menorah Park Rating Scale was used to yield four levels of engagement. The Philadelphia Geriatric Affect Rating Scale was used to identify four affect states. Dementia severity was ascertained by PAS-CIS scores mapped onto the Global Deterioration Scale. RESULTS: Increased levels of constructive engagement and positive affect were observed during participation in the Sensory Day sessions, relative to measures taken before the session. CONCLUSIONS: This novel approach to activity programming demonstrates that it is possible to provide group-based activities for residents with severe and very severe dementia which result in increased engagement and positive mood.
While working on valency lexicons for Czech and English, it was necessary to define treatment of multiword entities (MWEs) with the verb as the central lexical unit. Morphological, syntactic and semantic properties of such MWEs had to be formally specified in order to create lexicon entries and use them in treebank annotation. Such a formal specification has also been used for automated quality control of the annotation vs. the lexicon entries. We present a corpus-based study, concentrating on multilayer specification of verbal MWEs, their properties in Czech and English, and a comparison between the two languages using the parallel Czech-English Dependency Treebank (PCEDT). This comparison revealed interesting differences in the use of verbal MWEs in translation (discovering that such MWEs are actually rarely translated as MWEs, at least between Czech and English) as well as some inconsistencies in their annotation. Adding MWE-based checks should thus result in better quality control of future treebank/lexicon annotation. Since Czech and English are typologically different languages, we believe that our findings will also contribute to a better understanding of verbal MWEs and possibly their more unified treatment across languages. This work has been supported by the Grant No.
We deal with syntactic identification of occurrences of multiword expression (MWE) from an existing dictionary in a text corpus. The MWEs we identify can be of arbitrary length and can be interrupted in the surface sentence. We analyse and compare three approaches based on linguistic analysis at a varying level, ranging from surface word order to deep syntax. The evaluation is conducted using two corpora: the Prague Dependency Treebank and Czech National Corpus. We use the dictionary of multiword expressions SemLex, that was compiled by annotating the Prague Dependency Treebank and includes deep syntactic dependency trees of all MWEs. 1
Stanford Dependencies (SD) provide a functional characterization of the gram-matical relations in syntactic parse-trees. The SD representation is useful for parser evaluation, for downstream applications, and, ultimately, for natural language un-derstanding, however, the design of SD fo-cuses on structurally-marked relations and under-represents morphosyntactic realiza-tion patterns observed in Morphologically Rich Languages (MRLs). We present a novel extension of SD, called Unified-SD (U-SD), which unifies the annotation of structurally- and morphologically-marked relations via an inheritance hierarchy. We create a new resource composed of U-SD-annotated constituency and dependency treebanks for the MRL Modern Hebrew, and present two systems that can automat-ically predict U-SD annotations, for gold segmented input as well as raw texts, with high baseline accuracy. 1
Even though the quality of unsupervised dependency parsers grows, they often fail in recognition of very basic dependencies. In this paper, we exploit a prior knowledge of STOP-probabilities (whether a given word has any children in a given direction), which is obtained from a large raw corpus using the reducibility principle. By incorporating this knowledge into Dependency Model with Valence, we managed to considerably outperform the state-of-theart results in terms of average attachment score over 20 treebanks from CoNLL 2006 and 2007 shared tasks.
Emotions in music are conveyed by a variety of acoustic cues. Notably, the positive association between sound intensity and arousal has particular biological relevance. However, although amplitude normalization is a common procedure used to control for intensity in music psychology research, direct comparisons between emotional ratings of original and amplitude-normalized musical excerpts are lacking. In this study, 30 nonmusicians retrospectively rated the subjective arousal and pleasantness induced by 84 six-second classical music excerpts, and an additional 30 nonmusicians rated the same excerpts normalized for amplitude. Following the cue-redundancy and Brunswik lens models of acoustic communication, we hypothesized that arousal and pleasantness ratings would be similar for both versions of the excerpts, and that arousal could be predicted effectively by other acoustic cues besides intensity. Although the difference in mean arousal and pleasantness ratings between original and amplitude-normalized excerpts correlated significantly with the amplitude adjustment, ratings for both sets of excerpts were highly correlated and shared a similar range of values, thus validating the use of amplitude normalization in music emotion research. Two acoustic parameters, spectral flux and spectral entropy, accounted for 65% of the variance in arousal ratings for both sets, indicating that spectral features can effectively predict arousal. Additionally, we confirmed that amplitude-normalized excerpts were adequately matched for loudness. Overall, the results corroborate our hypotheses and support the cue-redundancy and Brunswik lens models.
Please note: This article is in Greek. Compiling a dialectal dictionary: the “Syntychies” lexical database: This paper introduces the reader to the issues of making an online dialectal dictionary, presenting some of the matters that have arisen while producing a lexical database of the Cypriot Greek dialect. Most problems related to the selection of data and compilation of lemmas were caused by the great variation in orthographic and/or morphological representation of Cypriot word forms. The database which has been created as part of the “Syntychies” research program is available on the website http://lexcy.library.ucy.ac.cy. The choices that have been adopted in this database after a lexical analysis of a large amount of data outline a framework for compiling other dialectical dictionaries of Greek.
The Spanish norms for pictures in shows 15 to 20 of the International Affective Picture System (IAPS) are reported in this paper. Participants were 811 undergraduate university students (521 women), who rated the valence, arousal, and dominance of 358 pictures. The correlations between the North-American and the Spanish ratings were all highly significant and, like in the United States, the picture distribution in the bidimensional affective space, defined by the ratings on affective valence and arousal, displayed the typical boomerang shape. Our data also corroborated gender differences in aversive pictures found in North-Americans. These results are fully consistent with those obtained in the first and second part of the Spanish adaptation, and demonstrate that the standardization of IAPS in our country has been successful. Finally, our data confirmed the cross-cultural differences found in arousal and dominance ratings: Spanish participants tended to assign higher arousal and lower dominance scores to the pictures, as a whole, than North-Americans. These data support the general cultural stereotypes that exist for these countries and suggest that the IAPS might be a reliably index of cultural differences in emotional disposition.
A process for the design and manufacture of 3D tactile textures with predefined affective properties was developed. Twenty four tactile textures were manufactured. Texture measures from the domain of machine vision were used to characterize the digital representations of the tactile textures. To obtain affective ratings, the textures were touched, unseen, by 107 participants who scored them against natural, warm, elegant, rough, simple, and like, on a semantic differential scale. The texture measures were correlated with the participants' affective ratings using a novel feature subset evaluation method and a partial least squares genetic algorithm. Six measures were identified that are significantly correlated with human responses and are unlikely to have occurred by chance. Regression equations were used to select 48 new tactile textures that had been synthesized using mixing algorithms and which were likely to score highly against the six adjectives when touched by participants. The new textures were manufactured and rated by participants. It was found that the regression equations gave excellent predictive ability. The principal contribution of the work is the demonstration of a process, using machine vision methods and rapid prototyping, which can be used to make new tactile textures with predefined affective properties.
OBJECTIVES: This study investigates an overall autonomic hypoactivity reflecting hypoarousal as important aetiological factor in ADHD at baseline during rest and in response towards stimuli. In addition, effects of methylphenidate (MPH) are examined. We further assessed whether this hypoarousal is a stable characteristic or ameliorated by arousing emotional stimuli. METHODS: Boys with ADHD were examined with (n = 35) or without MPH (n = 45) and compared with healthy boys (n = 22) regarding skin conductance level (SCL) during rest and skin conductance responses (SCRs) as well as valence and arousal ratings in response to positive, neutral, and negative pictures. RESULTS: ADHD children without MPH were characterized by reduced baseline SCL and overall reduced SCRs. ADHD children with MPH never differed from control children. All groups displayed normal valence and arousal ratings of the stimuli and enhanced SCRs to emotional in comparison to neutral pictures. CONCLUSIONS: This is the first study to unravel (1) a general autonomic hypoactivity in ADHD children at baseline and in response to low arousing neutral and highly arousing emotional stimuli, and (2) hints that MPH normalizes this hypoactivity. Results contribute to the understanding of ADHD aetiology and MPH functionality, and are consistent with the cognitive-energetic model of ADHD.