Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Using statistical approaches beside the traditional methods of natural language processing could significantly improve both the quality and performance of several natural language processing (NLP) tasks. The effective usage of these approaches is subject to the availability of the informative, accurate and detailed corpora on which the learners are trained. This article introduces a bootstrapping method for developing annotated corpora based on a complex and rich linguistically motivated elementary structure called supertag. To this end, a hybrid method for supertagging is proposed that combines both of the generative and discriminative methods of supertagging. The method was applied on a subset of Wall Street Journal (WSJ) in order to annotate its sentences with a set of linguistically motivated elementary structures of the English XTAG grammar that is using a lexicalised tree-adjoining grammar formalism. The empirical results confirm that the bootstrapping method provides a satisfactory way for annotating the English sentences with the mentioned structures. The experiments show that the method could automatically annotate about 20% of WSJ with the accuracy of F-measure about 80% of which is particularly 12% higher than the F-measure of the XTAG Treebank automatically generated from the approach proposed by Basirat and Faili [(2013). Bridge the gap between statistical and hand-crafted grammars. Computer Speech and Language, 27, 1085–1104].
Dutch is well-known for its verb clusters, i.e. constructions in which multiple verbs group together. This dissertation presents the most influential analyses of verb clusters in descriptive and generative syntax. It discusses phenomena that are typically related to cluster formation, such as the occurrence of an infinitive where one expects a past participle (i.e. Infinitivus Pro Participio or the IPP effect), word order variation, and the interruption of clusters by non-verbal material. Furthermore, this dissertation investigates how a corpus-based study can shed new light on the current syntactic theories with respect to cluster formation. For the corpus study, syntactically annotated corpora or treebanks are used, since they allow for the empirical investigation of Dutch syntax beyond the lexical level. The observations from the treebanks with regard to the set of clustering verbs, the word order variation in verb clusters, and the instances of cluster interruption are compared to the literature. Special attention goes out to constructions containing te-infinitives, as it is not always trivial to decide whether they are part of the verb cluster or not. Based on the results of the corpus study, a novel analysis of verb clusters is proposed in the framework of Head-driven Phrase Structure Grammar (HPSG). It is demonstrated that this analysis deals more adequately with verb clusters than previous HPSG approaches. An important consequence of the new analysis is that it not only deals with genuine verb clusters, but also accounts for ambiguous constructions. In addition, it extends to the analysis of other phenomena, such as adposition stranding.
This paper describes the strategies devised in order to convert the DiCoInfo, Dictionnaire fondamental de l’informatique et de l’Internet, a specialized lexical database, into a learners’ dictionary. Our main goal is to obtain a user-oriented dictionary (i.e. that meets specific user needs). Firstly, we defined the types of users towards which our dictionary is targeted: translation students are our first intended users. Then we determined the use situations and the functions of our dictionary: it should provide assistance in communicative and cognitive situations (Tarp, 2008). We made several changes to adapt the data categories of the DiCoInfo to these functions and user needs. In addition, we simplified the presentation: layout, display of data categories, access to data and addition of multimedia. In this user-oriented version, the data is presented in such a way that users who do not have a background in linguistics can easily interpret the contents of the data categories. Finally, different technologies were integrated in the process and hopefully contribute to make the new version even more accessible.
This article compiles a list of lemmas of the second class weak verbs of Old English by using the latest version of the lexical database Nerthus, which incorporates the texts of the Dictionary of Old English Corpus. Out of all the inflecional endings, the most distinctive have been selected for lemmatization: the infinitive, the inflected infinitive, the present participle, the past participle, the second person present indicative singular, the present indicative plural, the present subjunctive singular, the first and third person of preterite indicative singular, the second person of the preterite indicative singular, the preterite indicative plural and the preterite subjunctive plural. When it is necessary to regularize, normalization is restricted to correspondences based on dialectal and diachronic variation. The analysis turns out a total of 1,064 lemmas of weak verbs from the second class.
<p>The Tromsø Old Russian and OCS Treebank (TOROT, nestor.uit.no)1 is, along with its parent treebank, the PROIEL corpus (foni.uio.no), the only existing treebank of Old Church Slavonic (OCS), Old East Slavic and Middle Russian texts. There are other tagged resources, such as the Old Russian subcorpus of the Russian National Corpus2 and the Manuskript corpus,3 but none of them, to our knowledge, currently provide syntactic annotation. \n<p>The TOROT presently contains approximately 160,000 word tokens of fully annotated OCS (Codex Marianus4 and Codex Suprasliensis), 85,000 word tokens of fully annotated Kiev-era Old East Slavic, and 60,000 word tokens of fully annotated 15th–17th-century Middle Russian. In addition, it contains the Codex Zographensis with automatic and partially hand-corrected morphological annotation and lemmatisation (sections of the Gospels missing in the Codex Marianus also have full syntactic annotation), and the PROIEL version of the Greek Gospels, with which the Codex Marianus and the Codex Zographensis are both aligned at token level (automatically, then hand-corrected).
<h3>Introduction</h3><br> English News Text Treebank: Penn Treebank Revised was developed by the Linguistic Data Consortium (LDC) with funding through a gift from Google Inc. It consists of a combination of automated and manual revisions of the <a href="../../../LDC99T42">Penn Treebank</a> annotation of Wall Street Journal (WSJ) stories. The data is comprised of 1,203,648 word-level tokens in 49,191 sentence-level tokens -- in all 2,312 of the original Penn Treebank WSJ files. <br> <h3>Data</h3><br> This release includes revised tokenization, part-of-speech, and syntactic treebank annotation intended to bring the full WSJ treebank section into compliance with the agreed-upon policies and updates implemented for current English treebank annotation specifications at LDC. Examples include English Web Treebank (<a href="../../../LDC2012T13">LDC2012T13</a>), OntoNotes (<a href="../../../LDC2013T19">LDC2013T19</a>), and English translation treebanks such as English Translation Treebank: An-Nahar Newswire (<a href="../../../LDC2012T02">LDC2012T02</a>). English Treebank Supplemental Guidelines are included in this release. <br> <h3>Samples</h3><br> Please view this <a href="desc/addenda/LDC2015T13.tree.txt">treebank</a> and <a href="desc/addenda/LDC2015T13.txt">tokenized</a> samples. <br> <h3>Updates</h3><br> None at this time. </br> Portions © 1987-1989 Dow Jones & Company, Inc., © 1999, 2015 Trustees of the University of Pennsylvania
Cet article présente le repérage des connecteurs de discours dans le corpus « French Treebank » (FTB) déjà annoté pour la morpho-syntaxe. C’est la première étape dans l’annotation discursive complète de ce corpus. Il s’agit de projeter sur le corpus les éléments répertoriés dans LexConn, lexique des connecteurs du français, et de filtrer les occurrences de ces éléments qui n’ont pas un emploi discursif mais par exemple un emploi d’adverbe de manière ou de préposition introduisant un complément sous-catégorisé. Plus de 10 000 connecteurs ont été identifiés.
In this methodological article I discuss the advantages of using parsed corpora (treebanks) for historical linguistics research. I argue that even apparently simple, descriptive questions such as how often different word orders occur cannot be adequately answered without the use of treebanks, which make the underlying theoretical assumptions explicit. I demonstrate this concretely with a case study from word order in participle clauses in the New Testament.
Text-based sentiment analysis is a growing research field in affective computing, driven by both commercial applications and academic interest. Continuous dimensional representations, such as valence-arousal (VA) space, can represent the affective state more precisely than discrete effective representations. In building dimensional sentiment applications, affective lexicons with valence-arousal ratings are useful resources but are still very rare. Therefore, recent studies have investigated the automatic development of VA lexicons using linear regression techniques. One of the major limitations of linear regression is the under-fitting problem which can cause a poor fit between the algorithm and the training data. To tackle this problem, this study proposes the use of a locally weighted linear regression (LWLR) model to predict the valence-arousal ratings of affective words. The locally weighted method performs a regression around the point of interest using only training data that are "local" to that point, and thus can reduce the impact of noise from unrelated training data. Experimental results show that the proposed method achieved better performance for VA word prediction.
La nueva base de datos del Proyecto Nerthus, llamada The Grid, fue presentada por Martin Arista en una conferencia dictada en la Universidad de Sheffield en 2013. The Grid consiste en cinco tablas relacionadas: la base de datos lexica Nerthus, una concordancia por fragmentos, una concordancia por palabras, un indice y un indice inverso de The Dictionary of Old English Corpus. The Grid no esta basado en formas de diccionario, sino en atestiguaciones textuales. De todas las lineas de investigacion posibles que esta nueva organizacion de la base de datos ofrece, este trabajo se hace cargo de la lematizacion de las formas textuales. La razon es que un corpus morfologicamente anotado de ingles antiguo es una asignatura pendiente de esta disciplina. La informacion morfologica solo esta disponible para las letras A-G, las cuales ya han sido publicadas por el The Dictionary of Old English, pero no existe, o no es tan facil de encontrar en los diccionarios actuales para las letras H-Y. El proposito de la investigacion es proporcionar un inventario de lemas de verbos fuertes basados en la evidencia textual que viene proporcionada por el Dictionary of Old English Corpus. Respecto al Proyecto Nerthus, esta tesis intenta desarrollar un sistema de busquedas basado en sucesivas busquedas, de manera que las formas mas transparentes sean etiquetadas antes que otras formas mas opacas. La restriccion del ambito de analisis a los verbos fuertes se basa en dos razones. La primera es que el sistema de verbos fuertes en ingles antiguo desempena un papel central en la derivacion y el desarrollo del lexico. Por otra parte, los verbos fuertes, caracterizados por la apofonia, o ablaut, pueden ser buscados no solo por la terminacion flexiva, pero tambien por la vocal radical, lo que contribuye a refinar el sistema de busquedas. El punto de partida de esta investigacion es que la labor de lematizacion se puede hacer en parte automaticamente y en parte manualmente. La informacion contenida en la base de datos, junto con las funcionalidades de Filemaker, pueden maximizar la parte automatica del analisis y minimizar la revision manual. La metodologia incluye tres pasos principales: la recopilacion de un corpus de verbos fuertes que se adapte al analisis, la identificacion de las formas flexivas, y la definicion de codigos de busqueda automatica. La lista de verbos fuertes se ha tomado de la lista de referencia de verbos fuertes del Proyecto Nerthus, que se basa en las siete clases de Campbell (1987) y Hogg and Fulk (2011), y en las subclases de Krygier (1994). Para la identificacion de las formas flexivas relevantes, los verbos fuertes que no han sido derivados, han sido derivados en el infinitivo, presente de indicativo, preterito de indicativo, presente de subjuntivo, preterito de subjuntivo, e imperativo, todos ellos en singular y plural. Para las busquedas en la base de datos lexica, esta tesis propone un sistema de cuatro codigos de busqueda sucesivos que estan disenados especificamente para buscar determinadas formas verbales. Aparte del inventario de verbos fuertes, en las conclusiones se presentan resultados en dos areas. Primero, esta tesis puede responder de manera motivada la cuestion de los limites en la automatizacion en el analisis morfologico. En segundo lugar, que tesis arroja luz sobre la cuestion de la regularizacion de la ortografia caracteristica del trabajo lexicografico o normalizacion.
International audience
This research explores the cultural and linguistic strategies of immigrant youth to negotiate inclusion/exclusion, including language discrimination in Vancouver, Canada. My theoretical framework draws upon the Arendtian notions of ‘public space’, and ‘action and speech’ as well as Bourdieu’s concepts of ‘symbolic violence’ and ‘habitus’. My methodology is a critical qualitative approach. Fourteen immigrant youth, aged 15–25, were involved in this research. The findings of this study indicate that unlike second-generation immigrants, first-generation immigrant youth face cultural and linguistic challenges. Non-recognition of youths’ distinct linguistic and social capitals, the imposition of official languages and the regulation of the education and language market according to the dominant linguistic norms include forms of discrimination against Turkish minority youth in Canada. Taken together, the findings suggest that immigrant youths’ cultural and linguistic experiences of inclusion and exclusion cannot be dissociated from the wider politics of the nation-state, popular hegemony and social inequalities in the host society.
Agzul Nextar ad nesleḍ inaw aflutlay n yimusnawen n tesnalsit n tmetti, I waken ad nẓer d acu I d tamuɣli- nsen, amek ara tili tutlayt tamezdayt di tmaziɣt akked tagnut n tira. Azal n usentel –agi yella-d imi atas n ubeddel I -d yeḍran di tsertit tasnalsit (aẓayeṛ, amud).Tamukris nneɣ d tagi: D acu I d tulmisin n tutlayt tamezdayt ɣer yimusnawen n tesnalsit(ismazaɣen)? D acuten i yisekkilen akked tiriɣtan ixtaren i tira n tutlayt tamezdayt? Amek i- d bnan ismazaɣen tamuɣli-nsen? Ɣef wacu I reṣṣan iskukal-nsen? Abstarct We chose to study the position of socilinguists in relation to the written language in Berber standard, because today the Amazigh language, whose use was reserved for daily communication, appropriates other areas such the audio-visual fields, teaching, scientific and political discourse, etc. Given the variation that characterizes and technological advances which it faces, the choice of a could serve as a lingua is required. Standardization is therefore urgent that any language promotion begins with the promotion of a linguistic standard (Sauzet; 2002). Therefore, it should ask the following question: What (s) (s) language (s), graphic (s) and spelling (s) advocate for linguists Berber? What types of arguments they are using to justify their choice? Key concepts: Language planning, epilinguistique speech, coding language, linguistic norm, standardization
The scent of blood is potentially one of the most fundamental and survival-relevant olfactory cues in humans. This experiment tests the first human parameters of perceptual threshold and emotional ratings in men and women of an artificially simulated smell of fresh blood in contact with the skin. We hypothesize that this scent of blood, with its association with injury, danger, death, and nutrition will be a critical cue activating fundamental motivational systems relating to either predatory approach behavior or prey-like withdrawal behavior, or both. The results show that perceptual thresholds are unimodally distributed for both sexes, with women being more sensitive. Furthermore, both women and men's emotional responses to simulated blood scent divide strongly into positive and negative valence ratings, with negative ratings in women having a strong arousal component. For women, this split is related to the phase of their menstrual cycle and oral contraception (OC). Future research will investigate whether this split in both genders is context-dependent or trait-like.
Along with the increasing development of language resources - i.e., new lexicons, lexical databases, corpora, treebanks - the need for their efficient interlinking is growing. With such a linking, one can easily benefit from all their properties and information. Considering the convergence of resources, universal lexicographic formats are frequently discussed. In the present thesis, we investigate and analyse methods of interlinking language resources automatically. We introduce a system for interlinking lexicons (such as VALLEX, PDT-Vallex, FrameNet or SemLex) that offer information on syntactic properties of their entries. The system is automated and can be used repeatedly with newer versions of lexicons under development. We also design a method for identification of multiword expressions in a parsed text based on syntactic information from the SemLex lexicon. An output that verifies feasibility of the used methods is, among others, the mapping between the VALLEX and the PDT-Vallex lexicons, resulting in tens of thousands of annotated treebank sentences from the PDT and the PCEDT treebanks added into VALLEX. Powered by TCPDF (www.tcpdf.org)
Abstract Relations between perceiving and knowing are well-worn problems that become visceral encounters with doubt and ambiguity in ‘mixed-reality’ environments. Locative narrative situates participants within stories where existent places function as the setting. Experiential confusion, between what is talked of as real and as imagined, is an often-reported phenomenon. Classical pragmatisms, and more broadly the writings of William James, understand the functioning of the body to be for the production of action, from which flows a naturalistic epistemology. for James, a thought’s reference to an object occurs in the medium of an ‘experienceable environment’ and is a condition of it being known; what something is known-as is how it functions in a particular context and the consequences that follow. Contemporary pragmatists express varying positions on the function of representation in perception at different levels of cognitive awareness, and the extent to which intentionality is derivative of linguistic norms. In the locative narrative iOS application The Lost Index No.1 – Landscape with Figures strategies of directing participant attention, movement, cognitive tasks and propositional content are used to guide the interpretation of events. The complex environment that is created plays with the multi-stability of perception and the ‘multi-stability meaning’ between terms, resulting in ambiguity and an enhanced flexibility of interpretation.
&lt;p&gt;We define a dynamic oracle for the Covington non-projective dependency parser. This is not only the first dynamic oracle that supports arbitrary non-projectivity, but also considerably more efficient (O(n)) than the only existing oracle with restricted non-projectivity support. Experiments show that training with the dynamic oracle significantly improves parsing accuracy over the static oracle baseline on a wide range of treebanks.&lt;/p&gt;
Alexithymia is believed to involve deficits in emotion processing and imagery ability. Previous findings suggest that it is especially related to deficits in processing the arousal dimension of emotion, and that discordance may exist between self-report and physiological responses to emotional stimuli in alexithymia. The current study used a well-established emotional imagery paradigm to examine emotion processing deficits and discordance in participants (N = 86) selected based on their extreme scores on the Toronto Alexithymia Scale-20. Physiological (skin conductance, heart rate, and corrugator and zygomaticus electromyographic responses) and self-report (valence, arousal ratings) responses were monitored during imagery of anger, fear, joy, and neutral scenes and emotionally neutral high arousal (action) scenes. Results from regression analyses indicated that alexithymia was largely unrelated to responses on valence-based measures (facial electromyography, valence ratings), but that it was related to arousal-based measures. Specifically, alexithymia was related to higher heart rate during neutral and lower heart rate during fear imagery. Alexithymia did not predict differential responses to action versus neutral imagery, suggesting specificity of deficits to emotional contexts. Evidence for discordance between physiological responses and self-report in alexithymia was obtained from within-person analyses using multilevel modeling. Results are consistent with the idea that alexithymic deficits are specific to processing emotional arousal, and suggest difficulties with parasympathetic control and emotion regulation. Alexithymia is also associated with discordance between self-reported emotional experience and physiological response to emotion, consistent with prior evidence.
BACKGROUND: When humans observe other people's emotions they not only can relate but also experience similar affective states. This capability is seen as a precondition for helping and other prosocial behaviors. Our study aims to quantify the influence of help-related picture content on subjectively experienced affect. It also assesses the impact of different scales on the way people rate their emotional state. METHODS: The participants (N=242) of this study were shown stimuli with help-related content. In the first subset, half the drawings depicted a child or a bird needing help to reach a simple goal. The other drawings depicted situations where the goal was achieved. The second subset showed adults either actively helping a child or as passive bystanders. We created control conditions by including pictures of the adults on their own. Participants were asked to report their affective responses to the stimuli using two types of 9-point scales. For one half of the pictures, scales of arousal (calm to excited) and of bipolar valence (unhappy to happy) were employed; for the other half, unipolar scales of pleasantness and unpleasantness (strong to absent) were used. RESULTS: Even non-dramatic depictions of simple need-of-help situations were rated systematically lower in valence, higher in arousal, less pleasant and more unpleasant than corresponding pictures with the child or bird not needing help. The presence of a child and adult together increased pleasantness ratings compared to pictures in which they were depicted alone. Arousal was lower for pictures showing only an adult than for those including a child. Depictions of active helping were rated similarly to pictures showing a passive adult bystander, when the need-of-help was resolved. Aggregated unipolar pleasantness and unpleasantness ratings accounted well for arousal and even better for bipolar valence ratings and for content effects on them. CONCLUSION: This is the first study to report upon the meaningful impact of harmless need-of-help content on self-reported emotional experience. It provides the basis for further investigating the links between subjective emotional experience and active prosocial behavior. It also builds upon recent findings on the correspondence between emotional ratings on bipolar and unipolar scales.
Discourse relations may be either mononuclear or multi-nuclear. A mononuclear relation holds a nucleus and a satellite, the nucleus usually reflects the intention focus of the discourse, while the satellite represents supportive information. Although there are several previous works focused on analyzing discourse relations, only limited works focused on recognizing nuclearity between discourse units. In this paper, we propose a discourse unit nuclearity classifier in Chinese Discourse Treebank (CDTB). Our classifier considers the context of the two arguments, part of speech information, as well as word pair information as the features. The evaluations in CDTB corpus shows the effectiveness of our methodology and creates a baseline of 53.21% in accuracy.
OBJECTIVE: To examine the impact of plain packaging of cigarettes with enhanced graphic health warnings on adolescents' perceptions of pack image and perceived brand differences. METHODS: Cross-sectional school-based surveys conducted in 2011 (prior to introduction of new cigarette packaging) and in 2013 (7-12 months afterwards). Students aged 12-17 years (2011 n=6338; 2013 n=5915) indicated whether they had seen a cigarette pack in previous 6 months. Students rated the character of four popular cigarette brands, indicated level of agreement regarding differences between brands in ease of smoking, quitting, addictiveness, harmfulness and look of pack; and indicated positive and negative perceptions of pack image. Changes in responses of students seeing cigarette packs in the previous 6 months (2011: 60%; 2013: 65%) were examined. RESULTS: Positive character ratings for each brand reduced significantly between 2011 and 2013. Changes were found for four of five statements reflecting brand differences. Significantly fewer students in 2013 than 2011 agreed that 'some brands have better looking packs than others' (2011: 43%; 2013: 25%, p<0.001), with larger decreases found among smokers (interaction p<0.001). Packs were rated less positively and more negatively in 2013 than in 2011 (p<0.001). The decrease in positive image ratings was greater among smokers. CONCLUSIONS: The introduction of standardised packaging has reduced the appeal of cigarette packs. Further research could determine if continued exposure to standardised packs creates more uncertainty or disagreement regarding brand differences in ease of smoking and quitting, perceived addictiveness and harms.
This paper explores the problem of parsing Chinese long sentences. Inspired by human sentence processing, a second-stage parsing method, referred as main structure parsing in this paper, are proposed to improve the pars-ing performance as well as maintaining its high accuracy and efficiency on Chinese long sentences. Three different methods have at-tempted in this paper and the result shows that the best performance comes from the method using Chinese comma as the bounda-ry of the sub- sentence. According to our ex-periment about testing on the Chinese de-pendency Treebank 1.0 data, it improves long dependency accuracy by around 6.0 % than the baseline parser and 3.2 % than the previ-ous best model. 1
This paper describes the submitted discourse parsing system of the natural language group of Soochow University (SoNLP-DP) to the CoNLL 2015 shared task. Our System classifies discourse relations into explicit and non-explicit relations and uses a pipeline platform to conduct every subtask to form an end-toend shallow discourse parser in the Penn Discourse Treebank (PDTB). Our system is evaluated on the CoNLL-2015 Shared Task closed track and achieves the 18.51% in F1-measure on the official blind test set.
This paper explores the interaction between eventive information and morpho-syntax based on Chinese VV compounds. Chinese VV Compounds’ identical morpho-syntactic structure represents different event relations between the two component words and the correct interpretation of the meaning of these compounds relies on the prediction on their event relations. Without overt syntactic clues, we propose that ontology-based conceptual classification can be used to predict the event relation between the two component words. Compounding is the most productive way to research multi-word expressions in Mandarin Chinese. A Mandarin VV compound can be classified according to the eventive relation between two simplex verbs, which specifies how the eventive meanings of the two simplex verbs combine to form the meaning of the compound. The way in which two events combine with each other depends upon their event types, and the three types of eventive relations that we deal with in this paper are coordinate, modificational, and resultative. Using an ontology-based prediction approach, we hypothesized that the eventive relations could be predicted by the conceptual classification of the two simplex verbs’ event types. First, we utilized SUMO and Sinica BOW to classify each simplex verb. Next, the correlation between the ontology-based classification of each verb position and each eventive type was scored using a manually tagged lexical database and a training set was established. Finally, we encoded the ontological information of each VV compound in a 3-tuple based on these correlation scores. This 3-tuple was represented as a three-dimensional vector and was used to predict the eventive type of the new VV compounds. The results of our findings show that the classification experiments on event relation of unknown VV compounds can be reliably predicted based on the ontological classification of their component words.
We here present the development and validation of the Verbal Affective Memory Test-24 (VAMT-24). First, we ensured face validity by selecting 24 words reliably perceived as positive, negative or neutral, respectively, according to healthy Danish adults' valence ratings of 210 common and non-taboo words. Second, we studied the test's psychometric properties in healthy adults. Finally, we investigated whether individuals diagnosed with Seasonal Affective Disorder (SAD) differed from healthy controls on seasonal changes in affective recall. Recall rates were internally consistent and reliable and converged satisfactorily with established non-affective verbal tests. Immediate recall (IMR) for positive words exceeded IMR for negative words in the healthy sample. Relatedly, individuals with SAD showed a significantly larger decrease in positive recall from summer to winter than healthy controls. Furthermore, larger seasonal decreases in positive recall significantly predicted larger increases in depressive symptoms. Retest reliability was satisfactory, rs ≥.77. In conclusion, VAMT-24 is more thoroughly developed and validated than existing verbal affective memory tests and showed satisfactory psychometric properties. VAMT-24 seems especially sensitive to measuring positive verbal recall bias, perhaps due to the application of common, non-taboo words. Based on the psychometric and clinical results, we recommend VAMT-24 for international translations and studies of affective memory.
About half of the discourse relations annotated in Penn Discourse Treebank (Prasad et al., 2008) are not explicitly marked using a discourse connective. But we do not have extensive theories of when or why a discourse relation is marked explicitly or when the connective is omitted. Asr and Demberg (2012a) have suggested an information-theoretic perspective according to which discourse connectives are more likely to be omitted when they are marking a relation that is expected or predictable. This account is based on the Uniform Information Density theory (Levy and Jaeger, 2007), which suggests that speakers choose among alternative formulations that are allowed in their language the ones that achieve a roughly uniform rate of information transmission. Optional discourse markers should thus be omitted if they would lead to a trough in information density, and be inserted in order to avoid peaks in information density. We here test this hypothesis by observing how far a specific cue, negation in any form, affects the discourse relations that can be predicted to hold in a text, and how the presence of this cue in turn affects the use of explicit discourse connectives.
We describe a technique to minimize weighted tree automata (WTA), a powerful formalisms that subsumes probabilistic context-free grammars (PCFGs) and latent-variable PCFGs. Our method relies on a singular value decomposition of the underlying Hankel matrix defined by the WTA. Our main theoretical result is an efficient algorithm for computing the SVD of an infinite Hankel matrix implicitly represented as a WTA. We provide an analysis of the approximation error induced by the minimization, and we evaluate our method on real-world data originating in newswire treebank. We show that the model achieves lower perplexity than previous methods for PCFG minimization, and also is much more stable due to the absence of local optima.
Given the concentration of economic growth and power in science fields and the current levels of racial stratification in schooling, this study examined (1) the effects of race on students’ connectedness to science and career aspirations, (2) the extent to which these effects were moderated by school racial composition and racialized tracking, and (3) the differences in modeling effects using separate variables for race and gender (i.e., White, Black, Hispanic, female) versus race/gender (e.g., White female, Black male, etc.). Using the lens of racial formation theory, this study situated access to science knowledge as a racial project, conferring and denying access to resources along racial lines. Reviews of the literature on science self-efficacy, identity, engagement, and career aspirations revealed an under-emphasis on school institutional factors, such as racial composition and racialized tracking (which are important in sociological literature), as shaping student outcomes. The study analyzed data from the nationally representative High School Longitudinal Study that surveyed students in 2009 during their freshman year in high school and again in 2012 during most students’ junior year (n = 6,998). Affective ratings (in self-efficacy, identity, engagement) and career aspirations for students measured in 2012 were examined as dependent variables and a variable for racialized tracking was estimated given schools’ placement of students in advanced science coursework in 2012. Although school racial composition was not found to moderate race on outcome effects, primary analyses demonstrated that the presence of racialized tracking in the students’ schools did moderate these effects. Overall these results suggested that the student subgroups most often at a disadvantage compared to White students for the science outcomes studied were Hispanic males and females; Black students’ ratings and aspirations were largely on par or exceeded those of their White counterparts. In addition, results indicated that racialized tracking served to exacerbate gaps for Hispanic students and may also diminish career aspirations for Black students. Finally, while examining effects by race/gender did provide some additional insight and nuance in the interpretation of these results, there were clear instances where these more detailed analyses were not needed or may have obscured results that were clearer when aggregated by race. Given these results, implications for policy, practice, and future research are discussed.
In the past seven years, Language Research Institute of Inner Mongolia University has constructed a 500,000-word scale Mongolian dependency treebank. The syntactic treebank provides a favorable data platform for language research and information processing. In order to effectively use the treebank, we have designed and implemented a graphical syntactic information retrieval system based on the Mongolian dependency treebank. As an application system, this retrieval system offers search and statistical analysis on word, phrase, syntactic fragment and syntactic structure level.
The aim of this study was to investigate if and how temporal context influences subjective affective responses to emotional images. To do so, we examined whether the subjective evaluation of a target image is influenced by the valence of its preceding image, and/or its overall position in a sequence of images. Furthermore, we assessed if these potentially confounding contextual effects can be moderated by a common procedural control: randomized stimulus presentation. Four groups of participants evaluated the same set of 120 pictures from the International Affective System (IAPS) presented in four different sequences. Our data reveal strong effects of both aspects of temporal context in all presentation sequences, modified only slightly in their nature and magnitude. Furthermore, this was true for both valence and arousal ratings. Subjective ratings of negative target images were influenced by temporal context most strongly across all sequences. We also observed important gender differences: females expressed greater sensitivity to temporal-context effects and design manipulations relative to males, especially for negative images. Our results have important implications for future emotion research that employs normative picture stimuli, and contributes to our understanding of context effects in general.
With a dependency grammar, this study provides a unified method for calculating the syntactic complexity in linear and hierarchical dimensions. Two metrics, mean dependency distance (MDD) and mean hierarchical distance (MHD), one for each dimension, are adopted. Some results from the Czech-English dependency treebank are revealed: (1) Positive asymmetries in the distributions of the two metrics are observed in English and Czech, which indicates both languages prefer the minimalization of structural complexity in each dimension. (2) There are significantly positive correlations between sentence length (SL), MDD, and MHD. For longer sentences, English prefers to increase the MDD, while Czech tends to enhance the MHD. (3) A trade-off relationship of syntactic complexity in two dimensions is shown between the two languages. English tends to reduce the complexity of production in the hierarchical dimension, whereas Czech prefers to lessen the processing load in the linear dimension. (4) The threshold of the MDD2 and MHD2 in English and
We describe a technique to minimize weighted tree automata (WTA), a powerful formalisms that subsumes probabilistic context-free grammars (PCFGs) and latent-variable PCFGs. Our method relies on a singular value decomposition of the underlying Hankel matrix defined by the WTA. Our main theoretical result is an efficient algorithm for computing the SVD of an infinite Hankel matrix implicitly represented as a WTA. We provide an analysis of the approximation error induced by the minimization, and we evaluate our method on real-world data originating in newswire treebank. We show that the model achieves lower perplexity than previous methods for PCFG minimization, and also is much more stable due to the absence of local optima.
We examined the potential cost of practicing suppression of negative thoughts on subsequent performance in an unrelated task. Cues for previously suppressed and unsuppressed (baseline) responses in a think/no-think procedure were displayed as irrelevant flankers for neutral words to be judged for emotional valence. These critical flankers were homographs with one negative meaning denoted by their paired response during learning. Responses to the targets were delayed when suppression cues (compared with baseline cues and new negative homographs) were used as flankers, but only following direct-suppression instructions and not when benign substitutes had been provided to aid suppression. On a final recall test, suppression-induced forgetting following direct suppression and the flanker task was positively correlated with the flanker effect. Experiment 2 replicated these findings. Finally, valence ratings of neutral targets were influenced by the valence of the flankers but not by the prior role of the negative flankers.
We introduce interpolation of trained MSTParser models as a resource combination method for multi-source delexicalized parser transfer. We present both an unweighted method, as well as a variant in which each source model is weighted by the similarity of the source language to the target language. Evaluation on the HamleDT treebank collection shows that the weighted model interpolation performs comparably to weighted parse tree combination method, while being computationally much less demanding.
A positive change in perceived self-efficacy expectation is considered an important goal of exposure-based treatments. In line with this, several studies have demonstrated a positive association between self-efficacy and therapy outcome. While those studies primarily focused on changes in self-efficacy expectation that are achieved through therapy, the present study sought to examine whether changes in self-efficacy prior to treatment have an influence on therapy outcome. To this end, 48 healthy subjects completed a differential fear conditioning task. After the fear acquisition phase, half of the subjects received a positive verbal feedback aimed at increasing self-efficacy (experimental group) whereas others received no feedback (control group). Our results not only show that self-efficacy beliefs can be enhanced through verbal feedback but also point to an enhanced extinction of conditioned fear in the experimental group relative to the control group, evident on the implicit (skin conductance responses) and explicit (valence rating) level. Our results may have clinical implications for exposure-based treatments in anxiety disorders.
The purpose of this paper is to present an approach to create semi-automatically ontology from Arabic texts. The whole process is supervised by a linguistic expert. Our involvement in this project focused on a lexical ontology, taking as model the WordNet ontology, and as input source, the “Arabic verbs” of a contemporary monolingual dictionary () /mζjm Alγny/ in the form a lexical database. The verb, pivot of a sentence, is our goal in creating concepts, by adopting the synset as our meaning representation model. The Markov clustering algorithm of a graph, generated by the defining verbs, obtained from the transitive closure, allowed us to detect similar verbs and to identify as well, for a given verbal entry, all of its synonyms. A tool has been implemented, and experiments have been carried out to evaluate and show efficiency of the proposed approach.
Songs heard between the ages of 15 and 24 should be remembered better and have a stronger relationship to autobiographical memories when compared with music from other phases of life (“reminiscence bump effect”). Additionally, the proportion of music-evoked autobiographical memories (MEAMs) is at a maximum in these years of early adolescence and then declines up to the age of 60. In our study we tried both to replicate these important findings based on a German sample and to further investigate the influence of the affective characteristics of the songs on the frequency of participants’ autobiographical memories. In Experiment 1 a group of adults ( N = 48, M age = 67.1 years) listened to excerpts from 80, number-one, popular music hits from 1930 to 2010 and gave written self-reports on MEAMs. In Experiment 2 the affective characteristics were rated by another group of adults ( N = 22, M age = 66 years) and were used to predict the frequency of MEAMs. As a main result of Experiment 1, we confirmed the reminiscence bump and decline effect with a small effect size for the ratings of feelings evoked by the song and with a medium effect size for the song recognition performance of those songs released during the participants’ age range of 15 to 24 years. The total number of MEAMs was only marginally influenced by a memory bump and decline effect, and participants showed a significant proportion of MEAMs up to the fifth decade. Experiment 2 revealed that the affective ratings of the songs were unequally distributed over the two-dimensional emotion space unlike the average rate of MEAMs which was nearly equally distributed. In contrast to previous research, we therefore conclude that popular songs can be associated with autobiographical memory over five decades of life – independent of the affective character of the music.
With the development of internet, there are billions of short texts generated each day. However, the accuracy of large scale short text classification is poor due to the data sparseness. Traditional methods used to use external dataset to enrich the representation of document and solve the data sparsity problem. But external dataset which matches the specific short texts is hard to find. In this paper, we propose a framework to solve the data sparsity problem without using external dataset. Our framework deal with large scale short text by making the most of semantic similarity of words which learned from the training short texts. First, we learn word distributed representation and measure the word semantic similarity from the training short texts. Then, we propose a method which enrich the document representation by using the word semantic similarity information. At last, we build classifiers based on the enriched representation. We evaluate our framework on both the benchmark dataset(Standford Sentiment Treebank) and the large scale Chinese news title dataset which collected by ourselves. For the benchmark dataset, using our framework can improve 3% classification accuracy. The result we tested on the large scale Chinese news title dataset shows that our framework achieve better result with the increase of the training set size.
Pain catastrophising is an exaggerated cognitive attitude implemented during pain or when thinking about pain. Catastrophising was previously associated with increased pain severity, emotional distress and disability in chronic pain patients, and is also a contributing factor in the development of neuropathic pain. To investigate the neural basis of how pain catastrophising affects pain observed in others, we acquired EEG data in groups of participants with high (High-Cat) or low (Low-Cat) pain catastrophising scores during viewing of pain scenes and graphically matched pictures not depicting imminent pain. The High-Cat group attributed greater pain to both pain and non-pain pictures. Source dipole analysis of event-related potentials during picture viewing revealed activations in the left (PHGL) and right (PHGR) paraphippocampal gyri, rostral anterior (rACC) and posterior cingulate (PCC) cortices. The late source activity (600-1100 ms) in PHGL and PCC was augmented in High-Cat, relative to Low-Cat, participants. Conversely, greater source activity was observed in the Low-Cat group during the mid-latency window (280-450 ms) in the rACC and PCC. Low-Cat subjects demonstrated a significantly stronger correlation between source activity in PCC and pain and arousal ratings in the long latency window, relative to high pain catastrophisers. Results suggest augmented activation of limbic cortex and higher order pain processing cortical regions during the late processing period in high pain catastrophisers viewing both types of pictures. This pattern of cortical activations is consistent with the distorted and magnified cognitive appraisal of pain threats in high pain catastrophisers. In contrast, high pain catastrophising individuals exhibit a diminished response during the mid-latency period when attentional and top-down resources are ascribed to observed pain.
Syntactic-semantic treebank for domain ontology creationThis paper focuses on the creation of a domain treebank for the purposes of compiling a domain ontology. The domain treebank is viewed as a suitable resource for extracting of semantic relations from syntactic structures. First, the steps for ontology building are considered. Then, the processing over glossaries and standards is described with regard to their syntactic annotation. The utility of deriving semantic knowledge from the Treebank is also illustrated via the basic phrases. The idea is that the domain knowledge is represented in the domain data, but via treebanking more linguistic patterns can be extracted, which to be mapped to concepts and relations in a domain ontology.
INTRODUCTION: Although conceptual models of sexual functioning have suggested a major role for implicit cognitive processing in sexual functioning, this has thus far, only been investigated in women. AIM: The aim of this study was to investigate the role of implicit cognition in sexual functioning in men. METHODS: Men with (N = 29) and without sexual dysfunction (N = 31) were compared. MAIN OUTCOME MEASURES: Participants performed two single-target implicit association tests (ST-IAT), measuring the implicit association of visual erotic stimuli with attributes representing, respectively, valence ('liking') and motivation ('wanting'). Participants also rated the erotic pictures that were shown in the ST-IAT on the dimensions of valence, attractiveness, and sexual excitement to assess their explicit associations with these erotic stimuli. Participants completed the International Index of Erectile Functioning for a continuous measure of sexual functioning. RESULTS: Unexpectedly, compared with sexually functional men, sexually dysfunctional men were found to show stronger implicit associations of erotic stimuli with positive valence than with negative valence. Level of sexual functioning, however, was not predicted by explicit nor implicit associations. Level of sexual distress was predicted by explicit valence ratings, with positive ratings predicting higher levels of sexual distress. CONCLUSIONS: Men with and without sexual dysfunction differed significantly with regard to implicit liking. Research recommendations and implications are discussed.