Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Objective: To carry out the native assessment of International Affective Picture System(IAPS) among Chinese older adults.Methods:Altogether 116 Chinese older adults,including 51 male and 65 female,from three communities in Dalian City,aged from 60 to 80 years,rated 60 pictures(positive:25,neutral:12,negative:23) selected from the IAPS in terms of valence,arousal and dominance with Self-Assessment Manikin(SAM).The mean affective ratings were compared to the normative ratings of USA National Institute of Mental Health(NIMH).Result: Reliability analysis indicated that the affective ratings of our sample were stable and highly internally consistent.The affective ratings of Chinese older participants were strongly correlated with the normative ratings of NIMH(r=0.92,0.54 and 0.88 respectively for valence,arousal and dominance,P0.001).But paired t test showed there were still significant differences between the two samples.Chinese aged reported relatively higher arousal and dominance than NIMH sample for all pictures [(5.33±0.93)vs.(4.83±1.25),(5.60±1.20)vs.(5.19±1.21),P0.001],but lower valence than NIMH sample[(4.99±2.28)vs.(5.28±1.85),P=0.020].Male and female Chinese older participants showed similar emotional responses to most pictures.But female Chinese older participants reported higher valence than male ones(5.05±2.33/4.93±2.24,P0.05).The 60 pictures were distributed as shape in the two-dimensional affective space(valence-arousal).The association between valence and arousal was pronounced and linear for positive pictures(r=0.71,P0.001),but unpronounced for negative pictures,(r=-0.35,P0.05).Conclusion: IAPS is highly internationally accessible just as the expectation of its designers.However,considering about great differences in many aspects such as culture,social living and age between Chinese aged and NIMH sample,they may have different affective experiences to the same emotional stimuli.Therefore it is necessary to do some revisal before the IAPS is applied to Chinese aged.
We present the LFG PARSEBANKER, a comprehensive toolkit for interactive incremental construction of a treebank as a parsed corpus. This web-based toolkit offers an environment for batch and interactive parsing, versioning, inspection of structures, discriminant-based disambiguation, and statistics. It has recently been extended with a structural search facility.
OBJECTIVE: We aimed to study the neural processing of emotion-denoting words based on a circumplex model of affect, which posits that all emotions can be described as a linear combination of two neurophysiological dimensions, valence and arousal. Based on the circumplex model, we predicted a linear relationship between neural activity and incremental changes in these two affective dimensions. METHODS: Using functional magnetic resonance imaging, we assessed in 10 subjects the correlations of BOLD (blood oxygen level dependent) signal with ratings of valence and arousal during the presentation of emotion-denoting words. RESULTS: Valence ratings correlated positively with neural activity in the left insular cortex and inversely with neural activity in the right dorsolateral prefrontal and precuneus cortices. The absolute value of valence ratings (reflecting the positive and negative extremes of valence) correlated positively with neural activity in the left dorsolateral and medial prefrontal cortex (PFC), dorsal anterior cingulate cortex, posterior cingulate cortex, and right dorsal PFC, and inversely with neural activity in the left medial temporal cortex and right amygdala. Arousal ratings and neural activity correlated positively in the left parahippocampus and dorsal anterior cingulate cortex, and inversely in the left dorsolateral PFC and dorsal cerebellum. CONCLUSION: We found evidence for two neural networks subserving the affective dimensions of valence and arousal. These findings clarify inconsistencies from prior imaging studies of affect by suggesting that two underlying neurophysiological systems, valence and arousal, may subserve the processing of affective stimuli, consistent with the circumplex model of affect.
As a result of the widespread use of English in science and scholarship, there is an increasing need of reference tools which provide accurate information to non-nativeespecially junior-researchers on the correct use of lexico-grammatical patterns of nontechnical words when writing their scientific papers in English and on the conventionalized phraseological characteristics of the genre. Our aim is to present SciE-Lex, a lexical database which provides information to help Spanish researchers to write research papers in English accurately. Whereas there are specialized monolingual and bilingual dictionaries with specific terminological information, there is a shortage of reference tools supplying information on the correct use of syntactic and collocational patterns of nontechnical words in the scientific register and on the conventionalized phraseological characteristics of the genre. Based on the analysis of a 3+ million word corpus of scientific English, in its first stage, SciE-Lex displays information on: word class, morphological variants, equivalent(s) in Spanish, patterns of occurrence, list of collocations, examples of real use, and notes to clarify usage. In a second stage we plan to include lexical bundles, compositional recurrent sequences of words, since several studies have confirmed the difficulties that learners have with them. Further research will provide SciE-Lex with information about the distribution of lexical bundles across the different sections and/or moves of the academic research article as well as their function in discourse.
This paper presented an experiment on semantic role labeling by using SVM.This experiment was based on Chinese PropBank 5.0,which consisted of 1 652 sentences.The role-labeling set of this experiment included subject,object,indirect object,time and location.It used two-phase classification method with eight features,including path,phrase type,etc.For the small scaled training set,the experiment on testing set could reach the accuracy of 89.73% and the recall of 91.26% for semantic role labeling.Results highlight the effectiveness and efficiency of proposed approach for shallow semantic parsing of Chinese.
In this paper, we describe our work on building a parallel treebank for a less studied and typologically dissimilar language pair, namely Swedish and Turkish. The treebank is a balanced syntactically annotated corpus containing both fiction and technical documents. In total, it consists of approximately 160,000 tokens in Swedish and 145,000 in Turkish. The texts are linguistically annotated using different layers from part of speech tags and morphological features to dependency annotation. Each layer is automatically processed by using basic language resources for the involved languages. The sentences and words are aligned, and partly manually corrected. We create the treebank by reusing and adjusting existing tools for the automatic annotation, alignment, and their correction and visualization. The treebank was developed within the project Supporting research environment for minor languages aiming at to create representative language resources for language pairs dissimilar in language structure. Therefore, efforts are put on developing a general method for formatting and annotation procedure, as well as using tools that can be applied to other language pairs easily. 1.
In this paper, the high-level prosodic patterns of prosodic word (PW), prosodic phrase (PPh) and breath group/prosodic phrase group (BG/PG) for syllable pitch-level and duration are explored using an automatic joint prosody labeling and modeling method. Experimental results on a treebank speech corpus showed that the explored high-level prosodic patterns not only matched well with our a priori knowledge about Mandarin prosody, but also conformed well to other previous studies. They can therefore be integrated to form a meaningful Mandarin prosody hierarchy.
We present a robust parser which is trained on a treebank of ungrammatical sentences. The treebank is created automatically by modifying Penn treebank sentences so that they contain one or more syntactic errors. We evaluate an existing Penn-treebank-trained parser on the ungrammatical treebank to see how it reacts to noise in the form of grammatical errors. We re-train this parser on the training section of the ungrammatical treebank, leading to an significantly improved performance on the ungrammatical test sets. We show how a classifier can be used to prevent performance degradation on the original grammatical data.
BACKGROUND: Despite the highly acclaimed psychometric features of a 360-degree assessment in the fields of economics, military, and education, there has been increased interest in developing 360-degree instruments to assess competencies in graduate medical education only in the past recent years. Most of the effort to date, however, has focused on developing instruments and testing their reliability and feasibility. Insufficient attention has gone into issues of construct validity and particularly understanding the underlying constructs on which the instruments are based as well as the phenomena that affect ratings. PURPOSE: In preparation for developing a 360-degree assessment instrument, we explored variations in evaluators' opinion type of a competent resident and offer observation about evaluator's professional background and opinions. METHOD: Evaluators from two residency programs ranked 36 opinion statements, using a relative-ranking model, based on their opinion of a competent resident. By-person factor analysis was used to structure opinion types. RESULTS: Factor analysis of 156 responses identified four factors interpreted as four different opinion types of a competent resident: (a) altruistic, compassionate healer (n = 42 evaluators), (b) scientifically grounded clinician (n = 30), (c) holistic, humanistic clinician (n = 62), and (d) patient-focused, health manager (n = 31). Although 72% of nurses/respiratory therapist evaluators expressed type C, 28% expressed other types just as often. Only 14% of evaluator physicians expressed type D, and the remainders were evenly split among the other types. CONCLUSIONS: Our evaluators in 360-degree system expressed four opinion types of a competent resident. The individual opinion and not professional background influences the characteristics an evaluator values in a competent resident. We propose that these values will have an impact on competency assessment and should be taken into account in a 360-degree assessment.
This paper discusses a framework for development of bilingual and multilingual comprehension assistants and presents a prototype implementation of an English-Bulgarian comprehension assistant. The framework is based on the application of advanced graphical user interface techniques, WordNet and compatible lexical databases as well as a series of NLP preprocessing tasks, including POS-tagging, lemmatisation, multiword expressions recognition and word sense disambiguation. The aim of this framework is to speed up the process of dictionary look-up, to offer enhanced look-up functionalities and to perform a context-sensitive narrowing-down of the set of translation alternatives proposed to the user.
With the advent of the Internet, billions of images are now freely available online and constitute a dense sampling of the visual world. Using a variety of non-parametric methods, we explore this world with the aid of a large dataset of 79,302,017 images collected from the Internet. Motivated by psychophysical results showing the remarkable tolerance of the human visual system to degradations in image resolution, the images in the dataset are stored as 32 x 32 color images. Each image is loosely labeled with one of the 75,062 non-abstract nouns in English, as listed in the Wordnet lexical database. Hence the image database gives a comprehensive coverage of all object categories and scenes. The semantic information from Wordnet can be used in conjunction with nearest-neighbor methods to perform object classification over a range of semantic levels minimizing the effects of labeling noise. For certain classes that are particularly prevalent in the dataset, such as people, we are able to demonstrate a recognition performance comparable to class-specific Viola-Jones style detectors.
A Floresta Sintá(c)tica tem como objetivo criar e disponibilizar um corpus sintaticamente anotado. Neste artigo, são apresentados dois novos materiais do projeto: Selva (300 mil palavras e parcialmente revisto) e Amazônia (3.8 milhões de palavras, não revisto). Para lidar com um material tão grande e variado foi construída a interface Milhafre. O artigo mostra, ainda, como vem sendo enfrentado o desafio de compatibilizar, de uma lado, o usuário lingüista, que pode ter um perfil muito heterogêneo e, em geral, pouca familiaridade determinadas formalizações mais utilizadas em informática e, de outro, um único modelo de anotação sintática, freqüentemente pouco conhecido do lado “lingüístico não-computacional” e uma interface de acesso e manipulação de corpora capaz de lidar com um objeto tão complexo como a língua. Palavras-chave: árvores sintáticas, corpus anotado, corpus revisto, busca em corpora.
We present a description of a new resource (Prague Dependency Treebank of Spoken Language) being created for English and Czech to be used for the task of speech understanding, broad natural language analysis for dialog systems and other speech-related tasks, including speech editing. The resources we have created so far contain audio and a standard transcription of spontaneous speech, but as a novel layer, we add an edited (ldquoreconstructedrdquo) version of the spoken utterances. These edits go beyond the scope of current speech reconstruction efforts in that we allow, on top of the usual deletions of speech artifacts, fillers, etc. also for word modifications, insertions and word order changes. We have used both monologue and dialogue recordings in English and Czech to verify the feasibility of such transcription. We have also assessed the quality of the resulting annotation since the relative freedom of the editing raises an issue of what a ldquocorrectrdquo annotation is.
In the paper we describe the results in the development of the tools for handling diverse multilingual lexical resources such as monolingual and multilingual dictionaries, terminological dictionaries, complex lexicographic databases or WordNet semantic networks. All the presented tools are based on the Dictionary Editor and Browser (DEB) platform which uses standard XML formats. In this direction we strive to standardization of the lexical resources and also their interoperability. All the presented tools are freely available. We summarize the basic features of the DEB platform as a whole and then concentrate on four applications: DEBDict (a general dictionary browser), DEBTerm (multilingual terminological dictionary editor), PRALED (Czech Lexical Database system) and Visual Browser (graphical semantic network browser).
A web-based collaborative environment including on-line authoring tools that is managed by a central database was developed in collaboration with several countries including Peru, Bolivia, and the United States. The application involved developing a linguistics database and eLearning environment for documenting, preserving, and promoting language training for Aymara, a language indigenous to Peru and Bolivia. The database, an ontology management system called Lyra, incorporates all elements of the language (dialogues, phrase patterns, phrases, words, and morphemes) as well as cultural multimedia resources (images and sound recordings). The organization of the database enables a high level of integration among language elements and cultural resources. Authoring tools are used by experts in the Aymara language to build the linguistic database. These tools are accessible on-line as part of the collaborative environment using standard web browsers incorporating the Java plug-in. The eLearning student interface is a web-based program written in Flash. The Flash program automatically interprets and formats data objects retrieved from the database in XML format. The student interface is presented in Spanish and English. A web service architecture is used to publish the database on-line so that it can be accessed and utilized by other application programs in a variety of formats
Graph-based and transition-based approaches to dependency parsing adopt very different views of the problem, each view having its own strengths and limitations. We study both approaches under the framework of beam-search. By developing a graph-based and a transition-based dependency parser, we show that a beam-search decoder is a competitive choice for both methods. More importantly, we propose a beam-search-based parser that combines both graph-based and transition-based parsing into a single system for training and decoding, showing that it outperforms both the pure graph-based and the pure transition-based parsers. Testing on the English and Chinese Penn Treebank data, the combined system gave state-of-the-art accuracies of 92.1% and 86.2%, respectively.
The reflection of the category of the comic in literary speech is under investigation in the article. The basis of the taxonomic description of morphological means in the Russian language, offered by the author, consists of the following ways of creation of the comic: the usage of homonymy and contiguous phenomena to it, play up of the meanings of the same linguistic unit, repetition of the word in different grammatical forms (polyptot), divergence from the linguistic norms.
We present the STYX system, which is designed as an electronic corpus-based exercise book of Czech morphology and syntax with sentences directly selected from the Prague Dependency Treebank, the largest annotated corpus of the Czech language. The exercise book offers complex sentence processing with respect to both morphological and syntactic phenomena, i. e. the exercises allow students of basic and secondary schools to practice classifying parts of speech and particular morphological categories of words and in the parsing of sentences and classifying the syntactic functions of words. The corpus-based exercise book presents a novel usage of annotated corpora outside their original context.
We previously observed robust activation in the hippocampal region in response to novel valenced stimuli during an fMRI recognition memory paradigm in healthy individuals across the lifespan. In this study, we compared activation during the memory task in elderly controls (EC) and individuals with mild cognitive impairment (MCI). In 23 right-handed participants (15 EC, 8 MCI; 16M/7F; mean age=71.6), neural activity was compared for novel and previously learned (familiar) items. During encoding, participants viewed 10 B/W pictures of baby and elderly faces with happy/sad expressions repeated in a block 6x, alternating with blocks of 10 circles. Participants indicated whether each face was happy or sad. After a 20-minute consolidation period, the 10 encoded faces were presented 4x each intermixed with 40 new faces (half happy, half sad). Participants identified “new” or previously learned (“old”) faces. Random effects group analysis was performed using SPM5 to compare activity in response to novel vs familiar faces. Outside the scanner, participants viewed 40 new faces (half sad, half happy), 10 faces from the encoding scan, and 80 novel faces from the recognition scan. They indicated whether each face was “new” or “old” and rated valence and arousal of each face. EC showed significant activity in right fusiform and hippocampal regions in response to novel faces compared to previously learned faces (MNI coordinates FF: 40, -58, -14, T=8.07; pcorrected=.0001; HC: 22, -12, -16; T=5.60, puncorrected=.048). Significant activations were also observed in occipital and frontal cortices. A similar pattern was found in response to baby faces but did not hold for elderly faces, suggesting that arousal is important. MCI did not show increased activity in response to novel versus familiar faces in predicted regions. Groups did not significantly differ in overall reaction time/accuracy during encoding, valence/arousal ratings or accuracy during the post-scan task, or on the Florida Affect Battery, suggesting MCIs' affective perception was intact. MCIs were slower and less accurate during recognition scans. Lack of encoding-associated activity in MCIs suggests that continued studies of the functional correlates of the emotional-memory enhancement effect in the early stage of Alzheimer's disease are warranted.
Morphological processes in Semitic languages deliver space-delimited words which introduce multiple, distinct, syntactic units into the structure of the input sentence. These words are in turn highly ambiguous, breaking the assumption underlying most parsers that the yield of a tree for a given sentence is known in advance. Here we propose a single joint model for performing both morphological segmentation and syntactic disambiguation which bypasses the associated circularity. Using a treebank grammar, a data-driven lexicon, and a linguistically motivated unknown-tokens handling technique our model outperforms previous pipelined, integrated or factorized systems for Hebrew morphological and syntactic processing, yielding an error reduction of 12% over the best published results so far. 1
Melatonin in elderly patients Rixt F. Riemersma-van der Lek, Dick F. Swab, Jos Twisk, Elly M. Hol, Witte J.G. Hoogendijk and Eus Van Someren* Sleep and Cognition Laboratory, Netherlands Institute for Neuroscience, Amsterdam, The Netherlands (Received 10 March 2007; final version received 29 January 2008) Long-term combined light and melatonin treatment improve sleep, cognition and mood in demented elderly patients. A good night’s sleep sustains cognitive performance. Disturbed sleep, a frequent decisive factor in caregiver burden and institutionalisation of Alzheimer patients, may thus augment their characteristic impairments. We hypothesized that a possible reversible lack of activation of the circadian clock could contribute to sleep problems, and performed the first controlled human study on the effect of prolonged combined stimulation with light and melatonin. During a 3.5 year double-blind placebo-controlled randomized follow-up study, 189 elderly patients received daily supplementation of the circadian synchronisers light (+ lux, whole-day), and/or melatonin (2.5 mg). Half-yearly assessments were made of actigraphic sleep–wake rhythm estimates, cognition (MMSE), and non-cognitive symptoms. Combined light and melatonin treatment improved nocturnal restlessness by 8 + 3% per year (p 5 0.01), resulting in increased sleep duration and efficiency and a more pronounced 24-hour amplitude. Light improved cognition by 0.9 + 0.4 MMSE points or 5% (p 1⁄4 0.04). Light ameliorated depressive symptoms by 19% (Cornell Scale for Depression in Dementia). Melatonin ameliorated the worsening of psychiatric symptoms that occurred in subjects about to drop out of the study due to nursing home placement or death (Questionnaire format of the Neuropsychiatric Inventory). A negative effect of melatonin on affect (Philadelphia Geriatric Centre Affect Rating Scale) was counteracted in combination with bright light. Combined treatment also attenuated aggressive behaviour (Cohen-Mansfield Agitation Index). Both light and melatonin treatment enhanced the nocturnal rise of the 24-hour saliva melatonin rhythm, as measured in the absence of melatonin gifts. Melatonin treatment however also resulted in an increased daytime melatonin level, which was associated with an attenuation of diurnal activity. This first study on long-term stimulation of the human circadian timing system showed that improvement of the sleep–wake rhythm contributed to attenuation of cognitive The complete description of this study is available in: Riemersma-van der Lek et al. 2008. JAMA 299:2642–2655. *Corresponding author. Email: e.van.someren@nin.knaw.nl Biological Rhythm Research Vol. 40, No. 1, February 2009, 83–84 ISSN 0929-1016 print/ISSN 1744-4179 online! 2009 Taylor & Francis DOI: 10.1080/09291010802067155 http://www.informaworld.com D ow nl oa de d by [V rij e U ni ve rs ite it A m ste rd am ] a t 0 3: 36 2 0 A pr il 20 15 decline, with an affect exceeding that achieved with acetylcholinesterase inhibitors. Light improved non-cognitive symptoms, whereas melatonin may aggravate withdrawal behaviour and should – for long-term treatment in demented elderly – preferably be given in a lower dosage and in combination with bright light. 84 R.F. Riemersma-van der Lek et al. D ow nl oa de d by [V rij e U ni ve rs ite it A m ste rd am ] a t 0 3: 36 2 0 A pr il 20 15
The dual process theory proposes that evaluative conditioning is a form of learning distinct from Pavlovian conditioning and that it displays different functional characteristics such as not being subject to modulation. However, when assessed online as opposed to post-experimentally, modulation of evaluative conditioning by context change has been found in a contingency reversal procedure. Reversal of evaluative learning was found to be faster when trained in a different context rather than in the original training context. The present study addressed the question whether context change or instructions would affect the rate of reversal of evaluative learning and whether reversal learning would accelerate across repetitions. A picture-picture paradigm was used to expose participants to CS-US pairs and contingency was reversed three times during the experiment. Participants were required to provide online causal judgements and valence ratings after each set of 10 training trials. Context change, but not instructions, displayed a trend in affecting reversal of evaluative learning with participants displaying faster learning on trials immediately subsequent to contingency reversal. Instructions affected the reversal of contingency judgements. There was no evidence of acceleration across repetitions for either measure or manipulation.
We present the second version of the Penn Discourse Treebank, PDTB-2.0, describing its lexically-grounded annotations of discourse relations and their two abstract object arguments over the 1 million word Wall Street Journal corpus. We describe all aspects of the annotation, including (a) the argument structure of discourse relations, (b) the sense annotation of the relations, and (c) the attribution of discourse relations and each of their arguments. We list the differences between PDTB-1.0 and PDTB-2.0. We present representative statistics for several aspects of the annotation in the corpus. 1.
ProPOSEL is a prototype prosody and PoS (part-of-speech) English lexicon for Language Engineering, derived from the following language resources: the computer-usable dictionary CUVPlus, the CELEX-2 database, the Carnegie-Mellon Pronouncing Dictionary, and the BNC, LOB and Penn Treebank PoS-tagged corpora. The lexicon is designed for the target application of prosodic phrase break prediction but is also relevant to other machine learning and language engineering tasks. It supplements the existing record structure for wordform entries in CUVPlus with syntactic annotations from rival PoS-tagging schemes, mapped to fields for default closed and open-class word categories and for lexical stress patterns representing the rhythmic structure of wordforms and interpreted as potential new text-based features for automatic phrase break classifiers. The current version of the lexicon comes as a textfile of 104052 separate entries and is intended for distribution with the Natural Language ToolKit; it is therefore accompanied by supporting Python software for manipulating the data so that it can be used for Natural Language Processing (NLP) and corpus-based research in speech synthesis and speech recognition. 1. ProPOSEL: Derivation and Rationale A pronunciation lexicon is an integral part of the front-end NLP module in a generic Text-to-Speech (TTS) synthesis system and constitutes a natural way of giving such a system both prosodic and syntactic insights into input text. For English, three such resources- originally developed for Automatic Speech Recognition (ASR) and listing words and their phonetic transcriptions- are widely used: CELEX-2 (Baayen et al, 1996); PRONLEX (Kingsbury et
This article examines the constitution of ‘Japanese women’s language,’ or joseego, as a prescriptive linguistic norm for women through an analysis of scholarly and media representations of linguistic femininity in contemporary Japan. We distinguish two kinds of interrelated norms as constituting linguistic femininity: norms centered around general stylistic features such as politeness, gentleness, and refinement (the first-order norms) and those specificying particular linguistic forms including phonological, morphological, and lexical features (the second-order norms). Our analysis shows that both scholarly and media representations tend to share the dominant ideology of feminine speech in terms of general stylistic features and that they both link those features to specific linguistic forms in terms of Standard Japanese, but that the media representations are relatively more flexible in this linkage than the former and allows more room for contestation and the negotiation of alternative femininities. Through this analysis, we discuss the complex indexical process in which linguistic forms are ideologically linked to femininity as well as the tenuous nature of this linkage.
The article is devoted to psychological aspects of communicative influence on professional contact of militia officer, using a district militia officer functions as an example reflecting peculiarities of law-enforcement organs activities as a whole. The role and a place of communicative influence on the process of goals achievement by a militia officer are point out. The results of psychological research are given, confirming the necessity of psycholinguistic norms and rules use by a law-enforcement organs member in order to fulfill his functions successfully
What’s the best way to assess the performance of a semantic component in an NLP system? Tradition in NLP evaluation tells us that comparing output against a gold standard is a good idea. To define a gold standard, one first needs to decide on the representation language, and in many cases a first-order language seems a good compromise between expressive power and efficiency. Secondly, one needs to decide how to represent the various semantic phenomena, in particular the depth of analysis of quantification, plurals, eventualities, thematic roles, scope, anaphora, presupposition, ellipsis, comparatives, superlatives, tense, aspect, and time-expressions. Hence it will be hard to come up with an annotation scheme unless one permits different level of semantic granularity. The alternative is a theory-neutral black-box type evaluation where we just look at how systems react on various inputs. For this approach, we can consider the well-known task of recognising textual entailment, or the lesser-known task of textual model checking. The disadvantage of black-box methods is that it is difficult to come up with natural data that cover specific semantic phenomena. 1. Evaluating Meaning Formal methods for the analysis of the meaning of natural language expressions have long been restricted to the ivory tower built by semanticists, logicians, and philosophers of language. It was only in exceptional cases that they made their way directly into open domain NLP tools. Recently, this situation has changed. Thanks to the development of treebanks (large collections of texts annotated with syntactic structures), robust statistical parsers trained on such treebanks, and the development of large-scale semantic lexica, we now have at our disposal systems that are able to produce formal semantic representations achieving
The late positive potential (LPP) is a sustained positive deflection in the event-related potential that is larger following the presentation of emotional compared to neutral visual stimuli. Recent studies have indicated that the magnitude of the LPP is sensitive to emotion regulation strategies such as reappraisal, which involves generating an alternate interpretation of emotional stimuli so that they are less negative. It is unclear, however, whether reappraisal-related reductions in the LPP reflect reduced emotional processing or increased cognitive demands following reappraisal instructions. In the present study, we sought to examine whether a more or less negative description preceding the presentation of unpleasant images would similarly modulate the LPP. The LPP was recorded from 26 subjects as they viewed unpleasant and neutral International Affective Picture System images. All participants heard a brief description of the upcoming picture; prior to unpleasant images, this description was either more neutral or more negative. Following the more neutral description, the magnitude of the LPP, unpleasant ratings, and arousal ratings were all reliably reduced. These results indicate that changes in narrative are sufficient to modulate the electrocortical response to the initial viewing of emotional pictures, and are discussed in terms of recent studies on reappraisal and emotion regulation.
The paper presents a tool assisting manual annotation of linguistic data developed at the Department of Computational linguistics, IBL-BAS. Chooser is a general-purpose modular application for corpus annotation based on the principles of commonality and reusability of the created resources, language and theory independence, extendibility and user-friendliness. These features have been achieved through a powerful abstract architecture within the Model-View-Controller paradigm that is easily tailored to task-specific requirements and readily extendable to new applications. The tool is to a considerable extent independent of data format and representation and produces outputs that are largely consistent with existing standards. The annotated data are therefore reusable in tasks requiring different levels of annotation and are accessible to external applications. The tool incorporates edit functions, pass and arrangement strategies that facilitate annotators ’ work. The relevant module produces tree-structured and graph-based representations in respective annotation modes. Another valuable feature of the application is concurrent access by multiple users and centralised storage of lexical resources underlying annotation schemata, as well as of annotations, including frequency of selection, updates in the lexical database, etc. Chooser has been successfully applied to a number of tasks – POS tagging, WS annotation, syntactic annotation. 1.
Wordnets are lexical databases in which words are organized into clusters based on their meanings, and they are linked to each other through different semantic and lexical relations. The first wordnet called the Princeton WordNet was created for English, which were followed by various wordnets created within the framework of the EuroWordNet and BalkaNet projects, among others. Here we focus on the development of wordnets in general and of the Hungarian WordNet (HuWN). The process of constructing HuWn is illustrated by examples, some language-specific and language-independent problems encountered during the construction process are discussed, and then basic statistical data on HuWN are presented as well. Finally, two subontologies of HuWN, namely, the financial domain ontology and the legal domain ontology are also presented, and possible applications of WordNets are outlined.
We describe the annotation of multiword expressions and multiword named entities in the Prague Dependency Treebank. This paper includes some statistics of data and inter-annotator agreement. We also present an easy way to search and view the annotation, even if it is closely connected with deep syntactic treebank.
Research on insight—the phenomenon of suddenly solving an apparently intransigent problem—has been hampered because stimulus problems have been few, ad hoc, heterogeneous, and difficult to solve. Responding to the need for a larger pool of problems of a similar type and of varying level of difficulty, we report an experiment testing the validity of rebuses as insight problems. A rebus combines verbal and visual clues to a common phrase, such as PAINS (“growing pains”). Solving a rebus requires breaking implicit assumptions of normal reading, similar to the restructuring required in insight. We hypothesized that, the more implicit assumptions are involved, the more difficult the solution. The results of a two-part experiment supported the hypothesis, with participants solving more problems involving one assumption than they did problems involving two or more. Also, rebus performance correlated significantly with self-rated insight and with scores on remote associates, but not with general verbal ability. The findings suggest that rebus puzzles may be a useful source of theoretically grounded insight problems.
This paper discusses a new approach to conceptual design. The presented methodology is based on the structure of meanings in the design process. The search and evaluation of meanings form the foun-dations of developing this structure of meanings. In order to facilitate the use and operation of the mean-ings, the WordNet lexical database is used. An ex-isting visualization of WordNet is used for the pro-cess of meaning search. The WordNet::Similarity software for the measure of the relatedness of mean-ings in this database is the basic tool used for the evaluation process. The concept of similarity is con-cerned with the degree of interconnections between different meanings. Such search and evaluation tech-niques are later on incorporated into our methodol-ogy of the structure of meanings to support the de-sign process. The measures of relatedness of mean-ings are developed as a convergence criterion for ap-plication in the evaluation processes. Further on, our methodology for the structure of meanings is used to construct meanings in a given example of shape. All the steps of the design methodology, including the search and evaluation processes involved in develop-ing of the structure of the meanings, are elucidated. This design example is discussed to clarify possible implications of the design methodology. Finally, the paper presents directions for developing and further extensions of the proposed design methodology.
Mediation analysis is widely used in the social sciences. Despite the popularity of mediation models, few researchers have used graphical methods, other than structural path diagrams, to represent their models. Plots of the mediated effect can help a researcher better understand the results of the analysis and convey these results to others. This article presents a method for creating and interpreting plots of the mediated effect for a variety of mediation models, including models with (1) a dichotomous independent variable, (2) a continuous independent variable, and (3) an interaction between an independent variable and the mediating variable. An empirical example is then presented to illustrate these plots. Sample code for creating plots of the mediated effect in R and SAS is also included, and may be downloaded from www.psychonomic.org/archive.
Abstract: The Web 2.0 maximizes Internet concept of encouraging its users to cooperate effectively for offer of virtual services and content organization. Among various potentialities of Web 2.0, folksonomy appears as a result of free attribution of tags to Web's resources by user himself. Folksonomies describe Web's resources; however, they aren't integrated in metadata in general. In order for them to be intelligible by machines and therefore used in Semantics Web context, they have to be automatically allocated to specific metadata elements. There are many metadata patterns. The focus of this investigation will be Dublin Core (DC) which is a gathering of metadata for description of electronic resources and which has been adopted by Institutional Repositories as a way of standardization and interoperability. We propose an investigation which intends to identify of metadata originated from folksonomies and integrate them in a DC Ontology extended so as to allow that values reported in tags may be conveniently gathered by protocol for metadata harvesting, specifically Open Archives Initiative - Protocol for Metadata Harvesting (OAI-PMH). This paper will present results of pilot study developed in beginning of investigation as well as metadata preliminarily defined. Metadata may be defined as a group of for description of resources [1]. There are many standards of metadata, however, in repository context; we can point out Dublin Core Metadata Element Set (DCMES) or simply Dublin Core (DC) which is a metadata pattern for description of electronic resources. This standard is well diffused, used globally and on a broad scale due to some factors: a) it was created specifically for description of electronic elements; b) it has an initiative which is responsible for its development, maintenance and spreading, Dublin Core Metadata Initiative (DCMI); c) it is group of metadata used for protocol Open Archives Initiative - Protocol for Metadata Harvesting (OAI-PMH), a mechanism for data transfer between digital repositories. The insertion of metadata in repositories may be done by authors themselves, professionals who mediate deposit or of final users. The more active participation of users in construction and organization of Internet contents is result of evolution of technologies used in Web, so-called Web 2.0, it is 'the network as platform, spanning all connected devices; Web 2.0 applications are those that make most of intrinsic advantages of that platform: delivering software as a continually-updated service that gets better more people use it, consuming and remixing data from multiple sources, including individual users, while providing their own data and services in a form that allows remixing by others, creating network effects through an architecture of participation, and going beyond page metaphor of Web 1.0 to deliver rich user experiences.' [2]. Among possibilities of Web 2.0 Folksonomy comes up as the result of personal free tagging of and objects (anything with an URL) for one's own retrieval. The tagging is done in a social environment (shared and open to others). The act of tagging is done by person consuming information [3]. The tags which make up a folksonomy would be key-words, categories or metadata [4]. In this brief definition of tag, we can notice that when attributed by users they can represent different roles. In a study [5][6] following roles are pointed out: Identifying What (or Who) it is About, Identifying What it Is, Identifying Who Owns It, Refining Categories, Identifying Qualities or Characteristics, Self Reference and Task Organizing. In another study, Kinds of Tags (KoT), which compared tags with DC metadata elements, authors observed that there are some tags which cannot be inserted in any of already existing and therefore, concluded that other metadata may be defined in order to include descriptions arising from folksonomies. Some probable which were identified: Action_Towards_Resource, To_Be_Used_In, Rate e Depth [7][8]. The KoT is being developed in partnership with following universities: Universidade do Minho (Portugal), University of Bologna (Italy), UKOLN (United Kingdom), Universidad Carlos III (Spain), La Trobe University (Australia) and Universit? Libr? de Bruxelles -Facult? de Philosophie et Lettres (Belgium) and has objective of verifying how tags derived from folksonomies can be normalized aiming at their interoperability with metadata standards, specifically DC. Summing up, metadata are groups of for description of digital resources, holding different standards, among them DC which is adopted by Repositories as basis for protocol for metadata harvesting (o OAI-PMH). In Web 2.0 context, folksonomies arise, which are result of Web resource tagging by its own users. Tags are a complementary form of description which expresses user's view of resource being used. It can be observed through preliminary results of KoT project that current of description defined in DCMI Metadata Terms do not include all descriptive attributed by resource users by means of these tags. In context shown, giving continuity to analysis resulting from KoT project, we propose an investigation which aims at identifying metadata derived from folksonomies and integrate them in a DC Ontology extended so as to enable that values reported in tags may be conveniently gathered by protocol for metadata harvesting. Being so, we intend to develop a qualitative approach research and answer following questions: Q1 - Which metadata are necessary to contain folksonomy values?; Q2 - Which metadata should be created and which is their relation with already existing ones in Dublin Core?; Q3 - Which codification schemes should be used and what is their relation to already recommended by DCMI?; Q4 - Which Ontology related to DC already enable access to previously established conceptualizations? Q5 - Accomplishing what is stipulated in DCAM, what is extension of DC Ontology which should be made available openly? The procedures are divided in four stages: 1) Analysing tags contained in KoT project dataset- at this stage we will analyse all tags in relation to resources to which they have been attributed. Complementarily, to settle doubts, it will be necessary to turn to lexical resources (dictionaries, encyclopaedias, Word Net, Wikipedia, etc) and to analyse tags in relation to its users to understand functionality of tag attributed as a metadata element. At this stage a pilot study will be developed to refine methodology proposed to verify if variants proposed for grouping and analysing tags are adequate to identify probable new metadata elements which could be extracted from folksonomies. 2) Propose complementary metadata to DC - Establishing description originated from folksonomies based on DC standard, DCAM model, ISO Standard 15836-2003 and NISO Standard Z39.85-2007 norms. At this stage we intend to propose and/or qualifiers complementary to DC. 3) Forming an Ontology - Here we intend to fulfil Integration of DC Ontologies with and/or qualifiers derived from folksonomies. The ontology will be created from Prot?g? tool and coded in OWL. 4) Validation of proposal - carried out by scientific community as methodology and results obtained will be presented in relevant events and scientific magazines and by DCMI Social Tagging community through investigations via online questionnaires and workshops proposed to community. It is intended that results of research may provide support so that applications based on Artificial Intelligence permit automation of harvesting processes including description provided from folksonomies. This paper will present results of pilot study (that is being finalized) alongside with preliminary results of first research stage: tag analysis. This stage will be done in following phases: a) Analysis and grouping of tags in their variant forms; b) Analysis of tags in relation to DC metadata and its qualifiers. The preliminary results of KoT point to possible proposal of some metadata or element refinements to DCMI. Those terms will potentially accommodate tags that currently do not have a metadata holder. The results of this research will therefore allow to determinate if KoT preliminary findings are verified and in which extension. The final paper will conclude with this discussion.
Studies examining factors that influence when words are learned typically investigate one lexical category or a small set of words. We provide the first evaluation of the relation between input frequency and age of acquisition for a large sample of words. The MacArthur-Bates Communicative Development Inventory provides norming data on age of acquisition for 562 individual words collected from the parents of children aged 0; 8 to 2; 6. The CHILDES database provides estimates of frequency with which parents use these words with their children (age: 0; 7-7; 5; mean age: 36 months). For production, across all words higher parental frequency is associated with later acquisition. Within lexical categories, however, higher frequency is related to earlier acquisition. For comprehension, parental frequency correlates significantly with the age of acquisition only for common nouns. Frequency effects change with development. Thus, frequency impacts vocabulary acquisition in a complex interaction with category, modality and developmental stage.
Annotated list of dependency bigrams occurring in the PDT more than five times and having part-of-speech patterns that can possibly form a collocation. Each bigram is assigned to one of the six MWE categories by three annotators.
Predicting Word-Naming and Lexical Decision Times from a Semantic Space Model Brendan T. Johns (johns4@indiana.edu) Department of Psychological and Brain Sciences, 1101 E. Tenth St. Bloomington, In 47405 USA Michael N. Jones (jonesmn@indiana.edu) Department of Psychological and Brain Sciences, 1101 E. Tenth St. Bloomington, In 47405 USA Abstract organization of semantic memory. In a lexical decision task, a letter string is presented and the participant provides a speeded response of whether the string is a word or not. In a naming task, the participant’s task is to name the presented word aloud as quickly as possible. Both measures produce an index of a word’s identification latency. Orthographic and phonological factors are certainly large components of both LDT and NT, but semantics plays a significant role as well, and co-occurrence models have yet to be extended to predicting reaction time variance for these single-word identification tasks. Modeling of retrieval times is usually done by looking for the best environmental correlates of LDT and NT (Adelman & Brown, 2008). Some of the most influential models of retrieval times are based upon word frequency. Word frequency (WF) has been used to drive many different types of models, including serial-searched rank frequency models (Murray & Forster, 2004), threshold activation models (Coltheart, et al., 2001), and connectionist models (Seidenberg & McClelland, 1989). However, recent evidence suggests that word frequency may not drive retrieval times but, rather, the causal factor is a word’s contextual diversity (Adelman, Brown, & Quesada, 2006; Adelman & Brown, 2008). Contextual diversity (CD) is the number of different contexts that a word appears in, and is based on the rational analysis of memory (Anderson & Milson, 1989), particularly the principle of likely need (PLN). PLN states that the more unique contexts a word appears in, the more likely the word will be needed in any future context. Hence, a word with a high CD should be faster to retrieve under this principle. A word’s CD value is typically computed by simply counting the number of different documents in which it appears across a text corpus. This measure has been shown to be a better predictor of LDT and NT than WF (Adelman, et al., 2006). However, operationalizing CD as the number of documents in which a word occurs may not be a fair instantiation of PLN. A word that appears in many documents may have a high WF, but it should have a low CD if those documents are highly redundant, as is the case with words that belong to a popular discourse topic for which many documents exist. It is the number of different contexts and the uniqueness of contexts that determines a word’s likely need. This calls for a measure of CD that considers the semantic uniqueness of documents that a word appears in. Based on PLN, it is reasonable to assume that if a word appears in a context it has never before occurred in, We propose a method to derive predictions for single-word retrieval times from a semantic space model trained on text corpora. In Experiment 1 we present a large corpus analysis demonstrating that it is the number of unique semantic contexts a word appears in across language, rather than simply the number of contexts or the frequency of the word, that is the most salient predictor of lexical decision and naming times. In Experiment 2, we develop a co-occurrence learning model that weights new contextual uses of a word based on fit to what currently exists in the word’s memory representation, and demonstrate this model’s superiority in fitting the human data compared to models built using information about the word’s frequency or number of contexts. Finally, in Experiment 3 we find that building lexical representations using semantic distinctiveness naturally produces a better-organized semantic space to make predictions for semantic similarity between words. Keywords: Co-occurrence model; Lexical-decision; LSA; Contextual distinctiveness Introduction The last decade has seen remarkable progress with co- occurrence models of lexical semantics (e.g., Lund & Burgess, 1996; Landauer & Dumais, 1997). These models learn semantic representations for words by observing lexical co-occurrence patterns across a large text corpus, typically representing the words in a high-dimensional semantic space. This approach provides both an account of the semantic representation for words and an account of the learning mechanisms humans use to build and organize semantic memory. Co-occurrence models have seen considerable success at accounting for data in a wide variety of semantic tasks, including TOEFL synonyms (Landauer & Dumais, 1997), semantic similarity ratings and exemplar categorization (Jones & Mewhort, 2007), and free association norms (Griffiths, Steyvers, & Tenenbaum, To date, all applications of co-occurrence models have been to semantic similarity between two words or two documents. The standard prediction of semantic similarity in these models is some measure of the angle between two vectors. However, co-occurrence models should, in theory, contain sufficient information in the magnitude of their representations to make predictions about single word retrieval as well. Lexical decision time (LDT) and word naming time (NT) are both important variables that offer insight into the
There is an increasing interest in multimodal communication as suggested by several national and international projects (ISLE, HUMAINE, SIMILAR, CHIL, AMI, CALO, VACE, CALLAS), the attention devoted to the topic by well-known institutions and organizations (the National Institute of Standards and Technology, the Linguistic Data Consortium), and the success of conferences related to multimodal communication (ICMI, IVA, Gesture, Measuring Behavior, Nordic Symposium on Multimodal Communication, LREC Workshops on Multimodal Corpora).
We have constructed a large scale and detailed database of lexical types in Japanese from a treebank that includes detailed linguistic information. The database helps treebank annotators and grammar developers to share precise knowledge about the grammatical status of words that constitute the treebank, allowing for consistent large-scale treebanking and grammar development. In addition, it clarifies what lexical types are needed for precise Japanese NLP on the basis of the treebank. In this paper, we report on the motivation and methodology of the database construction.
This article presents a Web-based tool for the creation of divergent-thinking and open-ended creativity tasks. A Java program generates HTML forms with PHP scripting that run an Alternate Uses Task and/or open-ended response items. Researchers may specify their own instructions, objects, and time limits, or use default settings. Participants can also be prompted to select their best responses to the Alternate Uses Task (Silvia et al., 2008). Minimal programming knowledge is required. The program runs on any server, and responses are recorded in a standard MySQL database. Responses can be scored using the consensual assessment technique (Amabile, 1996) or Torrance’s (1998) traditional scoring method. Adoption of this Web-based tool should facilitate creativity research across cultures and access to eminent creators. The Creative Task Creator may be downloaded from the Psychonomic Society’s Archive of Norms, Stimuli, and Data, www.psychonomic.org/archive.
We survey the evaluation methodology adopted in information extraction (IE), as defined in a few different efforts applying machine learning (ML) to IE. We identify a number of critical issues that hamper comparison of the results obtained by different researchers. Some of these issues are common to other NLP-related tasks: e.g., the difficulty of exactly identifying the effects on performance of the data (sample selection and sample size), of the domain theory (features selected), and of algorithm parameter settings. Some issues are specific to IE: how leniently to assess inexact identification of filler boundaries, the possibility of multiple fillers for a slot, and how the counting is performed. We argue that, when specifying an IE task, these issues should be explicitly addressed, and a number of methodological characteristics should be clearly defined. To empirically verify the practical impact of the issues mentioned above, we perform a survey of the results of different algorithms when applied to a few standard datasets. The survey shows a serious lack of consensus on these issues, which makes it difficult to draw firm conclusions on a comparative evaluation of the algorithms. Our aim is to elaborate a clear and detailed experimental methodology and propose it to the IE community. Widespread agreement on this proposal should lead to future IE comparative evaluations that are fair and reliable. To demonstrate the way the methodology is to be applied we have organized and run a comparative evaluation of ML-based IE systems (the Pascal Challenge on ML-based IE) where the principles described in this article are put into practice. In this article we describe the proposed methodology and its motivations. The Pascal evaluation is then described and its results presented.
While large-scale corpora and various corpus query tools have long been recognized as essential language resources, the value of word association norms as language resources has been largely overlooked. This paper conducts some initial comparisons of the lexical relationships observed within Japanese collocation data extracted from a large corpus using the Japanese language version of the Sketch Engine (SkE) tool (Srdanović et al., 2008) and the relationships found within Japanese word association sets taken from the large-scale Japanese Word Association Database (JWAD) under ongoing construction by Joyce (2005, 2007). The comparison results indicate that while some relationships are common to both linguistic resources, many lexical relationships are only observed in one resource. These findings suggest that both resources are necessary in order to more adequately cover the diverse range of lexical relationships. Finally, the paper reflects briefly on the implementation of association-based word-search strategies into electronic dictionaries proposed by Zock and Bilac (2004) and Zock (2006).
Statistical parsing of noun phrase (NP) structure has been hampered by a lack of goldstandard data. This is a significant problem for CCGbank, where binary branching NP derivations are often incorrect, a result of the automatic conversion from the Penn Treebank. We correct these errors in CCGbank using a gold-standard corpus of NP structure, resulting in a much more accurate corpus. We also implement novel NER features that generalise the lexical information needed to parse NPs and provide important semantic information. Finally, evaluating against DepBank demonstrates the effectiveness of our modified corpus and novel features, with an increase in parser performance of 1.51%. 1