Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
This paper presents Thai syntactic resource: Thai CG treebank, a categorial approach of language resources. Since there are very few Thai syntactic resources, we designed to create treebank based on CG formalism. Thai corpus was parsed with existing CG syntactic dictionary and LALR parser. The correct parsed trees were collected as preliminary CG treebank. It consists of 50,346 trees from 27,239 utterances. Trees can be split into three grammatical types. There are 12,876 sentential trees, 13,728 noun phrasal trees, and 18,342 verb phrasal trees. There are 17,847 utterances that obtain one tree, and an average tree per an utterance is 1.85.
In this paper semantic classes of Czech verbs are presented as they are obtained from the lexical database VerbaLex that has recently been built at the NLP Centre FI MU. At the moment we have in VerbaLex 82 semantic classes covering 10,482 Czech verb lemmata and 19,556 verb valency frames. We discuss the criteria for establishing semantic classes: the most important one is grouping verbs according to their senses. The second one exploits relations between semantic classes of Czech verbs and semantic roles and subcategorization features as they are used in VerbaLex valency frames. We also touch on the issue of the ontology that could be used to describe the meanings of the verbs in the semantic classes. The semantic classification of Czech verbs can be extended for other languages via Interlingual Index (ILI) existing in WordNets and it can be used in the various applications in the NLP area (machine translation, syntactic analysis, semantic search, information extraction and others).
Most treebank work in the past has focused on European and Asian languages. The Wikipedia Treebank page lists treebanks (or treebank projects) for about 20 modern European languages (ranging from Basque to Swedish), five Asian languages (Chinese, Japanese, Hindi, Korean, Thai), two ancient languages (Greek and Latin), plus Arabic and Hebrew. Almost no treebanking work has been done on African or American indigenous languages.1 In the past we have explored parallel treebanks for English, German and Swedish [7]. Now we would like to explore to what extent our tools and guidelines will work when we include a very different language, Quechua, for which only few NLP resources exist. Since Quechua is spoken in Latin America, Spanish as parallel language is a natural choice. We have first compiled a parallel corpus Quechua Spanish. We have then stepwise analyzed and annotated the Quechua and the Spanish texts. For Spanish we have used the treebanking guidelines developed by [8]. As for Quechua there were no such guidelines so that we had to experiment with finding the appropriate grammar formalism and develop our own guidelines. In this paper we describe the characteristics of Quechua and our steps towards its morphological and syntactic annotation. We argue for Role and Reference Grammar as a suitable grammar formalism. We briefly describe how we annotated the parallel Spanish texts and demonstrate how we plan to align the Quechua with the Spanish trees.
In this paper, we show that by integrating existing NLP techniques and Semantic Web tools in a novel way, we can provide a valuable contribution to the solution of the knowledge acquisition bottleneck problem. NLP techniques to create a domain ontology on the basis of an open domain corpus have been combined with Semantic Web tools. More specifically, Watson and Prompt have been employed to enhance the kick-o ontology while Cornetto, a lexical database for Dutch, has been adopted to establish a link between the concepts and their Dutch lexicalization. The lexicalized ontology constitutes the basis for the cross-language retrieval of learning objects within the LT4eL eLearning project.
The aim of the present paper was to study heart rate changes during a video stimulation depicting two actors (male and female) producing dynamic facial expressions of happiness, sadness, and a neutral expression. We measured ballistocardiographic emotion-related heart rate responses with an unobtrusive measurement device called the EMFi chair. Ratings of subjective responses to the video stimuli were also collected. The results showed that the video stimuli evoked significantly different ratings of emotional valence and arousal. Heart rate decelerated in response to all stimuli and the deceleration was the strongest during negative stimulation. Furthermore, stimuli from the male actor evoked significantly larger arousal ratings and heart rate responses than the stimuli from the female actor. The results also showed differential responding between female and male participants. The present results support the hypothesis that heart rate decelerates in response to films depicting dynamic negative facial expressions. The present results also support the idea that the EMFi chair can be used to perceive emotional responses from people while they are interacting with technology.
In this paper, we would like to introduce a new approach to recover Vietnamese text's accents. Given a Vietnamese text in which accents are lost, our goal is to seek for a recovered text that yields a best lexical probability. Using a dynamic programming approach, we first build a model of language for Vietnamese as a lexical database which gives lexical probabilities to Vietnamese sentences. Second, we construct a map of literal translations of Vietnamese words to restrict our searching space. Finally, we apply dynamic programming as a searching engine to seek out the most probable sentence. We also use the co-occurrence graph to increase the accuracy of selection, the experimental results show that the average accuracy of our approach is about 93%-94%.
This paper describes a system that resolves prepositional phrase attachment ambiguity in English sentence process-ing. This attachment problem is ubiquitous in English text, and is widely known as a place where semantics determines syntactic form. The decision is made based on a four-tuple composed of the head verb of the verb phrase, the head noun of the noun phrase, and the preposition and head noun in the prepositional phrase. A corpus with known results, the Penn Treebank, is used for training and testing purposes. During training, known results are used to build a lattice of hierarchical categories taken from WordNet. These lattices are then compared to the novel lattices derived from the test four-tuples. The results of the system are 90.53% correct attachment decisions.
This study examined undergraduate, non-music majors’ familiarity with and preference \nfor Arabic music as compared to other world music. Several factors were examined to assess \ntheir effect on music preference including familiarity, musical characteristics, and student \ncharacteristics. Study participants included 203 undergraduate, non-music majors enrolled in six \nsections of music appreciation classes. Participants were divided into Caucasian and non- \nCaucasian groups ranging from 18 to 42 years of age. Music excerpts from Africa (Congo), Latin \nAmerica (Mexico), Asia (Japan), and the Middle East (Kuwait) were used as examples of \ndifferent world music. Arabic music was introduced as a new factor in this study that had not been explored in previous research. Knowing about students’ familiarity and preference for \nArabic music may help in understanding the ramifications of its inclusion in music programs, \nand the proper method of introducing it to the students in the classroom. Participants listened to the 12 musical excerpts and completed the WMFPT questionnaire. Results indicated that \nparticipants were not familiar with the world music excerpts, but did like the excerpts to a \nmoderate degree. Significant positive relationships were found between preference and \nfamiliarity, within preference ratings, and within familiarity ratings. The most influential musical characteristics in liking world music were rhythm, tempo, and timbre, with rhythm being the most influential. Participants’ background seems to have no significant relationship with either familiarity or preference. Results revealed that playing a musical instrument, musical training, and previous exposure to music of other cultures significantly affected preference and familiarity ratings.
Objective:To explore the characteristics of emotion cognitive processing and cognitive regulation in alexithymia.Methods:A total of 117 alexithymic subjects(TAS-20 scores ≥58)and 118 nonalexithymic subjects(TAS-20 scores ≤38)were selected with the Chinese version of 20-item Toronto Alexithymia Scale(TAS-20),and their scores on the Center for Epidemiologic Studies Depression Scale(CES-D)and Cognitive Emotion Regulation Questionnaire CERQ were compared.Then 51 alexithymic subjects and 54 nonalexithymic subjects were required to rate 120 affective pictures to three dimensions(valence,arousal and dominant).Results:(1)Compared with nonalexithymic group,alexithymic group got higher scores in negative coping dimension[(47.3±5.9) vs.(41.9±5.9),P0.001],while got lower scores in positive coping [(65.2±7.7) vs.(71.1±7.3),P0.001].(2)In valence rating,alexithymic group gave lower score for positive pictures[(7.0±1.0) vs.(7.7±1.0),P0.001]and higher score for negative pictures [(2.4±1.0) vs.(1.4±1.0),P0.001]than nonalexithymic group.In arousal dimension,alexithymic group gave lower scores to both positive pictures and negative pictures than nonalexithymic group[(6.3±1.2) vs.(6.8±1.1),(6.4±1.5) vs.(7.2±1.4);P0.01].Conclusion:Alexithymia subjects have deficits in emotion cognitive processing and emotion cognitive regulation.
In this paper we extend a shallow parser [6] with prepositional phrase attachment. Although the PP attachment task is a well-studied task in a discriminative learning context, it is mostly addressed in the context of artificial situations like the quadruple classification task [18] in which only two possible attachment sites, each time a noun or a verb, are possible. In this paper we provide a method to evaluate the task in a more natural situation, making it possible to compare the approach to full statistical parsing approaches. First, we show how to extract anchor-pp pairs from parse trees in the GENIA and WSJ treebanks. Next, we discuss the extension of the shallow parser with a PP-attacher. We compare the PP attachment module with a statistical full parsing approach [4] and analyze the results. More specifically, we investigate the domain adaptation properties of both approaches (in this case domain shifts between journalistic and medical language). Keywords prepositional phrase attachment, shallow parsing, machine learning of language 1
A novel series of trifluoromethyl-containing quinazoline derivatives with a variety of functional groups was designed, synthesized, and tested for their antitumor activity by following a pharmacophore hybridization strategy. Most of the 20 compounds displayed moderate to excellent antiproliferative activity against five different cell lines (PC3, LNCaP, K562, HeLa, and A549). After three rounds of screening and structural optimization, compound 10 b was identified as the most potent one, with IC<sub>50</sub> values of 3.02, 3.45, and 3.98 μM against PC3, LNCaP, and K562 cells, respectively, which were comparable to the effect of the positive control gefitinib. To further explore the mechanism of action of 10 b against cancer, experiments focusing on apoptosis induction, cell cycle arrest, and cell migration assay were conducted. The results showed that 10 b was able to induce apoptosis and prevent tumor cell migration, but had no effect on the cell cycle of tumor cells.
GLARF relations are generated from treebank and parses for English, Chinese and Japanese. Our evaluation of system output for these input types requires consideration of multiple correct answers.
Applying statistical parsers developed for English to languages with freer word-order has turned out to be harder than expected. This paper investigates the adequacy of different statistical parsing models for dealing with a (relatively) free word-order language. We show that the recently proposed Relational-Realizational (RR) model consistently outperforms state-of-the-art Head-Driven (HD) models on the Hebrew Treebank. Our analysis reveals a weakness of HD models: their intrinsic focus on configurational information. We conclude that the form-function separation ingrained in RR models makes them better suited for parsing nonconfigurational phenomena.
The creation of the Slovene Lexical Database is one of the central goals of a major Slovene lexicographic and human language technologies project (Communication in Slovene: http://www.slovenscina.eu). The database should, on the one hand, serve the needs of posterior compilation of monolingual and bilingual dictionaries, and, on the other hand, be robust enough to serve further human language technologies needs. The initial design phase of the database has been aimed at defining the guidelines and principles of the database compilation. In this period, several similar lexical database projects have been scrutinized. The paper presents a review of their features and then focuses on the semantic level of the description of the lexical unit. This level has two subcategories, namely the semantic indicator and the argument structure. The organizing principles of these two categories, especially the latter, are presented along with sample materials.
In this paper, we introduce an automatic method for classifying a given question using broad semantic categories in an existing lexical database (i.e., WordNet) as the class tagset. For this, we also constructed a large scale entity supersense database that contains over 1.5 million entities to the 25 WordNet lexicographer’s files (supersenses) from titles of Wikipedia entry. To show the usefulness of our work, we implement a simple redundancy-based system that takes the advantage of the large scale semantic database to perform question classification and named entity classification for open domain question answering. Experimental results show that the proposed method outperform the baseline of not using question classification. 關鍵詞: 自動問題回答,問題分類,辭彙語意資料庫,辭網,維基百科
The intensity and valence of 30 emotion terms, 30 events typical of those emotions, and 30 autobiographical memories cued by those emotions were each rated by different groups of 40 undergraduates. A vector model gave a consistently better account of the data than a circumplex model, both overall and in the absence of high-intensity, neutral valence stimuli. The Positive Activation - Negative Activation (PANA) model could be tested at high levels of activation, where it is identical to the vector model. The results replicated when ratings of arousal were used instead of ratings of intensity for the events and autobiographical memories. A reanalysis of word norms gave further support for the vector and PANA models by demonstrating that neutral valence, high-arousal ratings resulted from the averaging of individual positive and negative valence ratings. Thus, compared to a circumplex model, vector and PANA models provided overall better fits.
WordNet is an on-line lexical database of English language. Word meanings are represented by synonym sets (synsets). Each synset corresponds to one lexical concept. Different lexical and semantic relations link the synsets. An algorithm for automated forming Bulgarian synsets, corresponding to WordNet synsets, is presented. The three main steps of the algorithm are: I Automated improving of synsets on synonym dictionary that includes discovering synsets representing one concept; synsets which contain words for two or more different concepts; synsets with missed or incorrect synonyms placed in them, etc.; II. Finding a correspondence between lines of Englilish-Bulgarian Dictionary and WordNet synsets; III. Forming Bulgarian synsets coresponding to 55,000 WordNet concepts, using results from steps I. and II.
From the introduction: It is not always easy to define what a word means. We can choose between a variety of possibilities, from simply pointing at the correct object as we say its name to lengthy definitions in encyclopaedias, which can sometimes fill multiple pages. Although the former approach is pretty straightforward and is also very important for first language acquisition, it is obviously not a practical solution for defining the semantics of the whole lexicon. The latter approach is more widely accepted in this context, but it turns out that defining dictionary and encyclopaedia entries is not an easy task. In order to simplify the challenge of defining the meaning of words, it is of great advantage to organize the lexicon in a way that the structure in which the words are integrated gives us information about the meaning of the words by showing their relation to other words. These semantic relations are the focal point of this paper. In the first chapter, different ways to describe meaning will be discussed. It will become obvious why semantic relations are a very good instrument to organizing the lexicon. The second chapter deals with WordNet, an electronic lexical database which follows precisely this approach. We will examine the semantic relations which are used in WordNet and we will study the distinct characteristics of each of them. Furthermore, we will see which contribution is made by which relation to the organization of the lexicon. Finally, we will look at the downside of the fact that WordNet is a manually engineered network by examining the shortcomings of WordNet. In the third chapter, an alternative approach to linguistics is introduced. We will discuss the principles of corpus linguistics and, using the example of the British National Corpus, we will consider possibilities to extract semantic relations from language corpora which could help to overcome the deficiencies of the knowledge based approach. In the fourth chapter, I will describe a project the goal of which is to extend WordNet by findings from cognitive linguistics. Therefore, I will discuss the development process of a piece of software that has been programmed in the course of this thesis. Furthermore, the results from a small‐scale study using this software will be analysed and evaluated in order to check for the success of the project.
In this paper we describe an approach to target language modeling which is based on a large treebank. We assume a bag of bags as input for the target language gener-ation component, leaving it up to this com-ponent to decide upon word and phrase or-der. An experiment with Dutch as target language shows that this approach to can-didate translation reranking outperforms standard n-gram modeling, when measur-ing output quality with BLEU, NIST, and TER metrics. 1
Treebanks, as a quantitative extension of decades of syntactic theorizing, typically use annotation schemes with a small set of well-motivated phrasal categories.For constituency-based treebanks, these phrasal categories are selected to describe distributional regularities.These treebanks are often used as a data set for estimating Probabilistic Context Free Grammars (PCFGs) for parsing, but the phrasal category sets which are best for constituency description may be suboptimal for constituency parsing.Specifically, phrasal categories may exhibit a probabilistic bias towards different expansions in different parts of the overall tree, and there may be unanticipated but useful correlations between constituency annotation and other levels of linguistic annotation.In this thesis, the symbol-splitting technique of Johnson ( 1998) is extended to enrich syntactic categories with information about local syntactic context on the English Penn Treebank and the German Verbmobil II Treebank.The split symbols are then subjected to two different clustering techniques to preserve only relevant category distinctions, forming linguistically-motivated generalizations and assuaging data sparsity.The symbol-splitting and clustering techniques are then employed, on the Verbmobil treebank, to enrich syntactic categories with information about implicit prosodic break strength alone and then together with information about local context.Local syntactic context is found to be helpful on both treebanks examined.Experiments on the German Verbmobil II Treebank then show that information about Vita
In the article we present the compilation of a lexical database for Slovene, which is being undertaken within the framework of the project Communication in Slovene. We first shed light upon the concept from the theoretical point of view, and then position it within the context of concrete results in similar projects for other individual European languages. The dual purpose of the lexical database for Slovene, i.e., for dictionary applications and for natural language processing, determines the description of lexical units from three basic viewpoints: semantic, syntactic and collocational. On the basis of these points of departure we present the construction of the lexical database in terms of content levels and elucidate theoretical reflections on content solutions. Finally we present a tool for processing the lexicogrammatical profile of wordsWord Sketch, and describe the programme interface for the production of a lexical database.
Recent advances in research on affect have shown variations in individuals' affective circumplex, which are related to differences in arousal sensitivity (Blascovich, 1990; 1992; Feldman, 1995a). Given evidence suggesting that psychopaths are hyposensitive to arousal, they should therefore demonstrate an arousal-focus when processing emotional material. The present study was designed to assess this attentional bias. Participants were 41 offenders divided into psychopathic and nonpsychopathic groups according to their scores on the Psychopathy Checklist-Revised (PCL-R; Hare, 1991). Pleasure and arousal ratings and similarity judgments of emotional words and pictures of facial affect were employed to provide both objective and projective measures of the psychopath's interpretation of the material. Results indicated that psychopaths and nonpsychopaths did not differ in either their ratings or similarity judgments of the stimuli. However, when participants were further subdivided into sex offenders and non sex offenders, a pattern emerged that suggested that the psychopath's affective structure is indeed characterized by an arousal-focus, but only when faced with "intriguing" stimulation. Results are discussed in light of a new theory of psychopathy suggesting that psychopaths may be addicted to arousal, as represented in their lifestyles and their affective responses.
In this paper, we will describe some theoretical and practical issues raised during the construction of the Basque Dependency Treebank (BDT): the syntactic annotation of EPEC (Reference Corpus for the Processing of Basque). EPEC is a 300,000 word corpus of standard written Basque whose purpose is to be a training corpus for the development and improvement of several NLP (Natural Language Processing) tools for Basque. BDT will be the first corpus for the Basque language tagged at syntactic level. We will also present the dependency-based annotation hierarchy that we have established for the syntactic tagging. Decisions made during design of the annotation hierarchy are based on the description of Basque grammar made by Euskaltzaindia (Academy for the Basque Language). When describing dependency relations, we consider lexical units as syntactic heads. This will open up a way for us to work with semantics.
Broad-coverage annotated treebanks necessary to train parsers do not exist for many resource-poor languages. The wide availability of parallel text and accurate parsers in English has opened up the possibility of grammar induction through partial transfer across bitext. We consider generative and discriminative models for dependency grammar induction that use word-level alignments and a source language parser (English) to constrain the space of possible target trees. Unlike previous approaches, our framework does not require full projected parses, allowing partial, approximate transfer through linear expectation constraints on the space of distributions over trees. We consider several types of constraints that range from generic dependency conservation to language-specific annotation rules for auxiliary verb analysis. We evaluate our approach on Bulgarian and Spanish CoNLL shared task data and show that we consistently outperform unsupervised methods and can outperform supervised learning for limited training data.
Parallel treebanks provide a systematic way of expressing the structural relationships between source and target texts. In this paper, we present the general design principles behind the Copenhagen Dependency Treebanks, a set of parallel treebanks for Danish, English, German, Italian and Spanish with a unified annotation of morphology, syntax, discourse, and tranlational equivalence. Finally, we suggest some hypotheses about morphology and discourse, and describe how we plan to explore them empirically on the basis of the treebanks.
Parallel grammars and parallel treebanks can be a useful method for studying linguistic diversity and commonality. We use this approach to study how arguments to similar predicates are realized across languages. To that end, we formulate formal principles for aligning at phrase and word levels based on translational correspondences at predicate-argument level. A first version of a new tool for creating, storing, visualizing and searching treebank alignment at different levels has been constructed. 1
Enhanced sensitivity to information of negative (compared to positive) valence has an adaptive value, for example, by expediting the correct choice of avoidance behavior. However, previous evidence for such enhanced sensitivity has been inconclusive. Here we report a clear advantage for negative over positive words in categorizing them as emotional. In 3 experiments, participants classified briefly presented (33 ms or 22 ms) masked words as emotional or neutral. Categorization accuracy and valence-detection sensitivity were both higher for negative than for positive words. The results were not due to differences between emotion categories in either lexical frequency, extremeness of valence ratings, or arousal. These results conclusively establish enhanced sensitivity for negative over positive words, supporting the hypothesis that negative stimuli enjoy preferential access to perceptual processing.
In this thesis, I present arguments for a model of language acquisition with three characteristics. These are (1) Continuity in the abstract principles of Universal Grammar; (2) Lexical Learning, or the setting of syntactic parameters based upon the acquisition of morphology; and (3) Morpholexical Learning, which is the abstraction of morphological patterns and generalizations from a lexical database. Continuity accounts for what is invariant in language development. Lexical Learning accounts for what is languageparticular, and which therefore must be learned. Morpholexical Learning accounts for the sequence of developmental stages observed in child language data. The main goal of this thesis is to demonstrate that Morpholexical Learning, in conjunction with paradigmatic structure in the lexicon, provides a model for the acquisition of inflectional morphology. I demonstrate this proposal with data on the acquisition of subject-verb agreement morphology in German. In Chapter One, I present an introduction to the concerns and main proposals of this thesis. In Chapter Two, I motivate the existence of paradigmatic structure with three diachronic case studies. In Chapter Three, I return to the acquisitional debates introduced in Chapter One. I argue that the Continuity Hypothesis represents a preferable alternative to the Maturational Hypothesis. Next, I show that Lexical Learning of clausal representations is superior to the Lexical Projection Hypothesis and the Full Competence Hypothesis. I argue that Morpholexical Learning provides an answer to the Developmental Problem which Continuity and Lexical Learning create. In Chapter Four, I provide an extended case study of the acquisition of subjectverb agreement in German. The construction of word-specific paradigms during the stages under examination accounts for the pattern of agreement errors which German children produce. In Chapter Five, I continue the analysis of paradigm mixture begun in Chapter Two. The patterns of paradigm mixture attested Latin, German, and Icelandic are in essence identical to one another, which suggests universal principles of inflectional organization. In Chapter Six, I conclude the thesis with a sketch of how children develop from the "word-specific paradigm" stage to the "general paradigm stage".
In this paper we describe our participation at the EVALITA 2009 Con- stituency Parsing Task. We used the Berkeley Parser, obtaining the best F1, that is 78:73. This result corresponds to an increment of 15.85% with respect to the best result obtained at EVALITA 2007 by the Bikel's parser (F1 = 67:96). A further important advantage of the Berkeley parser is that it does not require any language adaptation in addition to the need of retraining it on the new treebank. For comparison, we also report the results obtained by the Bikel's parser on the 2009 treebank.
Linguistic database summaries in the sense of Yager (1982), further extended to an implementable form by Kacprzyk & Yager (2001) and Kacprzyk, Yager & Zadrozny (2000), are extremely simple natural language like statements exemplified by, for a personnel database, “most employees are young and well paid” (with some degree of truth). They have been implemented in business contexts (cf. Kacprzyk & Zadrozny,????, Kacprzyk, Wilbik and Zadrozny, 2006–2008). An effective and efficient way of their generation was proposed by Kacprzyk & Zadrozny (????) by using an interactive procedure based on Kacprzyk & Zadrozny's (????) fuzzy database queries with linguistic quantifiers. Moreover, in Kacprzyk & Zadrozny (???) the role of Zadeh's (???) protoform was shown and their use advocated. Though linguistic database summaries have a strong resemblance to natural language generation (NLG), this issue was never considered. In this paper we indicate some important issues that are common to linguistic database summarization and natural language generation, and propose some possible research directions.
We consider linguistic database summaries in the sense of Yager (1982), in an implementable form proposed by Kacprzyk & Yager (2001) and Kacprzyk, Yager & Zadrozny (2000), exemplified by, for a personnel database, “most employees are young and well paid” (with some degree of truth) and their extensions as a very general tool for a human consistent summarization of large data sets. We advocate the use of the concept of a protoform (prototypical form), vividly advocated by Zadeh and shown by Kacprzyk & Zadrozny (2005) as a general form of a linguistic data summary. Then, we present an extension of our interactive approach to fuzzy linguistic summaries, based on fuzzy logic and fuzzy database queries with linguistic quantifiers. We show how fuzzy queries are related to linguistic summaries, and that one can introduce a hierarchy of protoforms, or abstract summaries in the sense of latest Zadeh’s (2002) ideas meant mainly for increasing deduction capabilities of search engines. We show an implementation for the summarization of Web server logs.