Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
This article is about analysis of data obtained in repeated measures designs in psycholinguistics and related disciplines with items (words) nested within treatment (5 type of words). Statistics tested in a series of computer simulations are:F 1,F 2,F 1 &F 2,F′, minF′, plus two decision procedures, the one suggested by Forster and Dickinson (1976) and one suggested by the authors of this article. The most common test statistic,F 1 &F 2, turns out to be wrong, but all alternative statistics suggested in the literature have problems too. The two decision procedures perform much better, especially the new one, because it systematically takes into account the subject by treatment interaction and the degree of word variability.
We modified the traditional (verbal) digit span task for administration via computer and the Internet. This online version collects data on the floor and ceiling of a subject’s span capacity, rather than generating a rough estimate of capacity based on a 50% success rate, as the traditional version does. We compared the two versions within adult subjects in two cohorts: college-age normal readers and college-age reading-disabled readers. To explore the reliability of the online version as a research tool, we employed the Bland-Altman approach to examine agreement between instruments. The online version yielded spans similar to those yielded by the traditional version, tending toward smaller values at the high end and larger values at the low end of span sizes, in the typical readers. It differentiated between the better and poorer readers reliably, and to the same extent as does the verbal version. The online version of the digit span task is comparable to the traditional version in assessing verbal span capacity; the code with which to implement the task is available at www.psychonomic.org/archive.
The Implicit Association Test (IAT; Greenwald, McGhee, & Schwartz, 1998) is one of the most widely used tools for assessing implicit attitudes. To date, most IAT experiments have been run using Inquisit, a PC-based program. In the present article, we describe a method for conducting IAT experiments using PsyScope, a free, downloadable, Macintosh-based program (see Bonatti, n.d., for the OS X version; Cohen, MacWhinney, Flatt, & Provost, 1993, for the OS 9 version). In addition, we explain how data can be imported into SPSS for analysis. Preliminary results indicate that, in comparison with the PC version of the IAT, the Macintosh version provides similar sensitivity in measuring implicit self-esteem. Our PsyScope script and SPSS syntax may be downloaded from www.psychonomic.org/archive.
Hearing loss is a confounding variable that is rarely addressed in behavioral research despite its prevalence across the life span. Currently, the most common method of experimental control over hearing acuity is through self report of perceived impairment. We argue that this technique may lack sensitivity and that researchers should more commonly utilize standardized hearing screening procedures. Distinctive patterns of hearing loss are reviewed with attention to populations that commonly participate in behavioral research. We explain standard techniques for conducting pure tone hearing screening using a conventional portable audiometer and outline a procedure for how researchers can modify a conventional laptop computer for audiometric screening when a standard audiometer is unavailable. We offer a sample hearing screening program that researchers may use toward the development of their own protocol. This program is freely available for download at www psychonomic.org/archive.
Query expansion(QE) has been proved to be one of effective methods for improving the performance of the information retrieval(IR) system.Therefore,a new fuzzy QE method based on synonymy thesaurus is proposed,and the synonymy thesaurus is built based on the famous lexical database WordNet.In the synonymy thesaurus,the similarity between the synonyms is,which is obtained by Tanimoto coefficient.By using this synonymy thesaurus,query expansion can be done well.Then the fuzzy QE method is introduced into the document information retrieval system together with the modified vector space model.The experimental results show that the developed information retrieval system has got more effective performance than before by using the fuzzy query expansion method.One feature of the proposed information retrieval model is that it can be treated as one of simple semantic models.Another feature is that the expansion degree is controllable based on different thresholds.
A lack of surveillance system infrastructure in the Asia-Pacific region is seen as hindering the global control of rapidly spreading infectious diseases such as the recent avian H5N1 epidemic. As part of improving surveillance in the region, the BioCaster project aims to develop a system based on text mining for automatically monitoring Internet news and other online sources in several regional languages. At the heart of the system is an application ontology which serves the dual purpose of enabling advanced searches on the mined facts and of allowing the system to make intelligent inferences for assessing the priority of events. However, it became clear early on in the project that existing classification schemes did not have the necessary language coverage or semantic specificity for our needs. In this article we present an overview of our needs and explore in detail the rationale and methods for developing a new conceptual structure and multilingual terminological resource that focusses on priority pathogens and the diseases they cause. The ontology is made freely available as an online database and downloadable OWL file.
In this paper, we report on the role of the Urdu grammar in the Parallel Grammar (ParGram) project (Butt, M., King, T. H., Niño, M.-E., & Segond, F. (1999). A grammar writer’s cookbook. CSLI Publications; Butt, M., Dyvik, H., King, T. H., Masuichi, H., & Rohrer, C. (2002). ‘The parallel grammar project’. In: Proceedings of COLING 2002, Workshop on grammar engineering and evaluation, pp. 1–7). The Urdu grammar was able to take advantage of standards in analyses set by the original grammars in order to speed development. However, novel constructions, such as correlatives and extensive complex predicates, resulted in expansions of the analysis feature space as well as extensions to the underlying parsing platform. These improvements are now available to all the project grammars.
The empirical investigation of human gesture stands at the center of multiple research disciplines, and various gesture annotation schemes exist, with varying degrees of precision and required annotation effort. We present a gesture annotation scheme for the specific purpose of automatically generating and animating character-specific hand/arm gestures, but with potential general value. We focus on how to capture temporal structure and locational information with relatively little annotation effort. The scheme is evaluated in terms of how accurately it captures the original gestures by re-creating those gestures on an animated character using the annotated data. This paper presents our scheme in detail and compares it to other approaches.
This work is about multimodal and expressive synthesis on virtual agents, based on the analysis of actions performed by human users. As input we consider the image sequence of the recorded human behavior. Computer vision and image processing techniques are incorporated in order to detect cues needed for expressivity features extraction. The multimodality of the approach lies in the fact that both facial and gestural aspects of the user’s behavior are analyzed and processed. The mimicry consists of perception, interpretation, planning and animation of the expressions shown by the human, resulting not in an exact duplicate rather than an expressive model of the user’s original behavior.
Word-fragment completion is a frequently used test in implicit memory research. A database of 196 Spanish fragments was recently published (Dasí, Soler, & Ruiz, 2004) in which the fragments were described for indices, such as difficulty, familiarity, frequency, number of meanings, and so on (www.psychonomic.org/archive). In this work, a new index, thepriming index, is described for the same 196 fragments. This index is calculated for each fragment by subtracting the difficulty index (the proportion of correct completion when the fragment is not studied) from the proportion of correct completion when the fragment is studied, and it means the capacity of an item to be primed. In order to determine whether the new index performs well, we have replicated some important experimental effects, such as priming, frequency, and difficulty. We consider that this index can help researchers to better select stimuli for fragment-completion tasks.
The aim of this paper is to investigate whether a treebank grammar can be used to automatically classify and annotate German phrases contained in a MT lexicon. Phrases from the lexicon appear in their citation form and may differ structurally from the phrase tokens found in the corpus. We describe the grammar extraction process for a formalism called Tree-Generating Binary Grammar and evaluate the performance of subsets of the obtained grammar on a set of four types of lexical phrases.
We argue in favor of the the use of labeled directed graph to represent various types of linguistic structures, and illustrate how this allows one to view NLP tasks as graph transformations. We present a general method for learning such transformations from an annotated corpus and describe experiments with two applications of the method: identification of non-local depenencies (using Penn Treebank data) and semantic role labeling (using Proposition
This paper investigates new design options for the feature space of a dependency parser. We focus on one of the simplest and most efficient architectures, based on a deterministic shift-reduce algorithm, trained with the perceptron. By adopting second-order feature maps, the primal form of the perceptron produces models with comparable accuracy to more complex architectures, with no need for approximations. Further gains in accuracy are obtained by designing features for parsing extracted from semantic annotations generated by a tagger. We provide experimental evaluations on the Penn Treebank.
This paper proposes a novel Chinese syntactic parsing model based on semantic class, which is a variant of normal lexicalized statistical model. It attempts to make use of the syntactic and semantic similarity between Chinese words and then produces a more knowledgeable estimate of the probability of grammar rules. A simple but effective unsupervised method is designed to determine the proper semantic class of given words. Semantic class is used to improve the performance of parsing model. We evaluate our methods on the widely used Penn Chinese Treebank. Experimental results show that it outperforms a famous lexicalized model significantly on appropriate semantic class levels.
REVIEWS 781 Hedin, Tora. Changing Identities:Language Variation on Czech Television. Acta universitatis Stockholmeinsis. Stockholm Slavic Studies, 29. Stockholm University, Stockholm, 2005. xv + 219 pp. Tables. Figures. Illustrations. Notes. Bibliography. Appendix. Index. SEK 282.00 (paperback). This is a published version of the author's doctoral thesis,written inEnglish. The study examines indetail language variation inCzech television discourse, primarily between January 1997 and September 2003. It is a thoroughly researched and well informed work, with reference to all the important corpus-based studies of the spoken language? Kucera, Kravcisinova & Bednaf ova, and Bayer and Maglione, as well as toGammelgaard and Bermel (who both draw on dialogue in literary texts)? and to over seventy television broadcasts. The quantitative analysis is confined to a more manageable corpus of fifteen programmes, comprising 24,000 words, which are aimed at a range of different audiences and contain a mixture of prepared and unpre pared speech. Most of the examples cited have been well documented and analysed in other linguistic environments elsewhere, but comparatively little attention has hitherto been paid to themilieu of Czech television. The chapter on the historical background to theCzech language situation and the linguistic debate over the coexistence of Standard Czech and Common Czech (and their variant and intermediate forms) serves to contextualize the discussion of the empirical data. Chapter three provides a concise overview of language and the mass media, with some particularly helpful comments for the lesswell initiated on concepts such as the facade, the conversational framework and television as flow and image. It is to the author's advantage here that she is able to draw on a number of important studies in Scandinavian languages which are linguistically inaccessible tomost English- and Czech-speaking scholars. While the main part of the study employs a quantitative synchronic approach, the most engaging and thought provoking section of the book is arguably the qualitative diachronic analysis presented in chapter six. The explanation and exemplification of differences in language use in Czech television discourse before and after 1989 make for especially interesting reading and would provide the basis for a more detailed investigation in their own right.The final chapter on code-switching is similarly worthwhile, although this reviewer would have liked to see greater consideration given to the linguistic constraints on different types of code-mixing (both intrasentential and intersentential). The difficultyfor the author of thisbook is the sheer breadth of the subject matter and the range of the variables affecting television participants' choice of language variety.While a macrolinguistic study of this typemay provide a sound statistical basis for frequency-based analysis of usage, it can at best only offer tentative generalized suggestions as to the role played by the inter relationship between social, geographical, cultural and personal factors, variables such as age, professional status and gender, and the specific context of each television programme. It isperhaps a shame that the author did not correlate her main findings with the data from the Prague Spoken Corpus and the Brno Spoken Corpus of theCzech National Corpus with a view to defining more clearly any regional differences in usage. The fact thatCzech 782 SEER, 85, 4, OCTOBER 2007 television is largely Prague-based may have a more profound bearing on the use of language than has generally been appreciated. Not surprisingly, this study has many of the strengths and some of the weaknesses of a typical doctoral thesis. It offers a comprehensive summary and evaluation of existing research and provides very useful cross-references. It also highlights the complexity of language usage in a linguistic settingwhere stylisticallyand functionally divergent forms coexist, and where theprestigious 'standard' variant is not the spoken norm. Most importantly, it offers new statistical information to add to the existing body of data on morphological, phonological and lexical variation, and to substantiate claims that language choice always depends to a significant extent on the purpose of the dialogue and the formality of the situation.However, minor problems with editing and proof-reading detract from the overall quality of thework. Furthermore, the selection of television broadcasts inevitably contains a degree of subjectivity and is not indicative of the speech of the population as a whole. Finally, itwould appear that a...
This paper describes a method for automatic acquisition of wide-coverage treebank-based deep linguistic resources for Japanese, as part of a project on treebankbased induction of multilingual resources in the framework of Lexical-Functional Grammar (LFG). We automatically annotate LFG f-structure functional equations (i.e. labelled dependencies) to the Kyoto Text Corpus version 4.0 (KTC4) (Kurohashi and Nagao 1997) and the output of of Kurohashi-Nagao Parser (KNP) (Kurohashi and Nagao 1998), a dependency parser for Japanese. The original KTC4 and KNP provide unlabelled dependencies. Our method also includes zero pronoun identification. The performance of the f-structure annotation algorithm with zero-pronoun identification for KTC4 is evaluated against a manually-corrected Gold Standard of 500 sentences randomly chosen from KTC4 and results in a pred-only dependency f-score of 94.72%. The parsing experiments on KNP output yield a pred-only dependency f-score of 82.08%.
A method of data collection is presented that unites the efficiency of mass testing with the ease of instant electronic data collection that is typical of computer-based experiments run on individual participants. A wireless response system (WRS), originally designed as a teaching tool, is used to replicate three classic and robust effects from the memory literature (effects of false memory, levels of processing, and word frequency). It is shown that for these types of experimental designs, data can be collected more efficiently (in both time and effort) with the WRS method than through traditional mass- and individual-testing methods alone. The advantages and limitations of WRSs for use in mass electronic data collection are discussed.
MLR, I02.3, 2007 849 them to a traditional published form as was done inEngler's critical edition. This gives readers amuch more immediate sense of connection with Saussure's thought, as well as providing some inkling ofwhat it is like towork with theoriginal manuscript materials. The textproper ends on page 240, with 87 pages given over to the bibliography of secondary literatureon Saussure since I970. It isa very fulland useful listingcovering awide range ofworks coming from linguists and literarycritics, and while some omis sions are inevitable, nothing can detract from the fact that this volume marks a true watershed in thedevelopment of Saussurean studies in theEnglish-speaking world. UNIVERSITY OF EDINBURGH JOHN E. JOSEPH La Langue, lestyle,lesens: etudesoffertes aAnne-Marie Garagnon. By CLAIRE BADIOU MONFERRAN, FRiED1ERIC CALAS, JULIEN PIAT, and CHRISTELLE REGGIANI. Paris: L'Improviste. 2005. 383pp. E28. ISBN978-2-913764-27-9. The titleof thevolume, a collection of articles inhonour ofAnne-Marie Garagnon, indicates the threemain themes explored. There are fourmain sections which reflect the interestsboth of the dedicatee and of her colleagues and formerpupils. The first section, entitled 'La langue entre histoire et systeme', isparticularly rich and inter esting. It examines what the editors call 'l'heritage conceptuel etmethodologique du xxe siecle' (p. 7), opening up interesting questions in the history of French linguis tic thought and of the French language, and in linguistic theory and ideologies of language. The second part, 'Faits de langue et faitsde style', includes articles which discuss and analyse particular lexical and semantic features of the French language from the Renaissance to the present day. The third part, 'Effets de style et effets de discours', offers insights into current trends in stylistic analysis, which inFrance is particularly marked by theories of 'enonciation'. Notable in the final part, en titled 'Stylistique et hermeneutique des formes', are the articles byGeorges Molinie and Thomas Clerc, which aim to present the epistemological basis ofAnne-Marie Garagnon's stylistic approach; other contributors explore certain stylisticmotifs in literary texts ranging fromLa Princesse de Cleves to Proust. It is impossible in the space of a review to do justice to all the articles and I propose to focus on a num ber of papers which seem to reflectwell the overall quality of the volume. Gilles Siouffi's excellent article, for example, reviews the notion of 'langue classique' and considers how a stylistic and grammatical analysis of its features can enable us better tounderstand it.He challenges the paradoxical notion of classical French as at once forming the norm forcontemporary French and being considered as 'radicalement autre' (p. i). A number of contributions treat the analysis of proper names fromdif ferentperspectives. Delphine Denis provides an interesting discussion of the status of the proper name in the seventeenth century. She shows how in this period there gradually emerges a theoryof proper nouns 'plus ou moins rapprochee du nom com mun' (p. 33). She notes how thegrammars and volumes ofobservations on theFrench language elaborate an opposition between common and proper nouns, but, with the possible exception of thePort-Royal logic,never discuss thecomplex nature ofproper nouns or the nature of theirmeaning. Adopting a stylistic approach, Van Dung Le Flanchec uses an analysis ofMaurice Sceve's 'Mon Orphee' toquestion modern ana lyses of proper names in terms of antonomasia, which should in her view rather be considered a stylisticquestion. Adopting a similar stylistic approach, Claire Badiou Monferran looks at different types of anaphora in a corpus of seventeenth-century nouvelles galantes; once again she sees the repeated use of anaphoric demonstrative pronouns as being a stylistic traitrather than an inherent feature of theFrench of the 85o Reviews period. Notable among the articleswhich adopt an approach associated with theories of 'enonciation' is the article by Frederic Calas. Starting froman analysis of an extract from Marivaux's Yournaux and fromLesage's Gil Blas de Santillane, he demonstrates thatepanorthosis isnot simply a rhetorical figurebut also gives thediscourse an ironic effect,therebycriticizing 'l'hypocrisiemondaine' (p. 247). In short, thevolume offers insights into a wide range of perspectives which are currently being applied to the study of French language and literature.What theyperhaps all have in common is the desire to reflecton 'legeste hermeneutique' (p. i). UNIVERSITY...
Understanding user interests from text documents can provide support to personalized information recommendation services. Typically, these services automatically infer the user profile, a structured model of the user interests, from documents that were already deemed relevant by the user. Traditional keyword-based approaches are unable to capture the semantics of the user interests. This work proposes the integration of linguistic knowledge in the process of learning semantic user profiles that capture concepts concerning user interests. The proposed strategy consists of two steps. The first one is based on a word sense disambiguation technique that exploits the lexical database WordNet to select, among all the possible meanings (senses) of a polysemous word, the correct one. In the second step, a naïve Bayes approach learns semantic sensebased user profiles as binary text classifiers (userlikes and user-dislikes) from disambiguated documents. Experiments have been conducted to compare the performance obtained by keyword-based profiles to that obtained by sense-based profiles. Both the classification accuracy and the effectiveness of the ranking imposed by the two different kinds of profile on the documents to be recommended have been considered. The main outcome is that the classification accuracy is increased with no improvement on the ranking. The conclusion is that the integration of linguistic knowledge in the learning process improves the classification of those documents whose classification score is close to the likes / dislikes threshold (the items for which the classification is highly uncertain). 1
We present two methods to address the problem of sparsity in the FrameNet lexical database. The first method is based on the idea that a word that belongs to a frame is ``similar'' to the other words in that frame. We measure the similarity using a WordNet-based variant of the Lesk metric. The second method uses the sequence of synsets in WordNet hypernym trees as feature vectors that can be used to train a classifier to determine whether a word belongs to a frame or not. The extended dictionary produced by the second method was used in a system for FrameNet-based semantic analysis and gave an improvement in recall. We believe that the methods are useful for bootstrapping FrameNets for new languages.
This paper presents recent extensions to Poliqarp, an open source tool for indexing and searching morphosyntactically annotated corpora, which turn it into a tool for indexing and searching certain kinds of treebanks, complementary to existing treebank search engines. In particular, the paper discusses the motivation for such a new tool, the extended query syntax of Poliqarp and implementation and efficiency issues.
This paper reports about our efforts in creating a tri-lingual parallel treebank. The focal points are consistency checking and all aspects of sub-sentential alignment. We discuss the alignment guidelines, the importance of quality checks, and special alignment problems. Then we look at alignment algorithms and alignment visualization tools and we compare our own TreeAligner with other alignment tools. Our constituent structure treebanks contain just over 1,000 sentences and around 18,000 tokens in each language.
The linguistic quality of a parallel treebank depends crucially on the parallelism between the source and target language annotations. We propose a linguistic notion of translation units and a quantitative measure of parallelism for parallel dependency treebanks, and demonstrate how the proposed translation units and parallelism measure can be used to compute transfer rules, spot annotation errors, and compare different annotation schemes with respect to each other. The proposal is evaluated on the 100,000 word Copenhagen Danish-English Dependency Treebank.
Recent studies focussed on the question whether less-congurational languages like German are harder to parse than English, or whether the lower parsing scores are an \nartefact of treebank encoding schemes and data structures, as claimed by K¨ubler et al. (2006). This claim is based on the assumption that PARSEVAL metrics fully reflect parse quality across treebank encoding schemes. In this paper we present new experiments to test this claim. We use the \nPARSEVAL metric, the Leaf-Ancestor metric as well as a dependency-based evaluation, and present novel approaches measuring the effect of controlled error insertion on treebank trees and parser output. We also provide extensive past-parsing crosstreebank conversion. The results of the experiments show that, contrary to K¨ubler et al. (2006), the question whether or not German is harder to parse than English remains undecided.
This paper presents a method to automatically acquire wide-coverage, robust, probabilistic Lexical-Functional Grammar resources for Chinese from the Penn Chinese Treebank (CTB). Our starting point is the earlier, proofof- \nconcept work of (Burke et al., 2004) on automatic f-structure annotation, LFG grammar acquisition and parsing for Chinese using the CTB version 2 (CTB2). We substantially extend and improve on this earlier research as regards coverage, robustness, quality and fine-grainedness of the resulting LFG resources. We achieve this through (i) improved LFG analyses for a number of core Chinese phenomena; (ii) a new automatic f-structure annotation architecture which involves an intermediate dependency representation; (iii) scaling the approach from 4.1K trees in CTB2 to 18.8K trees in CTB version 5.1 (CTB5.1) and (iv) developing a novel treebank-based approach to recovering non-local dependencies (NLDs) for Chinese parser output. Against a new 200-sentence good standard of manually constructed f-structures, the method achieves 96.00% f-score for f-structures automatically generated for the original CTB trees and 80.01%for NLD-recovered f-structures generated for the trees output by Bikel’s parser.
This paper describes how a treebank of ungrammatical \nsentences can be created from a treebank of well-formed sentences. The treebank creation procedure involves the automatic introduction of frequently occurring grammatical errors into the sentences in an existing treebank, and the minimal transformation of the analyses in the treebank so \nthat they describe the newly created ill-formed sentences. \nSuch a treebank can be used to test how well a parser is able to ignore grammatical errors in texts (as people can), and can be used to induce a grammar capable of analysing such sentences. This paper also demonstrates the first of these uses.
In this paper we describe the current state of a new Japanese lexical resource: the Hinoki treebank. The treebank is built from dictionary definitions, examples and news text, and uses an HPSG based Japanese grammar to encode both syntactic and semantic information. It is combined with an ontology based on the definition sentences to give a detailed sense level description of the most familiar 28,000 words of Japanese.
Whilst the degree to which a treebank subscribes to a specific linguistic theory limits the usefulness of the resource, the availability of more formats for the same resource plays a crucial role both in NLP and linguistics. Conversion tools and multi-format treebanks are useful for investigating portability of NLP systems and validity of annotation. Unfortunately, conversion is a quite complex task since it involves grammatical rules and linguistic knowledge to be incorporated into the converter program.
We present the Modified French Treebank (MFT), a completely revamped French Treebank, derived from the Paris 7 Tree-bank (P7T), which is cleaner, more co-herent, has several transformed structures, and introduces new linguistic analyses. To determine the effect of these changes, we investigate how theMFT fares in statistical parsing. Probabilistic parsers trained on the MFT training set (currently 3800 trees) already perform better than their counter-parts trained on five times the P7T data (18,548 trees), providing an extreme ex-ample of the importance of data quality over quantity in statistical parsing. More-over, regression analysis on the learning curve of parsers trained on the MFT lead to the prediction that parsers trained on the full projected 18,548 tree MFT training set will far outscore their counterparts trained on the full P7T. These analyses also show how problematic data can lead to problem-atic conclusions–in particular, we find that lexicalisation in the probabilistic parsing of French is probably not as crucial as was once thought (Arun and Keller (2005)).
Proceedings of the Sixth International Workshop on Treebanks and \nLinguistic Theories. \nEditors: Koenraad De Smedt, Jan Hajič and Sandra Kübler. \nNEALT Proceedings Series, Vol. 1 (2007), 1-6. \n© 2007 The editors and contributors. \nPublished by \nNorthern European Association for Language \nTechnology (NEALT) \nhttp://omilia.uio.no/nealt. \nElectronically published at \nTartu University Library (Estonia) \nhttp://hdl.handle.net/10062/4476.
International audience
The authors examined the relationship of hostility with (a) affective ratings of pictures and (b) state affects evoked by task-induced stress in 95 healthy men and women 22-37 years of age. Pictures were from the International Affective Picture System (IAPS; P. J. Lang, M. M. Bradley, & B. N. Cuthbert, 1999). Stressors included a startle task, mental arithmetic task, and choice-deadline reaction time task. The circumplex model of affect was used to structure the self-reported state affects. The authors found that hostility was associated with displeasure, high arousal, and low dominance ratings of IAPS pictures. Hostility was related to unpleasant affect and unactivated unpleasant affect during the experiment, and subscale paranoia was related to activated unpleasant affect. Findings suggest that participants scoring high on hostility are prone to negative emotional reactions.
Proceedings of the Sixth International Workshop on Treebanks and \nLinguistic Theories. \nEditors: Koenraad De Smedt, Jan Hajič and Sandra Kübler. \nNEALT Proceedings Series, Vol. 1 (2007), 163-174. \n© 2007 The editors and contributors. \nPublished by \nNorthern European Association for Language \nTechnology (NEALT) \nhttp://omilia.uio.no/nealt. \nElectronically published at \nTartu University Library (Estonia) \nhttp://hdl.handle.net/10062/4476.
Parsing unrestricted text is useful for many language technology applications but requires parsing methods that are both robust and efficient. MaltParser is a language-independent system for data-driven dependency parsing that can be used to induce a parser for a new language from a treebank sample in a simple yet flexible manner. Experimental evaluation confirms that MaltParser can achieve robust, efficient and accurate parsing for a wide range of languages without language-specific enhancements and with rather limited amounts of training data.
Abstract In this chapter, we discuss the development and use of picture stimuli incorporated in the International Affective Picture System (IAPS, pronounced “eye-aps”; Lang, Bradley, & Cuthbert, 2005), a large set of emotionally evocative color photographs that includes pleasure, arousal, and dominance ratings made by men and women. The IAPS is currently used in experimental investigations of emotion and attention worldwide, providing experimental control in the selection of emotional stimuli, facilitating the comparison of results across different studies, and encouraging replication within and across psychological and neuroscience research laboratories. Numerous studies in our laboratory over the past 15 years have explored subjective, psychophysiological, behavioral, and neurophysiological reactions when viewing these affective stimuli. Basic findings from these studies, which will be informative for researchers considering or using the IAPS stimuli, are briefly summarized in this chapter.
The interpretation of nominal compounds is one of the most difficult problems in natural language processing. This paper proposes a new model for the automatic classification of four coarse-grained semantic relations involved in Chinese compound nominalizations. In such a model, for a compound nominalization, its paraphrased syntactic role occurrences (PSRO) in a treebank are exploited to form feature vectors for supervised classifiers. To solve the problem of data sparseness, the World Wide Web is used to discover relational clusters and such clusters are employed to produce smoothed PSRO feature vectors for the compound nominalizations. The experimental results show that such a method is very effective.
The dependency relation is the most essential ingredient in a dependency-based theory of syntax. This paper presents some statistical findings on the dependency relation extracted from a Chinese dependency treebank. A sentence in the proposed treebank can easily be converted into a SSyntS graph in Meaning-Text Theory. The statistics on the dependency relation show that modifiers make up 55% of all dependencies and actants have a lower proportion of 45%. The paper demonstrates it is possible to extract from the treebank active and passive valence information of a word (or word class). The paper gives a formula to calculate the mean dependency distance (MDD) for a specific type of dependency relation in a language and obtains MDD of all dependency types in Chinese. These figures show that some dependencies tend to be much farther apart than others, and demonstrate that dependency distance tends to minimization and different dependency types have varying preference on the direction of dependency.
We present several improvements to unlexicalized parsing with hierarchically state-split PCFGs. First, we present a novel coarse-to-fine method in which a grammar’s own hierarchical projections are used for incremental pruning, including a method for efficiently computing projections of a grammar without a treebank. In our experiments, hierarchical pruning greatly accelerates parsing with no loss in empirical accuracy. Second, we compare various inference procedures for state-split PCFGs from the standpoint of risk minimization, paying particular attention to their practical tradeoffs. Finally, we present multilingual experiments which show that parsing with hierarchical state-splitting is fast and accurate in multiple languages and domains, even without any language-specific tuning. 1