Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Proceedings of the Ninth International Workshop \non Treebanks and Linguistic Theories. \nEditors: Markus Dickinson, Kaili Müürisep and Marco Passarotti. \nNEALT Proceedings Series, Vol. 9 (2010), 245-256. \n© 2010 The editors and contributors. \nPublished by \nNorthern European Association for Language \nTechnology (NEALT) \nhttp://omilia.uio.no/nealt. \nElectronically published at \nTartu University Library (Estonia) \nhttp://hdl.handle.net/10062/15891.
Excessive or addictive Internet use can be linked to different online activities, such as Internet gaming or cybersex. The usage of Internet pornography sites is one important facet of online sexual activity. The aim of the present work was to examine potential predictors of a tendency toward cybersex addiction in terms of subjective complaints in everyday life due to online sexual activities. We focused on the subjective evaluation of Internet pornographic material with respect to sexual arousal and emotional valence, as well as on psychological symptoms as potential predictors. We examined 89 heterosexual, male participants with an experimental task assessing subjective sexual arousal and emotional valence of Internet pornographic pictures. The Internet Addiction Test (IAT) and a modified version of the IAT for online sexual activities (IATsex), as well as several further questionnaires measuring psychological symptoms and facets of personality were also administered to the participants. Results indicate that self-reported problems in daily life linked to online sexual activities were predicted by subjective sexual arousal ratings of the pornographic material, global severity of psychological symptoms, and the number of sex applications used when being on Internet sex sites in daily life, while the time spent on Internet sex sites (minutes per day) did not significantly contribute to explanation of variance in IATsex score. Personality facets were not significantly correlated with the IATsex score. The study demonstrates the important role of subjective arousal and psychological symptoms as potential correlates of development or maintenance of excessive online sexual activity.
This article gives a survey of the main issues confronting the compilers of monolingual dictionaries in the age of the Internet. Among others, it discusses the relationship between a lexical database and a monolingual dictionary, the role of corpus evidence, historical principles in lexicography vs. synchronic principles, the instability of word meaning, the need for full vocabulary coverage, principles of definition writing, the role of dictionaries in society, and the need for dictionaries to give guidance on matters of disputed word usage. It concludes with some questions about the future of dictionary publishing. Keywords: Monolingual Dictionaries, Lexical Database, Dictionary Structure, Word Meaning, Meaning Change, Usage, Usage Notes, Historical Principles Of Lexicography, Synchronic Principles Of Lexicography, Register, Slang, Standard English, Vocabulary Coverage, Consistency Of Sets, Phraseology, Syntagmatic Patterns, Problems Of Compositionality, Linguistic Prescriptivism, Lexical Evidence **This article is an edited version of a plenary address delivered at the conference on 'Dictionaries, More Than Words', which took place at the Faculty of Social Sciences, University of Ljubljana, Ljubljana, Slovenia, 6 February 2009.
This paper explores whether and how a firm should adapt its strategy in view of consumer use of prior customer ratings. Specifically, we consider optimal pricing and whether the firm should offer an unexpected frill to early customers to enhance their product experiences. We show that if price history is unobserved by consumers, a forward-looking firm should always modify its strategy from single-period optimal one, but it may be optimal to do so by lowering price, by lowering price and offering frills, or by raising price and offering frills, depending on the market growth rate. Specifically, the last strategy becomes optimal when market growth rate is high enough. The results are similar when the price history is observed by consumers, except that no deviation from single-period profit maximization choices is optimal when market growth is low enough. We also analyze whether the firm should prefer that the price information be stated in or left out of consumer reviews. In addition, in considering the effects of consumer heterogeneity, we conclude that the optimal firm's effort to affect ratings is higher when the idiosyncratic part of consumer uncertainty is larger.
The paper describes an approach to expedite the process of manual annotation of a Hindi dependency treebank. We propose a way by which consistency among a set of manual annotators could be improved. Furthermore, we show that our setup can also prove useful for evaluating when an inexperienced annotator is ready to start participating in the production of the treebank. We test our approach on sample sets of data obtained from an ongoing work on creation of this treebank. The results asserting our proposal are reported in this paper. 1.
Proceedings of the Ninth International Workshop \non Treebanks and Linguistic Theories. \nEditors: Markus Dickinson, Kaili Müürisep and Marco Passarotti. \nNEALT Proceedings Series, Vol. 9 (2010), 151-162. \n© 2010 The editors and contributors. \nPublished by \nNorthern European Association for Language \nTechnology (NEALT) \nhttp://omilia.uio.no/nealt. \nElectronically published at \nTartu University Library (Estonia) \nhttp://hdl.handle.net/10062/15891.
Proceedings of the Ninth International Workshop \non Treebanks and Linguistic Theories. \nEditors: Markus Dickinson, Kaili Müürisep and Marco Passarotti. \nNEALT Proceedings Series, Vol. 9 (2010), 233-244. \n© 2010 The editors and contributors. \nPublished by \nNorthern European Association for Language \nTechnology (NEALT) \nhttp://omilia.uio.no/nealt. \nElectronically published at \nTartu University Library (Estonia) \nhttp://hdl.handle.net/10062/15891.
Language users are increasingly turning to electronic resources to address their lexical information needs, due to their convenience and their ability to simultaneously capture different facets of lexical knowledge in a single interface. In this paper, we discuss techniques to respond to a user’s lexical queries by providing multilingual and multimodal information, and facilitating navigating along different types of links. To this end, structured information from sources like WordNet, Wikipedia, Wiktionary, as well as Web services is linked and integrated to provide a multi-faceted yet consistent response to user queries. The meanings of words in many different languages are characterized by mapping them to appropriate WordNet sense identifiers and adding multilingual gloss descriptions as well as example sentences. Relationships are derived from WordNet and Wiktionary to allow users to discover semantically related words, etymologically related words, alternative spellings, as well as misspellings. Last but not least, images, audio recordings, and geographical maps extracted from Wikipedia and Wiktionary allow for a multimodal experience. 1.
Several studies have investigated the neural responses triggered by emotional pictures, but the specificity of the involved structures such as the amygdala or the ventral striatum is still under debate. Furthermore, only few studies examined the association of stimuli's valence and arousal and the underlying brain responses. Therefore, we investigated brain responses with functional magnetic resonance imaging of 17 healthy participants to pleasant and unpleasant affective pictures and afterwards assessed ratings of valence and arousal. As expected, unpleasant pictures strongly activated the right and left amygdala, the right hippocampus, and the medial occipital lobe, whereas pleasant pictures elicited significant activations in left occipital regions, and in parts of the medial temporal lobe. The direct comparison of unpleasant and pleasant pictures, which were comparable in arousal clearly indicated stronger amygdala activation in response to the unpleasant pictures. Most important, correlational analyses revealed on the one hand that the arousal of unpleasant pictures was significantly associated with activations in the right amygdala and the left caudate body. On the other hand, valence of pleasant pictures was significantly correlated with activations in the right caudate head, extending to the nucleus accumbens (NAcc) and the left dorsolateral prefrontal cortex. These findings support the notion that the amygdala is primarily involved in processing of unpleasant stimuli, particularly to more arousing unpleasant stimuli. Reward-related structures like the caudate and NAcc primarily respond to pleasant stimuli, the stronger the more positive the valence of these stimuli is.
In the absence of context, the process of listening to acoustic scenes results in deriving an explicit semantic description and an implicit assessment of its acoustic properties in terms of its affective value. In this work, we mainly exploit the relationship between context-free associations of audio clips containing unconstrained acoustic sources with their affective values for clustering. Using over two hundred clips from the BBC sound effects library, we present a novel, quantitative method to compare the clusters of audio clips obtained using its context-free description with the clusters obtained from their affective measures; namely valence, arousal and dominance. Our results indicate that comparing clusters across representations is a suitable approach to determine an appropriate number of clusters to index audio clips in an un-supervised manner. In this paper we present our findings and examples of the resulting clusters of audio clips.
The aim of this paper is twofold. We focus, on the one hand, on the task of dynamically annotating English compound nouns, and on the other hand we propose disambiguation methods and techniques which facilitate the annotation task. Both the aforementioned are part of a larger on-going effort which aims to create HPSG annotation for the texts from the Wall Street Journal (henceforward WSJ) sections of the Penn Treebank (henceforward PTB) with the help of a hand-written large-scale and wide-coverage grammar of English, the English Resource Grammar (henceforward ERG; Flickinger (2002)). As we show in this paper, such annotations are very rich linguistically, since apart from syntax they also incorporate semantics, which does not only ensure that the treebank is guaranteed to be a truly sharable, re-usable and multi-functional linguistic resource, but also calls for the necessity of a better disambiguation of the internal (syntactic) structure of larger units of words, such as compound nouns, since this has an impact on the representation of their meaning, which is of utmost interest if the linguistic annotation of a given corpus is to be further understood as the practice of adding interpretative linguistic information of the highest quality in order to give “added value ” to the corpus. 1.
In this paper, we present an on-going project aiming at extending the WordNet lexical database by encoding common sense featural knowledge elicited from language speakers. Such extension of WordNet is required in the framework of the STaRS.sys project, which has the goal of building tools for supporting the speech therapist during the preparation of exercises to be submitted to aphasic patients for rehabilitation purposes. We review some preliminary results and illustrate what extensions of the existing WordNet model are needed to accommodate for the encoding of commonsense (featural) knowledge. 1
Comparing with the traditional way of manually developing grammar based on linguistic theory, corpus-oriented grammar development is more promising. To develop HPSG grammar through the corpus-oriented way, a treebank is an indispensable part. This paper first compares existing Chinese treebanks and chooses one of them as the basic resource for HPSG grammar development. Then it proposes a new design of part-of-speech tags based on the assumption that it is not only simple enough to reduce ambiguity of morphological analysis as much as possible, but also rich enough for HPSG grammar development. Finally, it introduces some on-going work about utilizing a Chinese scientific paper treebank in HPSG grammar development.
Co-constructing communicative effectiveness is often challenging in English as a lingua franca (ELF): speakers have considerably less to go on in terms of shared expectations of cultural knowledge and linguistic norms. A university environment provides a convenient backdrop for sharing at least academic conventions – although these vary more than might be surmised from the uniform labelling of such event types. This paper looks into some discourse and lexicogrammatical features in academic ELF, using ELFA as the database. The data consists of spoken language, which provides direct access to the ways in which meanings are negotiated in ongoing discourse, and the speech events are typically polylogic. ELF discourse requires close cooperation from the participants, which is reflected in its enhanced explicitness among other things. The explicitation strategies speakers display facilitate mutual comprehensibility and contribute to social cohesion within the multi-participant groups. Such strategies also help overcome the potential problems participants might have in dealing with a variety of formal deviations from ordinary English as a native language (ENL). Most of the time ELF bears a very close resemblance to Standard English, but signs of incipient ELF-specific developments are also in evidence.
Proceedings of the Ninth International Workshop \non Treebanks and Linguistic Theories. \nEditors: Markus Dickinson, Kaili Müürisep and Marco Passarotti. \nNEALT Proceedings Series, Vol. 9 (2010), 5. \n© 2010 The editors and contributors. \nPublished by \nNorthern European Association for Language \nTechnology (NEALT) \nhttp://omilia.uio.no/nealt. \nElectronically published at \nTartu University Library (Estonia) \nhttp://hdl.handle.net/10062/15891
Theories proposing that how one thinks and feels is influenced by feedback from the body remain controversial. A central but untested prediction of many of these proposals is that how well individuals can perceive subtle bodily changes (interoception) determines the strength of the relationship between bodily reactions and cognitive-affective processing. In Study 1, we demonstrated that the more accurately participants could track their heartbeat, the stronger the observed link between their heart rate reactions and their subjective arousal (but not valence) ratings of emotional images. In Study 2, we found that increasing interoception ability either helped or hindered adaptive intuitive decision making, depending on whether the anticipatory bodily signals generated favored advantageous or disadvantageous choices. These findings identify both the generation and the perception of bodily responses as pivotal sources of variability in emotion experience and intuition, and offer strong supporting evidence for bodily feedback theories, suggesting that cognitive-affective processing does in significant part relate to "following the heart."
Parsing requires the quantitative information of grammatical functions of part of speec h.This paper studied grammatical functions of part of speech by using the Chinese Dependency Treebank based on the Probabilistic Valency Pattern Theor y.According to the frequency,we divided grammatical functions of verbs into principal function,secondary function and part functio n.From the aspect of quantitative analysis,we validated and complemented formers’ conclusion,which helped us to have a clearer understanding of grammatical functions of verb s.This paper is also a development of the Probabilistic Valency Pattern Theor y.
In this paper we describe the development of a schema for the annotation of attribution relations and present the first findings and some relevant issues concerning this phenomenon. Following the D-LTAG approach to discourse, we have developed a lexically anchored description of attribution, considering this relation, contrary to the approach in the PDTB, independently from other discourse relations. This approach has allowed us to deal with the phenomenon in a broader perspective than previous studies, reaching therefore a more accurate description of it and making it possible to raise some still unaddressed issues. Following this analysis, we propose an annotation schema and discuss the first results concerning its applicability. The schema has been applied to a pilot portion of the ISST corpus of Italian and represents the initial phase of a project aiming at the creation of an Italian Discourse Treebank. We believe this work will raise some awareness concerning the fundamental importance of attribution relations. The identification of the source has in fact strong implications for the attributed material. Moreover, it will make overt the complexity of a phenomenon for long underestimated. 1.
The Varro toolkit is a system for identifying and counting a major class of regularity in treebanks and annotated natural language data in the form of treestructures: frequently recurring unordered subtrees. This software has been designed for use in linguistics to be maximally applicable to actually existing treebanks and other stores of tree-structurable natural language data. It minimizes memory use so that moderately large treebanks are tractable on commonly available computer hardware. This article introduces condensed canonically ordered trees as a data structure for efficiently discovering frequently recurring unordered subtrees.
In this paper, we introduce our recent work on Chinese HPSG grammar development through treebank conversion. By manually defining grammatical constraints and anno-tation rules, we convert the bracketing trees in the Penn Chinese Treebank (CTB) to be an HPSG treebank. Then, a large-scale lexi-con is automatically extracted from the HPSG treebank. Experimental results on the CTB 6.0 show that a HPSG lexicon was successfully extracted with 97.24 % accu-racy; furthermore, the obtained lexicon achieved 98.51 % lexical coverage and 76.51 % sentential coverage for unseen text, which are comparable to the state-of-the-art works for English. 1
Parsing requires the quantitative information of syntactic functions of part of speech.Based on the Probabilistic Valency Pattern Theory,this paper studied syntactic functions of Chinese nouns by using the Chinese Dependency Treebank.According to the frequency,we divided syntactic functions of nouns into typical functions and atypical functions,and proposed the model of Correlated Markedness and probabilistic valency pattern of grammatical functions of nouns.From this quantitative analysis,we validated and complemented the conclusions of previous studies,which helped us to have a clearer understanding of syntactic functions of Chinese nouns and provided a reference for teaching Chinese as a second language.
The novel anthocyanins, malvidin 3-O-(6-O-(4-O-malonyl-alpha-rhamnopyranosyl)-beta-glucopyranoside)-5-O-beta-glucopyranoside (2), malvidin 3-O-(6-O-alpha-rhamnopyranosyl-beta-glucopyranoside)-5-O-(6-O-malonyl-beta-glucopyranoside) (3), malvidin 3-O-(6-O-(4-O-malonyl-alpha-rhamnopyranosyl)-beta-glucopyranoside)-5-O-(6-O-malonyl-beta-glucopyranoside) (4), malvidin 3-O-(6-O-(4-O-malonyl-alpha-rhamnopyranosyl)-beta-glucopyranoside) (5) and malvidin 3-O-(6-O-(Z)-p-coumaroyl-beta-glucopyranoside)-5-O-beta-glucopyranoside (6), in addition to the 3-O-(6-O-alpha-rhamnopyranosyl-beta-glucopyranoside)-5-O-beta-glucopyranoside (1) and the 3-O-(6-O-(E)-p-coumaroyl-beta-glucopyranoside)-5-O-beta-glucopyranoside (7) of malvidin have been isolated from purple leaves of Oxalis triangularis A. St.-Hil. In pigments 2, 4 and 5 a malonyl unit is linked to the rhamnose 4-position, which has not been reported previously for any anthocyanin before. The identifications were mainly based on 2D NMR spectroscopy and electrospray MS.
We present algorithms for higher-order dependency parsing that are “third-order” in the sense that they can evaluate substructures containing three dependencies, and “efficient ” in the sense that they require only O(n4) time. Importantly, our new parsers can utilize both sibling-style and grandchild-style interactions. We evaluate our parsers on the Penn Treebank and Prague Dependency Treebank, achieving unlabeled attachment scores of
Large-scale phrase structure treebank and dependency structure treebank are developed and interconverted for the purpose of syntactic analysis on true corpus. The head percolation table is constructed on modern Chinese dependency grammar by discussing the relationship between phrase structure and dependency structure based on Penn Chinese Treebank (CTB),and CTB from phrase structure is converted to dependency structure treebank using the head percolation table. 200 sentences are chosen from CTB randomly to evaluate the conversion performance. Precision of the conversion has attained 99.50%. The achieved dependency structure treebank can be used to analyze Chinese dependency relation.
In this paper, we offer broad insight into the underperformance of Arabic constituency parsing by analyzing the interplay of linguistic phenomena, annotation choices, and model design. First, we identify sources of syntactic ambiguity understudied in the existing parsing literature. Second, we show that although the Penn Arabic Treebank is similar to other treebanks in gross statistical terms, annotation consistency remains problematic. Third, we develop a human interpretable grammar that is competitive with a latent variable PCFG. Fourth, we show how to build better models for three different parsers. Finally, we show that in application settings, the absence of gold segmentation lowers parsing performance by 2–5 % F1. 1
Nouns are generally easier to learn than verbs (e.g., Bornstein, 2005; Bornstein et al., 2004; Gentner, 1982; Maguire, Hirsh-Pasek, & Golinkoff, 2006). Yet, verbs appear in children's earliest vocabularies, creating a seeming paradox. This paper examines one hypothesis about the difference between noun and verb acquisition. Perhaps the advantage nouns have is not a function of grammatical form class but rather related to a word's imageability. Here, word imageability ratings and form class (nouns and verbs) were correlated with age of acquisition according to the MacArthur-Bates Communicative Development Inventory (CDI) (Fenson et al., 1994). CDI age of acquisition was negatively correlated with words' imageability ratings. Further, a word's imageability contributes to the variance of the word's age of acquisition above and beyond form class, suggesting that at the beginning of word learning, imageability might be a driving factor.
A commonly held assumption is that processes underlying explicit and implicit memory are distinct. Recent evidence, however, suggests that they may interact more than previously believed. Using the remember-know procedure the current study examines the relation between recollection, a process thought to be exclusive to explicit memory, and performance on two implicit memory tasks, lexical decision and word stem completion. We found that, for both implicit tasks, words that were recollected were associated with greater priming effects than were words given a subsequent familiarity rating or words that had been studied but were not recognised (misses). Broadly, our results suggest that non-voluntary processes underlying explicit memory also benefit priming, a measure of implicit memory. More specifically, given that this benefit was due to a particular aspect of explicit memory (recollection), these results are consistent with some strength models of memory and with Moscovitch's (2008) proposal that recollection is a two-stage process, one rapid and unconscious and the other more effortful and conscious.
BACKGROUND: In this controlled postdiagnosis study, the authors examined various aspects of body image of breast cancer survivors in cross-sectional and longitudinal designs. METHODS: In 2004 and 2007 the Body Image Scale (BIS) was completed by the same 248 disease-free women who had been treated for stage II and III breast cancer between 1998 and 2002. "Poorer" body image was defined as greater than the 70th percentile (N=76 women) of the BIS scores in contrast to "better" body image (N=172 women). Breast cancer survivors were examined clinically in 2004, and their BIS scores were compared with the scores from an age-matched group of women from the general population. RESULTS: In this cross-sectional study, poorer body image in 2004 was associated significantly with modified radical mastectomy, undergoing or planning to undergo breast-reconstructive surgery, a change in clothing, poor physical and mental health, chronic fatigue, and reduced quality of life (QoL). In univariate analyses, most of these factors and manually planned radiotherapy were significant predictors of poorer body image in 2007. In multivariate analyses, manually planned radiotherapy, poor physical QoL and high BIS score in 2004 remained independent predictors of a poorer body image in 2007. Body image ratings were relatively stable from 2004 to 2007. Twenty-one percent of breast cancer survivors reported body image dissatisfaction, similar to the proportion of dissatisfaction in controls. CONCLUSIONS: In this cross-sectional analysis, body image in breast cancer survivors was associated with the types of surgery and radiotherapy and with mental distress, reduced health, and impaired QoL. Body image ratings were relatively stable over time, and the antecedent body image score was a strong predictor of body image at follow-up. Body image in breast cancer survivors differed very little from that in controls.
In this paper, we propose a novel selftraining strategy for parsing which is based on Treebank conversion (SSPTC). In SSPTC, we make full use of the strong points of Treebank conversion and self-training, and offset their weaknesses with each other. To provide good parse selection strategies which are needed in self-training, we score the automatically generated parse trees with parse trees in source Treebank as a reference. To maintain the constituency between source Treebank and conversion Treebank which is needed in Treebank conversion, we get the conversion trees with the help of self-training. In our experiments, SSPTC strategy is utilized to parse Tsinghua Chinese Treebank with the help of Penn Chinese Treebank. The results significantly outperform the baseline parser. 1
We investigate a number of approaches to generating Stanford Dependencies, a widely used semantically-oriented dependency representation. We examine algorithms specifically designed for dependency parsing (Nivre, Nivre Eager, Covington, Eisner, and RelEx) as well as dependencies extracted from constituent parse trees created by phrase structure parsers (Charniak, Charniak-Johnson, Bikel, Berkeley and Stanford). We found that phrase structure parsers systematically outperform algorithms designed specifically for dependency parsing. The most accurate method for generating dependencies is the Charniak-Johnson reranking parser, with 89 % (labeled) attachment F1 score. The fastest methods are Nivre, Nivre Eager, and Covington. When used with a linear classifier to make local parsing decisions, these methods can parse the entire Penn Treebank development set (section 22) in less than 10 seconds on an Intel Xeon E5520. However, this speed comes with a substantial drop in F1 score (about 76 % for labeled attachment) compared to competing methods. By tuning how much of the search space is explored by the Charniak-Johnson parser, we are able to arrive at a balanced configuration that is both fast and nearly as good as the most accurate approaches. 1.
Complications arise for standoff annotation when the annotation is not on the source text itself, but on a more abstract representation. This is particularly the case in a language such as Arabic with morphological and orthographic challenges, and we discuss various aspects of these issues in the context of the Arabic Treebank. The Standard Arabic Morphological Analyzer (SAMA) is closely integrated into the annotation workflow, as the basis for the abstraction between the explicit source text and the more abstract token representation. However, this integration with SAMA gives rise to various problems for the annotation workflow and for maintaining the link between the Treebank and SAMA. In this paper we discuss how we have overcome these problems with consistent and more precise categorization of all of the tokens for their relationship with SAMA. We also discuss how we have improved the creation of several distinct alternative forms of the tokens used in the syntactic trees. As a result, the Treebank provides a resource relating the different forms of the same underlying token with varying degrees of vocalization, in terms of how they relate (1) to each other, (2) to the syntactic structure, and (3) to the morphological analyzer. 1.
Semantic dependency analysis is practicable way to semantic analysis. This paper describes a Chinese semantic dependency analysis system using HowNet. The system takes sentences with phrase syntactic information as input. First, it determines the headword of each phrase to get the dependency structure of the sentence. Second, it takes the syntactic constituent as the basic unit of semantic labeling and determines the semantic relation through searching HowNet and Semantic Information Structure Library.We randomly extract 100 sentences from Penn Chinese Treebank as test data. There are totally 2783 pairs of word. The system determines the semantic relations of 2546 pairs.The labeling ratio is 91.5%.
This paper proposes a new cascade algorithm based on conditional random fields. The algorithm is applied to automatic recognition of Chinese verb-object collocation, and combined with a new sequence labeling of “ONIY”. Experiments compare identified results under two segmentations and part-of-speech tag sets. The comprehensive experimental results show that the best performance is 90.65% in F-score over Tsinghua Treebank, and 82.00% in F-score over the segmentation and part-of-speech tagging scheme of Peking University. Our experiments show that the proposed algorithm can greatly improve recognition accuracy of multi-nested collocation, and play a positive role on long distance collocation.
Ott, N. & R. Ziai (2010). Evaluating dependency parsing performance on german learner language. In M. Dickinson, K. Müürisep & M. Passarotti (eds.), Proceedings of the Ninth International Workshop on Treebanks and Linguistic Theories. Vol. 9 of NEALT Proceeding Series, 175–186.
Recognition of special linguistic patterns in a certain language is very helpful for many NLP applications such as information extraction, machine translation and parsing. State-of-the-arts syntax parsers are based on given grammar. The used grammar is context free and cannot discover complex patterns which contain multiple linguistic units. We propose an unsupervised method to automatically discover the complex linguistic patterns from a classically parsed corpus. A specialized and efficient algorithm is applied to mine the frequent subtrees in the forest and the found subtrees are formalized as the linguistic patterns. The approach is validated on the Penn Chinese Treebank with found linguistic patterns.
The paper presents an approach to valency frame extraction for Croatian verbs on basis of morphological and syntactic features of wordforms from syntactically annotated sentences. We have used a gold standard sample of approximately 1200 sentences and 30.000 tokens from the Croatian Dependency Treebank and a frame instance extraction algorithm. We extracted 936 verb frame instances for 424 different verbs – consisting of lemmas, morphosyntactic tags and syntactic functions of the encountered wordforms – and manually assigned tectogrammatical functors to their elements. Distributional properties are given in terms of co-occurrences for each of these features. The obtained results will serve for further development of valency frame extraction procedures.
We show that the standard beam-search algorithm can be used as an efficient decoder for the global linear model of Zhang and Clark (2008) for joint word segmentation and POS-tagging, achieving a significant speed improvement. Such decoding is enabled by: (1) separating full word features from partial word features so that feature templates can be instantiated incrementally, according to whether the current character is separated or appended; (2) deciding the POS-tag of a potential word when its first character is processed. Early-update is used with perceptron training so that the linear model gives a high score to a correct partial candidate as well as a full output. Effective scoring of partial structures allows the decoder to give high accuracy with a small beam-size of 16. In our 10-fold crossvalidation experiments with the Chinese Treebank, our system performed over 10 times as fast as Zhang and Clark (2008) with little accuracy loss. The accuracy of our system on the standard CTB 5 test was competitive with the best in the literature. 1
Language is not a system of signs that allow the exchange of information among individuals but also the means to engage in intersubjective relationship as well as to mark the identity of the speaker: the connection between language and identity is often so strong that a single feature of linguistic usage can be enough to identity someone’s belonging in a certain group. The linguistic matrix, that consists of being aware of how linguistic register expectations become linguistic norms, constitutes the basis of the community of speakers and society. Linguistic behaviour is exposed to distinct social dynamics and new linguistic habits and can be influenced by the sense of belonging, as perceived by the speakers, and the strength of the language (vitality and prestige). Pluralinguistic formations make allowance for the different ways of communication, adapting them to local customs, in this way avoiding harm, as well as introducing a positive bias, in populations so far relegated to minority status.
This paper describes a first test run of Sentitext, a sentiment analysis system under development. Unlike most existing systems, Sentitext is entirely based on linguistic knowledge and independent of any domain, using a wide coverage lexical database and lacking learning algorithms or classifiers, strictly speaking. Results on this first test, for which a collection of hotel review texts from Tripadvisor has been used, are extremely encouraging, given the high polarity hit rate. Keywords: Sentiment analysis, opinion mining, on-line user reviews. 1 Introduccion *
A new Chinese chunking algorithm is proposed based on Naive Bayes model and semantic features. Through the analysis of Chinese chunking task, Naive Bayes model that combines different types of features were applied for its rapid performance of training and test. Semantic features were utilized to further improve the accuracy. Experimental results on the Chinese chunking corpus of Chinese Penn Treebank show that the algorithm achieves impressive accuracy of 92.8% in terms of the F-score.
u-tokyo.ac.jp Several recent discourse parsers have employed fully-supervised machine learning approaches. These methods require human annotators to beforehand create an extensive training corpus, which is a time-consuming and costly process. On the other hand, unlabeled data is abundant and cheap to collect. In this paper, we propose a novel semi-supervised method for discourse relation classification based on the analysis of cooccurring features in unlabeled data, which is then taken into account for extending the feature vectors given to a classifier. Our experimental results on the RST Discourse Treebank corpus and Penn Discourse Treebank indicate that the proposed method brings a significant improvement in classification accuracy and macro-average F-score when small training datasets are used. For instance, with training sets of c.a. 1000 labeled instances, the proposed method brings improvements in accuracy and macro-average F-score up to 50% compared to a baseline classifier. We believe that the proposed method is a first step towards detecting low-occurrence relations, which is useful for domains with a lack of annotated data. 1
This article details a series of carefully designed experiments aiming at evaluating the influence of automatic pre-annotation on the manual part-of-speech annotation of a corpus, both from the quality and the time points of view, with a specific attention drawn to biases. For this purpose, we manually annotated parts of the Penn Treebank corpus (Marcus et al., 1993) under various experimental setups, either from scratch or using various pre-annotations. These experiments confirm and detail the gain in quality observed before (Marcus et al., 1993; Dandapat et al., 2009; Rehbein et al., 2009), while showing that biases do appear and should be taken into account. They finally demonstrate that even a not so accurate tagger can help improving annotation speed. 1
Whereas the main phonological changes between OJ and NJ took place during the EMJ period, it was during the LMJ period that most of the significant grammatical changes took place which transformed Japanese from its premodern to its contemporary shape in both morphology and syntax. The course and precise dating of some changes is difficult to trace through the written sources; they are mainly observable in the sources dating from the end of the period. It may be no coincidence that sweeping changes took place during a period of civil war and great social upheaval and change which also would have resulted in a relaxation of social and linguistic norms.
The contextual cueing effect (CC) refers to the phenomenon in which visual search performance is faster for targets appearing in previously exposed configurations than for targets appearing in new configurations. We investigated whether the learned configurations affect other types of response such as affective evaluation. Participants were asked to search T-target among rotated L-distractors. The mean reaction time showed a typical CC. Then, participants were asked to evaluate how much they like the repeated or new configurations (Experiment 1), or how much they feel the difficulty of target detection (Experiment 2). The results showed that both the liking and difficulty ratings for the repeated configurations were lower than those for the new configurations.
This article describes psycholinguistic lexical databases available in various languages, including English, Spanish and Portuguese. These lexical databases are important for researchers in Psycholinguistics and other related areas, providing a pool of experimental materials and allowing for an efficient process of selection of these experimental materials. The process of gathering statistics is slow, resulting in a small pool of materials in the short-term. The need to find an alternative method to gather limited or yet unavailable statistics for a specific language led us to consider gathering statistics from other languages and to compute their triangulation. Our aim was to automatize the computation of statistics such as Familiarity, Imageability, Age of Acquisition and Written Word Frequency for that specific language. We will describe the process of preparing this data and triangulating and comparing statistics for some languages in an attempt of finding a relationship between them. The results were analysed considering correlations between each statistic in each pair of languages and by computing the mean of absolute differences between each language’s values.