Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
FLEx lexical database XML export on 2013-04-30.
In this paper, we propose a method for au-tomatic clause boundary annotation in the Hindi Dependency Treebank. We show that the clausal information implicitly encoded in a dependency structure can be made explicit with no or less human interven-tion. We exercised the proposed approach on 16,000 sentences of Hindi Dependency Treebank. Our approach gives an accuracy of 94.44 % for clause boundary identifica-tion evaluated over 238 clauses. The resul-tant corpus has varied usages and can be utilized for developing a statistical clause boundary identifier. 1
Slovene Lexical Database was created between 2008 and 2012 and represents a comprehensive syntactic and semantic description of a selected set of Slovene words. The description was based exclusively on the analysis of reference corpora of Slovene. The database is structured as a network of interrelated semantic and syntactic information about a particular word. Semantic level represents the top level in the hierarchy with the lexical unit as its core element. This includes all senses of the headwrd, multi-word expressions and phraseological units. Each sense is described with a short semantic indicator and/or whole-sentence definition which includes typical syntactic environment of the headword with the relevant number, form and semantic types in a valency frame (semantic frame). These are also reflected in a number of syntactic structures and corresponding collocations. All the higher types of information are confirmed by a selection of corpus examples. Multi-word expressions and phraseological units are treated independently from particular senses of the headword and have their own internal structure which requires the same types of information as single-word entries or senses.
In this paper, we discuss our efforts to anno-tate nominals in the Hindi Treebank with the semantic property of animacy. Although the treebank already encodes lexical information at a number of levels such as morph and part of speech, the addition of animacy informa-tion seems promising given its relevance to varied linguistic phenomena. The suggestion is based on the theoretical and computational analysis of the property of animacy in the con-text of anaphora resolution, syntactic parsing, verb classification and argument differentia-tion. 1
Human communication relies on words—spoken or written labels for the concepts we intend to convey. These linguistic units map meanings onto forms that can be recognized and produced by others within a shared communication system. To make this possible, words are stored in long-term memory within what is often called the mental lexicon. This repository includes orthographic, phonological, morphological, and semantic information, and enables retrieval whenever comprehension or production demands it. The act of retrieving such information is what researchers describe as lexical access. In reading, the orthographic stimulus must be matched with its stored representation, just as the phonological form of the acoustic signal must be matched during speech comprehension. In production, by contrast, the intended meaning serves as the entry point, giving access to the phonological or orthographic form required for speech or writing. The term lexical access was first popularized in studies of visual word recognition, but its use has since expanded. In current literature, especially on word recognition, alternative terms such as lexical retrieval or lexical processing are often preferred, since access implies a discrete lexical entry that can be “looked up.” This assumption is at odds with many contemporary models, which favor distributed, parallel-activation accounts where sublexical units such as letters, phonemes, or morphemes contribute dynamically to recognition. In contrast, in word production the notion of access is less contentious, because selecting the correct lexical item from meaning necessarily requires pinpointing a specific representation. Research into lexical processing has focused on two main questions: the nature of the stored representations in the mental lexicon and the cognitive procedures through which they are retrieved. Much of this work has been conducted by cognitive psychologists, leading to an emphasis on mechanisms of retrieval rather than linguistic content. Moreover, explanations have focused on the cognitive rather than the neural level, though psycholinguistic theories are increasingly informed by neuroscience. Indeed, although the present article emphasizes cognitive perspectives—as suggested by its title—key findings on neural and electrophysiological correlates of lexical processing are also acknowledged, since they have substantially contributed to refining and constraining cognitive theories. The bulk of empirical research has concentrated on visual word recognition, not only because reading experiments offer precise control and measurement, but also because of their pedagogical and societal importance. Nevertheless, the same theoretical questions extend to spoken word recognition, speech production, and writing, each of which poses its own challenges for models of lexical access. In recent years, important advances have reshaped the field: large-scale megastudies and open-access lexical databases now allow researchers to examine the joint influence of multiple lexical and semantic variables, moving beyond traditional factorial designs. Computational modeling has also become more diverse, integrating Bayesian frameworks, hybrid connectionist approaches, and deep learning architectures, while empirical work has expanded to a wider range of languages and writing systems. The references selected throughout this article represent either foundational studies that shaped the field or recent contributions that capture the current state of debate, providing a framework for understanding how humans connect word forms to meanings in real time. Updated in September 2025 by Maria Fernández-López.
Information Structure (IS) determines the “communicative” segmentation of the meaning of an utterance, which makes it central to the semantics‐syntax‐ intonation interface and therefore also to NLP. Despite this relevance, IS has not received much attention in the context of the majority of the reference treebanks for data-driven NLP that already contain a semantic and syntactic layers of annotation. We present our work in progress on the annotation of the Penn TreeBank with the thematicity dimension of the IS as defined in the Meaning-Text Theory. We experiment with tagging and transitionbased parsing techniques. Especially the latter achieve acceptable accuracy with even very small training samples, which is promising for languages with scarce resources.
We investigate statistical dependency parsing of two closely related languages, Croatian and Serbian.As these two morphologically complex languages of relaxed word order are generally under-resourced -with the topic of dependency parsing still largely unaddressed, especially for Serbian -we make use of the two available dependency treebanks of Croatian to produce state-of-the-art parsing models for both languages.We observe parsing accuracy on four test sets from two domains.We give insight into overall parser performance for Croatian and Serbian, impact of preprocessing for lemmas and morphosyntactic tags and influence of selected morphosyntactic features on parsing accuracy.
In this paper, we provide a quantitative analysis of non-projective constructions attested in the Ancient Greek Dependency Treebank (AGDT). We consider the different types of formal constraints and metrics that have become standardized in the literature on non-projectivity (planarity, wellnestedness, gap-degree, edge-degree). We also discuss some of the linguistic factors that cause non-projective edges in Ancient Greek. Our results confirm the remarkable extension of non-projectivity in the AGDT, both in terms of quantitative incidence of non-projective nodes and for their complexity, which is not paralleled by the corpora of modern languages considered in the literature. At the same time, the usefulness of other constraint (especially well-nestedness) is confirmed by our researches. 1
In this paper, we investigate errors in syntax annotation with the Turku Dependency Treebank, a recently published treebank of Finnish, as study material. This treebank uses the Stanford Dependency scheme as its syntax representation, and its published data contains all data created in the full double annotation as well as timing information, both of which are necessary for this study.
Many countries use national-level surveys to capture student opinions about their university experiences. It is necessary to interpret survey results in an appropriate context to inform decision-making at many levels. To provide context to national survey outcomes, we describe patterns in the ratings of science and engineering subjects from the UK’s National Student Survey (NSS). New, robust statistical models describe relationships between the Overall Satisfaction’ rating and the preceding 21 core survey questions. Subjects exhibited consistent differences and ratings of “Teaching”, “Organisation” and “Support” were thematic predictors of “Overall Satisfaction” and the best single predictor was “The course was well designed and running smoothly”. General levels of satisfaction with feedback were low, but questions about feedback were ultimately the weakest predictors of “Overall Satisfaction”. The UK’s universities affiliated groupings revealed that more traditional “1994” and “Russell” groups over-performed in a model using the core 21 survey questions to predict “Overall Satisfaction”, in contrast to the under-performing newer universities in the Million+ and Alliance groups. Findings contribute to the debate about “level playing fields” for the interpretation of survey outcomes worldwide in terms of differences between subjects, institutional types and the questionnaire items.
We present an empirical study on constructing a Japanese constituent parser, which can output function labels to deal with more detailed syntactic information.Japanese syntactic parse trees are usually represented as unlabeled dependency structure between bunsetsu chunks, however, such expression is insufficient to uncover the syntactic information about distinction between complements and adjuncts and coordination structure, which is required for practical applications such as syntactic reordering of machine translation.We describe a preliminary effort on constructing a Japanese constituent parser by a Penn Treebank style treebank semi-automatically made from a dependency-based corpus.The evaluations show the parser trained on the treebank has comparable bracketing accuracy as conventional bunsetsu-based parsers, and can output such function labels as the grammatical role of the argument and the type of adnominal phrases.
This paper proposes a combined model for POS tagging, dependency parsing and co-reference resolution for Bulgarian — a pro-drop Slavic language with rich mor-phosyntax. We formulate an extension of the MSTParser algorithm that allows the simultaneous handling of the three tasks in a way that makes it possible for each task to benefit from the information available to the others, and conduct a set of experi-ments against a treebank of the Bulgarian language. The results indicate that the pro-posed joint model achieves state-of-the-art performance for POS tagging task, and outperforms the current pipeline solution. 1
Factors Related to Undergraduate Psychology Majors Learning Statistics Tamarah Faye Smith Doctor of Philosophy: Educational Psychology Major Advisor: Dr. Frank Farley The American Psychological Association (APA) has outlined goals for psychology undergraduates. These goals are aimed at several objectives including the need to build skills for interpreting and conducting psychological research (APA, 2007). These skills allow psychologists to conduct research that is covered in the media (Farley et al. 2009) and influences policy and law (Fischer, Stein & Heikkinen, 2009; Steinberg, Cauffman, Woolard, Graham & Banich, 2009a; Steinberg, Cauffman, Woolard, Graham & Banich, 2009b). One of the fundamental courses required for building these skills is statistics, a course that begins at the undergraduate level. Research has suggested that performance after completing statistics courses is weak for many students (Garfield, 2003; Hirsch & O'Donnell, 2001; Konold et al. 1993; Mulhern & Wylie, 2005; Schau & Mattern, 1997). The current study examined factors that may be related to performance on a statistical test. A sample of 231 students enrolled in or having already completed a statistics course for psychology majors completed a statistical skill questionnaire, built by the author, to measure performance with four APA outlined goals. To measure student attitudes the Survey of Attitudes Toward Statistics (SATS-36; Schau, 2003) was completed with adapted questions to measure perceived attitudes of peers and faculty toward statistics. Finally, questions pertaining to classroom techniques and content areas covered were assessed. Building off of social cognitive theory (SCT; Bandura, 1986) and expectancy-value theory (Eccles & Wigfield, 2002), it was expected that lower attitudes, such as low value and low interest, among the students and those perceived to be held by faculty and peers would be related to lower performance on the statistical test. A series of linear regressions were conducted and revealed no significant relationship between perceived faculty attitudes and performance. Students' own liking and positive affect ratings were positive predictors of performance indicating a gain of 3-4% on the statistical test. However, an interesting negative relationship emerged with respect to students' value of statistics and peer interest scores where performance on the statistical test decreased as value and peer interest increased. This may be demonstrating issues pertaining to the SATS-36 validity when measuring students' value as well as issues with the items created to measure perceived peer interest. The results of a factor analysis on perceived attitude measures for peers and faculty suggest that the need for more items is necessary, particularly for faculty attitudes. Finally, this study provides a first look at the performance of a sample of psychology students with APA goals for quantitative reasoning. Results showed that students performed best at reading basic descriptive statistics (M=74.5%), and worst when choosing statistical tests for a given research hypothesis (M=30%). Performance on questions pertaining to confidence intervals (M=38%) and discriminating between statistical and practical significance (M=39%) was also low. Future research can address limitations of this study by expanding the sample to include a broader range of psychology undergraduates and including additional items for measuring perceived attitudes. Other methodological approaches, such as experimental design and directly measuring faculty attitudes, should also be considered. Finally, further research and replication are necessary to determine if scores on the statistical test will continue to be low with other samples and varying question formats. These results can then be used to generate conversation about why and how students are, or are not, learning the appropriate quantitative skills.
The paper presents the process of constructing a publicly available treebank of public messages written in Croatian. The messages were collected from various electronic sources – e-mail, blog, Facebook and SMS – and published on the Zagreb Museum of Contemporary Art LED facade within the Babel art project. The project aimed to use the facade as an open-space blog or social interface for enabling citizens to publicly express their views. Construction and current state of the treebank is presented along with future work plans. A comparison of Babel Treebank with Croatian Dependency Treebank and SETimes.HR treebank regarding differing domains and annotation schemes is briefly sketched. The treebank is used as a test platform for introducing a new standard for syntactic annotation of Croatian texts. An experiment with morphosyntactic tagging and dependency parsing of the treebank is conducted, providing first insight to computational processing of non-standard text in Croatian.
This chapter deals with the main methodological issues underlying the building of the SciE-Lex lexical database and discusses and justifies the information included. SciE-Lex was initially conceived as a response to the lack of reference tools that can help scientists write scientific papers in phraseologically competent and native-like English. While there are a number of specialised dictionaries that include specific terminological information, there is a shortage of writing aids that provide information about the use of non-technical terms in scientific genres. SciE-Lex aims at filling this gap by focusing on the description of general terms in scientific English. This article describes the two stages in the building of the database, the first one including morphosyntactic and collocational information, and the second one focusing on phraseological information.
Syntactic parsing is an important technique in the natural language processing, yet Latvian is still lacking an efficient general coverage syntax parser. This paper reports on the first experiments on statistical syntactic parsing for Latvian — a highly inflective Indo-European language with a relatively free word order. We have induced a statistical parser from a small, non-balanced Latvian Treebank using the MaltParser toolkit and measured the unlabeled attachment score (UAS). As MaltParser is based on the dependency grammar approach, we have also developed a convertor from the hybrid dependency-based annotation model used in the Latvian Treebank to the pure dependency annotation model. We have obtained a promising 74.63 % UAS in 10-fold cross-validation using only ~2500 sentences. The results revealed that best results can be achieved using non-projective stack parsing algorithm with lazy arc adding strategy, but comparably good results can be achieved using projective parsing algorithms combined with appropriate projectiviziation preprocessing.
Recent developments in Natural Language Processing (NLP) are heading towards knowledge rich resources and technology. Integration of linguistically sound grammars, sophisticated machine learning settings and world knowledge background is possible given the availability of the appropriate resources: deep multilingual treebanks, representing detailed syntactic and semantic information; and vast quantities of world knowledge information encoded within ontologies and Linked Open Data datasets (LOD). Thus, the addition of world knowledge facts provides a substantial extension of the traditional semantic resources like WordNet, FrameNet and others. This extension comprises numerous types of Named Entities (Persons, Locations, Events, etc.), their properties (Person has a birthDate; birthPlace, etc.), relations between them (Person works for an Organization), events in which they participated (Person participated in war, etc.), and many other facts. This huge amount of structured knowledge can be considered the missing ingredient of the knowledgebased NLP of 80’s and the beginning of 90’s. The integration of world knowledge within language technology is defined as an ontology-to-text relation comprising different language and world knowledge in a common model. We assume that the lexicon is based on the ontology, i.e. the word senses are represented by concepts, relations or instances. The problem of lexical gaps is solved by allowing the storage of not only lexica, but also free phrases. The gaps in the ontology (a missing concept for a word sense) are solved by appropriate extensions of the ontology. The mapping is partial in the sense that both elements (the lexicon and the ontology) are artefacts and thus — they are never complete. The integration of the interlinked ontology and lexicon with the grammar theory, on the other hand, requires some additional and non-trivial reasoning over the world knowledge. We will discuss phenomena like selectional constraints, metonymy, regular polysemy, bridging relations, which live in the intersective areas between world facts and their language reflection. Thus, the actual text annotation on the basis of ontology-to-text relation requires the explication of additional knowledge like co-occurrence of conceptual information, discourse structure, etc. Such knowledge is mainly present in deeply processed language resources like HPSG-based (LFG-based) treebanks (RedWoods treebank, DeepBank, and others). The inherent characteristics of these language resources is their dynamic nature. They are constructed simultaneously with the development of a deep grammar in the corresponding linguistic formalism. The grammar is used to produce all potential analyses of the sentences within the treebank. The correct analyses are selected manually on the base of linguistic discriminators which would determine the correct linguistic production. The annotation process of the sentences provides feedback for the grammar writer to update the grammar. The life cycle of a dynamic language resource can be naturally supported by the semantic technology behind the ontology and LOD modeling the grammatical knowledge as well as the annotation knowledge; supporting the annotation process; reclassification after changes within the grammar; querying the available resources; exploitation in real applications. The addition of a LOD component to the system would facilitate the exchange of language resources created in this way and would support the access to the existing resources on the web.
This paper presents a reranking approach to combining constituent and dependency parsing, aimed at improving parsing performance on both sides. Most previous combination methods rely on complicated joint decoding to integrate graph- and transition-based dependency models. Instead, our approach makes use of a high-performance probabilistic context free grammar (PCFG) model to output k-best candidate constituent trees, and then a dependency parsing model to rerank the trees by their scores from both models, so as to get the most probable parse. Experimental results show that this reranking approach achieves the highest accuracy of constituent and dependency parsing on Chinese treebank (CTB5.1) and a comparable performance to the state of the art on English treebank (WSJ).
We present an automatic animacy classier for Dutch that can determine the animacy status of nouns | how alive the noun’s referent is (human, inanimate, etc.). Animacy is a semantic property that has been shown to play a role in human sentence processing, felicity and grammaticality. Although animacy is not marked explicitly in Dutch, we expect knowledge about animacy to be helpful for parsing, translation and other NLP tasks. Only a few animacy classiers and animacyannotated corpora exist internationally. For Dutch, animacy information is only available in the Cornetto lexical-semantic database. We augment this lexical information with context information from the Dutch Lassy Large treebank, to create training data for an animacy classier that uses a novel kind of context features. We use the k-nearest neighbour algorithm with distributional lexical features, e.g. how frequently the noun occurs as a subject of the verb ‘to think’ in a corpus, to decide on the (predominant) animacy class. The size of the Lassy Large corpus makes this possible, and the high level of detail these word association features provide, results in accurate Dutch-language animacy classication.
In this paper we describe the expansion of probabilistic context free grammar for Urdu language, we did some experiments in probabilistic context free grammar for Urdu language, extraction of CFG rules through The Penn Treebank, evaluation of PCFG from CFG. This PCFG is further useful for parsing the Urdu Sentences. The Tree-bank based grammar is the best technique for building up the PCFG for language through some easy and understandable steps as compared to theoretical. A PCFG can be used to estimate a number of useful probabilities concerning a sentence and its parse-tree(s). The resulting Penn Treebank based PCFG is widely used in natural language processing, speech recognition, and integrated spoken language systems as well as in theoretical linguistics.
Body image disturbances are core symptoms of eating disorders (EDs). Recent evidence suggests that changes in body image may occur prior to ED onset and are not restricted to in-vivo exposure (e.g. mirror image), but also evident during presentation of abstract cues such as body shape and weight-related words. In the present study startle modulation, heart rate and subjective evaluations were examined during reading of body words and neutral words in 41 student female volunteers screened for risk of EDs. The aim was to determine if responses to body words are attributable to a general negativity bias regardless of ED risk or if activated, ED relevant negative body schemas facilitate priming of defensive responses. Heart rate and word ratings differed between body words and neutral words in the whole female sample, supporting a general processing bias for body weight and shape-related concepts in young women regardless of ED risk. Startle modulation was specifically related to eating disorder symptoms, as was indicated by significant positive correlations with self-reported body dissatisfaction. These results emphasize the relevance of examining body schema representations as a function of ED risk across different levels of responding. Peripheral-physiological measures such as the startle reflex could possibly be used as predictors of females' risk for developing EDs in the future.
Abstract In this article we present some statistical data on the distribution of parts of speech and dependency relations in a large manually annotated Hungarian Treebank, the Szeged Dependency Treebank. We hypothesize that the domain of the text influences the distribution of the above elements, thus we pay special attention to differences between domains. We present the characteristic rank-frequency distributions of parts of speech and dependency relations in Hungarian and analyse the domain similarities and differences among sub-corpora as regards the above distributions. Our results reveal that the computer and newspaper texts are most similar to each other while the domains literature and compositions also exhibit some similarities. On the other hand, the business news and the law sub-corpora are unique, both having their own characteristics.
The Digital World encounters rapid development nowadays, especially through the proliferation of social media in Indonesia. Twitter has become one of social media with expanded users within every sectors of society. There are so many part both individual as well as organization/enterprise which utilize twitter as tool for communication, business, customer relation, and other activities. Through the twitter's ever-expanding users with those particular purposes, the precise method to effectively and efficiently analyzing opinion-contained sentences become crucially needed. Therefore this research made for method analyzing through lexical based and model based approaches by machine learning to classify opinion-contained tweets using those 2 methods. The tested machine learning method are Support Vector Machine (SVM), Maximum Entropy (ME), Multinomial Naive Bayes (MNB), and k-Nearest Neighbor (k-NN). Based on the test outcome, lexical based approach highly depended on lexical database which became opinion classification matrix. Whilst machine learning approach can produce better accuracy due to its capability in new training data modeling based on outcome model. However, machine learning model based approach depends on various factors in analyzing sentiment.
The investigation of gender differences in emotion has attracted much attention given the potential ramifications on our understanding of sexual differences in disorders involving emotion dysregulation. Yet, research on content-specific gender differences across adulthood in emotional responding is lacking. The aims of the present study were twofold. First, we sought to investigate to what extent gender differences in the self-reported emotional experience are content specific. Second, we sought to determine whether gender differences are stable across the adult lifespan. We assessed valence and arousal ratings of 14 picture series, each of a different content, in 94 men and 118 women aged 20 to 81. Compared to women, men reacted more positively to erotic images, whereas women rated low-arousing pleasant family scenes and landscapes as particularly positive. Women displayed a disposition to respond with greater defensive activation (i.e., more negative valence and higher arousal), in particular to the most arousing unpleasant contents. Importantly, significant interactions between gender and age were not found for any single content. This study makes a novel contribution by showing that gender differences in the affective experiences in response to different contents persist across the adult lifespan. These findings support the “stability hypothesis” of gender differences across age.
of a dissertation at the University of Miami. Dissertation supervised by Professor Alexandra L. Quittner No. of pages in text. (70) Objective: CF is a progressive, life-shortening disease treated primarily with palliative medications. Among the consequences of the disease’s progression and daily treatments are discomfort and pain. Previous studies have suggested that pain is common in patients with CF; however, little is known about the factors associated with this pain or its impact on clinical outcomes. The relationships between pain and health outcomes over time, such as adherence and health-related quality of life (HRQOL), are largely unknown in this population. The purpose of this study was to systematically assess pain in adolescents with CF and evaluate its associations with adherence, social support, and HRQOL over a six-month period. Methods: The current study is part of a multi-center NIH SBIR Phase II randomized, controlled trial. The sample consisted of 95 participants, with a mean age of 15.69. Participants completed a battery of measures during three consecutive clinic visits, approximately 3 months apart. Participants also completed an online pain diary for the 6 days following each clinic visit. Diaries assessed pain intensity, location, duration, affective rating, and coping responses. Results: Overall, 73% of participants completed one or more of the online pain diaries across these time points, with 44% of the sample completing all six diaries. Pain was reported by 74.5% of participants. Of those who experienced pain, intensity was generally mild. Daily pain ratings, as assessed by online diaries, were highly variable within participants. Path analyses indicated that worse treatment adherence and poor social functioning were directly related higher pain and ultimately related to worse HRQOL. Conclusions: These results indicated that pain is common in adolescents with CF and that it interferes significantly with HRQOL. Treatment adherence appears to be particularly predictive of pain in this population. Regular assessment of pain and HRQOL is recommended.
In this dissertation, the robustness of the relationship between the lexical frequency of phonotactic patterns and word-acceptability is examined for words of Amharic, an understudied Semitic language. The patterns under investigation span the whole verb root and include both under-represented and over-represented consonant distributions in the lexicon. A state-of-the-art probabilistic model, the Maximum Entropy phonotactic learner, is used to acquire a phonotactic grammar from the input (the lexicon) and the predictions of that grammar are compared with the results of two Amharic nonce-word rating tasks designed specifically to investigate a range of consonantal phonotactic patterns. The first task investigates consonant co-occurrence patterns (homorganic consonants, identical consonants, and fricatives). In the Amharic verb lexicon, identical consonants are under- represented in some locations and over-represented in others whereas homorganic consonants and fricatives (a previously unknown pattern independently acquired by the model) are under-represented. The phonotactic learner successfully learned the under-represented patterns and the comparison between the model predictions and the experimental results show evidence for a relationship between lexical frequency and word acceptability for under -representation. However, speaker judgements show no preference for over-representation. The second task examines the distribution of single consonants within the verb root with respect to under-representation, over- representation and positional restrictions. Evidence for a relationship between lexical frequency and phonotactic probability was observed for both under-represented and over-represented consonants, but tied to a particular location. The correlation between speaker judgments and model predictions is low for this task, due in part to the way the model deals with over-representation. This investigation demonstrates not only that word acceptability is influenced by phonotactic probability for both under-represented and over-represented patterns, but also that probabilistic models can be used to investigate the phonotactics of a language, even in the absence of speaker judgement data. These models can therefore be used to assess the phonotactics of languages where experimental data is difficult to obtain and broaden our knowledge of phonotactic typology
This file contains the guidelines, which were employed to administer the Self-Assessment Manikin (SAM) [1] and the productivity questionnaire to the participants of the experiment reported in [2]. The guidelines have been written by following the technical manual by Lang et al. [3]. This document will be updated while preparing the camera-ready version of the paper. The camera-ready version of the paper will be made available on Figshare as well (the publisher permits that) [1] Bradley, L.: Measuring emotion: the self-assessment semantic differential. Journal of Behavior Therapy and Experimental Psychiatry. 25, 1, 49–59 (1994).<br>[2] Graziotin, D., Wang, X., & Abrahamsson, P.. Are Happy Developers more Productive? The Correlation of Affective States of Software Developers and their self-assessed Productivity. In Product-Focused Software Process Improvement,. Paphos, Cyprus. Springer Verlag. In Press (2013).<br>[3] Lang, P.J. et al.: International affective picture system (IAPS): Technical manual and affective ratings. Gainesville FL NIMH Center for the study of emotion and attention University of Florida. Technical Report A–6 (1999).
A new approach to lexicographic work, in which the lexicographer is seen more as a validator of the choices made by computer, was recently envisaged by Rundell and Kilgarriff (2011). In this paper, we describe an experiment using such an approach during the creation of Slovene Lexical Database (Gantar, Krek, 2011). The corpus data, i.e. grammatical relations, collocations, examples, and grammatical labels, were automatically extracted from 1,18-billion-word Gigafida corpus of Slovene. The evaluation of the extracted data consisted of making a comparison between the time spent writing a manual entry and a (semi)-automatic entry, and identifying potential improvements in the extraction algorithm and in the presentation of data. An important finding was that the automatic approach was far more effective than the manual approach, without any significant loss of information. Based on our experience, we would propose a slightly revised version of the approach envisaged by Rundell and Kilgarriff in which the validation of data is left to lower-level linguists or crowd-sourcing, whereas high-level tasks such as meaning description remain the domain of lexicographers. Such an approach indeed reduces the scope of lexicographer’s work, however it also results in the ability of bringing the content to the users more quickly.
The gaze cueing effect involves the rapid orientation of attention to follow the gaze direction of another person. Previous studies reported reciprocal influences between social variables and the gaze cueing effect, with modulation of gaze cueing by social features of face stimuli and modulation of the observer's social judgements from the validity of the gaze cues themselves. However, it remains unclear which social dimensions can affect-and be affected by-gaze cues. We used computer-averaged prototype face-like images with high and low levels of perceived trustworthiness and dominance to investigate the impact of these two fundamental social impression dimensions on the gaze cueing effect. Moreover, by varying the proportions of valid and invalid gaze cues across three experiments, we assessed whether gaze cueing influences observers' impressions of dominance and trustworthiness through incidental learning. Bayesian statistical analyses provided clear evidence that the gaze cueing effect was not modulated by facial social trait impressions (Experiments 1-3). However, there was uncertain evidence of incidental learning of social evaluations following the gaze cueing task. A decrease in perceived trustworthiness for non-cooperative low dominance faces (Experiment 2) and an increase in dominance ratings for faces whose gaze behaviour contradicted expectations (Experiment 3) appeared, but further research is needed to clarify these effects. Thus, this study confirms that attentional shifts triggered by gaze direction involve a robust and relatively automatic process, which could nonetheless influence social impressions depending on perceived traits and the gaze behaviour of faces providing the cues.
Unsupervised dependency parsing is acquiring great relevance in the area of Natural Language Processing due to the increasing number of utterances that become available on the Internet. Most current works are based on Depen- dency Model with Valence (DMV) (12) or Extended Valence Grammars (EVGs) (11), in both cases the dependencies between words are modeled by using a fixed structure of automata. We present a framework for unsupervised induction of dependency structures based on CYK parsing that uses a simple rewriting tech- niques of the training material. Our model is implemented by means of a k-best CYK parser, an inductor for Probabilistic Bilexical Grammars (PBGs) (8) and a simple technique that rewrites the treebank from k trees with their probabilities. An important contribution of our work is that the framework accepts any existing algorithm for automata induction making the automata structure fully modifiable. Our experiments showed that, it is the training size that influences parameteriza- tion in a predictable manner. Such flexibility produced good performance results in 8 different languages, in some cases comparable to the state-of-the-art ones.
Lexical resources such as WordNet and VerbNet are widely used in a multitude of NLP tasks, as are annotated corpora such as treebanks. Often, the resources are used as-is, without question or examination. This practice risks missing significant performance gains and even entire techniques. This paper addresses the importance of resource quality through the lens of a challenging NLP task: detecting selectional preference violations. We present DAVID, a simple, lexical resource-based preference violation detector. With asis lexical resources, DAVID achieves an F1-measure of just 28.27%. When the resource entries and parser outputs for a small sample are corrected, however, the F1-measure on that sample jumps from 40 % to 61.54%, and performance on other examples rises, suggesting that the algorithm becomes practical given refined resources. More broadly, this paper shows that resource quality matters tremendously, sometimes even more than algorithmic improvements. 1
Financial literacy education has traditionally never been a major subject area in most public schools in Maryland. With the unanticipated struggling of the United States of America (USA) Economy, Maryland (one of the richest states in the USA and its residents, especially of its major urban city, Baltimore, is facing severe family and personal financial crisis. This unanticipated financial crisis characterized by declining family and personal savings, mortgage defaults and foreclosures, and tenant evictions due to rent defaults is motivating educators, business entities, and politicians to develop education strategies to educate children from future financial crisis. Although there is an overwhelming consensus for providing financial education as a major curriculum to our schoolchildren, the practicability of most curriculums is obscure. To develop a practical approach to teaching financial literacy in elementary schools, an assessment using a Grid Familiarity Rating Chart (Very Familiar -Not Familiar) of basic financial and economics concepts was assigned to 210 students (4th -5th graders) in four inner city schools. Students received information to place a check mark by each concept if they are very familiar or familiar with the concept. Results revealed that 97% of the students were very familiar or familiar with the basic financial concepts as compared to 49% of the students who were very familiar or familiar with basic economic concepts. Even when students were encouraged to explain their very familiar or familiar concepts, 80% provided explanations that were accurate or almost accurate for basic financial concepts as compared to 20% accuracy for explanations with basic economic concepts. Conclusively, the most basic economic concepts (scarcity, choice, and opportunity cost), which are essential decision making tools for most financial decisions should be simplified and included in financial education lesson plans.
An MT-oriented system using Conditional Random Fields (CRFs) is presented to identify English Prepositional Phrases (PPs) within business domain. For the purpose of English-Chinese Machine Translation (MT), we, under the guidance of the theory of Syntactic Functional Grammar (SFG), refine PP function chunks into four types instead of the binary attachment. In order to improve the identification of these chunk types, we revise the Penn Treebank tagset with four major changes being made. A small size of 998k English annotated corpus in business domain is semi-automatically built based on our new tagset employing the Maximum Entropy model. Experiments show that our system achieves an accuracy of 88.45%, higher than other reported approaches. The adjustments made in the PP chunk types and POS tagset give rise to 4.11%, 4.25% and 4.15% increase in the precision, recall and F-score respectively.
La segmentation d'un texte en Unites Discursives Minimales (UDM) a pour but de decouper le texte en segments qui ne se chevauchent pas. Ces segments sont ensuite relies entre eux afin de construire la structure discursive d'un texte. La plupart des approches existantes utilisent une analyse syntaxique extensive. Malheureusement, certaines langues ne disposent pas d'analyseur syntaxique robuste. Dans cet article, nous etudions la faisabilite de la segmentation discursive de textes arabes en nous basant sur une approche d'apprentissage supervisee qui predit les UDM et les UDM imbriques. La performance de notre segmentation a ete evaluee sur deux genres de corpus: des textes de livres de l'enseignement secondaire et des textes du corpus Arabic Treebank. Nous montrons que la combinaison de traits typographiques, morphologiques et lexicaux permet une bonne reconnaissance des bornes de segments. De plus, nous montrons que l'ajout de traits syntaxiques n'ameliore pas les performances de notre segmentation.
The Stuttgart-Tubingen Tag Set (STTS) (Schiller et al., 1995) has long been established as a quasi-standard for part-of-speech (POS) tagging of German. It has been used, with minor modifications, for the annotation of three German newspaper treebanks, the NEGRA treebank (Skut et al., 1997), the TiGer treebank (Brants et al., 2002) and the TuBa-D/Z (Telljohann et al., 2004). One major drawback, however, is the lack of tags for the analysis of language phenomena from domains other than the newspaper domain. A case in point is spoken language, which displays a wide range of phenomena which do not (or only very rarely) occur in newspaper text.
Dependency parsing has gained more and more interest in natural language processing in recent years due to its simplicity and general applicability for diverse languages. Previous work demonstrates that part-of-speech (POS) is an indispensable feature in dependency parsing since pure lexical features suffer from serious data sparseness problem. However, due to little morphological changes, Chinese POS tagging has proven to be much more challenging than morphology-richer languages such as English (94% vs. 97% on POS tagging accuracy). This leads to severe error propagation for Chinese dependency parsing. Our experiments show that parsing accuracy drops by about 6% when replacing manual POS tags of the input sentence with automatic ones generated by a state-of-the-art statistical POS tagger. To address this issue, this paper proposes a solution by jointly optimizing POS tagging and dependency parsing in a unique model. We propose for our joint models several dynamic programming based decoding algorithms which can incorporate rich POS tagging and syntactic features. Then we present an effective pruning strategy to reduce the search space of candidate POS tags, leading to significant improvement of parsing speed. Experimental results on two Chinese data sets, i.e. Penn Chinese Treebank 5.1 and Penn Chinese Treebank 7, demonstrate that our joint models significantly improve both the state-of-the-art tagging and parsing accuracies. Detailed analysis shows that the joint method can help resolve syntax-sensitive POS ambiguities <formula formulatype="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex Notation="TeX">$\{{\ssr{NN}},{\ssr{VV}}\}$</tex> </formula>. In return, the POS tags become more reliable and helpful for parsing since the syntactic features are used in POS tagging. This is the fundamental reason for the performance improvement.
Light verb constructions (LVCs) are verb and noun combinations in which the verb has lost its meaning to some degree and the noun is used in one of its original senses. They often share their syntactic pattern with other constructions (e.g. verbobject pairs) thus LVC detection can be viewed as classifying certain syntactic patterns as light verb constructions or not. In this paper, we explore a novel way to detect LVCs in texts: we apply a dependency parser to carry out the task. We present our experiments on a Hungarian treebank, which has been manually annotated for dependency relations and light verb constructions. Our results outperformed those achieved by state-of-the-art techniques for Hungarian LVC detection, especially due to the high precision and the treatment of long-distance dependencies.
In this work, we present a data-driven method to enhance syntax trees with additional dependencies as defined in the wellknown Stanford Dependencies scheme, so as to give more information about the structure of the sentence. This hybrid method utilizes both machine learning and a rule-based approach, and achieves a performance of 93.1 % in F1-score, as evaluated using an existing treebank of Finnish. The resulting tool will be integrated into an existing Finnish parser and made publicly available at the address
Signal Detection models as well as the Two-High-Threshold model (2HTM) have been used successfully as measurement models in recognition tasks to disentangle memory performance and response biases. A popular method in recognition memory is to elicit confidence judgements about the presumed old/new status of an item, allowing for the easy construction of ROCs. Since the 2HTM assumes fewer latent memory states than response options are available in confidence ratings, the 2HTM has to be extended by a mapping function which models individual rating scale usage. Unpublished data from 2 experiments in Bröder and Schütz (2009) validate the core memory parameters of the model, and 3 new experiments show that the response mapping parameters are selectively affected by manipulations intended to affect rating scale use, and this is independent of overall old/new bias. Comparisons with SDT show that both models behave similarly, a case that highlights the notion that both modelling approaches can be valuable (and complementary) elements in a researcher's toolbox.
The Treebanks as the sets of syntactically annotated sentences, are the most widely used language resource in the application of Natural Language Processing. The occurrence of errors in the automatically created Treebanks is one of the main obstacles limiting the using of these resources in the real world applications. This paper aims to introduce an statistical method for diminishing the amount of errors occurred in a specific English LTAG-Treebank proposed in Basirat and Faili (2013). The problem has been formulated as a classification problem and has been tackled by using several classifiers. The experiments show that by using this approach, about 95% of the errors could be detected and more than 77% of them could successfully be corrected in the case of using Adaboost classifier. In addition, it has been shown that the new treebank could reach a high of 76% F-measure which is 8% higher than the original treebank.
Major depressive disorder (MDD) is associated with risk for chronic pain, but the mechanisms contributing to the MDD and pain relationship are unclear. To examine whether disrupted emotional modulation of pain might contribute, this study assessed emotional processing and emotional modulation of pain in healthy controls and unmedicated persons with MDD (14 MDD, 14 controls). Emotionally charged pictures (erotica, neutral, mutilation) were presented in 4 blocks. Two blocks assessed physiological-emotional reactions (pleasure/arousal ratings, corrugator electromyography (EMG), startle modulation, skin conductance) in the absence of pain and 2 blocks assessed emotional modulation of pain and the nociceptive flexion reflex (NFR, a physiological measure of spinal nociception) evoked by suprathreshold electric stimulations. Results indicated pictures generally evoked the intended emotional responses; erotic pictures elicited pleasure, subjective arousal, and smaller startle magnitudes, whereas mutilation pictures elicited displeasure, corrugator EMG activation, and subjective/physiological arousal. However, emotional processing was partially disrupted in MDD, as evidenced by a blunted pleasure response to erotica and a failure to modulate startle according to a valence linear trend. Furthermore, emotional modulation of pain was observed in controls but not MDD, even though there were no group differences in NFR threshold or emotional modulation of NFR. Together, these results suggest supraspinal processes associated with emotion processing and emotional modulation of pain may be disrupted in MDD, but brain to spinal cord processes that modulate spinal nociception are intact. Thus, emotional modulation of pain deficits may be a phenotypic marker for future pain risk in MDD.