Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
In contemporary South Korean society, there is a strong emphasis on cultural homogeneity and, simultaneously, the development of English proficiency as a human resource. Since language is inextricably linked to identity, bilingual learners from English speaking countries may feel pressure to conform to Korean cultural and linguistic norms, leading to negative identity practices that discourage the use of English. Like the “model minority” stereotype which has been assigned to Asian learners in the United States, the pervasive belief that learners from English speaking countries are highly proficient in English may have adverse effects on students who do not meet the conceptualized standard. To explore educational problems associated with the English-Korean bilingual learner, a case study was conducted on an American-Korean elementary school student. Results revealed that the learner avoided speaking English in public, learning English in formal contexts, and talking about American ethnic traditions, which has resulted in significant deficiencies in English pronunciation and literacy. The avoidance of explicit instruction appears to have precluded the development of cognitive and metacognitive strategies useful in overcoming language deficiencies in an English as a Foreign Language (EFL) context. Recommendations for educational reform have been suggested.
In this work, we present a data-driven method to enhance syntax trees with additional dependencies as defined in the wellknown Stanford Dependencies scheme, so as to give more information about the structure of the sentence. This hybrid method utilizes both machine learning and a rule-based approach, and achieves a performance of 93.1 % in F1-score, as evaluated using an existing treebank of Finnish. The resulting tool will be integrated into an existing Finnish parser and made publicly available at the address
Reliably coded treebanks are a goldmine for linguistics research. Answering a typical research question involves: (a) querying a treebank to extract sentences containing the feature to be investigated, (b) recognizing and keeping track of characteristics that determine the way in which the linguistic feature is encoded, and (c) using statistics to find out which (combination) of these characteristics determines the outcome of the linguistic feature. While sufficient tools are available for steps (a) and (c) in this process, step (b) has not received much attention yet. This paper describes how the programs “Cesax” and “CorpusStudio” can be used jointly to construct a “corpus research database”, a database that contains the sentences of interest selected in step (a), as well as user-definable pre-calculated characteristics for step (b).
Affective response may dominate users' reactions to the synthesized tactile sensations that are proliferating in today's handheld and gaming devices, yet it is largely unmeasured, modeled or characterized. A better understanding of user perception will aid the design of tactile behavior that engages touch, with an experience that satisfies rather than intrudes. We measured 30 subjects' affective response to vibrations varying in rhythm and frequency, then examined how differences in demographic, everyday use of touch, and tactile processing abilities contribute to variations in affective response. To this end, we developed five affective and sensory rating scales and two tactile performance tasks, and also employed a published `Need for Touch' (NFT) questionnaire. Subjects' ratings, aggregated, showed significant correlations among the five scales and significant effect of the signal content (rhythm and frequency). Ratings varied considerably among subjects, but this variation did not coincide with demographic, NFT score or tactile task performance. The linkages found among the rating scales confirm this as a promising approach. The next step towards a comprehensive picture of individuals' patterns of affective response to tactile sensations entails pruning, integration and redundancy reduction of these scales, then their formal validation.
We have implemented a rule-based prototype of a Spanish-to-Cuzco Quechua MT system enhanced through the addition of statistical components. The greatest difficulty during the translation process is to generate the correct Quechua verb form in subordinated clauses. The prototype has several rules that decide which verb form should be used in a given context. However, matching the context in order to apply the correct rule depends crucially on the parsing quality of the Spanish input. As the form of the subordinated verb depends heavily on the conjunction in the subordinated Spanish clause and the semantics of the main verb, we extracted this information from two treebanks and trained different classifiers on this data. We tested the best classifier on a set of 4 texts, increasing the correct subordinated verb forms from 80% to 89%.
We present a reformulation of the word pair features typically used for the task of disambiguating implicit relations in the Penn Discourse Treebank. Our word pair features achieve significantly higher performance than the previous formulation when evaluated without additional features. In addition, we present results for a full system using additional features which achieves close to state of the art performance without resorting to gold syntactic parses or to context outside the relation.
People believe that women are more emotionally intense than men, but the scientific evidence is equivocal. In this study, we tested the novel hypothesis that men and women differ in the neural correlates of affective experience, rather than in the intensity of neural activity, with women being more internally (interoceptively) focused and men being more externally (visually) focused. Adult men (n = 17) and women (n = 17) completed a functional magnetic resonance imaging study while viewing affectively potent images and rating their moment-to-moment feelings of subjective arousal. We found that men and women do not differ overall in their intensity of moment-to-moment affective experiences when viewing evocative images, but instead, as predicted, women showed a greater association between the momentary arousal ratings and neural responses in the anterior insula cortex, which represents bodily sensations, whereas men showed stronger correlations between their momentary arousal ratings and neural responses in the visual cortex. Men also showed enhanced functional connectivity between the dorsal anterior insula cortex and the dorsal anterior cingulate cortex, which constitutes the circuitry involved with regulating shifts of attention to the world. These results demonstrate that the same affective experience is realized differently in different people, such that women's feelings are relatively more self-focused, whereas men's feelings are relatively more world-focused.
This paper investigates the effect of the label bias problem of maximum entropy Markov models for part-of-speech tagging, a typical sequence prediction task in natural language processing. This problem has been underexploited and underappreciated. The investigation reveals useful information about the entropy of local transition probability distributions of the tagging model which enables us to exploit and quantify the label bias effect of part-of-speech tagging. Experiments on a Vietnamese treebank and on a French treebank show a significant effect of the label bias problem in both of the languages.
This study examines the impact of online word of mouth (WOM) and expert reviews on movies' box office revenues, both in the U.S. domestic market and in the international markets. Using a sample of 169 movies released in 2008, the study discovered that the frequency of online WOM and the valence rating of expert reviews were significant factors for box office outcomes in the domestic market. The study also found that only the frequency of online WOM was a significant factor in the international markets. The findings suggest that online WOM and expert reviews play a critical role in moviegoers' consumption behavior in the age of the Internet and social media.
We describe a novel approach to detecting empty categories (EC) as represented in de-pendency trees as well as a new metric for measuring EC detection accuracy. The new metric takes into account not only the position and type of an EC, but also the head it is a dependent of in a dependency tree. We also introduce a variety of new features that are more suited for this approach. Tested on a sub-set of the Chinese Treebank, our system im-proved significantly over the best previously reported results even when evaluated with this more stringent metric. 1
Large-scale linguistically annotated cor-pora have played a crucial role in advanc-ing the state of the art of key natural lan-guage technologies such as syntactic, se-mantic and discourse analyzers, and they serve as training data as well as evaluation benchmarks. Up till now, however, most of the evaluation has been done on mono-lithic corpora such as the Penn Treebank, the Proposition Bank. As a result, it is still unclear how the state-of-the-art analyzers perform in general on data from a vari-ety of genres or domains. The completion of the OntoNotes corpus, a large-scale, multi-genre, multilingual corpus manually annotated with syntactic, semantic and discourse information, makes it possible to perform such an evaluation. This paper presents an analysis of the performance of publicly available, state-of-the-art tools on all layers and languages in the OntoNotes v5.0 corpus. This should set the bench-mark for future development of various NLP components in syntax and semantics, and possibly encourage research towards an integrated system that makes use of the various layers jointly to improve overall performance. 1
We present a classification model that predicts the presence or omission of a lexical connective between two clauses, based upon linguistic features of the clauses and the type of discourse relation holding between them.The model is trained on a set of high frequency relations extracted from the Penn Discourse Treebank and achieves an accuracy of 86.6%.Analysis of the results reveals that the most informative features relate to the discourse dependencies between sequences of coherence relations in the text.We also present results of an experiment that provides insight into the nature and difficulty of the task.
ABSTRACT Online travel reviews are emerging as a powerful source of information affecting tourists' pre-purchase evaluation of a hotel organization. This trend has highlighted the need for a greater understanding of the impact of online reviews on consumer attitudes and behaviors. In view of this need, we investigate the influence of online hotel reviews on consumers' attributions of service quality and firms' ability to control service delivery. An experimental design was used to examine the effects of four independent variables: framing; valence; ratings; and target. The results suggest that in reviews evaluating a hotel, remarks related to core services are more likely to induce positive service quality attributions. Recent reviews affect customers' attributions of controllability for service delivery, with negative reviews exerting an unfavorable influence on consumers' perceptions. The findings highlight the importance of managing the core service and the need for managers to act promptly in addressing customer service problems.
status: Published
The ventromedial prefrontal cortex (vmPFC) plays a critical role in processing appetitive stimuli. Recent investigations have shown that reward value signals in the vmPFC can be altered by emotion regulation processes; however, to what extent the processing of positive emotion relies on neural regions implicated in reward processing is unclear. Here, we investigated the effects of emotion regulation on the valuation of emotionally evocative images. Two independent experimental samples of human participants performed a cognitive reappraisal task while undergoing fMRI. The experience of positive emotions activated the vmPFC, whereas the regulation of positive emotions led to relative decreases in vmPFC activation. During the experience of positive emotions, vmPFC activation tracked participants' own subjective ratings of the valence of stimuli. Furthermore, vmPFC activation also tracked normative valence ratings of the stimuli when participants were asked to experience their emotions, but not when asked to regulate them. A separate analysis of the predictive power of vmPFC on behavior indicated that even after accounting for normative stimulus ratings and condition, increased signal in the vmPFC was associated with more positive valence ratings. These results suggest that the vmPFC encodes a domain-general value signal that tracks the value of not only external rewards, but also emotional stimuli.
We present “CDG Lab”, an integrated environment for development of dependency grammars and treebanks. It uses the Categorial Dependency Grammars (CDG) as a formal model of dependency grammars. CDG are very expressive. They generate unlimited dependency structures, are analyzed in polynomial time and are conservatively extendable by regular type expressions without loss of parsing efficiency. Due to these features, they are well adapted to definition of large scale grammars. CDG Lab supports the analysis of correctness of treebanks developed in parallel with evolving grammars.
Recent experimental studies have examined GOODNESS IS BRIGHTNESS and a host of other primary metaphors. However, complex mappings such as INTELLIGENCE IS BRIGHTNESS have been largely ignored, nor has there been any attempt to distinguish their effects from those of primary metaphors such as GOODNESS IS BRIGHTNESS. The current study assesses both the nonprimary metaphoric mapping INTELLIGENCE IS BRIGHTNESS and the well-documented primary metaphor GOODNESS IS BRIGHTNESS in a visual priming task. The study finds that a bright background encourages photos of faces to be rated as both more intelligent and wellintentioned, though the background does not significantly affect either attribute alone. This suggests that two metaphors with the same source domain can reinforce each other. The study also underscores the difficulty in assessing a non-primary mapping in isolation from other factors.
peer reviewed
The present paper focuses on ways in which the pragmatic (functional) meaning that arises from various contextual features, known in corpus linguistics as semantic prosody, can become an integral part of lexicographical descriptions as they are represented in the Slovene Lexical Database (SLD). This is particularly important for the treatment of phraseology and idiomatics. First, the theoretical background is provided, with the focus on the prototype theory and its practical implications for monolingual lexicography. A parallel is drawn with the model of meaning analysis in the SLD. The second part begins with a brief introduction to semantic prosody and continues with an analysis of monolingual meaning descriptions in the SLD against a number of authentic corpus examples, investigating how their pragmatic components have been identified. The analysis of corpus data shows that pragmatics is an important contributor to the process of sense discrimination in works of lexical and lexicographic relevance.
With the growing interest in statistical parsing, special attention has recently been devoted to the problem of comparing different treebanks to assess which languages or domains are more difficult to parse relative to a given model. A common methodology for comparing parsing difficulty across treebanks is based on the use of the standard labeled precision and recall measures. As an alternative, in this article we propose an information-theoretic measure, called the expected conditional cross-entropy (ECC). One important advantage with respect to standard performance measures is that ECC can be directly expressed as a function of the parameters of the model. We evaluate ECC across several treebanks for English, French, German, and Italian, and show that ECC is an effective measure of parsing difficulty, with an increase in ECC always accompanied by a degradation in parsing accuracy.
National audience
This paper discusses the extension of a sys-tem developed for automatic discovery of tree-bank annotation inconsistencies over an entire corpus to the particular case of evaluation of inter-annotator agreement. This system makes for a more informative IAA evaluation than other systems because it pinpoints the incon-sistencies and groups them by their structural types. We evaluate the system on two corpora- (1) a corpus of English web text, and (2) a corpus of Modern British English. 1
This paper presents our preliminary conclusions as part of an ongoing effort to construct a new dependency representation framework for Turkish.We aim for this new framework to accommodate the highly agglutinative morphology of Turkish as well as to allow the annotation of unedited web data, and shape our decisions around these considerations.In this paper, we firstly describe a novel syntactic representation for morphosyntactic sub-word units (namely inflectional groups (IGs) in Turkish) which allows inter-IG relations to be discerned with perfect accuracy without having to hide lexical information.Secondly, we investigate alternative annotation schemes for coordination structures and present a better scheme (nearly 11% increase in recall scores) than the one in Turkish Treebank (Oflazer et al., 2003) for both parsing accuracies and compatibility for colloquial language.
Although several syntactically annotated corpora (or treebanks) exist for Dutch, they are seldomly used for descriptive linguistic research because there are no easy-to-use exploitation tools available.This demonstration paper describes GrETEL, a linguistic search engine (http:// nederbooms.ccl.kuleuven.be/eng/gretel)that enables non-technical users to consult treebanks in a user-friendly way.Instead of a formal search expression, a natural language example is used as input to the system, allowing users to search for similar constructions as the example they provide.In the first version of GrETEL, only written Dutch (LASSY) was included.Based on user requests we have now included the Spoken Dutch Corpus (CGN) as well.
We present, here, our analysis of systematic divergences in parallel English-Hindi dependency treebanks based on the Computational Paninian Grammar (CPG) framework. Study of structural divergences in parallel treebanks not only helps in developing larger treebanks automatically, but can also be useful for many NLP applications such as data-driven machine translation (MT) systems. Given that the two treebanks are based on the same grammatical model, a study of divergences in them could be of advantage to such tasks, along with making it more interesting to study how and where they diverge. We consider two parallel trees divergent based on differences in constructions, relations marked, frequency of annotation labels and tree depth. Some interesting instances of structural divergences in the treebanks have been discussed in the course of this paper. We also present our task of alignment of the two treebanks, wherein we talk about our extraction of divergent structures in the trees, and discuss the results of this exercise. 1
This paper introduces an advanced, efficient approach for rule based English to Bengali (E2B) machine translation (MT), where Penn-Treebank parts of speech (PoS) tags, HMM (Hidden Markov Model) Tagger is used.Fuzzy-If-Then-Rule approach is used to select the lemma from rule-based-knowledge. The proposed E2B-MT has been tested through F-Score measurement, and the accuracy is more than eighty percent.
This paper investigates the appropriateness of using lexical cohesion analysis to assess Chinese readability. In addition to term frequency features, we derive features from the result of lexical chaining to capture the lexical cohesive information, where E-HowNet lexical database is used to compute semantic similarity between nouns with high word frequency. Classification models for assessing readability of Chinese text are learned from the features using support vector machines. We select articles from textbooks of elementary schools to train and test the classification models. The experiments compare the prediction results of different sets of features.
В статье рассматривается проблема нормы и нормативного подхода к языку в диахроническом плане.Определяется специфика нормативного похода к языковым средствам в различных лингвистических традициях и выявляются основные характеристики лингвистической нормы.В статье указывается, что на каждом этапе развития языка складываются свои нормы как резуль
The paper presents the process of constructing a publicly available treebank of public messages written in Croatian. The messages were collected from various electronic sources – e-mail, blog, Facebook and SMS – and published on the Zagreb Museum of Contemporary Art LED facade within the Babel art project. The project aimed to use the facade as an open-space blog or social interface for enabling citizens to publicly express their views. Construction and current state of the treebank is presented along with future work plans. A comparison of Babel Treebank with Croatian Dependency Treebank and SETimes.HR treebank regarding differing domains and annotation schemes is briefly sketched. The treebank is used as a test platform for introducing a new standard for syntactic annotation of Croatian texts. An experiment with morphosyntactic tagging and dependency parsing of the treebank is conducted, providing first insight to computational processing of non-standard text in Croatian.
This paper presents a reranking approach to combining constituent and dependency parsing, aimed at improving parsing performance on both sides. Most previous combination methods rely on complicated joint decoding to integrate graph- and transition-based dependency models. Instead, our approach makes use of a high-performance probabilistic context free grammar (PCFG) model to output k-best candidate constituent trees, and then a dependency parsing model to rerank the trees by their scores from both models, so as to get the most probable parse. Experimental results show that this reranking approach achieves the highest accuracy of constituent and dependency parsing on Chinese treebank (CTB5.1) and a comparable performance to the state of the art on English treebank (WSJ).
The problem of Vietnamese syntactic parsing, especially constituency parsing, has recently been tackled by several research groups. A common effort of the Vietnamese language processing community has allowed the creation of VietTreebank, a reference parsed corpus containing about 10,000 sentences for the constituency parsing task. In this paper, we present our work to build a reference treebank, based on VietTreebank, for the dependency parsing task, which has not yet been very well studied for Vietnamese. First we define a dependency label set by adapting the dependency schema developed by the NLP group at Stanford university and taking into account the particularities of Vietnamese grammar. Then we propose an algorithm to convert a constituency treebank to a dependency one. The algorithm is tested on a set of 100 sentences of VietTreebank corpus and gives very good results. Finally, we carry out an experiment on Vietnamese dependency parsing using MaltParser tool and the dependency treebank converted from VietTreebank.
Data-driven research in linguistics typically involves the processes of data annotation, data visualization and identification of relevant patterns. We describe our experience in incorporating these processes at an undergraduate course on language information technology. Students collectively annotated the syntactic structures of a set of Classical Chinese poems; the resulting treebank was put on a platform for corpus search and visualization; finally, using this platform, students investigated research questions about the text of the treebank. 1
В статье рассматривается проблема нормы и нормативного подхода к языку в диахроническом плане.Определяется специфика нормативного похода к языковым средствам в различных лингвистических традициях и выявляются основные характеристики лингвистической нормы.В статье указывается, что на каждом этапе развития языка складываются свои нормы как резуль
We describe Abstract Meaning Representation (AMR), a semantic representation language in which we are writing down the meanings of thousands of English sentences. We hope that a sembank of simple, whole-sentence semantic structures will spur new work in statistical natural language understanding and generation, like the Penn Treebank encouraged work on statistical parsing. This paper gives an overview of AMR and tools associated with it.
The latter part of the lifespan is commonly associated with a decline of cognitive functions, but also with changes in emotional responding.To explore the effect of age on processing of emotional stimuli, we used a two-task design.In a stimulus-rating task, we investigated the emotional responses to 15 different schematic facial emotional stimuli (one neutral, seven positive, seven negative) on Arousal, Valence and Potency measures in 20 younger (21-32 yrs, M=26, SD=3.7) and 20 older (65-81 yrs, M=72, SD=4.9) participants.In a visual attention task, we used the same 15 stimuli in a visual search paradigm to investigate differences between younger and older participants in how the emotional properties of these emotional stimuli influence visual attention.The results from the stimulus-rating task showed significantly reduced range in responses to emotional stimuli in the older compared to the younger group.This difference was found on both emotional Arousal and Potency measures, but not on emotional Valence measures; indicating an age-related flattening of affect on two of the three emotional key dimensions.The results from the visual search task showed -apart from the general extension of response latencies in older -no general emotion-related differences between how emotional stimuli influences attention in the younger and older groups.Analysis of the relationships between attention and emotion measures showed that higher ratings on Arousal and Potency were associated with both shorter reaction times and fewer errors in the attention task.This correlation was age-independent, indicating a similar influence from emotional Arousal on detection of angry faces in younger and older adults.
Aiming at the area of machine translation applications,this paper conduct research on the construction of Chinese Sentence-Category Dependency Treebank(CSCDT) based on the theory of hierarchical network of concepts.Conceptual category tagset and sentence-category relation tagset for the treebank are presented also with the example tree of CSCDT.
A time-windowing feature extraction approach based on time-frequency (TF) analysis is adopted here to investigate the time-course of the discrimination between musical appraisal electroencephalogram (EEG) responses, under the parameter of familiarity. An EEG data set, formed by the responses of nine subjects during music listening, along with self-reported ratings of liking and familiarity, is used. Features are extracted from the beta (13-30 Hz) and gamma (30-49 Hz) EEG bands in time windows of various lengths, by employing three TF distributions (spectrogram, Hilbert-Huang spectrum, and Zhao-Atlas-Marks transform). Subsequently, two classifiers (k-NN and SVM) are used to classify feature vectors in two categories, i.e., "likeâ and "dislike,â under three cases of familiarity, i.e., regardless of familiarity (LD), familiar music (LDF), and unfamiliar music (LDUF). Key findings show that best classification accuracy (CA) is higher and it is achieved earlier in the LDF case {91.02 ± 1.45% (7.5-10.5 s)} as compared to the LDUF case {87.10 ± 1.84% (10-15 s)}. Additionally, best CAs in LDF and LDUF cases are higher as compared to the general LD case {85.28 ± 0.77%}. The latter results, along with neurophysiological correlates, are further discussed in the context of the existing literature on the time-course of music-induced affective responses and the role of familiarity.
In this paper, we propose a method for au-tomatic clause boundary annotation in the Hindi Dependency Treebank. We show that the clausal information implicitly encoded in a dependency structure can be made explicit with no or less human interven-tion. We exercised the proposed approach on 16,000 sentences of Hindi Dependency Treebank. Our approach gives an accuracy of 94.44 % for clause boundary identifica-tion evaluated over 238 clauses. The resul-tant corpus has varied usages and can be utilized for developing a statistical clause boundary identifier. 1
Slovene Lexical Database was created between 2008 and 2012 and represents a comprehensive syntactic and semantic description of a selected set of Slovene words. The description was based exclusively on the analysis of reference corpora of Slovene. The database is structured as a network of interrelated semantic and syntactic information about a particular word. Semantic level represents the top level in the hierarchy with the lexical unit as its core element. This includes all senses of the headwrd, multi-word expressions and phraseological units. Each sense is described with a short semantic indicator and/or whole-sentence definition which includes typical syntactic environment of the headword with the relevant number, form and semantic types in a valency frame (semantic frame). These are also reflected in a number of syntactic structures and corresponding collocations. All the higher types of information are confirmed by a selection of corpus examples. Multi-word expressions and phraseological units are treated independently from particular senses of the headword and have their own internal structure which requires the same types of information as single-word entries or senses.