Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Using neural networks to estimate the probabilities of word sequences has shown significant promise for statistical language modeling. Typical modeling methods include multi-layer neural networks, log-bilinear networks and recurrent neural networks, etc. In this paper, we propose the temporal kernel neural network language model, a variant of models mentioned above. This model explicitly captures long-term dependencies of words with exponential kernel, where the memory of history is decayed exponentially. Additionally, several sentences with variable lengths as a mini-batch are efficiently implemented for speeding up. Experimental results show that the proposed model is very competitive to the recurrent neural network language model and obtains the lower perplexity of 111.6 (more than 10% reduction) than the state-of-the-art results reported in the standard Penn Treebank Corpus. We further apply this model to Wall Street Journal speech recognition task, and observe significant improvements in word error rate.
This paper proposes a combined model for POS tagging, dependency parsing and co-reference resolution for Bulgarian — a pro-drop Slavic language with rich mor-phosyntax. We formulate an extension of the MSTParser algorithm that allows the simultaneous handling of the three tasks in a way that makes it possible for each task to benefit from the information available to the others, and conduct a set of experi-ments against a treebank of the Bulgarian language. The results indicate that the pro-posed joint model achieves state-of-the-art performance for POS tagging task, and outperforms the current pipeline solution. 1
It has been observed that the inclusion of morphosyntactic information in dependency treebanks is crucial to obtain high results in dependency parsing for some languages. In this paper we explore in depth to what extent it is useful to include morphological features, and the impact of diverse morphosyntactic annotations on statistical dependency parsing of Spanish. For this, we give a detailed analysis of the results of over 80 experiments performed with MaltParser through the application of MaltOptimizer. Our goal is to isolate configurations of morphosyntactic features which would allow for optimizing the parsing of Spanish texts, and to evaluate the impact that each feature has, independently and in combination with others. 1
We present an empirical study on constructing a Japanese constituent parser, which can output function labels to deal with more detailed syntactic information.Japanese syntactic parse trees are usually represented as unlabeled dependency structure between bunsetsu chunks, however, such expression is insufficient to uncover the syntactic information about distinction between complements and adjuncts and coordination structure, which is required for practical applications such as syntactic reordering of machine translation.We describe a preliminary effort on constructing a Japanese constituent parser by a Penn Treebank style treebank semi-automatically made from a dependency-based corpus.The evaluations show the parser trained on the treebank has comparable bracketing accuracy as conventional bunsetsu-based parsers, and can output such function labels as the grammatical role of the argument and the type of adnominal phrases.
Many countries use national-level surveys to capture student opinions about their university experiences. It is necessary to interpret survey results in an appropriate context to inform decision-making at many levels. To provide context to national survey outcomes, we describe patterns in the ratings of science and engineering subjects from the UK’s National Student Survey (NSS). New, robust statistical models describe relationships between the Overall Satisfaction’ rating and the preceding 21 core survey questions. Subjects exhibited consistent differences and ratings of “Teaching”, “Organisation” and “Support” were thematic predictors of “Overall Satisfaction” and the best single predictor was “The course was well designed and running smoothly”. General levels of satisfaction with feedback were low, but questions about feedback were ultimately the weakest predictors of “Overall Satisfaction”. The UK’s universities affiliated groupings revealed that more traditional “1994” and “Russell” groups over-performed in a model using the core 21 survey questions to predict “Overall Satisfaction”, in contrast to the under-performing newer universities in the Million+ and Alliance groups. Findings contribute to the debate about “level playing fields” for the interpretation of survey outcomes worldwide in terms of differences between subjects, institutional types and the questionnaire items.
In this paper, we investigate errors in syntax annotation with the Turku Dependency Treebank, a recently published treebank of Finnish, as study material. This treebank uses the Stanford Dependency scheme as its syntax representation, and its published data contains all data created in the full double annotation as well as timing information, both of which are necessary for this study.
In this paper, we provide a quantitative analysis of non-projective constructions attested in the Ancient Greek Dependency Treebank (AGDT). We consider the different types of formal constraints and metrics that have become standardized in the literature on non-projectivity (planarity, wellnestedness, gap-degree, edge-degree). We also discuss some of the linguistic factors that cause non-projective edges in Ancient Greek. Our results confirm the remarkable extension of non-projectivity in the AGDT, both in terms of quantitative incidence of non-projective nodes and for their complexity, which is not paralleled by the corpora of modern languages considered in the literature. At the same time, the usefulness of other constraint (especially well-nestedness) is confirmed by our researches. 1
We investigate statistical dependency parsing of two closely related languages, Croatian and Serbian.As these two morphologically complex languages of relaxed word order are generally under-resourced -with the topic of dependency parsing still largely unaddressed, especially for Serbian -we make use of the two available dependency treebanks of Croatian to produce state-of-the-art parsing models for both languages.We observe parsing accuracy on four test sets from two domains.We give insight into overall parser performance for Croatian and Serbian, impact of preprocessing for lemmas and morphosyntactic tags and influence of selected morphosyntactic features on parsing accuracy.
Information Structure (IS) determines the “communicative” segmentation of the meaning of an utterance, which makes it central to the semantics‐syntax‐ intonation interface and therefore also to NLP. Despite this relevance, IS has not received much attention in the context of the majority of the reference treebanks for data-driven NLP that already contain a semantic and syntactic layers of annotation. We present our work in progress on the annotation of the Penn TreeBank with the thematicity dimension of the IS as defined in the Meaning-Text Theory. We experiment with tagging and transitionbased parsing techniques. Especially the latter achieve acceptable accuracy with even very small training samples, which is promising for languages with scarce resources.
In this paper, we discuss our efforts to anno-tate nominals in the Hindi Treebank with the semantic property of animacy. Although the treebank already encodes lexical information at a number of levels such as morph and part of speech, the addition of animacy informa-tion seems promising given its relevance to varied linguistic phenomena. The suggestion is based on the theoretical and computational analysis of the property of animacy in the con-text of anaphora resolution, syntactic parsing, verb classification and argument differentia-tion. 1
Slovene Lexical Database was created between 2008 and 2012 and represents a comprehensive syntactic and semantic description of a selected set of Slovene words. The description was based exclusively on the analysis of reference corpora of Slovene. The database is structured as a network of interrelated semantic and syntactic information about a particular word. Semantic level represents the top level in the hierarchy with the lexical unit as its core element. This includes all senses of the headwrd, multi-word expressions and phraseological units. Each sense is described with a short semantic indicator and/or whole-sentence definition which includes typical syntactic environment of the headword with the relevant number, form and semantic types in a valency frame (semantic frame). These are also reflected in a number of syntactic structures and corresponding collocations. All the higher types of information are confirmed by a selection of corpus examples. Multi-word expressions and phraseological units are treated independently from particular senses of the headword and have their own internal structure which requires the same types of information as single-word entries or senses.
In this paper, we propose a method for au-tomatic clause boundary annotation in the Hindi Dependency Treebank. We show that the clausal information implicitly encoded in a dependency structure can be made explicit with no or less human interven-tion. We exercised the proposed approach on 16,000 sentences of Hindi Dependency Treebank. Our approach gives an accuracy of 94.44 % for clause boundary identifica-tion evaluated over 238 clauses. The resul-tant corpus has varied usages and can be utilized for developing a statistical clause boundary identifier. 1
Nous présenterons les différentes couches d'annotation du treebank Rhapsodie, un corpus de français parlé richement annoté. Le corpus contient plusieurs niveaux de segmentation indépendants: en unités illocutoires pour la macrosyntaxe, en unités rectionnelles pour la microsyntaxe, en périodes, paquets intonatifs et groupes accentuels pour la prosodie. Les unités rectionnelles sont analysées en dépendance, avec un traitement fin des phénomènes d'entassements (coordination, reformulation, négo...
Aiming at the area of machine translation applications,this paper conduct research on the construction of Chinese Sentence-Category Dependency Treebank(CSCDT) based on the theory of hierarchical network of concepts.Conceptual category tagset and sentence-category relation tagset for the treebank are presented also with the example tree of CSCDT.
В статье рассматривается проблема нормы и нормативного подхода к языку в диахроническом плане.Определяется специфика нормативного похода к языковым средствам в различных лингвистических традициях и выявляются основные характеристики лингвистической нормы.В статье указывается, что на каждом этапе развития языка складываются свои нормы как резуль
This paper investigates the appropriateness of using lexical cohesion analysis to assess Chinese readability. In addition to term frequency features, we derive features from the result of lexical chaining to capture the lexical cohesive information, where E-HowNet lexical database is used to compute semantic similarity between nouns with high word frequency. Classification models for assessing readability of Chinese text are learned from the features using support vector machines. We select articles from textbooks of elementary schools to train and test the classification models. The experiments compare the prediction results of different sets of features.
This paper introduces an advanced, efficient approach for rule based English to Bengali (E2B) machine translation (MT), where Penn-Treebank parts of speech (PoS) tags, HMM (Hidden Markov Model) Tagger is used.Fuzzy-If-Then-Rule approach is used to select the lemma from rule-based-knowledge. The proposed E2B-MT has been tested through F-Score measurement, and the accuracy is more than eighty percent.
We present, here, our analysis of systematic divergences in parallel English-Hindi dependency treebanks based on the Computational Paninian Grammar (CPG) framework. Study of structural divergences in parallel treebanks not only helps in developing larger treebanks automatically, but can also be useful for many NLP applications such as data-driven machine translation (MT) systems. Given that the two treebanks are based on the same grammatical model, a study of divergences in them could be of advantage to such tasks, along with making it more interesting to study how and where they diverge. We consider two parallel trees divergent based on differences in constructions, relations marked, frequency of annotation labels and tree depth. Some interesting instances of structural divergences in the treebanks have been discussed in the course of this paper. We also present our task of alignment of the two treebanks, wherein we talk about our extraction of divergent structures in the trees, and discuss the results of this exercise. 1
Although several syntactically annotated corpora (or treebanks) exist for Dutch, they are seldomly used for descriptive linguistic research because there are no easy-to-use exploitation tools available.This demonstration paper describes GrETEL, a linguistic search engine (http:// nederbooms.ccl.kuleuven.be/eng/gretel)that enables non-technical users to consult treebanks in a user-friendly way.Instead of a formal search expression, a natural language example is used as input to the system, allowing users to search for similar constructions as the example they provide.In the first version of GrETEL, only written Dutch (LASSY) was included.Based on user requests we have now included the Spoken Dutch Corpus (CGN) as well.
This paper presents our preliminary conclusions as part of an ongoing effort to construct a new dependency representation framework for Turkish.We aim for this new framework to accommodate the highly agglutinative morphology of Turkish as well as to allow the annotation of unedited web data, and shape our decisions around these considerations.In this paper, we firstly describe a novel syntactic representation for morphosyntactic sub-word units (namely inflectional groups (IGs) in Turkish) which allows inter-IG relations to be discerned with perfect accuracy without having to hide lexical information.Secondly, we investigate alternative annotation schemes for coordination structures and present a better scheme (nearly 11% increase in recall scores) than the one in Turkish Treebank (Oflazer et al., 2003) for both parsing accuracies and compatibility for colloquial language.
This paper discusses the extension of a sys-tem developed for automatic discovery of tree-bank annotation inconsistencies over an entire corpus to the particular case of evaluation of inter-annotator agreement. This system makes for a more informative IAA evaluation than other systems because it pinpoints the incon-sistencies and groups them by their structural types. We evaluate the system on two corpora- (1) a corpus of English web text, and (2) a corpus of Modern British English. 1
National audience
With the growing interest in statistical parsing, special attention has recently been devoted to the problem of comparing different treebanks to assess which languages or domains are more difficult to parse relative to a given model. A common methodology for comparing parsing difficulty across treebanks is based on the use of the standard labeled precision and recall measures. As an alternative, in this article we propose an information-theoretic measure, called the expected conditional cross-entropy (ECC). One important advantage with respect to standard performance measures is that ECC can be directly expressed as a function of the parameters of the model. We evaluate ECC across several treebanks for English, French, German, and Italian, and show that ECC is an effective measure of parsing difficulty, with an increase in ECC always accompanied by a degradation in parsing accuracy.
The present paper focuses on ways in which the pragmatic (functional) meaning that arises from various contextual features, known in corpus linguistics as semantic prosody, can become an integral part of lexicographical descriptions as they are represented in the Slovene Lexical Database (SLD). This is particularly important for the treatment of phraseology and idiomatics. First, the theoretical background is provided, with the focus on the prototype theory and its practical implications for monolingual lexicography. A parallel is drawn with the model of meaning analysis in the SLD. The second part begins with a brief introduction to semantic prosody and continues with an analysis of monolingual meaning descriptions in the SLD against a number of authentic corpus examples, investigating how their pragmatic components have been identified. The analysis of corpus data shows that pragmatics is an important contributor to the process of sense discrimination in works of lexical and lexicographic relevance.
peer reviewed
ABSTRACT Online travel reviews are emerging as a powerful source of information affecting tourists' pre-purchase evaluation of a hotel organization. This trend has highlighted the need for a greater understanding of the impact of online reviews on consumer attitudes and behaviors. In view of this need, we investigate the influence of online hotel reviews on consumers' attributions of service quality and firms' ability to control service delivery. An experimental design was used to examine the effects of four independent variables: framing; valence; ratings; and target. The results suggest that in reviews evaluating a hotel, remarks related to core services are more likely to induce positive service quality attributions. Recent reviews affect customers' attributions of controllability for service delivery, with negative reviews exerting an unfavorable influence on consumers' perceptions. The findings highlight the importance of managing the core service and the need for managers to act promptly in addressing customer service problems.
Large-scale linguistically annotated cor-pora have played a crucial role in advanc-ing the state of the art of key natural lan-guage technologies such as syntactic, se-mantic and discourse analyzers, and they serve as training data as well as evaluation benchmarks. Up till now, however, most of the evaluation has been done on mono-lithic corpora such as the Penn Treebank, the Proposition Bank. As a result, it is still unclear how the state-of-the-art analyzers perform in general on data from a vari-ety of genres or domains. The completion of the OntoNotes corpus, a large-scale, multi-genre, multilingual corpus manually annotated with syntactic, semantic and discourse information, makes it possible to perform such an evaluation. This paper presents an analysis of the performance of publicly available, state-of-the-art tools on all layers and languages in the OntoNotes v5.0 corpus. This should set the bench-mark for future development of various NLP components in syntax and semantics, and possibly encourage research towards an integrated system that makes use of the various layers jointly to improve overall performance. 1
This study examines the impact of online word of mouth (WOM) and expert reviews on movies' box office revenues, both in the U.S. domestic market and in the international markets. Using a sample of 169 movies released in 2008, the study discovered that the frequency of online WOM and the valence rating of expert reviews were significant factors for box office outcomes in the domestic market. The study also found that only the frequency of online WOM was a significant factor in the international markets. The findings suggest that online WOM and expert reviews play a critical role in moviegoers' consumption behavior in the age of the Internet and social media.
This paper investigates the effect of the label bias problem of maximum entropy Markov models for part-of-speech tagging, a typical sequence prediction task in natural language processing. This problem has been underexploited and underappreciated. The investigation reveals useful information about the entropy of local transition probability distributions of the tagging model which enables us to exploit and quantify the label bias effect of part-of-speech tagging. Experiments on a Vietnamese treebank and on a French treebank show a significant effect of the label bias problem in both of the languages.
People believe that women are more emotionally intense than men, but the scientific evidence is equivocal. In this study, we tested the novel hypothesis that men and women differ in the neural correlates of affective experience, rather than in the intensity of neural activity, with women being more internally (interoceptively) focused and men being more externally (visually) focused. Adult men (n = 17) and women (n = 17) completed a functional magnetic resonance imaging study while viewing affectively potent images and rating their moment-to-moment feelings of subjective arousal. We found that men and women do not differ overall in their intensity of moment-to-moment affective experiences when viewing evocative images, but instead, as predicted, women showed a greater association between the momentary arousal ratings and neural responses in the anterior insula cortex, which represents bodily sensations, whereas men showed stronger correlations between their momentary arousal ratings and neural responses in the visual cortex. Men also showed enhanced functional connectivity between the dorsal anterior insula cortex and the dorsal anterior cingulate cortex, which constitutes the circuitry involved with regulating shifts of attention to the world. These results demonstrate that the same affective experience is realized differently in different people, such that women's feelings are relatively more self-focused, whereas men's feelings are relatively more world-focused.
We present a reformulation of the word pair features typically used for the task of disambiguating implicit relations in the Penn Discourse Treebank. Our word pair features achieve significantly higher performance than the previous formulation when evaluated without additional features. In addition, we present results for a full system using additional features which achieves close to state of the art performance without resorting to gold syntactic parses or to context outside the relation.
We present an affective text analysis model that can directly estimate and combine affective ratings of multi-word terms, with application to the problem of sentence polarity/semantic orientation detection. Starting from a hierarchical compositional method for generating sentence ratings, we expand the model by adding multi-word terms that can capture non-compositional semantics. The method operates similarly to a bigram language model, using bigram terms or backing off to unigrams based on a (degree of) compositionality criterion. The affective ratings for n-gram terms of different orders are estimated via a corpus-based method using distributional semantic similarity metrics between unseen words and a set of seed words. N-gram ratings are then combined into sentence ratings via simple algebraic formulas. The proposed framework produces state-of-the-art results for word-level tasks in English and German and the sentence-level news headlines classification SemEval'07-Task14 task. The inclusion of bigram terms to the model provides significant performance improvement, even if no term selection is applied.
We present a novel method ("waste") for the segmentation of text into tokens and sentences. Our approach makes use of a Hidden Markov Model for the detection of segment boundaries. Model parameters can be estimated from pre-segmented text which is widely available in the form of treebanks or aligned multi-lingual corpora. We formally define the waste boundary detection model and evaluate the system's performance on corpora from various languages as well as a small corpus of computer-mediated communication.
Psycholinguistic research shows that key properties of the human sentence processor are incrementality, connectedness (partial structures contain no unattached nodes), and prediction (upcoming syntactic structure is anticipated). There is currently no broad-coverage parsing model with these properties, however. In this article, we present the first broad-coverage probabilistic parser for PLTAG, a variant of TAG that supports all three requirements. We train our parser on a TAG-transformed version of the Penn Treebank and show that it achieves performance comparable to existing TAG parsers that are incremental but not predictive. We also use our PLTAG model to predict human reading times, demonstrating a better fit on the Dundee eye-tracking corpus than a standard surprisal model.
Large-scale unlabeled data contains abundant lexical information for NLP tasks such as Chinese word segmentation and POS tagging.This work extracted high-dimensional distributional lexical information from a largescale unlabeled Chinese corpus.An auto-encoder then performed the unsupervised dimension reduction.The learned low-dimensional lexicon features were used as new lexical features for a joint Chinese word segmentation and POS tagging task.Experiments on the Chinese Treebank 5corpus showed that the additional lexicon features improve the performance and are better than those features learned by using the principal component analysis and the k-means algorithm.
Grote verzamelingen van vertaalde teksten – zogenaamde parallelle corpora - worden vaak automatisch op zins- en woordniveau gealigneerd om automatische vertaalsystemen op te trainen. Soms voegt men ook automatisch syntactische bomen aan de zinnen toe om meer taalkundige informatie eruit te kunnen halen. Als die bomen aan beide kanten verschijnen en de boomknopen ook worden gealigneerd, is er sprake van een parallelle treebank. De beste vertaalsystemen zijn bijna of helemaal puur statistisch, maar in recente jaren ontstond er een grotere nadruk op de integratie van meer taalkundig gemotiveerde data, waaronder ook het gebruik van parallel treebanks. Ze zijn echter alleen op een zeer grote schaal bruikbaar, omdat er door zo een systeem veel te leren is van hoe een taal typisch naar een andere moet worden omgezet. Daarom onderzoeken we technieken om automatisch de boomknopen accuraat te aligneren. Een bijkomend motief is het feit dat parallel treebanks ook voor andere applicaties bruikbaar zijn en als taalbronnen zelf van wetenschappelijk belang zijn. Het hele proces van het aligneren van knopen noemen wij tree alignment. Wij vinden dat een combinatie van statistiche en regelgebaseerde technieken met relatief weinig trainingsgegevens en weinig features zeer accurate alignments kan produceren. Ten slotte vinden we dat, wanneer wij alignments die relatief heel veel knopen aligneren – al zijn sommigen soms fout – op een syntactisch gebaseerde systeem toepassen, dat tot verbeterde automatische vertaling leidt, in vergelijking met hetzelfde systeem die op minder maar meer accurate alignments getrained is.
BACKGROUND: Patients can make valuable contributions towards promoting the safety of their health care. Health care professionals (HCPs) could play an important role in encouraging patient involvement in safety-relevant behaviours. However, to date factors that determine HCPs' attitudes towards patient participation in this area remain largely unexplored. OBJECTIVE: To investigate predictors of HCPs' attitudes towards patient involvement in safety-relevant behaviours. DESIGN: A 22-item cross-sectional fractional factorial survey that assessed HCPs' attitudes towards patient involvement in relation to two error scenarios relating to hand hygiene and medication safety. SETTING: Four hospitals in London PARTICIPANTS: Two hundred sixteen HCPs (116 doctors; 100 nurses) aged between 21 and 60 years (mean: 32): 129 female. OUTCOME MEASURES: Approval of patient's behaviour, HCP response to the patient, anticipated effects on the patient-HCP relationship, support for being asked as a HCP, affective rating response to the vignettes. RESULTS: HCPs elicited more favourable attitudes towards patients intervening about a medication error than about hand sanitation. Across vignettes and error scenarios, the strongest predictors of attitudes were how the patient intervened and how the HCP responded to the patient's behaviour. With regard to HCP characteristics, doctors viewed patients intervening less favourably than nurses. CONCLUSIONS: HCPs perceive patients intervening about a potential error less favourably if the patient's behaviour is confrontational in nature or if the HCP responds to the patient intervening in a discouraging manner. In particular, if a HCP responds negatively to the patient (irrespective of whether an error actually occurred), this is perceived as having negative effects on the HCP-patient relationship.
Dependency analysis relies on morphosyntactic evidence, as well as semantic evidence. In some cases, however, morphosyntactic evidence seems to be in conflict with semantic evidence. For this reason dependency grammar theories, annotation guidelines and tree-to-dependency conversion schemes often differ in how they analyze various syntactic constructions. Most experiments for which constituent-based treebanks such as the Penn Treebank are converted into dependency treebanks rely blindly on one of four-five widely used tree-to-dependency conversion schemes. This paper evaluates the down-stream effect of choice of conversion scheme, showing that it has dramatic impact on end results. 1