Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
Abstract This short reply seeks to clarify the concept of linguistic norm circles and to correct some misunderstandings of it implicit in Sealey & Carter's response. It also reinforces some doubts over their version of the linguistic system. Norm circles, it argues, provide an important part of the explanation for linguistic practices, but always in conjunction with other interacting causal powers.
It has recently been shown that different NLP models can be effectively combined using dual decomposition.In this paper we demonstrate that PCFG-LA parsing models are suitable for combination in this way.We experiment with the different models which result from alternative methods of extracting a grammar from a treebank (retaining or discarding function labels, left binarization versus right binarization) and achieve a labeled Parseval F-score of 92.4 on Wall Street Journal Section 23 -this represents an absolute improvement of 0.7 and an error reduction rate of 7% over a strong PCFG-LA product-model baseline.Although we experiment only with binarization and function labels in this study, there is much scope for applying this approach to other grammar extraction strategies.
Tree Substitution Grammar rules form a large and expressive class of features capable of representing syntactic and lexical patterns that provide evidence of an author’s native language. However, this class of features can be applied to any general constituent based model of grammar and previous work has done little to explore these options, relying primarily on the common Penn Treebank annotation standard. In this work we contrast the performance of syntactic features for Native Language Indentification using five different formalisms. The use of different formalisms captures complementary information from second language data, and can be used in combination to yield classification performance superior to any formalism taken on its own. 1
Recognizing negative emotions is impaired in Huntington’s disease (HD), while the subjective emotional involvement in affective pictures images is not completely known in these patients. We aimed to further evaluate emotional reaction in early HD patients, by means of subjective arousal and valence rating and EEG changes induced by International Affective Pictures (IAPS). We recruited 16 consecutive genetically confirmed HD outpatients, and 16 sex and age matched controls. Eighty-four color slides, 28 pleasant, 28 unpleasant and 28 neutral images, in random presentation, were chosen from the International Affective Picture System. EEG was recorded by 32 scalp electrodes. HD patients judged negative and positive affective images respectively more unpleasant and pleasant than controls. They also exhibited higher arousal for pictures, independently from their affective content. However, in HD patients we observed a reduced positivity in the 400-700 m.sec intervals during unpleasant pictures viewing. Present findings may suggest that emotional impact related to affective images is preserved in HD, but it coexists with impairment of late cortical processing following unpleasant stimuli.
We present an automatic animacy classifier for Dutch that can determine the animacy status of nouns — how alive the noun’s referent is (human, inanimate, etc.). Animacy is a semantic property that has been shown to play a role in human sentence processing, felicity and grammaticality. Although animacy is not marked explicitly in Dutch, we expect knowledge about animacy to be helpful for parsing, translation and other NLP tasks. Only a few animacy classifiers and animacy-annotated corpora exist internationally. For Dutch, animacy information is only available in the Cornetto lexical-semantic database. We augment this lexical information with context information from the Dutch Lassy Large treebank, to create training data for an animacy classifier that uses a novel kind of context features. We use the k-nearest neighbour algorithm with distributional lexical features, e.g. how fre-quently the noun occurs as a subject of the verb ‘to think ’ in a corpus, to decide on the (pre-dominant) animacy class. The size of the Lassy Large corpus makes this possible, and the high level of detail these word association features provide, results in accurate Dutch-language animacy classification. 1.
We can determine whether two texts are paraphrases of each other by finding out the extent to which the texts are similar. The typical lexical matching technique works by matching the sequence of tokens between the texts to recognize paraphrases, and fails when different words are used to convey the same meaning. We can improve this simple method by combining lexical with syntactic or semantic representations of the input texts. The present work makes use of syntactical information in the texts and computes the similarity between them using word similarity measures based on WordNet and lexical databases. The texts are converted into a unified semantic structural model through which the semantic similarity of the texts is obtained. An approach is presented to assess the semantic similarity and the results of applying this approach is evaluated using the Microsoft Research Paraphrase (MSRP) Corpus.
Abstract This paper explores the textual-linguistic norms evident in the translation of culturally specific material in a sample of translated South African children's books in Afrikaans and English, with a view to investigating the tensions between domestication and foreignisation, particularly as related to different types of books, such as primers, local picture books, and international picture books. A detailed qualitative textual analysis of micro-level translation choices relating to cultural orientation is presented, comparing the 21 translations in the sample with their source texts, and comparing subsamples of different types of books with one another. The analysis suggests the prevalence of hybrid translation strategies that orient translated texts in multiple cultural directions, but also indicates potentially significant differences in this regard between different types of books, with translations of international picture books tending towards greater use of domesticating strategies, despite their generally culturally generic background.
This paper describes our integration efforts in two Northern European language infrastructures. Specifically, this work has been a collaboration between the META-NORD team at the University of Bergen and the INESS project, a large treebanking infrastructure project in Norway, in developing and documenting two complex resources, as well as making these accessible to the R&D community.
BACKGROUND: Patients can make valuable contributions towards promoting the safety of their health care. Health care professionals (HCPs) could play an important role in encouraging patient involvement in safety-relevant behaviours. However, to date factors that determine HCPs' attitudes towards patient participation in this area remain largely unexplored. OBJECTIVE: To investigate predictors of HCPs' attitudes towards patient involvement in safety-relevant behaviours. DESIGN: A 22-item cross-sectional fractional factorial survey that assessed HCPs' attitudes towards patient involvement in relation to two error scenarios relating to hand hygiene and medication safety. SETTING: Four hospitals in London PARTICIPANTS: Two hundred sixteen HCPs (116 doctors; 100 nurses) aged between 21 and 60 years (mean: 32): 129 female. OUTCOME MEASURES: Approval of patient's behaviour, HCP response to the patient, anticipated effects on the patient-HCP relationship, support for being asked as a HCP, affective rating response to the vignettes. RESULTS: HCPs elicited more favourable attitudes towards patients intervening about a medication error than about hand sanitation. Across vignettes and error scenarios, the strongest predictors of attitudes were how the patient intervened and how the HCP responded to the patient's behaviour. With regard to HCP characteristics, doctors viewed patients intervening less favourably than nurses. CONCLUSIONS: HCPs perceive patients intervening about a potential error less favourably if the patient's behaviour is confrontational in nature or if the HCP responds to the patient intervening in a discouraging manner. In particular, if a HCP responds negatively to the patient (irrespective of whether an error actually occurred), this is perceived as having negative effects on the HCP-patient relationship.
BACKGROUND: In Alzheimer's disease (AD), some patients present with cognitive impairment other than episodic memory disturbances. We evaluated whether occurrence of posterior atrophy (PA) and medial temporal lobe atrophy (MTA) could account for differences in cognitive domains affected. METHODS: In 329 patients with AD, we assessed five cognitive domains: memory, language, visuospatial functioning, executive functioning, and attention. Magnetic resonance imaging (MRI) was rated visually for the presence of MTA and PA. Two-way analyses of variance were performed with MTA and PA as independent variables, and cognitive domains as dependent variables. Gender, age, and education were covariates. As PA is often encountered in younger patients, analyses were repeated after stratification for age of onset (early onset, ≤65 years). RESULTS: The mean age of the participants was 67 years, 175 (53%) were female, and the mean Mini-Mental State Examination (score±standard deviation) was 20±5 points. Based on dichotomized magnetic resonance imaging ratings, 84 patients (26%) had MTA and PA, 98 (30%) had MTA, 57 (17%) had PA, and 90 (27%) had neither. MTA was associated with worse performance on memory, language, and attention (all, P<.05), and PA was associated with worse performance on visuospatial and executive functioning (both, P<.05). Stratification for age showed in patients with late-onset AD (n=173) associations between MTA and impairment on memory, language, visuospatial functioning, and attention (all, P<.05); in early-onset AD (n=156), patients with PA tended to perform worse on visuospatial functioning. CONCLUSIONS: Regional atrophy is related to impairment in specific cognitive domains in AD. The prevalence of PA in a large set of patients with AD and its association with cognitive functioning provides support for the usefulness of this visual rating scale in the diagnostic evaluation of AD.
Large-scale unlabeled data contains abundant lexical information for NLP tasks such as Chinese word segmentation and POS tagging.This work extracted high-dimensional distributional lexical information from a largescale unlabeled Chinese corpus.An auto-encoder then performed the unsupervised dimension reduction.The learned low-dimensional lexicon features were used as new lexical features for a joint Chinese word segmentation and POS tagging task.Experiments on the Chinese Treebank 5corpus showed that the additional lexicon features improve the performance and are better than those features learned by using the principal component analysis and the k-means algorithm.
Nous présenterons les différentes couches d'annotation du treebank Rhapsodie, un corpus de français parlé richement annoté. Le corpus contient plusieurs niveaux de segmentation indépendants: en unités illocutoires pour la macrosyntaxe, en unités rectionnelles pour la microsyntaxe, en périodes, paquets intonatifs et groupes accentuels pour la prosodie. Les unités rectionnelles sont analysées en dépendance, avec un traitement fin des phénomènes d'entassements (coordination, reformulation, négo...
We propose a new variant of Tree-Adjoining Grammar that allows adjunction of full wrapping trees but still bears only context-free expressivity. We provide a transformation to context-free form, and a further reduction in probabilistic model size through factorization and pooling of parameters. This collapsed context-free form is used to implement efficient grammar estimation and parsing algorithms. We perform parsing experiments the Penn Treebank and draw comparisons to Tree-Substitution Grammars and between different variations in probabilistic model design. Examination of the most probable derivations reveals examples of the linguistically relevant structure that our variant makes possible. 1
ABSTRACT. In this article, we describe our research on wide-coverage semantics for Frenchlanguage texts and on its application to produce detailed semantic descriptions of itineraries. Using a categorial grammar semi-automatically extracted from the French Treebank and a manually constructed semantic lexicon, the resulting parser computes discourse representation structures representing the meaning of arbitrary text. The main goal of this paper is to apply and specialize this general framework of wide-coverage semantics to the spatial and temporal organization of the Itipy corpus — a set of 19th century texts discussing voyages through the Pyrenees mountains. The implemented system gives satisfying results and opens the door to the integration with specialized extensions, such as a separate module computing the discourse relations between the textual units. RÉSUMÉ. Dans cet article, nous donnons une description de notre recherche sur la sémantique à large couverture pour le français et, plus précisément, de la façon dont ces expressions sémantiques sont utilisées pour donner l’interprétation spatio-temporelle des itinéraires. En utilisant une grammaire catégorielle extraite semi-automatiquement du French Treebank et un lexique sémantique construit manuellement, l’analyseur convertit les analyses syntaxiques en DRS (discourse representation structures). Le but principal de cet article est l’application et la spécialisation de cette méthodologie générale au corpus Itipy contenant des récits de voyage du 19e siècle. La chaîne de traitement complète donne des résultats tout à fait satisfaisants et laisse l’opportunité d’y ajouter des extensions spécialisées, comme un composant qui calculerait les relations discursives entre les parties du texte.
Participants viewed dynamic facial expressions that moved from a neutral expression to varying degrees of angry, happy, or sad or from these emotionally expressive faces to neutral.A contrast effect was observed for expressions that moved to a neutral state. That is, a neutral expression that began as angry was rated as having a mildly positive expression, whereas the same neutral expression was rated as negatively valenced when it began with a smile. In Experiment 2, static expressions presented sequentially elicited contrast effects, but they were weaker than those following dynamic expressions. Experiment 3 assessed a broad range of facial movements across varying degrees of angry and happy expressions. We observed momentum effects for movements that ended at mildly expressive points (25% and 50% expressive). For such movements, affect ratings were higher, as if the perceived expression moved beyond their endpoint. Experiment 4 assessed sad facial expressions and found both contrast and momentum effects for dynamic expressions to and from sad faces. These findings demonstrate new and potent contextual influences on dynamic facial expressions and highlight the importance of facial movements in social-emotional communication.
Empty elements (EEs) play a critical role in Chinese syntactic, semantic and discourse analysis. Previous studies employ a language-independent sentence-level approach to EE recovery, by casting it as a linear tagging or structured parsing problem. In comparison, this paper proposes a clause-level hybrid approach to address specific problems in Chinese EE recovery, which recovers EEs in Chinese language from the clause perspective and integrates the advantages of both linear tagging and structured parsing. In particular, a comma disambiguation method is employed to improve syntactic parsing and help determine clauses in Chinese. In this way, the noise introduced by sentence-level syntactic parsing and multiple EEs in the same position of a linear sentence can be well addressed. Evaluation on Chinese Treebank 6.0 shows the significant performance improvement of our clause-level hybrid approach over the state-of-the-art sentence-level baselines, and its great impact on a state-of-the-art Chinese syntactic parser.
Online content analysis employs algorithmic methods to identify entities in unstructured text. Both machine learning and knowledge-base approaches lie at the foundation of contemporary named entities extraction systems. However, the progress in deploying these approaches on web-scale has been been hampered by the computational cost of NLP over massive text corpora. We present SpeedRead (SR), a named entity recognition pipeline that runs at least 10 times faster than Stanford NLP pipeline. This pipeline consists of a high performance Penn Treebank- compliant tokenizer, close to state-of-art part-of-speech (POS) tagger and knowledge-based named entity recognizer.
This article considers dictionaries as lexical information / knowledge sources to be derived from a deeper, underlying, lexical database. These dictionary-tokens or -instantiations are inter alia specified by the users' needs. As a case in point of such a derivation meeting the needs of a multilingual society, a bidirectional bilingual learner dictionary is presented. Specific tools, such as editors with reversal function, and models, such as the hub-and-spoke model, are discussed as means to function within the lexicographical infrastructure of a multilingual society.
El objetivo del trabajo consiste en reutilizar el Treebank de dependencias EPEC-DEP (BDT) para construir el gold standard de la sintaxis superficial del euskera. El paso basico consiste en el estudio comparativo de los dos formalismos aplicados sobre el mismo corpus: el formalismo de la Gramatica de Restricciones (Constraint Grammar, CG) y la Gramatica de Dependencias (Dependency Grammar, DP). Como resultado de dicho estudio hemos establecido los criterios linguisticos necesarios para derivar la funciones sintacticas en estilo CG. Dichos criterios han sido implementados y evaluados, asi en el 75% de los casos somos capaces de derivar automaticamente las funciones sintacticas para construir el gold standard.
Four different patterns of biased ratings of facial expressions of emotions have been found in socially anxious participants: higher negative ratings of (1) negative, (2) neutral, and (3) positive facial expressions than nonanxious controls. As a fourth pattern, some studies have found no group differences in ratings of facial expressions of emotion. However, these studies usually employed valence and arousal ratings that arguably may be less able to reflect processing of social information. We examined the relationship between social anxiety and face ratings for perceived trustworthiness given that trustworthiness is an inherently socially relevant construct. Improving on earlier analytical strategies, we evaluated the four previously found result patterns using a Bayesian approach. Ninety-eight undergraduates rated 198 face stimuli on perceived trustworthiness. Subsequently, participants completed social anxiety questionnaires to assess the severity of social fears. Bayesian modeling indicated that the probability that social anxiety did not influence judgments of trustworthiness had at least three times more empirical support in our sample than assuming any kind of negative interpretation bias in social anxiety. We concluded that the deviant interpretation of facial trustworthiness is not a relevant aspect in social anxiety.
Part-of-speech (POS) tagging is a fundamental task in natural language processing (NLP). It provides useful information for many other NLP tasks, including word sense disambiguation, text chunking, named entity recognition, syntactic parsing, semantic role labeling, and semantic parsing. In this paper, we present a new method for Vietnamese POS tagging using dual decomposition. We show how dual decomposition can be used to integrate a word-based model and a syllable-based model to yield a more powerful model for tagging Vietnamese sentences. We also describe experiments on the Viet Treebank corpus, a large annotated corpus for Vietnamese POS tagging. Experimental results show that our model using dual decomposition outperforms both word-based and syllable-based models.
This paper investigates the impact on French dependency parsing of lexical generalization methods beyond lemmatization and morphological analysis. A distributional thesaurus is created from a large text corpus and used for distributional clustering and WordNet automatic sense ranking. The standard approach for lexical generalization in parsing is to map a word to a single generalized class, either replacing the word with the class or adding a new feature for the class. We use a richer framework that allows for probabilistic generalization, with a word represented as a probability distribution over a space of generalized classes: lemmas, clusters, or synsets. Probabilistic lexical information is introduced into parser feature vectors by modifying the weights of lexical features. We obtain improvements in parsing accuracy with some lexical generalization configurations in experiments run on the French Treebank and two out-of-domain treebanks, with slightly better performance for the probabilistic lexical generalization approach compared to the standard single-mapping approach. 1
The state is the primary provider of education in Singapore. The rise of the global knowledge economy, however, has generated an education industry that is worth about USD2.2 trillion per year globally. The Government of Singapore therefore revamped its higher-education sector so that Higher Private Education Organisations (HPEOs) can offer more flexible transnational education to attract fee-paying, overseas students and increase the education industry’s contribution to the national GDP. The private-education industry is currently facing competition from more established markets such as the United States, United Kingdom and Australia. Locally, HPEOs need to upgrade their service standard in order to meet the stringent registration requirements imposed by Singapore’s regulators and, at the same time, compete with each other in a high-cost environment. HPEOs therefore need to gain a better understanding of the relations among variables such as service quality, price satisfaction, image rating, overall satisfaction, repurchase intention and positive word of mouth. A better understanding will allow HPEOs to improve their marketing efforts in order to obtain competitive advantages. Prior research only focused on the correlation of variables and failed to take into account the inter-relationship among two or more variables. The measurement of these variables is often dependent on geographical factors, types of services and types of stakeholders. It is therefore the objective of this study to examine the measurement and relationships among service quality, price satisfaction, image rating, overall satisfaction, repurchase intention and positive word of mouth in the context of HPEOs in Singapore. A survey involving 554 participants from local HPEOs was conducted over a period of three years. Analysis of the data shows that attributes such as image rating and service quality are unique and HPEOs have to customize these attributes to meet the needs and wants of their students. This study found that service quality, price satisfaction, image rating, overall satisfaction, repurchase intention and positive word of mouth are all positively correlated. It was also found that some factors (e.g., overall satisfaction and repurchase intention) are mediators of the relationships between other variables.
Although disgust propensity (DP) has been implicated in the development of some anxiety disorders, the mechanism that may account for this association has not been fully elucidated. The present study examined the extent to which the potentiation of learned aversion might be one such mechanism. Participants (n = 103) were randomized to one of two evaluative conditioning (EC) paradigms consisting of 12 reinforced conditioned stimulus (CS+) pairings of the word "part" (condition one) or "some" (condition two) with 12 aversive unconditioned stimulus (US) images, and 12 pairings of the CS- word "cylinder" with 12 neutral images. Participants then completed measures of DP and trait anxiety and provided subjective affective ratings for the aversive US. The findings revealed that participants experienced significantly more disgust, anxiety, anger, sadness, and less happiness toward their respective CS+. In contrast, participants experienced significantly more happiness toward the CS-. Examination of the magnitude of evaluative change to the CS+ revealed the strongest effect for disgust. DP, but not trait anxiety, also predicted a greater increase in disgust, anger, and anxiety in response to the CS+ relative to the CS-. Furthermore, the association between DP and greater disgust, anger, and anxiety in response to the CS+ was mediated by more intense negative affective responding to the US among those higher in DP. The implication of these findings for better understanding how DP may confer risk for anxiety-related psychopathology is discussed.
The aim of the work is to profit the existing dependency Treebank EPEC-DEP (BDT) in order to build the gold standard for the surface syntax of Basque. As basic step, we make a comparative study of both formalisms, the Constraint Grammar formalism (CG) and the Dependency Grammar (DP) that have been applied on the corpus. As a result, we establish some criteria that will serve us to derive automatically the CG style syntactic function tags. Those criteria were implemented and evaluated; as a result, in the 75 % of the cases we are able to derive the CG style syntactic function tags for building the gold standard.
Extracting the users expected information from a large text collection based on some query is the aim of a Information Retrieval (IR) system. Now a days Assamese Digital documents are increasing at a huge rate and to collect the information efficiently from them we are in need of an Assamese IR system for retrieving documents. Comparing query and document term on lexical level and the shorter length query implies the problem like word mismatch. Adding additional term with user's query means expanding the query can improve the IR system's performance. The electronic lexical database, WordNet can help identifying the synonymous expressions and linguistic entities that are semantically similar with the input query. Here we present an Assamese IR system based on vector space model and show our result by considering the query vector as the original user's query and the query is extended using the Assamese WordNet synsets.
This article uses semi-supervised Expectation Maximization (EM) to learn lexico-syntactic dependencies, i.e. associations between words and the structures that occur with them. Due to Zipfian distributions in language, such dependencies are extremely sparse in labelled data, and unlabelled data are the only source for learning them. Specifically, we learn sparse lexical parameters of a generative parsing model (a Probabilistic Context-Free Grammar, PCFG) that is initially estimated over the Penn Treebank. Our lexical parameters are similar to supertags—they are fine-grained, and encode complex structural information at the pre-terminal level. Our goal is to use unlabelled data to learn these for words that are rare or unseen in the labelled data. We get large error reductions (up to 17.5%) in parsing ambiguous structures associated with unseen verbs, the most important case of learning lexico-structural dependencies, resulting in a statistically significant improvement in labelled bracketing score of the treebank PCFG. Our semi-supervised method incorporates structural and lexical priors from the labelled data to guide estimation from unlabelled data, and is the first successful use of semi-supervised EM to improve a generative structured model already trained over large labelled data. The method scales well to larger amounts of unlabelled data, and also gives substantial error reductions (up to 11.5%) for models trained on smaller amounts of labelled data, making it relevant to low-resource languages with small treebanks as well.
The Digital World encounters rapid development nowadays, especially through the proliferation of social media in Indonesia. Twitter has become one of social media with expanded users within every sectors of society. There are so many part both individual as well as organization/enterprise which utilize twitter as tool for communication, business, customer relation, and other activities. Through the twitter's ever-expanding users with those particular purposes, the precise method to effectively and efficiently analyzing opinion-contained sentences become crucially needed. Therefore this research made for method analyzing through lexical based and model based approaches by machine learning to classify opinion-contained tweets using those 2 methods. The tested machine learning method are Support Vector Machine (SVM), Maximum Entropy (ME), Multinomial Naive Bayes (MNB), and k-Nearest Neighbor (k-NN). Based on the test outcome, lexical based approach highly depended on lexical database which became opinion classification matrix. Whilst machine learning approach can produce better accuracy due to its capability in new training data modeling based on outcome model. However, machine learning model based approach depends on various factors in analyzing sentiment.
Abstract In this article we present some statistical data on the distribution of parts of speech and dependency relations in a large manually annotated Hungarian Treebank, the Szeged Dependency Treebank. We hypothesize that the domain of the text influences the distribution of the above elements, thus we pay special attention to differences between domains. We present the characteristic rank-frequency distributions of parts of speech and dependency relations in Hungarian and analyse the domain similarities and differences among sub-corpora as regards the above distributions. Our results reveal that the computer and newspaper texts are most similar to each other while the domains literature and compositions also exhibit some similarities. On the other hand, the business news and the law sub-corpora are unique, both having their own characteristics.
Body image disturbances are core symptoms of eating disorders (EDs). Recent evidence suggests that changes in body image may occur prior to ED onset and are not restricted to in-vivo exposure (e.g. mirror image), but also evident during presentation of abstract cues such as body shape and weight-related words. In the present study startle modulation, heart rate and subjective evaluations were examined during reading of body words and neutral words in 41 student female volunteers screened for risk of EDs. The aim was to determine if responses to body words are attributable to a general negativity bias regardless of ED risk or if activated, ED relevant negative body schemas facilitate priming of defensive responses. Heart rate and word ratings differed between body words and neutral words in the whole female sample, supporting a general processing bias for body weight and shape-related concepts in young women regardless of ED risk. Startle modulation was specifically related to eating disorder symptoms, as was indicated by significant positive correlations with self-reported body dissatisfaction. These results emphasize the relevance of examining body schema representations as a function of ED risk across different levels of responding. Peripheral-physiological measures such as the startle reflex could possibly be used as predictors of females' risk for developing EDs in the future.
Syntactic parsing is an important technique in the natural language processing, yet Latvian is still lacking an efficient general coverage syntax parser. This paper reports on the first experiments on statistical syntactic parsing for Latvian — a highly inflective Indo-European language with a relatively free word order. We have induced a statistical parser from a small, non-balanced Latvian Treebank using the MaltParser toolkit and measured the unlabeled attachment score (UAS). As MaltParser is based on the dependency grammar approach, we have also developed a convertor from the hybrid dependency-based annotation model used in the Latvian Treebank to the pure dependency annotation model. We have obtained a promising 74.63 % UAS in 10-fold cross-validation using only ~2500 sentences. The results revealed that best results can be achieved using non-projective stack parsing algorithm with lazy arc adding strategy, but comparably good results can be achieved using projective parsing algorithms combined with appropriate projectiviziation preprocessing.
This chapter deals with the main methodological issues underlying the building of the SciE-Lex lexical database and discusses and justifies the information included. SciE-Lex was initially conceived as a response to the lack of reference tools that can help scientists write scientific papers in phraseologically competent and native-like English. While there are a number of specialised dictionaries that include specific terminological information, there is a shortage of writing aids that provide information about the use of non-technical terms in scientific genres. SciE-Lex aims at filling this gap by focusing on the description of general terms in scientific English. This article describes the two stages in the building of the database, the first one including morphosyntactic and collocational information, and the second one focusing on phraseological information.
Turkish is an agglutinative language with rich morphology-syntax interactions. As an extension of this property, the Turkish Treebank is designed to represent sublexical dependencies, which brings extra challenges to parsing raw text. In this work, we use a joint POS tagging and parsing approach to parse Turkish raw text, and we show it outperforms a pipeline approach. Then we experiment with incorporating morphological feature prediction into the joint system. Our results show statistically significant improvements with the joint systems and achieve the state-ofthe-art accuracy for Turkish dependency parsing.
Exploiting data from a parallel treebank recently developed for Italian, English and French, the paper discusses issues related to the development of a dependency-based alignment system.We focus on the alignment of linguistic expressions and constructions which are structurally different in the languages that have to be aligned, and on how to deal with them using dependency rather than constituency.In order to analyze in particular the shifts related to syntactic structure, we present a selection of cases where a dependencybased and a constituency-based representation has been applied and compared.
In emotional speech research, it has been suggested that loudness, along with other prosodic features, may be an important cue in communicating high activation affects. In earlier studies, we found different voice quality stimuli to be consistently associated with certain affective states. In these stimuli, as in typical human productions, the different voice qualities entailed differences in loudness. To examine the extent to which the loudness differences among these voice qualities might influence the affective coloring they impart, two experiments were conducted with the synthesized stimuli, in which loudness was systematically manipulated. Experiment 1 used stimuli with distinct voice quality features including intrinsic loudness variations and stimuli where voice quality (modal voice) was kept constant, but loudness was modified to match the non-modal qualities. If loudness is the principal determinant in affect cueing for different voice qualities, there should be little or no difference in the responses to the two sets of stimuli. In Experiment 2, the stimuli included distinct voice quality features but all had equal loudness to test the hypothesis that equalizing the perceived loudness of different voice quality stimuli will have relatively little impact on affective ratings. The results suggest that loudness variation on its own is relatively ineffective whereas variation in voice quality is essential to the expression of affect. In Experiment 1, stimuli incorporating distinct voice quality features consistently obtained higher ratings than the modal voice stimuli with varied loudness. In Experiment 2, non-modal voice quality stimuli proved potent in affect cueing even with loudness differences equalized. Although loudness per se does not seem to be the major determinant of perceived affect, it can contribute positively to affect cueing: when combined with a tense or modal voice quality, increased loudness can enhance signaling of high activation states.
This paper describes our approaches to Na-tive Language Identification (NLI) for the NLI shared task 2013. NLI as a sub area of au-thor profiling focuses on identifying the first language of an author given a text in his sec-ond language. Researchers have reported sev-eral sets of features that have achieved rel-atively good performance in this task. The type of features used in such works are: lex-ical, syntactic and stylistic features, depen-dency parsers, psycholinguistic features and grammatical errors. In our approaches, we se-lected lexical and syntactic features based on n-grams of characters, words, Penn TreeBank (PTB) and Universal Parts Of Speech (POS) tagsets, and perplexity values of character of n-grams to build four different models. We also combine all the four models using an en-semble based approach to get the final result. We evaluated our approach over a set of 11 na-tive languages reaching 75 % accuracy. 1
At present, discourse parsing is an important research topic. Rhetorical Structure Theory (RST) is one of the most popular approaches in this field. In general, discourse parsing includes three stages: discourse segmentation, discourse relations detection and building up rhetorical trees. Different strategies are used when developing discourse parsers. One of the strategies to detect discourse relations is based on symbolic rules that take into account linguistic clues, such as discourse markers. Nevertheless, some discourse markers are ambiguous, that is, they can indicate more than one discourse relation. This fact constitutes a problem when assigning discourse relations automatically. In this paper, a symbolic approach to detect and solve discourse markers ambiguity in Spanish is developed. First, we detect ambiguous discourse markers, using the training corpus of the RST Spanish Treebank. Second, we extract linguistic contexts for these markers. Third, we design linguistic rules to solve the ambiguity of discourse markers. Fourth, we evaluate the rules, using the test corpus of the RST Spanish Treebank. Our approach outperforms the baseline created following the methodology of the state of the art. Therefore, we consider that the results obtained in our experiments are representative and constitute the first step towards the disambiguation of discourse markers senses in Spanish. However, there is room for improvement and the main limitations of the approach are presented. In the future, the rules will be integrated in a discourse parser for Spanish, and several related applications will be developed (automatic summarization and information extraction, among others).
For years observational techniques along with other methods have sought to explore the relationships of couple interactional exchanges to marital quality and longevity. However, many of the previous methodological procedures used might be inadequate at capturing the influential micro-dimensional nuances of interpartner couple affective stability and reciprocity. This study explored the dyadic patterns in 23 married couples' continuous affect ratings during two communication episodes. Multilevel modeling was used to assess the structure in the stability of one's own affect and the influence of partner affect over 3-, 6-, and 9-second time lags. Implications regarding the use of nested models to explore patterns of actor and partner effects are discussed.
We present a comparative study of transition-, graph- and PCFG-based models aimed at illuminating more precisely the likely contribution of CFGs in improving Chinese dependency parsing accuracy, especially by combining heterogeneous models. Inspired by the impact of a constituency grammar on dependency parsing, we propose several strategies to acquire pseudo CFGs only from dependency annotations. Compared to linguistic grammars learned from rich phrase-structure treebanks, well designed pseudo grammars achieve similar parsing accuracy and have equivalent contributions to parser ensemble. Moreover, pseudo grammars increase the diversity of base models; therefore, together with all other models, further improve system combination. Based on automatic POS tagging, our final model achieves a UAS of 87.23%, resulting in a significant improvement of the state of the art.
This paper describes our submission for SemEval2013 Task 2: Sentiment Analysis in Twitter. For the limited data condition we use a lexicon-based model. The model uses an affective lexicon automaticallygeneratedfrom a very large corpus of raw web data. Statistics are calculated over the word and bigram affective ratings and used as features of a Naive Bayes tree model. For the unconstrained data scenario we combine the lexicon-based model with a classifier built on maximum entropy language models and trained on a large external dataset. The two models are fused at the posterior level to produce a final output. The approach proved successful, reaching rankings of 9th and 4th in the twitter sentiment analysis constrained and unconstrained scenario respectively, despite using only lexical features.