Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
This article focuses on the variability of one of the subtypes of multi-word expressions, namely those consisting of a verb and a particle or a verb and its complement(s). We build on evidence from Estonian, an agglutinative language with free word order, analysing the behaviour of verbal multi-word expressions (opaque and transparent idioms, support verb constructions and particle verbs). Using this data we analyse such phenomena as the order of the components of a multi-word expression, lexical substitution and morphosyntactic flexibility.
This article details a series of carefully designed experiments aiming at evaluating the influence of automatic pre-annotation on the manual part-of-speech annotation of a corpus, both from the quality and the time points of view, with a specific attention drawn to biases. For this purpose, we manually annotated parts of the Penn Treebank corpus (Marcus et al., 1993) under various experimental setups, either from scratch or using various pre-annotations. These experiments confirm and detail the gain in quality observed before (Marcus et al., 1993; Dandapat et al., 2009; Rehbein et al., 2009), while showing that biases do appear and should be taken into account. They finally demonstrate that even a not so accurate tagger can help improving annotation speed. 1
This paper presents a method for the automatic detection and correction of malapropism errors found in documents using the WordNet lexical database, a search engine (Google) and a paronyms dictionary. The malapropisms detection is based on the evaluation of the cohesion of the local context using the search engine, while the correction is done using the whole text cohesion evaluated in terms of lexical chains built using the linguistic ontology. The correction candidates, which are taken from the paronyms dictionary, are evaluated versus the local and the whole text cohesion in order to find the best candidate that is chosen for replacement. The testing methods of the application are presented, along with the obtained results.
RATIONALE: Acute tryptophan depletion (ATD) decreases levels of central serotonin. ATD thus enables the cognitive effects of serotonin to be studied, with implications for the understanding of psychiatric conditions, including depression. OBJECTIVE: To determine the role of serotonin in conscious (explicit) and unconscious/incidental processing of emotional information. MATERIALS AND METHODS: A randomized, double-blind, cross-over design was used with 15 healthy female participants. Subjective mood was recorded at baseline and after 4 h, when participants performed an explicit emotional face processing task, and a task eliciting unconscious processing of emotionally aversive and neutral images presented subliminally using backward masking. RESULTS: ATD was associated with a robust reduction in plasma tryptophan at 4 h but had no effect on mood or autonomic physiology. ATD was associated with significantly lower attractiveness ratings for happy faces and attenuation of intensity/arousal ratings of angry faces. ATD also reduced overall reaction times on the unconscious perception task, but there was no interaction with emotional content of masked stimuli. ATD did not affect breakthrough perception (accuracy in identification) of masked images. CONCLUSIONS: ATD attenuates the attractiveness of positive faces and the negative intensity of threatening faces, suggesting that serotonin contributes specifically to the appraisal of the social salience of both positive and negative salient social emotional cues. We found no evidence that serotonin affects unconscious processing of negative emotional stimuli. These novel findings implicate serotonin in conscious aspects of active social and behavioural engagement and extend knowledge regarding the effects of ATD on emotional perception.
The article analyzes 97 elementary schoolbooks in Buenos Aires to determine which social representations about linguistic norm underlie in these school materials. The paper reviews -especially in the defi nitions of categories, and exercises and activities- the concepts of linguistic variety, standard language and español neutro. Based on these variables, this article sees the possible repercussions in social representations that students and teachers can develop from point of view of the publishing companies.
Motion pictures are so typical experience goods that consumers tend to look for more credible information. Hence, movie audiences consider movie viewers` reviews more important than the information provided by the film distributor. Recently many portal sites allow consumers to post their reviews and opinions so that other people check the number of consumer reviews and scores before going to the theater. There are a few previous researches studying the electronic word of mouth(eWOM) effect in the movie industry. They found that the volume of eWOM influenced the revenue of the movie significantly but the valence of eWOM did not affect it much (Liu 2006). The goal of our research is also to investigate the eWOM effects in general. But our research is different from the previous studies in several aspects. First, we study the eWOM effect in Korean movie industry. In other words, we would like to check whether we can generalize the results of the previous research across countries. The similar econometric models are applied to Korean movie data that include 746,282 consumer reviews on 439 movies. Our results show that both the valence(RATING) and the volume(LNMSG) of the eWOM influence weekly movie revenues. This result is different from the previous research findings that the volume only influences the revenue. We conjectured that the difference of self construal between Asian and American culture may explain this difference (Kitayama 1991). Asians including Koreans have more interdependent self construal than American, so that they are easily affected by other people`s thought and suggestion. Hence, the valence of the eWOM affects Koreans` choice of the movie. Second, we find the critical defect of the previous eWOM models and, hence, attempt to correct it. The previous eWOM model assumes that the volume of eWOM (LNMSG) is an independent variable affecting the movie revenue (LNREV). However, the revenue can influence the volume of the eWOM. We think that treating the volume of eWOM as an independent variable a priori is too restrictive. In order to remedy this problem, we employed a simultaneous equation in which the movie revenue and the volume of the eWOM can affect each other. That is, our eWOM model assumes that the revenue (LNREV) and the volume of eWOM (LNMSG) have endogenous relationship where they influence each other. The results from this simultaneous equation model showed that the movie revenue and the eWOM volume interact each other. The movie revenue influences the eWOM volume for the entire 8 weeks. The reverse effect is more complex. Both the volume and the valence of eWOM affect the revenue in the first week, but only the volume affect the revenue for the rest of the weeks. In the first week, consumers may be curious about the movie and look for various kinds of information they can trust, so that they use the both the quantity and quality of consumer reviews. But from the second week, the quality of the eWOM only affects the movie revenue, implying that the review ratings are more important than the number of reviews. Third, our results show that the ratings by professional critics (CRATING) had negative effect to the weekly movie revenue (LNREV). Professional critics often give low ratings to the blockbuster movies that do not have much cinematic quality. Experienced audiences who watch the movie for fun do not trust the professionals` ratings and, hence, tend to go for the low-rated movies by them. In summary, applied to the Korean movie ratings data and employing a simultaneous model, our results are different from the previous eWOM studies: 1) Koreans (or Asians) care about the others` evaluation quality more than quantity, 2) The volume of eWOM is not the cause but the result of the revenue, 3) Professional reviews can give the negative effect to the movie revenue.
This paper proposes a dependency parsing method that uses bilingual constraints to improve the accuracy of parsing bilingual texts (bitexts). In our method, a targetside tree fragment that corresponds to a source-side tree fragment is identified via word alignment and mapping rules that are automatically learned. Then it is verified by checking the subtree list that is collected from large scale automatically parsed data on the target side. Our method, thus, requires gold standard trees only on the source side of a bilingual corpus in the training phase, unlike the joint parsing model, which requires gold standard trees on the both sides. Compared to the reordering constraint model, which requires the same training data as ours, our method achieved higher accuracy because of richer bilingual constraints. Experiments on the translated portion of the Chinese Treebank show that our system outperforms monolingual parsers by 2.93 points for Chinese and 1.64 points for English. 1
We propose an effective approach to automatically identify predicate heads in Chinese sentences based on statistical pre-processing and rule-based post-processing. In the preprocessing stage, the maximal noun phrases in a sentence are recognized and replaced by “NP ” labels to simplify the sentence structure. Then a CRF model is trained to recognize the predicate heads of this simplified sentence. In the post-processing stage, a rule base is built according to the grammatical features of predicate heads. It is then utilized to correct the preliminary recognition results. Experimental results show that our approach is feasible and effective, and its accuracy achieves 89.14 % on Tsinghua Chinese Treebank. 1
Due to idiosyncrasies in their syntax, semantics or frequency, Multiword Expressions (MWEs) have received special attention from the NLP community, as the methods and techniques developed for the treatment of simplex words are not necessarily suitable for them. This is certainly the case for the automatic acquisition of MWEs from corpora. A lot of effort has been directed to the task of automatically identifying them, with considerable success. In this paper, we propose an approach for the identification of MWEs in a multilingual context, as a by-product of a word alignment process, that not only deals with the identification of possible MWE candidates, but also associates some multiword expressions with semantics. The results obtained indicate the feasibility and low costs in terms of tools and resources demanded by this approach, which could, for example, facilitate and speed up lexicographic work.
SpatialML is an annotation scheme for marking up references to places in natural language. It covers both named and nominal references to places, grounding them where possible with geo-coordinates, and characterizes relationships among places in terms of a region calculus. A freely available annotation editor has been developed for SpatialML, along with several annotated corpora. Inter-annotator agreement on SpatialML extents is 91.3 F-measure on a corpus of SpatialML-annotated ACE documents released by the Linguistic Data Consortium. Disambiguation agreement on geo-coordinates on ACE is 87.93 F-measure. An automatic tagger for SpatialML extents scores 86.9 F on ACE, while a disambiguator scores 93.0 F on it. Results are also presented for two other corpora. In adapting the extent tagger to new domains, merging the training data from the ACE corpus with annotated data in the new domain provides the best performance.
u-tokyo.ac.jp Several recent discourse parsers have employed fully-supervised machine learning approaches. These methods require human annotators to beforehand create an extensive training corpus, which is a time-consuming and costly process. On the other hand, unlabeled data is abundant and cheap to collect. In this paper, we propose a novel semi-supervised method for discourse relation classification based on the analysis of cooccurring features in unlabeled data, which is then taken into account for extending the feature vectors given to a classifier. Our experimental results on the RST Discourse Treebank corpus and Penn Discourse Treebank indicate that the proposed method brings a significant improvement in classification accuracy and macro-average F-score when small training datasets are used. For instance, with training sets of c.a. 1000 labeled instances, the proposed method brings improvements in accuracy and macro-average F-score up to 50% compared to a baseline classifier. We believe that the proposed method is a first step towards detecting low-occurrence relations, which is useful for domains with a lack of annotated data. 1
This study investigates how the duration of the stem vowel of regular and irregular English verbs is modulated by tense (present and past), regularity, lexical frequency, gang size of the vocalic alternation, imageability ratings, and vowel quality. The vocalic durations of 48 monosyllabic irregular verbs and 171 regular verbs were extracted from the Buckeye Corpus of spontaneous speech. A linear mixed effects regression model revealed that vowels of past tense forms tend to have longer durations than vowels of present tense forms, that vowels of words that are less imageable are realized with shorter durations, and that tense vowels are longer than lax vowels. Surprisingly, higher frequency irregular past tense forms were produced with longer vowels, contradicting Aylett and Turk [(2004); (2006)] and Bell et al. [(2003); (2009)]. Further, vowels were shorter for irregular verbs with larger vocalic alternation gangs, contradicting the predictions of Kuperman et al. [(2007)] but supporting hypotheses that units with a smaller information load have shorter durations. This pattern of results is interpreted as a consequence of the pressure for regularization during the production of irregular past tense forms.
Nivre’s method was improved by enhancing deterministic dependency parsing through application of a tree-based model. The model considers all words necessary for selection of parsing actions by including words in the form of trees. It chooses the most probable head candidate from among the trees and uses this candidate to select a parsing action. In an evaluation experiment using the Penn Treebank (WSJ section), the proposed model achieved higher accuracy than did previous deterministic models. Although the proposed model’s worst-case time complexity is O(n 2), the experimental results demonstrated an average parsing time not much slower than O(n). 1
The ANR Emotirob project aims at detecting emotions in an original application context: realizing an emotional companion robot for weakened children. This paper presents a system which aims at characterizing emotions by only considering linguistic content. It is based on the assumption that emotions can be compound: simple lexical words have an intrinsic emotional value, while verbal and adjectival predicates act as a function on the emotional values of their arguments. The paper describes the algorithm of compositional computation of the emotion and the lexical emotional norm used by this algorithm. A quantitative and qualitative analysis of the differences between system outputs and expert annotations is given, which shows satisfactory results, with a good detection of emotional valency in 90.0% of the test utterances.
Introduzione L’impiego di corpora elettronici nella ricerca linguistica si rivela partico larmente efficace quando ai dati raccolti ( raw data ) vengono aggiunte informazioni di carattere linguistico e funzionale che facilitano la loro in terpretazione da un punto di vista morfosintattico. Il risultato di tali pro cedure di analisi e rappresentato da Treebanks ovvero corpora annotati sintatticamente attraverso rappresentazioni ad albero. Solitamente distinta dalla procedura di mark-up , l’annotazione di un corpus realizzata con l’applicazione di sistemi computerizzati ( parsers ) che eseguono un’etichettatura delle parti del discorso ( POS tagging ) costituisce una procedura particolarmente efficace nel chiarire, ad esempio, casi di ambiguita grammaticale in cui un lessema, se privo di annotazione, puo essere interpretato sia come verbo (V) che come nome (N) (ad esempio work ). I parsers generalmente propongono descrizioni delle produzioni linguistiche di una comunita di parlanti basate su regole grammaticali estratte da un campione di dati e tale caratteristica talvolta non permette di eseguire l’analisi automatica anche di frasi non grammaticali . Si tratta, dunque, di sistemi di analisi basati su grammatiche di precisione che di stinguono principalmente cio che e grammaticale da cio che non lo e ba sandosi su un corpus necessariamente limitato di dati, in quanto la descri zione fa esclusivamente riferimento a campioni di frasi grammaticalmen te corrette di una lingua inseriti nel sistema. Tuttavia, la propensione dei parlanti a produrre costruzioni che presentano “deviazioni” dalla norma e il fatto che molti di essi usino una lingua straniera come strumento di co municazione comporta in se «that naturalistic ungrammatical sentences are of interest to linguists studying language production, language loss and language learning, and that the grammatical/ungrammatical distinc tion cannot therefore be completely dismissed» .
Summary With the celebration in 2008 of the 125th anniversary of the first publication of Olive Schreiner's novel, The Story of an African Farm, in 1873, the question of reliability of the text came up once again for review. This article accounts for the circumstances of the first printing in London with an inexperienced author as proofreader, without any existing standardisation or other lexical references to non-British usages particularly proto-Afrikaans, to consult, and the prevailing London publishing norms in control. Subsequent editions with numerous corrections by her hand, as well as by later editors, are mentioned, while the quest to establish a definitive edition is outlined, now that English South African usages incorporate many fringe language examples which have since become nativised into common usage. The article suggests that lax proofreading, on the one hand, together with scantily informed metropolitan standards of language outreach, on the other, have led to unfortunate errors being perpetuated, even in numerous scholarly spin-offs, despite the attempts of previous scholars to standardise the text to conform to present-day professional norms and conventions.
MLR, 105.3, 2010 889 between participants. The chapters in this book document the present concern about the state of the language and show how these concerns are really about the state of society, not of the language. University of Bristol Nils Langer Was istgutesDeutsch? Studien undMeinungen zum gepflegten Sprachgebrauch. Ed. by Armin Burkhardt. (Duden? Thema Deutsch, 8) Mannheim: Duden. 2007. 411pp. 25. ISBN 978-3-411-04213-5. It has been noted by many scholars that there is a singularly defensive attitude towards codified linguistic norms inGermany, memorably termed Sprachnormen frommigkeit by Peter von Polenz over twenty years ago, and the question of these norms has once again become a hotly debated issue inGermany, perhaps spread ing out from the insecurity caused by the unexpectedly contentious orthographic reform of 1996 and the concern about the effect of Anglicisms on the language, with puristic tendencies which had been suppressed since 1945 re-emerging after unification. This has taken the form of concern and discussion in themedia about the supposed decline of the language and increasing deviation fromwhat has tradi tionally been considered 'gutesDeutsch', and books such as those by Bastian Sick (e.g. Der Dativ istdem Genitiv sein Tod (Cologne: Kiepenheuer & Witsch, 2004)) lambasting 'bad' German have enjoyed immense (and thoroughly undeserved) success. The aim of this volume, as the editor makes clear in his introduction, is to provide a set of informed contributions to this debate from experts across a wide range ofGerman linguistic studies. A major problem, as the editor sees it,is the lack of connection between professional scholars of language and the general educated public, who often feel that such experts (especially those held responsible for the spelling reform) fail tounderstand their concerns and arewilling simply to observe the ongoing demise of the language from their ivory towers. Despite this aim, itmust be doubted whether the twenty-nine essays in this book can succeed in bridging this gap. They are divided into four sections entitled Rucksichten (three essays giving a historical perspective), Einsichten (nine essays on what can be considered good in terms of linguistic aspects such as pronunciation, grammar, and style),Hinsichten (twelve essays on theuse of the language in specific registers or genres), and Ansichten (five essays expressing views from a number of perspectives on the state of the language). In practice, all but one (that by Dieter E. Zimmer in the fourth section) are by precisely thekind of professional scholars who are held to lack concern for the dire condition of the language. In general, these essays are of an unusually consistent quality for an edited volume of this kind, and some are quite outstanding?notably those by Gottfried Kolde on Sprachpflege from 1945 to 1968, by Hans-Werner Eroms on what con stitutes 'good' grammar, and by Jiirgen Schiewe on Sprachkritik. The essays are informed and informative (and written in good German), and they are unanimous 890 Reviews thatwhat is good German' isGerman used appropriately for the topic, comprehen sibly,and imaginatively?in Schiewe's terms (p. 373) it isGerman characterized by Angemessenheit, Pragnanz, and Variation. Rudolf Hoberg's essay entitled 'Besseres Deutsch: Was kann und soil eine wissenschaftlich begrundete Sprachpflege tun' is a succinct and eminently sensible summary ofwhat the criteria should be forgood German, by an academic linguistwho despite the stereotype does care deeply about quality in language?but is not convinced thatGerman civilization as we know it will perish when people no longerwrite sentences like 'Die Blumen sturben sicher, wenn du sie nicht bald begossest'. It is unlikely that the book will win over the self-appointed guardians of lin guistic excellence and sundry other pedants to the view thatmodern German has immense vitality and is in no way in decline. But what James and Lesley Milroy in their Authority in Language, 3rd edn (London: Routledge, 1999) refer to as the complaint tradition'?the idea that the language of the present is in often unspecified ways 'bad' compared with the language of the past?has a long history across many languages, and the desire for stable linguistic norms appears very deep-rooted. In themain, the essays in this book can be recommended without hesitation, but, unfortunately, theywill probably neither...
This paper presents a new perspective on the origin and development of the Mary-merry-marry merger, the conditioned merger, or neutralization, of mid and low front vowels before /r/ in dialects of North American English. The city of Montreal, Quebec represents one of very few regions in which this merger has not taken hold, despite the fact that a near-complete merger is found in the nearby rural region of Quebec’s Eastern Townships. This paper attempts to shed light on this puzzling geographic distribution using data from archival interviews conducted with Eastern Townshippers born between 1895 and 1915. An acoustic analysis of the vowels before /r/ is presented and compared with data from recent studies of Montreal English. Acoustic analysis of the mean values of the first and second vowel formants shows a great deal of variation in these speakers’ productions of the historically low front vowel before /r/. In some tokens it is clearly merged with the mid vowel, while in others the two phonemes remain clearly distinct. Further, this variation is found both between speakers and in the speech of individuals themselves. Although not entirely homogenous, the speech community does appear to share general norms with regard to which words are or are not merged. These results demonstrate that the merger was not a lexically abrupt sound change. Rather, the results are consistent with a theory of sound change via lexical diffusion, which implies a much longer timeline for this change than previously assumed, suggesting its origins may go back many more generations. As such, it is suggested that the current geolinguistic pattern of the merger may be traced to the different settlement histories of Montreal and the Eastern Townships.
We show that the standard beam-search algorithm can be used as an efficient decoder for the global linear model of Zhang and Clark (2008) for joint word segmentation and POS-tagging, achieving a significant speed improvement. Such decoding is enabled by: (1) separating full word features from partial word features so that feature templates can be instantiated incrementally, according to whether the current character is separated or appended; (2) deciding the POS-tag of a potential word when its first character is processed. Early-update is used with perceptron training so that the linear model gives a high score to a correct partial candidate as well as a full output. Effective scoring of partial structures allows the decoder to give high accuracy with a small beam-size of 16. In our 10-fold crossvalidation experiments with the Chinese Treebank, our system performed over 10 times as fast as Zhang and Clark (2008) with little accuracy loss. The accuracy of our system on the standard CTB 5 test was competitive with the best in the literature. 1
In this paper, we argue for and demonstrate the use of Prolog as a tool to query annotated corpora. We present a case study based on the German TüBa-D/Z Treebank to show that flexible and efficient corpus querying can be started with a minimal amount of effort. We end this paper with a brief discussion of performance, that suggests that the approach is both fast enough and scalable. 1
Ott, N. & R. Ziai (2010). Evaluating dependency parsing performance on german learner language. In M. Dickinson, K. Müürisep & M. Passarotti (eds.), Proceedings of the Ninth International Workshop on Treebanks and Linguistic Theories. Vol. 9 of NEALT Proceeding Series, 175–186.
A commonly held assumption is that processes underlying explicit and implicit memory are distinct. Recent evidence, however, suggests that they may interact more than previously believed. Using the remember-know procedure the current study examines the relation between recollection, a process thought to be exclusive to explicit memory, and performance on two implicit memory tasks, lexical decision and word stem completion. We found that, for both implicit tasks, words that were recollected were associated with greater priming effects than were words given a subsequent familiarity rating or words that had been studied but were not recognised (misses). Broadly, our results suggest that non-voluntary processes underlying explicit memory also benefit priming, a measure of implicit memory. More specifically, given that this benefit was due to a particular aspect of explicit memory (recollection), these results are consistent with some strength models of memory and with Moscovitch's (2008) proposal that recollection is a two-stage process, one rapid and unconscious and the other more effortful and conscious.
Recognition of special linguistic patterns in a certain language is very helpful for many NLP applications such as information extraction, machine translation and parsing. State-of-the-arts syntax parsers are based on given grammar. The used grammar is context free and cannot discover complex patterns which contain multiple linguistic units. We propose an unsupervised method to automatically discover the complex linguistic patterns from a classically parsed corpus. A specialized and efficient algorithm is applied to mine the frequent subtrees in the forest and the found subtrees are formalized as the linguistic patterns. The approach is validated on the Penn Chinese Treebank with found linguistic patterns.
Co-constructing communicative effectiveness is often challenging in English as a lingua franca (ELF): speakers have considerably less to go on in terms of shared expectations of cultural knowledge and linguistic norms. A university environment provides a convenient backdrop for sharing at least academic conventions – although these vary more than might be surmised from the uniform labelling of such event types. This paper looks into some discourse and lexicogrammatical features in academic ELF, using ELFA as the database. The data consists of spoken language, which provides direct access to the ways in which meanings are negotiated in ongoing discourse, and the speech events are typically polylogic. ELF discourse requires close cooperation from the participants, which is reflected in its enhanced explicitness among other things. The explicitation strategies speakers display facilitate mutual comprehensibility and contribute to social cohesion within the multi-participant groups. Such strategies also help overcome the potential problems participants might have in dealing with a variety of formal deviations from ordinary English as a native language (ENL). Most of the time ELF bears a very close resemblance to Standard English, but signs of incipient ELF-specific developments are also in evidence.
We present the first evaluation of the utility of automatic evaluation metrics on surface realizations of Penn Treebank data. Using outputs of the OpenCCG and XLE realizers, along with ranked WordNet synonym substitutions, we collected a corpus of generated surface realizations. These outputs were then rated and post-edited by human annotators. We evaluated the realizations using seven automatic metrics, and analyzed correlations obtained between the human judgments and the automatic scores. In contrast to previous NLG meta-evaluations, we find that several of the metrics correlate moderately well with human judgments of both adequacy and fluency, with the TER family performing best overall. We also find that all of the metrics correctly predict more than half of the significant systemlevel differences, though none are correct in all cases. We conclude with a discussion of the implications for the utility of such metrics in evaluating generation in the presence of variation. A further result of our research is a corpus of post-edited realizations, which will be made available to the research community. 1
This paper shows that training a lexicalized parser on a lemmatized morphologically-rich treebank such as the French Treebank slightly improves parsing results. We also show that lemmatizing a similar in size subset of the English\nPenn Treebank has almost no effect on parsing performance with gold lemmas and leads to a small drop of performance when automatically assigned lemmas and POS tags are used. This highlights two facts: (i) lemmatization helps to reduce lexicon data-sparseness issues for French, (ii) it also makes the parsing process sensitive to correct assignment of POS tags to unknown words.
In this paper, we discuss our analysis and resulting new annotations of Penn Discourse Treebank (PDTB) data tagged as Concession. Concession arises whenever one of the two arguments creates an expectation, and the other ones denies it. In Natural Languages, typical discourse connectives conveying Concession are ‘but’, ‘although’, ‘nevertheless’, etc. Extending previous theoretical accounts, our corpus analysis reveals that concessive interpretations are due to different sources of expectation, each giving rise to critical inferences about the relationship of the involved eventualities. We identify four different sources of expectation: Causality, Implication, Correlation, and Implicature. The reliability of these categories is supported by a high inter-annotator agreement score, computed over a sample of one thousand tokens of explicit connectives annotated as Concession in PDTB. Following earlier work of (Hobbs, 1998) and (Davidson, 1967) notion of reification, we extend the logical account of Concession originally proposed in (Robaldo et al., 2008) to provide refined formal descriptions for the first three mentioned sources of expectations in Concessive relations. 1.
On the question of precisely what role common sense (or related datum like folk psychology, trust in pre-theoretic/intuitive judgments, etc.) should have in reigning in the possible excesses of our philosophical methods, the so-called 'continental' answer to this question, for the vast majority, would be "as little as possible", whereas the analytic answer for the vast majority would be "a reasonably central one". While this difference at the level of both rhetoric and meta-philosophy is sometimes -perhaps often -problematised by the actual philosophical practices of representative philosophers of either tradition, I will argue that this norm (and its absence) nonetheless continues to play an important justificatory role in relation to the use of some rather different methodological practices. In particular, many analytic philosophers not only explicitly invoke the value of common sense, but they also implicitly value it via techniques like conceptual analysis that want to explicate folk psychology and/or lay bare what is already embedded in the linguistic norms of a given culture, the widespread use of thought experiments and the way they function as 'intuition pumps', as well as the general aim to achieve 'reflective equilibrium' between our intuitions and reflective judgments in epistemology and political philosophy. Such methods, I will argue, enshrine a conservative, or, more positively, a modest understanding of the philosophical project in that it is invested in cohering with both a given body of knowledge and common sense. These methods are notably less perspicuous in continental philosophy. To bring some of the reasons why this might be so to the fore, this paper considers Deleuze's sustained attack on both good and common sense, which he argues are fundamental to the prevalence of a dogmatic image of thought. If Deleuze is right about this, and if the analytic tradition distils and perfects certain methods that are closely associated with this image of thought, then we have here a rather stark methodological contrast that calls for elaboration and evaluation.
Current automatic wrappers using DOM tree and visual properties of data records to extract the required information from the deep web generally have limitations such as the inability to check the similarity of tree structures accurately. Our study shows that data records located in the deep web do not only share similar visual properties and tree structures, but they are also related semantically in their contents. As such we are able to propose an ontological technique using existing lexical database for English (WordNet) for the extraction of data records from deep web pages. Wrappers designed based on ontological technique are able to reduce the number of potential data regions identified for data extraction, thus improve the data extraction accuracy. In this study, we use visual cue from the underlying browser rendering engine to locate and extract the relevant data region from the deep web by measuring the text and image sizes of data records. Experimental results show that our technique is robust and performs better than the existing state of the art wrappers. Unlike existing ontological based wrappers, our wrapper is domain independent and is able to extract wide range of data records with different structures.
BACKGROUND: Recent advances have been made in the application of cognitive training strategies as interventions for mental disorders. One novel approach, cognitive control training (CCT), uses computer-based exercises to chronically increase prefrontal cortex recruitment. Activation of prefrontal control mechanisms have specifically been identified with attenuation of emotional responses. However, it is unclear whether recruitment of prefrontal resources alone is operative in this regard, or whether prefrontal control is important only in the role of explicit emotion regulation. This study examined whether exposure to cognitive tasks before an emotional challenge attenuated the effects of the emotional challenge. AIMS: We investigated whether a single training session could alter participants' reactivity to subsequent emotional stimuli on two computer-based tasks as well as affect ratings made during the study. We hypothesized that individuals performing the Cognitive Control (CC) task as compared to those performing the Peripheral Vision (PV) comparison task would (1) report reduced negative affect following the mood induction and the emotion task, and (2) exhibit reduced reactivity (defined by lower affective ratings) to negative stimuli during both the reactivity and recovery phases of the emotion task and (3) show a reduced bias towards threatening information. METHOD: Fifty-nine healthy participants were randomized to complete CC tasks or PV, underwent a negative mood induction, and then made valence and arousal ratings for IAPS images, and completed an assessment of attentional bias. RESULTS: RESULTS indicated that a single-session of CC did not consistently alter participants' responses to either task. However, performance on the CC tasks was correlated on subsequent ratings of emotional images. CONCLUSIONS: While overall these results do not support the idea that affective responding is altered by making healthy volunteers use their prefrontal cortex before the affective task, they are discussed in the context of study design issues and future research directions.
Towards human consistent data driven decision support systems using verbalization of data mining results via linguistic data summaries We present how the conceptually and numerically simple concept of a fuzzy linguistic database summary can be a very powerful tool for gaining much insight into the essence of data that may be relevant for a business activity. The use of linguistic summaries provides tools for the verbalization of data analysis (mining) results which, in addition to the more commonly used visualization e.g. via a GUI, graphical user interface, can contribute to an increased human consistency and ease of use. The results (knowledge) derived are in a simple, easily comprehensible linguistic form which can be effectively and efficiently employed for supporting decision makers via the data driven decision support system paradigm. Two new relevant aspects of the analysis are also outlined which was first initiated by the authors. First, following Kacprzyk and Zadrożny [1] comments are given on an extremely relevant aspect of scalability of linguistic summarization of data, using their new concept of a conceptual scalability that is crucial for large applications. Second, following Kacprzyk and Zadrożny [2] it is further considered how linguistic data summarization is closely related to some types of solutions used in natural language generation (NLG), which can make it possible to use more and more effective and efficient tools and techniques developed in this another rapidly developing area. An application of a computer retailer is outlined.
In this paper authors will show how verb valency data, added to the Croatian dictionary in NooJ, enhances recognition of VP as well as NP and PP parts of a sentence. At the Department of Information Sciences two parallel PhD Theses were being developed. One is construction of a chunker for Croatian using NooJ and the other one is construction of Croatian verb valency lexicon (CRVLLEX). Here we combined the two projects by adding the data from CRVLLEX to the existing NooJ dictionary hoping thus to obtain better, improved Croatian chunker results. Our dictionary has over 36 000 entries of which 1 884 are verbs. Each verb is only marked by its category and the FLX. Additional data is being added from the CRVLLEX. Theoretic motivation behind the construction of CRVLLEX is Praha’s Dependency Treebank (PDT) but also good Czech verb valency dictionaries (VALLEX and Verbalex). So far, CRVLLEX has 1 739 verbs with 5 118 valency frames (approx. 3 frames per verb) and 173 syntactic-semantic classes. Each word entry in CRVLLEX contains headword lemma, reflexivity example (depending whether the verb is reflexive or not) and frame entry which is the part that describes the valency frame for each verb. Every verb also has the following attributes: aspect, frequency and form, which we hope will be of great importance in disambiguating parts of a sentence surrounding the verb.
This study utilized a unique contact lens method to examine the influence of right hemisphere processing on affective judgments and pain perception. Under conditions of unilateral visual stimulation, participants completed affective ratings of films and underwent a cold pressor pain task. The goal was to examine differences in affective judgments of films projected to one or the other hemisphere and to examine the influence of this unilateral visual stimulation on pain sensitivity. The results failed to replicate a previous finding of differential affective judgments under conditions of right and left visual stimulation. Nonetheless, in female subjects, unilateral visual stimulation was significantly associated with pain lateralization. This finding is discussed in terms of attentional processes in pain perception and limitations of the current lens design.
The article shows the wealth of colloquial language features in the city environment through the presence of texts in the city reality which are designed for collective receivers/recipients (eg. sign-board, information advertising, price labels, etc.). In the research, the components revealing descanting in the urban language (dialectal and sociolectal features) were found, as well as the associated evaluation of objects and phenomena, colloquiality or even familiarity of idea transfer, and free realisation of orthographic and stylistic norms. Urban texts bear testimony of frequent language taboo breaking in the original sphere as well as in the area violating tactfulness and politeness canons, up to violation of decency and modesty. In the thesis, the changes in the sphere of native words meaning (neologisms and neosemantisms) and examples of introducing allogenic lexemes (orientalisms) are discussed. The important feature of the examples analysed is ambiguity, present in the lexical area as well as in the global apprehension of the message, which could decide about the language game played with receivers.
Building up messages as a cognitive activity within the linguistic multi-level system is the result of the interaction between the various components of this system. Yet, this interactive process occurring in the language user's mind while encoding can vary from person to person. Likewise, it also differs in different recipients while decoding. This cognitively carried out difference in encoding can result in either unintended messages, though emptied out in mature and well-formed language constructions, or ill-formed ones. The present study aims at identifying the most frequent syntactic errors at the sentential level; and how immature or vague conceptualization manifests itself in the grammar-meaning relationship as reflected in the subjects' errors. Data were collected from the writings and the performance tests of mainly two Arab groups of third and fourth-year English majors in the course of Sociolinguistics at the University of Nizwa in the Sultanate of Oman. Errors were identified, classified and interpreted in terms of the underlying cognitive processes they went through during production. Errors were related to time-tense vague mapping, finite-nonfinite confusion, sentence-clause uncertainty, incorrect embedding, voice-related inaccuracy and verbless clauses or sentences. ********** Errors in second language performance are easy to observe, and undoubtedly serve as good indicators of a learner's level of second-language knowledge. Thus, they are an integral component of language learning. The phenomenon of error has received the interest of researchers, though traditionally has been regarded as the linguistic phenomena deviant from the language rules and standard usages, reflecting learners' deficiency in language competence. Hence, many teachers simply correct individual errors as they occur, giving little concern to identifying patterns of errors or to uncovering causes other than learner ignorance. However, interpreting what the cause of these errors has been a major concern for linguists, as well as classroom practitioners. In this respect, Steinberg and Sciarini (2006), state that the systematicity of most errors is the result of the application of certain strategies when relevant second-language knowledge is not available or incomplete. Conscious resort to any of these strategies is definitely a cognitive exercise, the results of which are various deviations of the linguistic norms, sentence patterns and rules that put the language components together to produce linguistic moulds embodying composites of thoughts. Presently, therefore, with the development of linguistics, applied linguistics, psychology and other relevant subjects, attitudes toward errors changed greatly. Instead of being problems to be overcome, errors are viewed as an evidence of the learners' stages in their target language (TL) development. It is through analyzing learner errors that errors are elevated from the status of undesirability to that of a guide to the inner working of the language learning process, as (Ellis, 1985, cited by Nassaji & Fotos, 2004, at http://faculty.ksu.edu.sa/dinaalsibai/ReviewGrammar.PDF) puts it. The term grammar is used in this paper to include the two language levels of syntax and morphology since some errors in one level cannot be explained in isolation from the other. For example, the past tense of bring is brought, a case of inflection under morphology, whereas the interrogative form Did bring them? of Ali brought them is a syntactic issue. Memory and logic, which are cognitive skills, play a crucial role in language learning. As Steinberg and Sciarini (2006: 34-35) state it, that while learning to identify the words of the language, formulating rules for their use, and relating speech to the environment and mind, the child employs an exceptional memory capacity. In this respect, a massive number of particular words, phrases and sentences have to be remembered in connection with the context, whether physical or mental, in which they took place. …
Current automatic wrappers using DOM tree and visual properties of data records to extract the required information from the search engine results pages generally have limitations such as the inability to check the similarity of tree structures accurately. Our study on the properties of data records shows that these data records located in search engine results pages are not only having similar visual properties and tree structures, but they are also related semantically in their contents. In this context, we propose an ontological technique using existing lexical database for English (WordNet) for the extraction of data records. We find that wrappers designed based on ontological technique are able to reduce the number of potential data regions to be extracted, thus they are able to improve the data extraction accuracy. We then use visual cue from the browser rendering engine to locate and extract the relevant data region from the web page by measuring the size of text and image of data records. Experimental results indicate that our technique is robust and performs better than the existing state of the art visual based wrappers.
This paper proposes an approach to improve graph-based dependency parsing by using decision history. We introduce a mechanism that considers short dependencies computed in the earlier stages of parsing to improve the accuracy of long dependencies in the later stages. This relies on the fact that short dependencies are generally more accurate than long dependencies in graph-based models and may be used as features to help parse long dependencies. The mechanism can easily be implemented by modifying a graphbased parsing model and introducing a set of new features. The experimental results show that our system achieves state-ofthe-art accuracy on the standard PTB test set for English and the standard Penn Chinese Treebank (CTB) test set for Chinese. 1
Investigating local linguistic norms to discover larger patterns of language behaviour has been standard practice in sociolinguistic study. Looking closely at socially salient variables reveals patterns that problematize accepted trajectories of variation as traditional and newly emerging sociolinguistic identities interact. This paper integrates findings from multiple complementary projects to describe the forces influencing the stopping of interdental fricatives (dis ting for this thing), a highly salient marker of Newfoundland English, in and around St. John’s, the province’s major city. In urbanizing communities multivariate analysis reveals variation patterns typical of dialect erosion: older men maintain traditional norms while younger women move toward the standard, especially in linguistically salient contexts. In the same communities, a timing-based approach finds that young women seem to be agentively inserting stopped forms, suggesting that they have adopted a system with fricatives as the default choice. When we contrast urban and rural communities and affiliations, we find a more complex pattern: style shifting is greatest among urban males and rural females. We posit that these seemingly divergent patterns result from efforts by speakers to position themselves within the local social landscape during a period of rapid social change.
This paper explores whether and how a firm should adapt its strategy in view of consumer use of prior customer ratings. Specifically, we consider optimal pricing and whether the firm should offer an unexpected frill to early customers to enhance their product experiences. We show that if price history is unobserved by consumers, a forward-looking firm should always modify its strategy from single-period optimal one, but it may be optimal to do so by lowering price, by lowering price and offering frills, or by raising price and offering frills, depending on the market growth rate. Specifically, the last strategy becomes optimal when market growth rate is high enough. The results are similar when the price history is observed by consumers, except that no deviation from single-period profit maximization choices is optimal when market growth is low enough. We also analyze whether the firm should prefer that the price information be stated in or left out of consumer reviews. In addition, in considering the effects of consumer heterogeneity, we conclude that the optimal firm's effort to affect ratings is higher when the idiosyncratic part of consumer uncertainty is larger.
Based on her study of the changes affecting the understanding of premarital intimate relationships in Ukrainian villages and cities during the period of mass modernization, the author argues that pre-modern sexual practices do not correlate precisely with modern sexual practices and thus cannot be described by current lexical definitions. The modern phrase sexual intercourse has no exact correspondence to such premarital practices as poliuvannia ("hunting") and prytula ("leaning against"). They consisted of non-penetrative (or incomplete penetrative) sexual activities, rather than sexual congress for the purpose of pleasure or reproduction. Unlike modern norms, the traditional culture allowed and even encouraged premarital intimate relationships, which were understood as a sign of healthy, successful maturation. Although premarital mixed-gender sleeping arrangements were tolerated in premodern villages but condemned in growing urban areas, the percentage of premarital births markedly increased in Ukrainian cities in the late nineteenth-early twentieth centuries. The author provides a brief statistical survey of premarital births.
This research is funded by NSF Grant OISE-0968369 to Ping Li and by a University Graduate Fellowship to Benjamin Zinszer. Snodgrass, J. G., & Vanderwart, M. (1980). A standardized set of 260 pictures: norms for name agreement, image agreement, familiarity, and visual complexity. Journal of experimental psychology. Human learning and memory, 6(2), 174-215. Malt, B. C., Sloman, S. A., & Gennari, S. P. (2003). Universality and language specificity in object naming. Journal of Memory and Language, 49(1), 20-42.
Systems for syntactically parsing sentences have long been recognized as a priority in Natural Language Processing. Statistics-based systems require large amounts of high quality syntactically parsed data. Using the XLE toolkit developed at PARC and the LFG Parsebanker interface developed at Bergen, the Parsebank Project at Powerset has generated a rapidly increasing volume of syntactically parsed data. By using these tools, we are able to leverage the LFG framework to provide richer analyses via both constituent (c-) and functional (f-) structures. Additionally, the Parsebanking Project uses source data from Wikipedia rather than source data limited to a specific genre, such as the Wall Street Journal. This paper outlines the process we used in creating a large-scale LFG-Based Parsebank to address many of the shortcomings of previously-created parse banks such as the Penn Treebank. While the Parsebank corpus is still in progress, preliminary results using the data in a variety of contexts already show promise.
PURPOSE: To evaluate traumatized bone marrow with a dual-energy (DE) computed tomographic (CT) virtual noncalcium technique. MATERIALS AND METHODS: In this prospective institutional review board-approved study, 21 patients with an acute knee trauma underwent DE CT and magnetic resonance (MR) imaging. A software application was used to virtually subtract calcium from the images. Presence of fractures was noted, and presence of bone bruise was rated on a four-point scale for six femoral and tibial regions by two radiologists. CT numbers were obtained in the same regions. Consensus reading of independently read MR images served as the reference standard. Image ratings and CT numbers were subjected to receiver operating characteristic curve analysis. RESULTS: After exclusion of 16 regions owing to artifacts, MR imaging revealed 59 bone bruises in the remaining 236 regions (19 of 114 femoral, 40 of 122 tibial). Fractures were present in eight patients. Visual rating revealed areas under the curve of 0.886 and 0.897 in the femur and 0.974 and 0.953 in the tibia for observers 1 and 2, respectively. For CT numbers, the respective areas under the curve were 0.922 and 0.974. If scores of 1 and 2 (strong or mild bone bruise) were counted as positive, sensitivities were 86.4% and 86.4% and specificities were 94.4% and 95.5% for observers 1 and 2, respectively. The kappa statistic demonstrated good to excellent agreement (femur, kappa = 0.78; tibia, kappa = 0.87). CONCLUSION: This DE CT virtual noncalcium technique can subtract calcium from cancellous bone, allowing bone marrow assessment and potentially making posttraumatic bone bruises of the knee detectable with CT.
In this paper, we present what we believe to be the first data-driven dependency parser for Urdu. The parser was trained and tuned using MaltParser system, a system for data-driven dependency parsing. The Urdu dependency treebank (UDT) is used for training and testing of the Urdu dependency parser, is also presented first time. The UDT contains corpus of 2853 sentences which are annotated at multiple levels such as part-of-speech (POS) level, chunk (phrase level) and dependency relations level. The UDT also contains information about the token counter, head of current token. The annotation is done manually to build UDT. Urdu Dependency Parsing system is evaluated by conducting a series of experiments. All experiments are performed using Maltparser default algorithm with different feature models. Initial, a base line simple feature model consisting word position, word, head and dependency relation is used for Urdu dependency parsing. Then feature model is enhanced by adding part-of-speech (POS) and chunk (Phrase level) information. The results of all parsing experiments are reported. The overall best labeled accuracy (LA) achieved 74.48 % and 90.14 % of unlabeled attachment score (UAS) is achieved. The error analysis is performed by comparing output data with treebank test data which manual parsed to analyze and classify the different types of errors produced by the parser. This is very useful to identify the future directions for future expansion of the treebank and for improving the parsing accuracy. 1.