Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
While the 2007-2010 financial crisis has hit a variety of countries asymmetrically, the case of Spain is particularly illustrative: this country experienced a pronounced housing bubble partly funded via spectacular developments in its securitization markets leading to looser credit standards and subsequent financial stability problems. We analyze the sequential deterioration of credit in this country considering rating changes in individual securitized deals and on balance sheet bank conditions. Using a sample of 20, 286 observations on securities and rating changes from 2000Q1 to 2010Q1 we build a model in which loan growth, on balancesheet credit quality and rating changes are estimated simultaneously. Our results suggest that loan growth significantly affects on balance-sheet loan performance with a lag of at least two years. Additionally, loan performance is found to lead rating changes with a lag of four quarters. Importantly, bank characteristics (in particular, observed solvency, cash flow generation and cost efficiency) also affect ratings considerably. Additionally, these other bank characteristics seem to have a higher weight in the rating changes of securities issued by savings banks as compared to those issued by commercial banks. JEL Classification: G21, G12
Data integration systems attempt to provide users with seamless and flexible access to information from multiple autonomous, distributed and heterogeneous data sources through a unified query interface. Besides data are continuously growing, maintained by different organizations and managed autonomously, querying data from heterogeneous data sources faces new challenges. As data integration has been automated, the ambiguity in concept interpretation also known as semantic heterogeneity has become one of the main obstacles to this process. Introduction of the Semantic Web Vision Ontologies WordNet ontology [3] is a large lexical database that is used in many schema matching algorithms to match schemas based on the semantics of attributes. In this paper ontology based semantic query reformulation technique is followed to improve the recall of the query. The reformulated query is optimized by removing disjunctive clauses in the query to reduce the computational cost of the semantic query execution. Experimental results show that the proposed optimization technique improves recall with minimal execution time.
Identifying factors that improve the assessment of athletes' psychological functioning is imperative to make proper return-to-play decisions following concussion. Prior research indicates that an individual's affect is related to symptom reporting. The present study examines two novel methods of affect assessment in college athletes at baseline participating in a sports-concussion management program. A total of 256 athletes completed a neuropsychological baseline battery with measurements of psychological symptoms (BDI-Fast Screen, Post-Concussion Symptom Scale, and ImPact Total Symptom Score) and a measure of affective memory bias (the Affective Verbal Learning Test; AVLT). Examiners completed an observation-based rating of affect. Multivariate analysis of variance and χ2 analyses were conducted to examine the effect of affect on symptom reports. Examiners' Affect Ratings were predictive of broad symptom reporting, while the performance based index of affect (Affective Verbal Learning Test, AVLT) was more predictive of depressive symptoms. These findings suggest that performance on the AVLT may be a useful indicator of self-reported depression in a collegiate athlete sample. Additionally, these results demonstrate that examiners' behavioral assessments of affect are important in the assessment of psychological functioning in athletes. Continued work should focus on developing objective measures that are sensitive and valid for the evaluation of outcomes from concussion.
Recent applications including the Semantic Web, Web ontology and XML have sparked a renewed interest on graph-structured databases. Among others, twig queries have been a popular tool for retrieving subgraphs from graph-structured databases. To optimize twig queries, selectivity estimation has been a crucial and classical step. However, the majority of existing works on selectivity estimation focuses on relational and tree data. In this paper, we investigate selectivity estimation of twig queries on possibly cyclic graph data. To facilitate selectivity estimation on cyclic graphs, we propose a matrix representation of graphs derived from prime labeling - a scheme for reachability queries on directed acyclic graphs. With this representation, we exploit the consecutive ones property (C1P) of matrices. As a consequence, a node is mapped to a point in a two-dimensional space whereas a query is mapped to multiple points. We adopt histograms for scalable selectivity estimation. We perform an extensive experimental evaluation on the proposed technique and show that our technique controls the estimation error under 1.3% on XMARK and DBLP, which is more accurate than previous techniques. On TREEBANK, we produce RMSE and NRMSE 6.8 times smaller than previous techniques.
UFAL). Abstract. Annotated corpora such as treebanks are important for the development of parsers, language applications as well as understanding of the language itself. Only very few languages possess these scarce resources. In this paper, we describe our eort in syntactically annotating a small corpora (600 sentences) of Tamil language. Our annotation is similar to Prague Dependency Treebank (PDT 2.0) and consists of 2 levels or layers: (i) morphological layer (m-layer) and (ii) analytical layer (a-layer). For both the layers, we introduce annotation schemes i.e. positional tagging for m-layer and dependency relations (and how dependency structures should be drawn) for a-layers. Finally, we evaluate our corpora in the tagging and parsing task using well known taggers and parsers and discuss some general issues in annotation for Tamil language.
We want to demonstrate on some selected linguistic issues that classical structural and functional linguistics even with its seemingly traditional approaches has something to offer to a formal description of language and its applications in natural language processing and to illustrate by a brief reference to Functional Generative Grammar (on the theoretical side of CL) and Prague Dependency Treebank (on the applicational side) a possible interaction between linguistics and CL.
Prepositional phrase (PP) consists of two parts which are a preposition as the leading part and a word or phrase as the tail part. In accordance with this fact, this paper proposes a new approach for identifying PP. In this method, PP identification is transformed into the collocation identification of preposition itself and the right boundary word. The Cascaded Conditional Random Fields (CCRFs) is used in this approach. With the Penn Chinese Treebank 5.1 as our experiment corpus, the F <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sub> rises to 94.63%. This approach obtains breakthrough in this specific field as the current F <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sub> is about 8.6% higher than any publicly published paper.
This paper is based on the assertions of Theory of Linguistic Variation and Change and proposes a discussion about possible actions of linguistic norms over two variable phenomena in Brazilian Portuguese: the position of clitic pronouns associated with a single verb, and the use of prepositions with verbal complements indicating a “goal/recipient”. We intend to describe each phenomenon by analysing data selected from newspapers from Sao Paulo and Rio Claro published between 1900 and 1915. We compare the descriptions and evaluate the role played by the standard and the common usage which is (already) perceptible in the ‘paulistas’ published pages of that period.
In this paper, we present a simple and effective fine-grained feature generation scheme for dependency parsing. We focus on the problem of grammar representation, introducing fine-grained features by splitting various POS tags to different degrees using HowNet hierarchical semantic knowledge. To prevent the oversplitting, we adopt a threshold-constrained bottomup strategy to merge the derived subcategories. We conduct the experiments on the Penn Chinese Treebank. The results show that, with the fine-grained features, we can improve the dependency parsing accuracies by 0.52 % (absolute) for the unlabeled first-order parser, and in the case of second-order parser, we can improve the dependency parsing accuracies by 0.61% (absolute). 1
The paper presents a preliminary research on possible relations between the syntactic structure and the polarity of a Czech sentence by means of the so-called sentiment analysis of a computer corpus. The main goal of sentiment analysis is the detection of a positive or negative polarity, or neutrality of a sentence (or, more broadly, a text). Most often this process takes place by looking for the polarity items, i.e. words or phrases inherently bearing positive or negative values. These words (phrases) are collected in the subjectivity lexicons and implemented into a computer corpus. However, when using sentences as the basic units to which sentiment analysis is applied, it is always important to look at their semantic and morphological analysis, since polarity items may be influenced by their morphological context. It is expected that some syntactic (and hypersyntactic) relations are useful for the identification of sentence polarity, such as negation, discourse relations or the level of embeddedness of the polarity item in the structure. Thus, we will propose such an analysis for a convenient source of data, the richly annotated Prague Dependency Treebank.
The short-term memory for the flavour of a wine in frequent but nonexpert drinkers was studied in an experiment comparing memory for wine under conditions where participants imagined and remembered a target wine with memory under conditions where participants carried out subsequent imaging and image rating tasks on the wine and on competing mental images. The results show that the cognitive operation of mental image rating produced enhanced memory when the operations were performed on the image of the target wine and somewhat impaired memory when operations were performed on an image of a competing flavour or operations performed on mental images in other domains. It is concluded that sensory traces of flavour may be effectively maintained in short-term memory when supported by appropriate encoding operations.
In this paper we try to present how information technologies as tools for the creation of digital bilingual dictionaries can help the preservation of natural languages. Natural languages are an outstanding part of human cultural values and for that reason they should be preserved as part of the world cultural heritage. We describe our work on the bilingual lexical database supporting the Bulgarian-Polish Online dictionary. The main software tools for the webpresentation of the dictionary are shortly described. We focus our special attention on the presentation of verbs, the richest from a specific characteristics viewpoint linguistic category in Bulgarian.
The Plain Meaning Rule is often assailed on the grounds that it is unprincipled—that it substitutes for careful analysis an interpreter's ad hoc and impressionistic intuition about the meaning of legal texts. But what if judges and lawyers had the means to test their intuitions about plain meaning systematically? Then initial linguistic impressions about the meaning of a legal text might be viewed as hypotheses to be tested, rather than determinative criteria upon which to base important decisions. There exists very little legal scholarship on corpus linguistics—the study of language function and use through large, electronic linguistic databases called corpora—and the role that corpus methods might play in legal interpretation. This omission becomes more and more striking as scholars and jurists (and even the United States Supreme Court) have found themselves persuaded by corpus-based arguments. This Article argues that the plain or ordinary meaning of a given term in a given context is an empirical matter that may be quantified through corpus-based methods. These methods, when applied to questions of legal ambiguity, present significant advantages over existing empirical approaches to plain meaning and over the prevailing intuition-based interpretive approach of many courts. Because large, sophisticated linguistic corpora are widely available and easy to use, and because corpus methods offer a more principled and systematic alternative to the impressionistic interpretation of legal texts, corpus linguistics may one day revolutionize the process of legal interpretation.
Using semi-supervised EM, we learn finegrained but sparse lexical parameters of a generative parsing model (a PCFG) initially estimated over the Penn Treebank. Our lexical parameters employ supertags, which encode complex structural information at the pre-terminal level, and are particularly sparse in labeled data – our goal is to learn these for words that are unseen or rare in the labeled data. In order to guide estimation from unlabeled data, we incorporate both structural and lexical priors from the labeled data. We get a large error reduction in parsing ambiguous structures associated with unseen verbs, the most important case of learning lexico-structural dependencies. We also obtain a statistically significant improvement in labeled bracketing score of the treebank PCFG, the first successful improvement via semi-supervised EM of a generative structured model already trained over large labeled data. 1
This preliminary study explored whether neurophysiological responses to visual stimuli, including attachment-related pictures, differed based on attachment status. Along with self-reported valence ratings and reaction times, recorded electroencephalographic (EEG) responses to a total of 100 images, 25 each of Positive, Negative, Neutral, and Personal (each participant's parents and child), were analyzed within and among three mothers with three attachment statuses (Dismissing, Preoccupied, and Secure), as judged by the Adult Attachment Interview (AAI). All three mothers gave their highest pleasantness ratings for Personal photographs. However, differences emerged when cross-region Alpha2 activation patterns in response to each picture type were compared amongst attachment categories. Alpha2 activation recorded during viewing of the participants' children's photographs was similar to viewing Negative pictures for mothers with insecure (Dismissing and Preoccupied) status; whereas the Alpha2 activation of the mother with Secure status towards photographs of her child was similar to Positive pictures. Different patterns of hemispheric asymmetry in Beta1 frequency when processing different picture types were also found. The mother with Dismissing status showed significantly stronger left-hemisphere Beta1 activation across all image types. The Preoccupied mother showed significantly stronger right-hemisphere Beta1 activation for all but the Neutral images, during which activation did not differ between the two hemispheres. The mother with Secure status showed significantly stronger Beta1 activation in the left hemisphere for all but parental Personal photos, during which activation did not differ between the two hemispheres. Implications from the current findings and future research possibilities are discussed.
A Morphological Analyzer and Generator are two crucial tools involving any Natural Language Processing of Dravidian Languages. The present paper discusses the improvization of the existing Morphological Analyzer and Generator for Tamil by defining and describing the relevant linguistic database required for the purpose of developing them. The implementation of an open source platform called Apertium to handle inflection as well as derivation for word level analysis and generation of Tamil is also discussed. The paper also presents the efficacy, coverage and speed of the module against the large corpora. The paper also draws inferences of the morphological categories in their inflection and problems in analysing them. I. INTRODUCTION A language like Tamil is regarded as morphologically rich wherein the words are formed of one or more stems/roots plus one or more suffixes. So the complexity of morphology requires a more sophisticated morphological analyzer and generator. A morphological analyzer is a computational tool to analyze word forms into their roots along with their constituent functional elements. The morphological generator is the reverse process of an analyzer i.e. from a given root and functional elements, it generates the well-formed word forms. The present attempt involves a practical adoption of lttoolbox for the Modern Standard Written Tamil in order to develop an improvised open source morphological analyzer and generator. The tool uses the computational algorithm called Finite State Transducers for one-pass analysis and generation, and the database is based on the morphological model called Word and Paradigm.
Nous presentons une architecture pour l’analyse syntaxique en deux etapes. Dans un premier temps un analyseur syntagmatique construit, pour chaque phrase, une liste d’analyses qui sont converties en arbres de dependances. Ces arbres sont ensuite reevalues par un reordonnanceur discriminant. Cette methode permet de prendre en compte des informations auxquelles l’analyseur n’a pas acces, en particulier des annotations fonctionnelles. Nous validons notre approche par une evaluation sur le corpus arbore de Paris 7. La seconde etape permet d’ameliorer significativement la qualite des analyses retournees, quelle que soit la metrique utilisee.
Syntactic structures have been good features for opinion analysis, but it is not easy to use them. To find these features by supervised learning methods, correct syntactic labels are indispensible. Two possible sources to acquire syntactic structures are parsing trees and dependency trees. For the annotation processing, parsing trees are more readable for annotators, while dependency trees are easier to use by programs. To use syntactic structures as features, this paper tried to annotate on human friendly materials and transform these annotations to the corresponding machine friendly materials. We annotated the gold answers of opinion syntactic structures on the parsing tree from Chinese Treebank, and then proposed methods to find their corresponding dependency relations on the dependency trees generated from the same sentence. With these relations, we could train a model to annotate opinion dependency relations automatically to provide an opinion dependency parser, which is language independent if language resources are incorporated. Experiment results show that the annotated syntactic structures and their corresponding dependency relations improve at least 8% of the performance of opinion analysis.
This paper examines how W.E.B. DuBois' concept of double consciousness influenced the interactions of 13 Black youth inside an after school Community Literacy Intervention Program (CLIP). Du Bois, a pre-eminent 20th century Black sociologist, used double consciousness as a lens to help explain social and psychological tensions that African Americans encounter while negotiating their identities in a societal context structured mainly upon dominant white cultural and linguistic norms and values. The authors provide a conceptual framework for understanding the interpretive processes that signify double consciousness which includes: surveying the context; assessing risks and identity consequences; articulating mainstream or race conscious reads, and bridging/or disengaging. Implications for pre-service teachers and particularly urban educators are discussed.
This paper presents a method for word sense disambiguation based on Lesk algorithm which uses lexical database WordNet as knowledge base. The word sense disambiguation is the process of automatically clarifying a meaning of a word in its context. In general, ontology means the meaning or analogous term. It can be interpreted by relating the word with other words in the sentence. This tool accepts English statement as input and gives best possible meaning of given word. Method is experimented with senseval-2 test data for lexical sample task. The results show the betterment over the original Lesk algorithm.
We consider a very simple, yet effective, approach to cross language adaptation of dependency parsers. We first remove lexical items from the treebanks and map part-of-speech tags into a common tagset. We then train a language model on tag sequences in otherwise unlabeled target data and rank labeled source data by perplexity per word of tag sequences from less similar to most similar to the target. We then train our target language parser on the most similar data points in the source labeled data. The strategy achieves much better results than a non-adapted baseline and stateof-the-art unsupervised dependency parsing, and results are comparable to more complex projection-based cross language adaptation algorithms. 1
The poor grammatical output of Machine Translation (MT) systems appeals syntax-based approaches within language modeling. However, previous studies showed that syntax-based language modeling using (Context-Free) Treebank Grammars was not very helpful in improving BLEU scores for Chinese-English machine translation. In this article we further study this issue in the context of Chinese-English syntax-based Statistical Machine Translation (SMT) where Synchronous Tree Substitution Grammars (STSGs) are utilized to model the translation process. In particular, we develop a Tree Substitution Grammar-based language model for syntax-based MT, and present three methods to efficiently integrate the proposed language model into MT decoding. In addition, we design a simple and effective method to adapt syntax-based language models for MT tasks. We demonstrate that the proposed methods are able to benefit a state-of-the-art syntax-based MT system. On the NIST Chinese-English MT evaluation corpora, we finally achieve an improvement of 0.6 BLEU points over the baseline.
AIM: To evaluate a standardised MRI acquisition protocol and a new image rating scale for disease severity in patients with progressive supranuclear palsy (PSP) and multiple systems atrophy (MSA) in a large multicentre study. METHODS: The MRI protocol consisted of two-dimensional sagittal and axial T1, axial PD, and axial and coronal T2 weighted acquisitions. The 32 item ordinal scale evaluated abnormalities within the basal ganglia and posterior fossa, blind to diagnosis. Among 760 patients in the study population (PSP = 362, MSA = 398), 627 had per protocol images (PSP = 297, MSA = 330). Intra-rater (n = 60) and inter-rater (n = 555) reliability were assessed through Cohen's statistic, and scale structure through principal component analysis (PCA) (n = 441). Internal consistency and reliability were checked. Discriminant and predictive validity of extracted factors and total scores were tested for disease severity as per clinical diagnosis. RESULTS: Intra-rater and inter-rater reliability were acceptable for 25 (78%) of the items scored (≥ 0.41). PCA revealed four meaningful clusters of covarying parameters (factor (F) F1: brainstem and cerebellum; F2: midbrain; F3: putamen; F4: other basal ganglia) with good to excellent internal consistency (Cronbach α 0.75-0.93) and moderate to excellent reliability (intraclass coefficient: F1: 0.92; F2: 0.79; F3: 0.71; F4: 0.49). The total score significantly discriminated for disease severity or diagnosis; factorial scores differentially discriminated for disease severity according to diagnosis (PSP: F1-F2; MSA: F2-F3). The total score was significantly related to survival in PSP (p<0.0007) or MSA (p<0.0005), indicating good predictive validity. CONCLUSIONS: The scale is suitable for use in the context of multicentre studies and can reliably and consistently measure MRI abnormalities in PSP and MSA. Clinical Trial Registration Number The study protocol was filed in the open clinical trial registry (http://www.clinicaltrials.gov) with ID No NCT00211224.
Flat noun phrase structure was, up until recently, the standard in annotation for the Penn Treebanks. With the recent addition of internal noun phrase annotation, dependency parsing and applications down the NLP pipeline are likely affected. Some machine translation systems, such as TectoMT, use deep syntax as a language transfer layer. It is proposed that changes to the noun phrase dependency parse will have a cascading effect down the NLP pipeline and in the end, improve machine translation output, even with a reduction in parser accuracy that the noun phrase structure might cause. This paper examines this noun phrase structure’s effect on dependency parsing, in English, with a maximum spanning tree parser and shows a 2.43%, 0.23 Bleu score, improvement for English to Czech machine translation. 1
PURPOSE: To report how youths, both with and without idiopathic scoliosis (IS), respond to questions about their self-image and perceptions of body shape. An additional purpose is to describe themes that emerged as important to youths with IS to better understand scoliosis from their perspective. METHODS: Descriptive qualitative and quantitative methods were utilized. Subject interviews were conducted, as part of a larger cognitive interviewing study on the Spinal Appearance Questionnaire, using a cross-sectional sample of 76 females between 8 and 16 years of age with IS and who were typically developing (TD), without scoliosis. RESULTS: IS and TD subjects revealed similar findings when asked what makes them look good versus their peers; self-image ratings were also positive. Predominant themes from open-ended responses include physical appearance, feelings, brace wear, and discomfort. CONCLUSION: Self-image and body shape did not differ significantly between groups. The identified themes warrant further exploration as they are significant and important to youth with scoliosis.
L’exploitation de corpus analyses syntaxiquement (ou corpus arbores) pour le public non specialiste n’est pas un probleme trivial. Si la communaute du TAL souhaite mettre a la disposition des chercheurs non-informaticiens des corpus comportant des annotations linguistiques complexes, elle doit imperativement developper des interfaces simples a manipuler mais permettant des recherches fines. Dans cette communication, nous presentons les modes de recherche « grand public » developpe(e)s dans le cadre du projet Scientext, qui met a disposition un corpus d’ecrits scientifiques interrogeable par partie textuelle, par partie du discours et par fonction syntaxique. Les modes simples sont decrits: un mode libre et guide, ou l’utilisateur selectionne lui-meme les elements de la requete, et un mode semantique, qui comporte des grammaires locales preetablies a l’aide des fonctions syntaxiques.
The conventional sequence labeling methods for Chinese word segmentation do not fully utilize the linguistic information, which restricts further improvements of the performance. Chinese morphology intensively investigates the constructions and usages of Chinese words, which is helpful to Chinese word segmentation. Furthermore, some word segmentation ambiguities cannot be resolved only by means of the lexical information, and the final disambiguations take place in the parsing process. In this paper, we propose a parsing-based Chinese word segmentation model, which can fully utilize the morphological and syntactic information. Experiments on Penn Chinese Treebank(CTB) 5.0 show that the proposed model obtains competitive performances as the CRFs-based model. To investigate the relationship between our parsing-based model and the CRFs-based model, a maximum entropy model based framework for integrating different knowledge sources is employed. The integrating model obtains an F-measure of 97.9, 25% in segmentation error rate reduction relative to the CRFs-based model, which indicates that the two models are complementary to each other.
OBJECTIVE: To determine the effectiveness of a modulated acoustic startle reflex paradigm with emotional imagery in studying physiological changes associated with emotional responses in persons with traumatic brain injury (TBI). SETTING: Outpatient rehabilitation hospital. PARTICIPANTS: Six individuals with moderate to severe TBI. Mean age was 32 years and mean years postinjury were 9.9. METHOD: The modulated acoustic startle reflex procedure involved imagery of emotional scripts (joy, anger, fear, and neutral) followed by a startle noise, versus startle noise alone (no script). MEASURES: Eyeblink and skin conductance response, subjective arousal and valence ratings of the scripts, and general anger questionnaire. RESULTS: Startle blink responses following anger imagery were significantly smaller than those following fear (P =.006) and neutral (P =.023) imagery. Skin conductance response did not change significantly based on the content of the scripts (P =.070). CONCLUSIONS: Large startle blink responses indicate avoidance of a stimulus. Our findings suggest that participants with TBI did not have an avoidant reaction to anger-inducing stimuli. Skin conductance response findings may imply arousal impairments. The modulated acoustic startle reflex was effective in measuring emotional responses; however, larger studies comparing persons with TBI with control groups are needed to further explore these findings.
This paper presents a probabilistic approach for POS tagging that combines HMMs and character language models being applied to Portuguese texts. In this approach, the emission probabilities for each hidden state in a HMM are estimated by a proper character language model. The tagger built has been trained and tested on Bosque, a subset of Floresta Sinta(c)tica treebank, reaching 96.2% accuracy with a 39-tag tagset and 92.0% with a 257-tag tagset extended with inflexion information.
Many recent experiments in the automatic classification of discourse relations have limited themselves to a small set of coarse categories. While there are eminent reasons to do so – annotation for a small sets of categories can be created more reliably, and possibly also be defined in a more clear – cut way than finer distinctions – it is an interesting question whether the finer-grained distinctions present in some annotated corpora can be reconstructed reliably. The present paper investigates the feasibility of such fine-grained tagging of discourse relations using data from the Penn Discourse Treebank. 1
OBJECTIVES: Dental phobia is currently classified as a specific phobia of the blood-injection-injury (BII) subtype. In another subtype, animal phobia, enhanced amplitudes of late event-related potentials have consistently been identified for patients during passive viewing of disorder-relevant pictures. However, this has not been shown for BII phobics, and studies with dental phobics are lacking. Findings on cardiac responses in BII phobia during exposure are heterogeneous, as some studies showed a diphasic pattern of heart rate acceleration and deceleration, whereas others observed pure acceleration. In contrast, heart rate increase has consistently been shown for dental phobics, resembling the reaction of animal phobics. Moreover, the BII subtype is characterized by elevated disgust reactivity whereas the role of habitual disgust proneness in dental phobia is unclear. METHODS: We recorded the electroencephalogram and the electrocardiogram from 18 dental phobic and 18 healthy women while they watched pictures depicting dental treatment, disgust, fear and neutral items. RESULTS: Phobics relative to controls showed an enhanced late positive potential (300-700 ms) and heart rate acceleration towards phobic material, reflecting motivated attention and fear. Affective ratings revealed that dental phobics experienced significantly higher levels of fear than disgust during exposure to phobia-relevant material. Patients' elevated habitual disgust proneness was restricted to specific domains, such as the oral incorporation of offensive objects. CONCLUSION: The psychophysiology of dental phobia resembles the fear-dominated subtypes of specific phobia reported in earlier studies. Future studies should continue to investigate whether the current classification of this disorder as BII phobia needs to be reconsidered.
Psycholinguistic studies on whether classifiers facilitate processing object-extracted relative clauses (RC) in Mandarin have often made use of a classifier mismatch-match configuration, wherein a preceding classifier mismatches the following RC-subject but matches the modified head noun. However, an examination of the Chinese Treebank corpus 5.0 shows this configuration rarely occurs. None of the 10 tokens of pre-RC classifiers conforms to the mismatch-match configuration in a real sense. Instead, either a dropped RC-subject or some intervening item successfully avoids anticipated lexical disruption effects induced by a mismatching classifier. The results of analysis suggest that the constructed examples used in previous psycholinguistic studies may not realistically test natural language processing procedures.
This paper introduces, an XML format developed to serialise the object model defined by the ISO Syntactic Annotation Framework SynAF. Basing on widespread best practices we adapt a popular XML format for syntactic annotations, TigerXML, with additional features to support a variety of syntactic phenomena including constituent and dependency structures, binding, and different node types such as compounds or empty elements. We also define interfaces to other formats and standards including the Morpho-syntactic Annotation Framework MAF and the ISOCat Data Category Registry. Finally a case study of the German Treebank TueBa-D/Z is presented, showcasing the handling of constituent structures, topological fields and coreference annotation.
Abstract The language of a speech community can only act as an identity marker for all of its speakers if linguistic norms are widely shared and if a minimal number of language varieties are spoken. This article examines briefly how a linguistic norm came to serve the whole of Iceland and how a situation of relative linguistic homogeneity was maintained for centuries. Sociolinguistic theory tells us that the speech community that we can reconstruct for early Iceland should lead to the establishment and maintenance of local norms. However, Iceland, arguably monodialectal, was certainly characterized by long-term linguistic homogeneity and remained a society where nucleated settlements barely formed over a thousand-year period. Scholars have argued that a mixture of dialects leveled shortly after the settlement of Iceland in the ninth century (Settlement). Studies show that dialect leveling requires dialect mixing, the convergence of people on one place, and sustained linguistic contact between the speakers. The settlement pattern of Iceland is indicative of population divergence (not convergence) and there is limited evidence of sustained contact. It is therefore proposed that the dialect leveling might be linked instead with significant population movements and social upheaval in mainland Scandinavia in the immediate pre-Viking period. The variety of Norse that was taken westward across the Atlantic might itself already have been the result of several earlier stages of mixing and koineization. It is only by combining linguistic, historical, and archaeological knowledge that this problem of how one linguistic norm came to serve the whole of Iceland can be understood.
Anaphora resolution is one of the most difficult tasks in NLP. The ability to identify non-referential pronouns before attempting an anaphora resolution task would be significant, since the system would not have to attempt resolving such pronouns and hence end up with fewer errors. In addition, the number of non-referential pronouns has been found to be non-trivial in many domains. The task of detecting non-referential pronouns could also be incorporated into a part-of-speech tagger or a parser, or treated as an initial step in semantic interpretation. In this article, I describe a machine learning method for identifying non-referential pronouns in an annotated subsegment of the Penn Arabic Treebank using three different feature settings. I achieve an accuracy of 97.22% with 52 different features extracted from a small window size of -5/+5 tokens surrounding each potentially non-referential pronoun.
Eighteenth-century language usage is markedly under-represented in the first two editions of the OED, whose quotations for this period were gathered almost entirely during the late nineteenth and early twentieth centuries. This article reviews some of the possible causes, characteristics and consequences of OED’s gap in eighteenth-century documentation and shows that female authors were particularly scanted. The role of quotations in the OED, as the evidential basis for the dictionary, is briefly considered, along with eighteenth-century (and Victorian/Edwardian) views on women and language, and the availability of female-authored texts for quotation by the lexicographers. The article reports sample reading in eighteenth-century female writers (especially Jean Adam, Penelope Aubin and Anna Seward), which shows that OED could easily have supplied its eighteenth-century deficiency from such authors, and that it often favoured distinctive usages in female-authored texts—innovative, eccentric or domestic vocabulary—rather than usage which exemplified linguistic norms (especially in poetry, where Seward’s case is examined). It also discusses revisions to the OED so far conducted in the third (ongoing) edition, and their implications for readers and editors of eighteenth-century texts.
Abstract The aim of computational semantics is to capture the meaning of natural language expressions in representations suitable for performing inferences, in the service of understanding human language in written or spoken form. First‐order logic is a good starting point, both from the representation and inference point of view. But even if one makes the choice of first‐order logic as representation language, this is not enough: the computational semanticist needs to make further decisions on how to model events, tense, modal contexts, anaphora and plural entities. Semantic representations are usually built on top of a syntactic analysis, using unification, techniques from the lambda‐calculus or linear logic, to do the book‐keeping of variable naming. Inference has many potential applications in computational semantics. One way to implement inference is using algorithms from automated deduction dedicated to first‐order logic, such as theorem proving and model building. Theorem proving can help in finding contradictions or checking for new information. Finite model building can be seen as a complementary inference task to theorem proving, and it often makes sense to use both procedures in parallel. The models produced by model generators for texts not only show that the text is contradiction‐free; they also can be used for disambiguation tasks and linking interpretation with the real world. To make interesting inferences, often additional background knowledge is required (not expressed in the analysed text or speech parts). This can be derived (and turned into first‐order logic) from raw text, semi‐structured databases or large‐scale lexical databases such as WordNet. Promising future research directions of computational semantics are investigating alternative representation and inference methods (using weaker variants of first‐order logic, reasoning with defaults), and developing evaluation methods measuring the semantic adequacy of systems and formalisms.
The dictionaries are among the wellknown tools for applications in everyday life, education, sciences, humanities, and human communication. The recent developments of information technologies contribute to the design and creation of new software tools with a wide range of applications, especially for natural language processing. The paper presents an online bilingual dictionary as a technological tool for applications in digital humanities, and describes the structure and content of the bilingual Lexical Database supporting a Bulgarian-Polish online dictionary. It focuses especially on the presentation of verbs, which form the richest from a specific characteristics viewpoint linguistic category in Bulgarian. The main software modules for webpresentation of this digital dictionary are also shortly described. 1
We introduce synchronous tree adjoining grammars (TAG) into tree-to-string transla-tion, which converts a source tree to a target string. Without reconstructing TAG deriva-tions explicitly, our rule extraction algo-rithm directly learns tree-to-string rules from aligned Treebank-style trees. As tree-to-string translation casts decoding as a tree parsing problem rather than parsing, the decoder still runs fast when adjoining is included. Less than 2 times slower, the adjoining tree-to-string system improves translation quality by +0.7 BLEU over the baseline system only al-lowing for tree substitution on NIST Chinese-English test sets. 1
Good Dictionary Examples or GDEX is a tool in the Sketch Engine designed to help lexicographers with identifying dictionary examples by ranking sentences according to how likely they are to be good candidates. The ranking is done automatically using various syntactic and lexical features. So far, only GDEX for English has been available. This paper presents the design and evaluation of Slovene GDEX, which was used for finding good examples for the new lexical database of Slovene, one of the activities in the Communication in Slovene project. Several different GDEX configurations were designed, evaluated and compared. The evaluation involved examining sentences of lemmas belonging to different word classes. Good sentences were logged for subsequent analysis with external data-mining software, WEKA. The observed behaviour was then used to adjust the parameters of the GDEX classifiers. We believe that the procedure of identifying features of good examples and their values, described in this paper, can be used for the development of GDEX for any language.
While most dialectological research so far focuses on phonetic and lexical phenomena, we use recent fieldwork in the domain of dialect syntax to guide the development of multidialectal natural language processing tools. In particular, we develop a set of rules that transform Standard German sentence structures into syntactically valid Swiss German sentence structures. These rules are sensitive to the dialect area, so that the dialects of more than 300 towns are covered. We evaluate the transformation rules on a Standard German treebank and obtain accuracy figures of 85% and above for most rules. We analyze the most frequent errors and discuss the benefit of these transformations for various natural language processing tasks. 1
Based on two syntactic dependency treebanks built with two different styles of Chinese, a statistical study is conducted regarding word-frequency and distributions. We extracted three grammatical words as the research objects and analyzed their network features, including all degree, out-degree, in-degree, all closeness, in-closeness, out-closeness and betweenness. Then these three nodes were removed from the networks. We recorded and compared the network features of the two original networks and the three networks from which one node is respectively removed, including the number of vertices, average degree, average path length, diameter, the number of isolated vertices, domain and density. The results show that all three function words are central nodes of the Chinese syntactic networks but have different status. Their influence to the overall structure is also quite different. The research not only provides a new method for the study about Chinese grammatical words but also provides a new way of thinking the node characteristics in the complex network.