Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
In this paper, we present a modular incremental statistical model for English full parsing. Unlike other full parsing approaches in which the analysis of the sentence is a uniform process, our model separates the full parsing into shallow parsing and sentence skeleton parsing. In shallow parsing, we finish POS tagging, Base NP identification, prepositional phrase attachment and subordinate clause identification. In skeleton parsing, we use a layered feature-oriented statistical method. Modularity possesses the advantage of solving different problems in parsing with corresponding mechanisms. Feature-oriented rule is able to express the complex lingual phenomena at the key point if needed. Evaluated on Penn Treebank corpus, we obtained 89.2% precision and 89.8% recall.
this paper I will discuss a framework for semantics which allows us to record truth-conditional and compositional analyses as dependency-style corpus annotations in a direct and fine-grained fashion. This method eliminates the need for a semantic representation formalism by decomposing semantic information into simple statements about (word or morpheme) tokens. A collection of such data would form a new kind of linguistic treebank. The main purpose of this article is to show that the present approach makes it possible to combine formal semantics and corpus-oriented study of language use in new and interesting ways. The methodology of this framework, which I call Token Dependency Semantics (TDS, Dahllf [4]), is in several respects different from the common one(s) in traditional formal semantics. TDS nevertheless delivers a fairly conventional (but ontologically restrained) analysis of truth-conditional meaning
XML and related W3C standards (XSLT, XML Schema, XPath, DOM, etc.) often take part in the linguistic data representation and interchange today but not as a direct way to implement efficient data manipulation software tools. The aim of the paper is to show that incorporation of XML and the standards that surround it can bring general applicability of the implemented system for various kinds of linguistic data and also easy extensibility of such systems. As a case study we present a designed and implemented system called DEB (Dictionary Editor and Browser) that is able to manage lexical data from dictionaries to lexical databases, semantic networks, and complex ontologies. The smart design also facilitates the connection to other linguistic tools such as corpus managers or morphological analysers.
The thesis is titled Mario Monteleone Electronic Dictionaries and Lexicography. Uses language with lexical databases provided concerning the relationship (and potential) between traditional lexicography and computational linguistics. In particular, Mr. Monteleone is trying to establish whether some of the lexicography analytical limits can be exceeded through the application of research methods in computational linguistics. To this end, Mr. Monteleone is conducting a detailed analysis of the two disciplines, of their traditional instruments, ie dictionaries paper and electronic dictionaries, in terms of their construction methods and their application areas. The result is a handbook for traditional lexicographers and computer focused on the identification and classification of linguistic data for the accomplishments of paper and electronic dictionaries.
Abstract. The problem of Prepositional Phrase (PP) attachment disambiguation consists in determining if a PP is part of a noun phrase, as in He sees the room with books, or an argument of a verb, as in He fills the room with books. Volk has proposed two variants of a method that queries an Internet search engine to find the most probable attachment variant. In this paper we apply the latest variant of Volk’s method to Spanish with several differences that allow us to attain a better performance close to that of statistical methods using treebanks. 1
Choosing the statistical model is the key problem in statistical parsing. Statistical model lies in the core of NLP parsing. This paper investigates 4 primary statistical parsing models, namely PCFG, history-based model, cascading parsing model and head-driven parsing model, and compares their performances in a 10000 Chinese treebank. The analysis based on the experiment were shown in the paper. The comparative study of these models can be exploited to build the practical and effective Chinese parser.
In den letzten Jahren ist die Zahl der verfgbaren linguistisch annotierten Korpora stndig gewachsen. Zu den bekanntesten gehren das Brown-Korpus, das Susanne-Korpus, die Penn-Treebank, das Negra-Korpus, das Tiger-Korpus und die im Zusam-
En este artículo estudiamos el problema de la estimación de gramáticas incontextuales \nestocásticas en formato general y su uso en un modelo de lenguaje híbrido. \nEn este trabajo se propone la estimación de una gramática incontextual estocástica usando \nuna nueva versión del algoritmo de Earley que permite manejar muestras parentizadas. El \nmodelo de lenguaje híbrido es definido como una combinación lineal de un modelo de ngramas \nbasado en palabras, que se utiliza para capturar las relaciones locales entre palabras, \ny una gramática estocástica, basada en categorías junto con una distribución de palabras en \ncategorías, que se utiliza para representar las relaciones a largo término entre estas categorías. Se han realizado experimentos usando el corpus UPenn Treebank. La evaluación \nde los modelos se ha realizado desde el punto de vista de la perplejidad de un conjunto \nde test, y desde el punto de vista de la tasa de errores por palabra en un experimento de \nreconocimiento automático del habla.
The aim of this article is to show that the role of legal terminology in juridical discourse can adequately be determined in a framework of a juridical textwork model only. This implies to analyse their role under pragmatic and intertextual aspects. It will be contended that neither semantics nor the wording of the law can be employed as a starting point for an adequate interpretation of legal language in public discourse. Of crucial importance is rather the question or how the legal text and social reality can combine to constitute the legal norm (which is more than just the legal text). This is illustrated with the example of German court decisions on the issue of sit-ins that were organised by the peace movement in the 1970s and 1980s to block the access to American army bases. It will be demonstrated that in the process of purling the coercion law in concrete normative terms through different courts, specific legal terms are semanticly modified and adjusted to specific language use, hence constituting semantic battles and linguistic norm conflicts in the juridical discourse.
Abstract This paper describes a framework for building story traces (compact global views of a narrative) and story projections (selections of key elements of a narrative) and their applications in text understanding and classification. Word and sense properties are extracted from documents using the WordNet lexical database enhanced with Prolog inference rules and a number of lexical transformations. Inference rules are based on navigation in various WordNet relation chains (hypernyms, meronyms, entailment and causality links, etc.) and derived relations expressed as Boolean combinations of node and edge properties used to direct the navigation. The resulting abstract story traces provide a compact view of the underlying narrative's key content elements and a means for automated indexing and classification of text collections. Ontology driven projections act as a kind of “semantic lenses” and provide a means to select a subset of a narrative whose key sense elements are subsumed by a set of concepts, predicates and properties expressing the focus of interest of a user. Finally, we discuss applications of these techniques in text understanding, classification of text collections and answering questions about a text.
Machine translation engines draw on various types of databases. This paper is concerned with Arabic as a source or target language, and focuses on lexical databases. The non-concatenative nature of Arabic morphology, the complex structure of Arabic word-forms, and the general use of vowel-free writing present a real challenge to NLP developers. We show here how and why a stem-grounded lexical database, the items of which are associated with grammar-lexis specifications – as opposed to a root-&-pattern database –, is motivated both linguistically and with regards to efficiency, economy and modularity. Arguments in favour of databases relying on stems associated with grammar-lexis specifications (such as DIINAR.1 or the Arabic dB under development at SYSTRAN), rather than on roots and patterns, are the following: (a) The latter include huge numbers of rule-generated word-forms, which do not actually appear in the language. (b) Rule-generated lemmas – as opposed to existing ones – are widely under-specified with regards to grammar-lexis relations. (c) In a Semitic language such as Arabic, the mapping of grammar-lexis specifications that need to be associated with every lexical entry of the database is decisive. (d) These specifications can only be included in a stem-based dB. Points (a) to (d) are crucial and in the context of machine translation involving Arabic.
El estudio experimental de la emoción requiere de estímulos que evoquen en una forma confiable reacciones psicológicas y fisiológicas que varien sistemáticamente sobre el rango de emociones de acuerdo a las dimensiones de valencia (agradable o desagradable), activación (excitado o calmado) y dominancia (alta y baja) (Lang, Bradley, Cuthbert, 1999). A pesar de que los correlatos neurales de las emociones básicas han sido investigados, la organización neural de las "emociones morales" en el cerebro humano no se conocen bien. El objetivo de la presente investigación fue obtener un grupo de estimulos diferenciados (fotografías) y caracterizarlos en términos de su valencia afectiva, activación, dominancia, y contenido moral, en una población mexicana. Se seleccionaron fotografías que representan escenas con una carga emocional amplia como violaciones morales (escenas de guerra, asaltos físicos, etc), escenas aversivas sin connotación moral (tumores, cuerpos mutilados) y escenas naturales (toallas, mesas, puertas, etc. ). Los sujetos evaluaron cada fotografía de acuerdo a su valencia, activación, dominancia y contenido moral (ausente o extremo). Para la evaluación, se utilizó la Escala Internacional Self-Assessment Maniki Affective Rating System desarrollada por Lang (1980). Se discute las implicaciones de los datos, para el estudio de las emociones y del juicio moral.
In this paper we will present work carried out lately on the 50,000 words Italian Spontaneous Speech Corpus called AVIP, under national project API, made available for free download from the website of the coordinator, the University of Naples. We will concentrate on the tuning of the parser for Italian which had been previously used to parse 100,000 words corpus of written Italian within the National Treebank initiative coordinated by ILC in Pisa. The parser receives as input the adequately transformed orthographic transcription of the dialogues making up the corpus, in which pauses, hesitations and other disfluencies have been turned into most likely corresponding punctiation marks, interjections or truncation of the word underlying the uttered segment.\nThe most interesting phenomenon we will discuss is without any doubts "overlapping", i.e. a speech event in which two people speak at the same time by uttering actual words or in some cases nonwords, when one of the speakers, usually the one which is not the current turntaker, interrupts the current speaker.\nThis phenomenon takes place at a certain point in time where it has to be anchored to the speech signal but in order to be fully parsed and subsequently semantically interpreted, it needs to be referred semantically to a following turn.
Natural language generation (NLG) is the task of formulating a fluent sequence of words in natural language to communicate information or ideas in applications like machine translation, human-computer dialogue, automatic summarization, and question-answering. Realization, a fundamental subtask of NLG, produces an individual sentence from a sentence plan specified in terms of linguistic relations between words and/or concepts. It involves determining the order of words, inserting function words like determiners and prepositions, performing morphological inflections, and ensuring grammaticality and agreement. An ultimate goal for natural language generation is to develop a large-scale, robust, general-purpose system. Two primary challenges are scaling up to broad coverage of syntax and producing high quality output. The irregularity of natural language makes it difficult to know how to combine linguistic primitives into fluent sentences. Also, the knowledge resources for making such a determination are time-consuming and labor-intensive to assemble, leading to a knowledge acquisition bottleneck. Evaluating whether a realizer performed appropriately is an additional challenge. There can often be more than one acceptable output, and no tools exist that can automatically assess grammaticality or fluency. This thesis takes an approach of using probabilistic models learned from text corpora to rank candidate sentences and output the most likely. It contributes (1) a symbolic mapping rule formalism and ruleset for mapping inputs to candidate outputs that achieves broad coverage through greater regularity, (2) a packed forest representation and efficient ranking algorithm that can manage the combinatorial growth in output candidates, and (3) an empirical evaluation of coverage, correctness, and the ability to handle underspecification. This evaluation is the first large-scale empirical evaluation of coverage and quality ever performed for sentence realization. The empirical evaluation is performed by automatically converting a set of 2400 hand-parsed sentences from the Penn Treebank corpus into system inputs, and then regenerating them using the system. The top-ranked output of the generator is compared to the original sentence. The results show better than 80% coverage of newspaper text and 94% precision (57% are exact matches) for almost fully-specified inputs, and the same coverage with 55% precision for minimally specified inputs.
This work deals with models used, or usable in the domain of Automatic Natural Language Processing, when one seeks a syntactic interpretation of a statement. This interpretation can be used as additional information for subsequent treatments, that can aim for instance at producing a semantic representation of the statement. It can also be used as a filter to select utterances belonging to a specific language, among several hypotheses, as done in Automatic Speech Recognition. As the syntactic interpretation of a statement is generally ambiguous with natural languages, the probabilisation of the space of syntactic trees can help in the analysis task: when several analyses are competing, one can then extract the most probable interpretation, or classify interpretations according to their probabilities. We are interested here in the probabilistic versions of Context-Free Grammars (PCFGs) and Substitution Tree Grammar (PTSGs). Syntactic treebanks, which as much as possible account for the language we wish to model, serve as the basis for defining the probabilistic parameters of such grammars. First, we exhibit in this thesis some drawbacks of the usual learning paradigms, due to the use of arbitrary heuristics (STSG DOP model), or to the use of learning criteria that consider these grammars as generative ones (creation of sentences from the grammar) rather than dedicated to analysis (creation of analyses from the sentence). In a second time, we propose new methods for training grammars, based on the traditional Maximum Entropy and Maximum Likelihood criteria. These criteria are instanciated so that they correspond to a syntactic analysis task rather than a language generation task. Specific training algorithms are necessary for their implementation, but traditional algorithms can cope with those models for the task of syntactic analysis. Lastly, we invest the problem of time complexity of syntactic analysis, which is a real issue for the effective use of PTSGs. We describe classes of PTSGs that allow the analysis of a sentence in polynomial complexity. We finally describe a method that enable the extraction of such a PTSG from the set of subtrees of a treebank. The PTSG produced by this method allows us to test our non-generative learning criterium on "realistic" data, and to give a statistical comparison between this criterium and the usual heuristic criterium in term of analysis performance.
Abstract. This paper explores the use of initial Stochastic Context-Free Grammars (SCFG) obtained from a treebank corpus for the learning of SCFG by means of estimation algorithms. A hybrid language model is defined as a combination of a word-based n-gram, which is used to capture the local relations between words, and a category-based SCFG with a word distribution into categories, which is defined to represent the long-term relations between these categories. Experiments on the UPenn Treebank corpus are reported. These experiments have been carried out in terms of the test set perplexity and the word error rate in a speech recognition experiment. 1
The paper presents a designed and implemented tool VisDic which implements the functionality of editing and viewing different lexical resources. It was developed in the Natural Language Processing Laboratory at the Faculty of Informatics, Masaryk University. The smart design and the accent on the usage of XML standards enable to manage lexical data from dictionaries to lexical databases, semantic networks, and complex ontologies. It also facilitates the connection to other linguistic tools such as corpus managers or morphological analyzers.
We explore learning prepositionalphrase attachment in Dutch, to use it as a filter in prosodic phrasing. From a syntactic treebank of spoken Dutch we extract instances of the attachment of prepositional phrases to either a governing verb or noun. Using cross-validated parameter and feature selection, we train two learning algorithms, IB1 and RIPPER, on making this distinction, based on unigram and bigram lexical features and a cooccurrence feature derived from WWW counts. We optimize the learning on noun attachment, since in a second stage we use the attachment decision for blocking the incorrect placement of phrase boundaries before prepositional phrases attached to the preceding noun. On noun attachment, IB1 attains an F-score of 82; RIPPER an F-score of 78. When used as a filter for prosodic phrasing, using attachment decisions from IB1 yields the best improvement on precision (by six points to 71) on phrase boundary placement.
Fiammetta Namer: Morphologic productivity, representativity and base complexity: the MoQuête system In this paper we propose to describe a corpus processor that can be used to prepare and create lexical databases for French, and whose exploitation is more particularly oriented towards morphological parsing. This processor, called MoQuête (Morphology & Queries), applies on a lexical database (LDB) whose realization is the results of the following steps: recover online corpora, tag them, lemmatize them, and submit them to derivational morphological parsing. Starting from the experiment described for dutch in (Krott & ah, 1999), we illustrate the MoQuête behavior with the presentation of a série of measures, performed from a 27 millions tokens LDB obtained from newspapers online corpora, and meant to evaluate the link between the quantitative morphologic productivity and the representativity of morphologic rules, in the frame of complex lexical units with a complex base.
Past research has found that individual differences in both attitudinal and situational variables may be associated with males’ likelihood of acquaintance rape (LAR). The present research was conducted to examine the predictive value of both attitudinal and situational factors on males’ likelihood of forcing a female acquaintance to have non-consensual sexual intercourse. In Study 1, male and female respondents (Rs) were presented with a scenario depicting a hypothetical sexual interaction between the respondent and a newly acquainted member of the opposite sex. As the encounter progressed from one sexual activity to the next, Rs made three ratings regarding their own and partner’s intent to engage in each successive activity. The scenario ended with the female refusing further activity and males’ affect ratings, adherence to attitudes conducive of rape, and LAR were measured. Males’ initial perceptions of female sexual intent (to later engage in sexual intercourse) best predicted LAR. Study 2 was conducted to examine the role of female sexual communication on perceptions of consent to sexual intercourse. Rs were presented with a scenario similar to that in Study 1, but at each stage were requested to rate the extent to which the female had consented to engage in each sexual activity. Males completed the same affect and attitudinal measures. The results of Study 2 again suggested that males’ initial perception of female consent to (later) engage in sexual intercourse best predicted LAR. The present research suggests that further investigation into the role of situational factors and males’ initial perception of sexual intent and consent in the aetiology of acquaintance rape is required.
Dealing with convergence in German speech islands in Russia, Brazil and the United states the article discusses the linguistic phenomena related to the notion of convergence from different vantage points including intralinguistic convergence (due to dialect-dialect contact), interlinguistic convergence (due to language-language contact), typological "convergence" (or intralinguistic change), pidginization, and cognitive processes of simplification. Most of the German speech islands are considered to be contracting - if not dying - varieties with respect to the reduction of their grammatical systems. Evidently, for a long time language contact (and sometimes variety contact) have severely increased. Linguistic norms have been weakened in terms of both norm certainty and norm loyalty thus giving way to processes similar to those common to pidgin languages. External induced changes are highly remarkable in all German speech islands. But the susceptibility for change and the ways of change are structured by systematical and typological constraints which probably turn out to be cognitive processes underlying quite "normal" linguistic change. This change is discussed as a subsequent process of "regularization" (of irregular forms), simplification (of rules) and loss of grammatical distinctions (and their compensation). The linguistic description of these interrelated processes is based on an integrated approach providing methodology from sociolinguistics, dialectology and research on language change, including the attempt to highlight the cognitive structures which furrow the line for internal simplifications under external pressure. Comparative speech island research seems to be a promising field of application for the description of the intermesh of these processes.
Abstract An elevated disgust sensitivity (DS) is considered to be a vulnerability factor for the development of a blood-injection-injury (BII) phobia. Within the present functional Magnetic Resonance Imaging (fMRI) study, 12 female BII phobics were scanned while viewing alternating blocks of 40 disgust-inducing, 40 fear-inducing, and 40 affectively neutral pictures. Each block lasted 60s and was repeated six times during the experiment. All scenes were phobia-irrelevant. Afterwards, the subjects gave affective ratings for the pictures and described their DS on a self-report measure for different areas (e.g., poor hygiene, unusual food, death/deformation). The responses were compared with those of 12 nonphobic females. The BII phobics showed a stronger occipital activation within the right cuneus and lingual gyrus during the first viewing of the disgusting pictures. Aside from this finding, which could be interpreted as reflecting increased attention, there was little evidence for a generally elevated DS in BII phobia. On the DS questionnaire, the patients had indicated a greater reactivity only for disorder-relevant contents (death/deformation). Further, both groups gave similar disgust ratings for the pictures and showed comparable brain-dynamic responses over all blocks of the disgust condition, which included the activation of both amygdalae and the left inferior frontal gyrus.
INTRODUCTION: Although evidence suggests that interpersonal psychotherapy may be an efficacious treatment for eating disorders, there is surprisingly little systematic knowledge about the interpersonal world of these patients. METHOD: SASB self-image ratings were used to explore interpersonal profiles in a large heterogeneous sample of eating disorders (N = 830), matched normal controls (N = 105) and a small group of controls with subclinical depression (N = 26). RESULTS: Eating disorder patients clearly presented with significantly more negative interpersonal profiles compared to controls. Within the eating disorder group, anorexics were characterized by high self-control, self-blame and self-attack. Patients with binge eating disorder expressed the least negative self-image, and were significantly more self-affirming than bulimics and less self-controlling than patients with atypical eating disorders. CONCLUSIONS: Eating disorder patients may have distinct interpersonal profiles that increase the risk of negative therapeutic reaction. Better knowledge of interpersonal processes in eating disorders may help to improve both diagnostic assessment and treatment.
We describe an algorithm for recovering non-local dependencies in syntactic dependency structures. The pattern-matching approach proposed by Johnson (2002) for a similar task for phrase structure trees is extended with machine learning techniques. The algorithm is essentially a classifier that predicts a non-local dependency given a connected fragment of a dependency structure and a set of structural features for this fragment. Evaluating the algorithm on the Penn Treebank shows an improvement of both precision and recall, compared to the results presented in (Johnson, 2002).
We investigate the performance of the Structured Language Model when one of its components is modeled by a connectionist model. Using a connectionist model and a distributed representation of the items in the history makes the component able to use much longer contexts than possible with currently used interpolated or backoff models, both because of the inherent capability of the connectionist model to fight the data sparseness problem, and because of the only sub-linear growth in the model size when increasing the context length. Experiments show significant improvement in perplexity and moderate reduction in word error rate over the baseline SLM results on the UPENN treebank and Wall Street Journal (WSJ) corpora respectively. The results also show that the probability distribution obtained by our model is much less correlated to regular N-grams than the baseline SLM model.
We have developed an example-based machine translation (EBMT) system that uses the World Wide Web for two different purposes: First, we populate the system's memory with translations gathered from rule-based MT systems located on the Web. The source strings input to these systems were extracted automatically from an extremely small subset of the rule types in the Penn-II Treebank. In subsequent stages, the source, target translation pairs obtained are automatically transformed into a series of resources that render the translation process more successful. Despite the fact that the output from on-line MT systems is often faulty, we demonstrate in a number of experiments that when used to seed the memories of an EBMT system, they can in fact prove useful in generating translations of high quality in a robust fashion. In addition, we demonstrate the relative gain of EBMT in comparison to on-line systems. Second, despite the perception that the documents available on the Web are of questionable quality, we demonstrate in contrast that such resources are extremely useful in automatically postediting translation candidates proposed by our system.
The primary purpose of this study was to assess the cross-cultural invariance of job performance ratings. A secondary purpose was to examine potential cross-cultural differences in correlates of performance ratings (i.e., ratee sex, age, tenure; supervisor's opportunity to observe ratee). Fast-food supervisors from Canada, South Korea, and Spain rated employees on their technical proficiency, customer service, and teamwork. Results show that these ratings demonstrate a basic level of measurement invariance, although the error variances of the ratings and pattern of construct variances and covariances were largely culture-specific. This suggests that supervisors across cultures may use and interpret the ratings similarly, but perceive differences in performance. Furthermore, age, tenure, and the supervisor's opportunity to observe the ratee were found to affect ratings differently across cultures. Overall, this study suggests that although job performance ratings are at least partially invariant across cultures, latent performance may not be, and we present some preliminary data as to why latent invariance may not exist.
BACKGROUND: Individual differences in neural circuitry that regulate emotional reactivity may be associated with alcoholism and antisocial personality disorder (ASPD), a common comorbid condition. The emotion-modulated startle reflex was used to investigate emotional reactivity among alcohol-dependent (AD) men with and without ASPD. METHODS: Sixty-two men were tested: (1) AD (n = 24), (2) AD-ASPD (n = 17), and (3) non-AD, non-ASPD controls (n = 21). Participants completed self-report instruments and clinical interviews and had eye-blink electromyograms measured in response to acoustic startle probes while viewing color photographs rated as affectively pleasant, neutral, and unpleasant. RESULTS: Startle blink magnitudes were larger during unpleasant as compared with pleasant slides for control and AD groups, resulting in significant linear trend effects (p < 0.001) and nonsignificant quadratic trend effects. In contrast, AD-ASPD did not show a significant difference in blink magnitude during unpleasant and pleasant slides and did not show a significant linear valence trend or quadratic trend effect (p > 0.6). Subjective valence and arousal ratings of the photographs were similar across groups. CONCLUSIONS: Adult male alcoholics with ASPD have abnormal emotional responsiveness to both pleasant and unpleasant stimuli relative to alcoholics without ASPD and to controls.
Treebanks are a valuable resource for the training of parsers that perform automatic annotation of unseen data. It has been shown that changes in the representation of linguistic annotation have an impact on the performance of a certain annotation task. We focus on the task of Topological Field Parsing for German using Probabilistic Context-Free Grammars in the present research. We investigate an iterative algorithm for tuning the label set of a given treebank to this task and show that the number of parses proposed by a context-free grammar is reduced considerably in addition to an increase in labeled precision and recall for the annotation of node labels. We also show that the optimal refinement can be achieved with a relatively small number of changes to the treebank. 1
Agents seeking to discover and compose needed Web services may face knowledge sharing interoperability problems due to differing ontologies. In practice, agents may not have a global consensus ontology that will facilitate knowledge sharing and integration of required services. We investigate a method for agents to develop local consensus ontologies to aid in the communication within a multi-agent system of business-tobusiness (B2B) agents. We compare variations of syntactic and semantic similarity matching to form local consensus ontologies with and without the use of a lexical database.
We use the grammatical relations (GRs) described in Carroll et al. (1998) to compare a number of parsing algorithms. A first ranking of the parsers is provided by comparing the extracted GRs to a gold standard GR annotation of 500 Susanne sentences: this required an implementation of GR extraction software for Penn Treebank style parsers. In addition, we perform an experiment using the extracted GRs as input to the Lappin and Leass (1994) anaphora resolution algorithm. This produces a second ranking of the parsers, and we investigate the number of errors that are caused by the incorrect 'GRs.
The article postulates a look at the problem of literary (general) language norms, considering the evolution of the linguistic system caused by the passage of time. In the diachronic view, the linguistic system of the language has a variant nature, wherein the level of intensification of variance is dependent on the depth of systemic structures. In diachronic research, one should accept the existence of a variant level of the superficial system and to a large degree the invariant level of the deep system. This assumption allows for discussion about an immanent norm, inherent to the ethnic, national language in an organic way, precisely connected with the tendency of the language to improve itself, and also about a textual norm being the reflection of the variant level of the superficial linguistic system. In the XVI century, in the Polish language, a reconstruction of the structure of the language as the result of extralinguistic factors such as print, in effect led to an intensification of normalization, particularly in printed texts. At this time, there came into existence so-called particular norms, i.e. a catalog of normative decisions characteristic for individual publishing houses, e.g. norms unique to the Publishing Houses of Hieronim Wietor and Florian Ungler.
We propose the use of Lexicalized Tree Adjoining Grammar (LTAG) as a source of features that are useful for reranking the output of a statistical parser. In this paper, we extend the notion of a tree kernel over arbitrary sub-trees of the parse to the derivation trees and derived trees provided by the LTAG formalism, and in addition, we extend the original definition of the tree kernel, making it more lexicalized and more compact. We use LTAG based features for the parse reranking task and obtain labeled recall and precision of 89.7%/90.0% on WSJ section 23 of Penn Treebank for sentences of length ≤ 100 words. Our results show that the use of LTAG based tree kernel gives rise to a 17% relative difference in f-score improvement over the use of a linear kernel without LTAG based features.
This article describes a corpus-based investigation of quantifier scope preferences. Following recent work on multimodular grammar frameworks in theoretical linguistics and a long history of combining multiple information sources in natural language processing, scope is treated as a distinct module of grammar from syntax. This module incorporates multiple sources of evidence regarding the most likely scope reading for a sentence and is entirely data-driven. The experiments discussed in this article evaluate the performance of our models in predicting the most likely scope reading for a particular sentence, using Penn Treebank data both with and without syntactic annotation. We wish to focus attention on the issue of determining scope preferences, which has largely been ignored in theoretical linguistics, and to explore different models of the interaction between syntax and quantifier scope.
The paper presents a maximum entropy Chinese character-based parser trained on the Chinese Treebank ("CTB" henceforth). Word-based parse trees in CTB are first converted into character-based trees, where word-level part-of-speech (POS) tags become constituent labels and character-level tags are derived from word-level POS tags. A maximum entropy parser is then trained on the character-based corpus. The parser does word-segmentation, POS-tagging and parsing in a unified framework. An average label F-measure 81.4% and word-segmentation F-measure 96.0% are achieved by the parser. Our results show that word-level POS tags can improve significantly word-segmentation, but higher-level syntactic strutures are of little use to word segmentation in the maximum entropy parser. A word-dictionary helps to improve both word-segmentation and parsing accuracy.
This paper presents a Java-based hyperbolic-style browser designed to render RDF files as structured ontological maps. The program was motivated by the need to browse the content of a web-accessible ontology server: WEB KB-2. The ontology server contains descriptions of over 74,500 object types derived from the WordNet 1.7 lexical database and can be accessed using RDF syntax. Such a structure creates complications for hyperbolic-style displays. In WEB KB-2 there are 140 stable ontology link types and a hyperbolic display needs to filter and iconify the view so different link relations can be distinguished in multi-link views. Our browsing tool, OntoRama, is therefore motivated by two possibly interfering aims: the first to display up to 10 times the number of nodes in a hyperbolic-style view than using a conventional graphics display; secondly, to render the ontology with multiple links comprehensible in that view.
Perceptions of social closeness and familiarity were assessed among 44 monozygotic (MZA) and 33 dizygotic (DZA) reunited twin pairs, and several individual twins and triplets. Significantly greater MZA than DZA closeness and familiarity were found. Closeness and familiarity ratings for co-twins exceeded those for nonbiological siblings with whom twins were raised. Correlations between perceptions of physical resemblance and social closeness and familiarity were positive and statistically significant. However, most correlations between social relatedness and contact time were non-significant. Associations between social relatedness and similarities in selected behavioral traits were also examined. The findings support various theoretical perspectives anticipating greater affiliation among close relatives than distant relatives.
In this paper we present a proposal to extend WordNet-like lexical databases by adding phrasets, i.e. sets of free combinations of words which are recurrently used to express a concept (let's call them recurrent free phrases). Phrasets are a useful source of information for different NLP tasks, and particularly in a multilingual environment to manage lexical gaps. Two experiments are presented to check the possibility of acquiring recurrent free phrases from dictionaries and corpora.
This paper investigates adapting a lexicalized probabilistic context-free grammar (PCFG) to a novel domain, using maximum a posteriori (MAP) estimation. The MAP framework is general enough to include some previous model adaptation approaches, such as corpus mixing in Gildea ( Other approaches falling within this framework are more effective. In contrast to the results in Gildea ( MAP adaptation can also be based on either supervised or unsupervised adaptation data. Even when no in-domain treebank is available, unsupervised techniques provide a substantial accuracy gain over unadapted grammars, as much as nearly 5% F-measure improvement.
We present a neural network method for inducing representations of parse histories and using these history representations to estimate the probabilities needed by a statistical left-corner parser. The resulting statistical parser achieves performance (89.1% F-measure) on the Penn Treebank which is only 0.6% below the best current parser for this task, despite using a smaller vocabulary size and less prior linguistic knowledge. Crucial to this success is the use of structurally determined soft biases in inducing the representation of the parse history, and no use of hard independence assumptions.