Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
This paper describes the grling-sdm system, which is asupervised probabilistic classifier that participated in the 1998SENSEVAL competition for word-sense disambiguation. This systemuses model search to select decomposable probability models describingthe dependencies among the feature variables.These types of models have been found to be advantageous in terms ofefficiency and representational power. Performance on the SENSEVALevaluation data is discussed.
In this article, I conduct a quantitative analysis of do absence in negative declaratives in the present tense in a dialect from the north-east of Scotland, Buckie. Analysis of nearly 800 contexts of use reveals that this variation is entirely conditioned by linguistic internal constraints. The most significant of these is person and number of the subject — 3rd person singular subjects and plural NPs have no do absence, while do is variable in the remaining pronouns. I argue that a syntactic explanation best accounts for this patterning of use. Where there is no overt - s inflection in the present tense (influenced by the “northern subject rule”), do is not obligatory in Buckie Scots. Frequency effects, lexical restrictions and processing constraints are called upon to account for the range of frequencies of do absence seen in the variable contexts. Lastly, there is no significant change in use of do across three generations of speakers, highlighting the community members’ relative immunity to prescriptive norms.
Spontaneous, conversational speech in probable dementia of Alzheimer type (DAT) participants and healthy older controls was analysed using eight linguistic measures. These were evaluated for their usefulness in discriminating between healthy and demented individuals. The measures were; noun rate, pronoun rate, verb rate, adjective rate, Clause-like Semantic Unit rate (all per 100 words), including three lexical richness measures; type token ratio (TTR), Brunt's Index (W) and Honor's statistic (R). Results suggest that these measures offer a sensitive method of assessing spontaneous speech output in DAT. Comparison between DAT and healthy older participants demonstrates that these measures discriminate well between these groups. This method shows promise as a diagnostic and prognostic tool, and as a measure for use in clinical trials. Further validation in a large sample of patient versus control "norms" in addition to evaluation in other types of dementia is considered.
An architecture for federating heterogeneousdictionary databases is described. It proposes acommon description language and query language toprovide for the exchange of information betweendatabases with different organizations, on differentplatforms and in different DBMSs. The common querylanguage has an SQL like structure. The first versionof the description language follows the TEI standardtag definitions for dictionaries with the expectationthat the description language will be expanded in thefuture. A practical implementation of the proposalsusing WWW technology for two multi-lingualdictionaries is described.
This paper describes the evaluation of a WSD method withinSENSEVAL. This method is based on Semantic Classification Trees (SCTs)and short context dependencies between nouns and verbs. The trainingprocedure creates a binary tree for each word to be disambiguated. SCTsare easy to implement and yield some promising results. The integrationof linguistic knowledge could lead to substantial improvement.
The Translational English Corpus (TEC) held at the Centre for Translation Studies at UMIST is a full-text, synchronic, general, monolingual, written, single, translational corpus of English. It is also a direct, multi-source-language, mono-translation-mode (written mode), mono-translation-method (human translation), largely into-mother-tongue, professional, published corpus. At the time of writing TEC represents four text categories: newspapers, biography, fiction, and inflight magazines. Hatim (1999) has argued that so far in translation studies an important distinction has been ignored. This distinction is between what is "in" and what is "of" the text. "In" refers to the language itself, which can be ana-lyzed through text analysis, while "of" refers to the text in its entirety, its overall effect in terms of ideology. In this paper, I will argue that TEC can indeed be a valuable, self-contained, single resource for studying precisely the "of" of translational language, that is, its ideological impact in the target language and culture. To this purpose, I will discuss some methodological issues as well as the methods for carrying out critical linguistic analy-sis. I will also suggest ways of investigating norms of lexical use in TEC and their possible ideological implications through the lexico-grammatical and collocational analysis of a set of key words relating to Europe in translated newspaper articles.
Language deviation is a kind of langUage form diverging from the language norm, In poetry, there are eight kinds of language deviation: lexical deviation.Phonological deviation. grammatical deviation. graphological deviation. semantic deviation.deviation of register. deviation of historical period and dialectal deviation. From the point of view of the aesthetic function, we analyze the language deviation of poetry in outer to appreciate the poems better.
A Classification Information Model is a pattern classification model.The model decides the proper class of an input instance by integrating individual decisions, each of which is made with each feature in the pattern.Each individual decision is weighted according to the distributional property of the feature deriving the decision. An individual decision and its weight are represented as classification information which is extracted from the training instances.In the word sense disambiguation based on the model, the proper sense of an input instance is determined by the weighted sum of whole individual decisions derived from the features contained in the instance.
The effects of lexical difficulty and talker variability on word recognition were examined in four groups of listeners: native English/normal hearing; native English/hearing impaired; non-native English/normal hearing; and non-native English/hearing impaired (hearing level matched to the native hearing impaired). Lexical difficulty was measured by the difference in performance to 75 lexically ‘‘easy’’ and ‘‘hard’’ words based on word frequency and Neighborhood Activation Theory [Luce and Pisoni (1998)]. The effect of talker variability was measured by the difference in performance between single and multiple talker (nine talkers) conditions. The familiarity of the 150 words was rated on a seven-point scale. An up–down adaptive procedure was used to determine the sound pressure level for 50% performance. Non-native listeners in both normal and hearing-impaired groups required a greater intensity for equal intelligibility than for the comparative native normal and hearing-impaired listeners. Results, however, showed significant effects of lexical difficulty and talker variability in all four groups. Structural equation modeling demonstrated that an auditory factor estimated by pure tone average, etc., accounts for four times more variance to performance than does a linguistic fluency factor measured by word familiarity ratings and native versus non-native status, however, the linguistic fluency factor is also essential to the model fit.
Age of acquisition (AoA) has been reported to be a predictor of the speed of reading words aloud (word naming) and lexical decision, with early-acquired words being responded to faster than later-acquired words in both tasks. All previous studies of AoA effects have, however, relied upon adult estimates of word learning age the validity of which it is easy to cast doubt upon. Using objective age of acquisition norms derived from children's naming data, this study shows that AoA effects do not depend upon the use of adult ratings. In addition to effects of real AoA, influences of word frequency and orthographic neighbourhood size were obtained in both word naming and lexical decision. Imageability affected lexical decision but not word naming, while the characteristics of the word's initial phoneme affected word naming but not lexical decision.
SENSEVAL set itself the task of evaluating automaticword sense disambiguation programs (see Kilgarriff andRosenzweig, this volume, for an overview of theframework and results). In order to do this, it wasnecessary to provide a `gold standard' dataset of `correct' answers. This paper will describe thelexicographic part of the process involved in creatingthat dataset. The primary objective was for a group oflexicographers to manually examine keywords in a largenumber of corpus contexts, and assign to each contexta sense-tag for the keyword, taken from the Hectordictionary. Corpus contexts also had to be manuallypart-of-speech (POS) tagged. Various observationsmade and insights gained by the lexicographers duringthis process will be presented, including a critiqueof the resources and the methodology.
1. Introduction Not every historical linguist embraces the idea of Chomsky's syntactocentrism with enthusiasm. It may be untimely to say unkind things about it, but there are syntactic problems which cannot be resolved satisfactorily only by formal operations. Under the current psycholinguistic views there seem to be some chances of recognizing the old conceptual world of the speaker and thus contributing to a more appropriate understanding of the writings he has left. Following chiefly Jackendoff's ideas expressed in The architecture of the language faculty (1997) -- yet with due respect for other linguistic and psycholinguistic orientations -- I will discuss grammatical relations which involve word order, thematic roles and word-formation (compounding) and which by structural standards prove so intractable. A common trait of them all is that they are structurally ambiguous and consequently differ in meanning, or that they are simply semantically opaque. 2. Word order An example of how weakly significant word order in Old English can be is the first part of the following sentence: 1) Storm oft holm gebringep, geofen in grimmum selum (Maxims I 112/50) which has been understood as either 'The sea often brings a storm, the ocean in stormy seasons' (Gordon 1954: 342) 'The sea often brings a storm' (Bosworth, entry gebringan) or 'The often brings forth a flood' (Reszkiewicz 1971:35) 'storm oft brings ocean into a furious condition' (Bosworth, entry soel) The interpretative difficulty lies in the fact that the functions of a grammatical subject and a grammatical object are not clearly transparent: the nouns and holm are both singular and each can agree with the finite form of the verb, gebringep, which as a two (or even three) argument requires a subject and an object. This brings up a question: which is which? Structurally speaking each can perform either function. They are both masculine, singular, of a-inflection of which nominative/accusative syncretism is a norm. Besides, there is no adjectival or pronominal modifier to help, neither can alliteration be helpful. Reszkiewicz searched for a clue to the functional identification in the position of the noun with regard to the and came to the conclusion that: Older Old English, especially poetry, lacked both the definite and the indefinite articles; the object often preceded the governing verb (Reszkiewicz 1971: 35). Although the grounds on which such a decision is reached are formally defens ible, empirically they are less so as they can be falsified by a sentence, also a gnomic verse, which reads: (2) Moegen mon sceal mid mete fedan (Maxims I 118/44) in which it is the subject man and not moegen which is closer to the finite form of the verb, sceal (moegen and man also show inflectional syncretism in this respect); this sententious saying means: 'One shall nourish strength with meat' (food) (Gordon 1954: 344) 'A man must feed strength with meat' (Bosworth, entry fedan) The proponents of either of the two meanings of the gnomic storm verse would probably try to persuade us that their views are compatible with the formal grammatical relations. But which of the meanings would satisfy the pragmatics of the discourse? Although the senses of particular lexical items are clear, a real cognitive image is still concealed. As a historical linguist I am more comfortable asking questions than answering them, so my glimpse into the Old English cognitive mind will be based on the possible, we now try to see, life as it would have been over a millenium of years ago. Since the conceptual structure of our example is not immediately predictable from the syntactic structure, nor is it found in the lexical structures, I will try to consider the language context first and then to search for similar uses of and holm. …
In this paper we present some observations concerning an experiment of (manual/automatic) semantic tagging of a small Italian corpus performed within the framework of the SENSEVAL/ROMANSEVAL initiative. Themain goal of the initiative was to set up a framework for evaluation of Word Sense Disambiguation systems (WSDS) through the comparative analysis of their performance on the same type of data. In this experiment there are two aspects which are of relevance: first, the preparation of the reference annotated corpus, and, second, the evaluation of the systems against it. In both aspects we are mainly interested here in the analysis of the linguistic side which can lead to a better understanding of the problem of semantic annotation of a corpus, be itmanual or automatic annotation. In particular, we will investigate, firstly, the reasons for disagreement between human annotators, secondly, some linguistically relevant aspects of the performance of the Italian WSDS and, finally, the lessons learned from the present experiment.
Pastiche is central to the resistant politics of Kathy Acker's writing--yet she would appear to agree with Fredric Jameson's influential critique of pastiche as "the wearing of a linguistic mask, speech in a dead language" (17). Her 1986 novel Don Quixote is all about having to speak "in a dead language" in the absence of a more "healthy" norm. It begins with the death of the protagonist, a female version of Cervantes's knight, who then goes on to narrate much of the subsequent story. Acker explains, "BEING DEAD, DON QUIXOTE COULD NO LONGER SPEAK. BEING BORN INTO AND PART OF A MALE WORLD, SHE HAD NO SPEECH OF HER OWN. ALL SHE COULD DO WAS READ MALE TEXTS WHICH WEREN'T HERS" (39). The novel then proceeds by plagiarism and pastiche, as Quixote goes on a quest--for a heterosexual love unsullied by patriarchal power relations--through fragments of numerous existing texts. Quixote rereads and pieces together a whole range of textual scraps, from Machiavelli's The Prince to a Godzilla movie. What becomes clear in her eccentric survey of (primarily) Western culture is that the lost, healthy linguistic norm is more than unhealthy for female readers--indeed, it is deadly.
Computer-driven systems for constructing composite faces of suspects (E-fit; Mac-a-Mug) have largely replaced mechanical systems (Photofit; the Identikit) in police use, yet little is known of their comparative effectiveness in rendering an accurate likeness. Participants (N = 24) constructed 2 of 4 familiar or unfamiliar faces, for one of which they used Photofit and for the other, E-fit. A likeness of each face was made first under target-absent conditions and then with photographs of the target present. The accuracy of the resulting composites was assessed by familiarity ratings, names elicited, and matching accuracy. The computer-driven system showed consistent superiority only when a familiar face was constructed in the presence of photographs; when participants worked from memory, E-fit was no better than Photofit. The implications of these findings for theories of face retrieval and the operational use of composites are discussed.
The present study used the picture perception paradigm to examine the extent to which three well-documented psychophysiological measures demonstrate consistency across time in response to emotional stimuli. The three measures were the eye-blink startle response and the activation in two facial muscle regions (zygomatic and corrugator). Twenty-seven young women were assessed on two occasions, 2 weeks apart. Whereas activation in the corrugator and zygomatic muscle regions demonstrated the predicted patterns at both assessments (with some attenuation in the zygomatic muscle regions), the startle response had limited consistency across the two assessments. The startle response revealed the predicted linear pattern of valence modulation during the first assessment. During the second assessment, startle magnitude response was a quadratic function of valence ratings and a linear function of arousal ratings. The unexpected pattern of startle response during the second session appeared to be related to the content of the pleasant slides, with action slides generating quadratic valence modulation and erotic slides continuing to exhibit the expected linear valence modulation.
We present a method for automatically detecting errors in a manually marked corpus using anomaly detection. Anomaly detection is a method for determining which elements of a large data set do not conform to the whole. This method fits a probability distribution over the data and applies a statistical test to detect anomalous elements. In the corpus error detection problem, anomalous elements are typically marking errors. We present the results of applying this method to the tagged portion of the Penn Treebank corpus.
The present study investigated the relationship between daily diary affect ratings and ambulatory cardiovascular activity in 117 male Vietnam combat veterans (61 with posttraumatic stress disorder [PTSD] and 56 without PTSD). Participants completed 12-14 hr of ambulatory monitoring and daily diary affect ratings. Compared with veterans without PTSD, veterans with PTSD reported higher negative affect and lower positive affect in daily diary ratings. No differences were detected for mean laboratory initial recordings or mean ambulatory heart rate (HR), systolic blood pressure (SBP), or diastolic blood pressure (DBP). However, compared with veterans without PTSD, veterans with PTSD demonstrated higher SBP and DBP variability and a higher proportion of HR activity (compared with initial recording values) during daily activity. There was a significant Time of Day x Group interaction for mean HR, with a trend for PTSD participants to maintain HR levels during evening hours.
This paper proposes a new error-driven HMM-based text chunk tagger with context-dependent lexicon. Compared with standard HMM-based tagger, this tagger uses a new Hidden Markov Modelling approach which incorporates more contextual information into a lexical entry. Moreover, an error-driven learning approach is adopted to decrease the memory requirement by keeping only positive lexical entries and makes it possible to further incorporate more context-dependent lexical entries. Experiments show that this technique achieves overall precision and recall rates of 93.40% and 93.95% for all chunk types, 93.60% and 94.64% for noun phrases, and 94.64% and 94.75% for verb phrases when trained on PENN WSJ TreeBank section 00-19 and tested on section 20-24, while 25-fold validation experiments of PENN WSJ TreeBank show overall precision and recall rates of 96.40% and 96.47% for all chunk types, 96.49% and 96.99% for noun phrases, and 97.13% and 97.36% for verb phrases.
In this paper, we present a neural-networks-based knowledge discovery and data mining (KDDM) methodology based on granular computing, neural computing, fuzzy computing, linguistic computing, and pattern recognition. The major issues include 1) how to make neural networks process both numerical and linguistic data in a data base, 2) how to convert fuzzy linguistic data into related numerical features, 3) how to use neural networks to do numerical-linguistic data fusion, 4) how to use neural networks to discover granular knowledge from numerical-linguistic data bases, and 5) how to use discovered granular knowledge to predict missing data. In order to answer the above concerns, a granular neural network (GNN) is designed to deal with numerical-linguistic data fusion and granular knowledge discovery in numerical-linguistic databases. From a data granulation point of view, the GNN can process granular data in a database. From a data fusion point of view, the GNN makes decisions based on different kinds of granular data. From a KDDM point of view, the GNN is able to learn internal granular relations between numerical-linguistic inputs and outputs, and predict new relations in a database. The GNN is also capable of greatly compressing low-level granular data to high-level granular knowledge with some compression error and a data compression rate. To do KDDM in huge data bases, parallel GNN and distributed GNN will be investigated in the future.
We present some novel machine learning techniques for the identification of subcategorization information for verbs in Czech. We compare three different statistical techniques applied to this problem. We show how the learning algorithm can be used to discover previously unknown subcategorization frames from the Czech Prague Dependency Treebank. The algorithm can then be used to label dependents of a verb in the Czech treebank as either arguments or adjuncts. Using our techniques, we are able to achieve 88% precision on unseen parsed text.
BOOK NOTICES 209 Linguistic databases. Ed. by John Nerbonne. (CSLI lecture notes 77.) Stanford, CA: CSLI, 1998. Pp. xxi, 243. The papers in this collection were originally presented at the 'Linguistic Databases' conference, University of Groningen, 23-24 March, 1995. Because ofthe almostproverbial rapidity with which information technology develops, the collection as a whole is dated already, but there is still much of interest to be found. Not all papers read at the conference are in this volume, but the papers cover a wide range of subjects, mostly practical in nature, not theoretical. After a clear and readable introduction by Nerbonne, the papers are presented in no particularorder, though the editor groups the papers in five main areas: syntactic corpora and databases, phonetic databases, applications in linguistic theory, applications, and extending basic technologies. The papers themselves are not presented according to this grouping, however, and at first sight the book appears rather disorganized. The wide variety of subjects can be deduced from the titles of the papers presented: 'Test suites for natural language processing', 'From annotated corpora to databases: The SgmlQL language', 'Markup of a test suite with SGML', 'An open systems approach for an acoustic-phonetic continuous speech database: The S_tools database-management system ', "The reading database of syllable structure',? database application for the generation of phonetic atlas maps', 'Swiss French polyphone and polyvar: Telephone speech databases to model inter- and intra-speaker variability', 'Investigating argument structure: The Russian nominalization database', "The use of a psycholinguistic database in the simplification of text for aphasie readers', "The computer learner corpus: A testbed for electronic EFL tools', 'Linking WordNet to a corpus query system', 'Multilingual data processing in the CELLAR environment '. The issue whether to use open free systems or closed proprietary systems is addressed in several papers. Some papers present applications developed both in open and closed systems. This is one area where developments have been going very fast, and nowadays freely available databases are often as capable as their commercial counterparts. Some of the applications presented in this collection are available from the Internet, and url's are often given. The collection can serve as a good introduction to the field for relative outsiders as ample references and links are given. The papers themselves vary greatly in subject matter so not all will be of interest to every reader. My particular favorite was 'From annotated corpora to databases: the SgmlQL language '. [BOUDEWUN REMPT.j Understanding phonology. By Carlos Gussenhoven and Haike Jacobs. (Understanding language series.) London: Arnold, 1998. Pp. xii, 286. This textbook is intended as an introduction to phonology aimed at 'students with little or no prior knowledge of linguistics' (back cover). As in many other textbooks, it uses exercises as a learning tool. Two types ofexercises are proposed. The ones identified by a key, 'intended as an expository aid' (xi), are provided with a solution in an appendix (though it is not always so much a clear cut answer as a guide for reflection, which is, to my view, a lot better). The ones identified by a dot are intended as practice material, and no solution is offered. I thought the idea of having two types of exercises a good one since it gives the reader the opportunity both for individual work and for discussion with others. Also, whenever it may apply, an optimality theoretic analysis is offered to describe a phonological process. Ch. 1, "The production of speech', is a basic introduction to phonology, phonetics, and phonation. Ch. 2,'Some typology: Sameness and difference', cleverly covers the universal and language specific aspects of phonological structures and typology. Ch. 3,'Making the form fit', addresses phonological grammar and adaptation by presenting the nativization of loan words in both the rules and the constraints approaches. Ch. 4, 'Underlying and surface representations', Ch. 5, 'Distinctive features', and Ch. 6, 'Ordered rules', deal with the basic notions of generative phonology within the SPE type formalism and introduce the reader to the school of linear phonology. Ch. 7,? case study: The diminutive suffix in Dutch', shows how these notions are applied. In Ch. 8, 'Levels of representation', Gussenhoven and Jacobs present an intermediate level of representation between the underlying representation and...
This paper describes the methodology that is being used to augment the Penn Treebank annotation with sense tags and other types of semantic information. Inspired by the results of SENSEVAL, and the high inter-annotator agreement that was achieved there, similar methods were used for a pilot study of 5000 words of running text from the Penn Treebank. Using the same techniques of allowing the annotators to discuss difficult tagging cases and to revise WordNet entries if necessary, comparable inter-annotator rates have been achieved. The criteria for determining appropriate revisions and ensuring clear sense distinctions are described. We are also using hand correction of automatic predicate argument structure information to provide additional thematic role labeling. 1.
Structure preserving grammar compaction (SPC) is a simple CFG compaction technique originally described in (van Genabith et al., 1999a, 1999b). It works by generalising category labels and in so doing plugs holes in the grammar. To date the method has been tested on small corpra only. In the present research we apply SPC to a large grammar extracted from the Penn Treebank and examine its effects on rule treebank grammar size and on rule accession rates (as an indicator of grammar completeness). 1 Introduction Tree banks and resources compiled from treebanks are potentially very useful in NLP. Grammars extracted from treebanks --- so called treebank grammars (Charniak, 1996) --- can form the basis of large coverage NLP systems. Such treebank grammars, however, can suffer from several shortcomings: they commonly feature a large number of flat, highly specific rules that may be rarely used, with ensuing costs for processing (load) under the grammar. Furthermore, the observ...
We present an implementation of a chart-based head-corner parsing algorithm for lexicalized Tree Adjoining Grammars. We report on some practical experiments where we parse 2250 sentences from the Wall Street Journal using this parser. In these experiments the parser is run without any statistical pruning; it produces all valid parses for each sentence in the form of a shared derivation forest. The parser uses a large Treebank Grammar with 6789 tree templates with about 120# 000 lexicalized trees. The results suggest that the observed complexity of parsing for LTAG is dominated by factors other than sentence length. 1. Motivation The particular experiments that we report on in this paper were chosen to discover certain facts about LTAG parsing in a practical setting. Specifically, we wanted to discover the importance of the worst-case results for LTAG parsing in practice. Let us take Schabes' Earleystyle TAG parsing algorithm (Schabes, 1994) which is the usual candidate for a practica...
Burnout was tested for in 754 mental health workers and related to self-image as assessed with Structural Analysis of Social Behavior (SASB, Benjamin 1974). A positive relation was found between burnout and negative self-image, and between the experience of personal accomplishment and positive self-image. Compared to self-image, gender, age and work setting did not explain any variance in burnout. Highly burned-out persons had a significantly more negative self-image than staff who had rated themselves as low burnout. Finally, the relation between self-image and burnout was studied in 210 subjects who had completed their self-image ratings one year before burnout was measured, with the same results: a negative self-image was related to higher burnout one year later. One general conclusion is that a tendency in staff to treat themselves in negative ways may function as a negative filter for coping with difficulties at work and thus be a risk factor for burnout.
REVIEWS Get access Christiane Fellbaum (ed.) WordNet: An Electronic Lexical Database. Cambridge, Mass.: MIT Press. 1998. xxii + 423 pages. ISBN 0-262-06197-X. $50. Geoffrey Sampson Geoffrey Sampson University of Sussex Search for other works by this author on: Oxford Academic Google Scholar International Journal of Lexicography, Volume 13, Issue 1, March 2000, Pages 54–59, https://doi.org/10.1093/ijl/13.1.54 Published: 01 March 2000
Traditional theories of finance posit that the pricing of securities in financial markets should be done according to the quality of their underlying technical fundamentals. However, research on financial markets has tended to indicate that factors other than technical fundamentals are often used by market participants to gauge the value of securities. This phenomenon may be quite prevalent in markets for initial public offerings (IPSs), where securities lack a financial history. The imagery and affect associated with securities can be a powerful basis upon which to judge their worth. Advanced business students in a securities analysis course were asked to evaluate a number of industry groups represented on the New York Stock Exchange in terms of a set of judgmental variables. After providing imagery and affective evaluations for each industry group, the participants judged the likelihood that they would invest in companies associated with each industry. Imagery and affective ratings were highly correlated with one another and with the likelihood of investing. Judgments of performance correlated poorly to moderately with actual market performance as measured by weighted average returns for the industry groups studied. The results suggest that imagery and affect are part of a coherent psychological framework for evaluating classes of securities, but that framework may have low validity for predicting performance.
This paper describes a hybrid proposal to combine n-grams and Stochastic Context-Free Grammars (SCFGs) for language modeling. A classical n-gram model is used to capture the local relations between words, while a stochastic grammatical model is considered to represent the long-term relations between syntactical structures. In order to define this grammatical model, which will be used on large-vocabulary complex tasks, a category-based SCFG and a probabilistic model of word distribution in the categories have been proposed. Methods for learning these stochastic models for complex tasks are described, and algorithms for computing the word transition probabilities are also presented. Finally, experiments using the Penn Treebank corpus improved by 30% the test set perplexity with regard to the classical n-gram models.
Data oriented parsing systems employ redundant stochastic tree substitution grammars (STSGs) to analyse natural language utterances on the basis of an annotated corpus (a treebank). An important component of such systems is the way in which the substitution probability of a parse tree fragment is estimated from its occurrences in the treebank. In the standard method for doing this, the probability of a fragment is directly correlated with its occurrence frequency in the collection of all fragments of all corpus trees. We show that this results in undesirable statistical biases. We therefore propose an alternative method, which estimates the substitution probability of a fragment as the probability that it has been involved in the derivation of a corpus tree. We show that this method has more plausible properties.
A BILINGUAL LEXICAL DATABASE FOR FRAME SEMANTICS Get access Thierry Fontenelle Thierry Fontenelle 19 Rue du Merschgrund (L-8373 Hobscheid, Luxembourg)University of Liège(B-4000 Liege, Belgium) (fontenel@pt_lu) Search for other works by this author on: Oxford Academic Google Scholar International Journal of Lexicography, Volume 13, Issue 4, December 2000, Pages 232–248, https://doi.org/10.1093/ijl/13.4.232 Published: 01 December 2000
I have been developing a computer program named RebLin. RebLin has special functions other than searching electronic dictionaries. In this paper, I will show how multi-functional RebLin expands the possibility of electronic dictionaries. First, I will compare RebLin with other computer programs which handle electronic dictionaries, analyzing how each of the software deal with inflected forms and derivatives. Then, I will explain how RebLin retrieves useful information out of other electronic resources such as those on the internet, digitized movies and a lexical database. Finally, I will analyze a log file which records words which users input in the search field of RebLin, and show how they make use of the various electronic resources which RebLin provides them with.
1.1 Notion of word....................................... 4 1.2 Tests of wordhood..................................... 5 1.3 Compatibility with other guidelines............................ 6
This dissertation describes a natural language processing research in the field of nominal compounds in general and technical English. The starting point for the studies presented was INTEX, a tool for automatic treatment of large corpora.<br />While analyzing the problem of large coverage listing and describing of compounds, we addressed the following issues:<br />1) Which methods of compound description should be used?<br />2) For what kind of applications is this description useful?<br />The first issue is treated in the context of electronic lexical databases such as they are admitted in the INTEX system. We analyze the inflectional morphology of compounds in French, English and Polish. We propose a method of automatic generation of their inflected forms. We describe the construction of two electronic dictionaries: one for general English compounds, and the other for simple and compound terms of the computer science technical English. We also present a library of finite-state automata and transducers for the recognition of English cardinal and ordinal numerals.<br />The utility of large coverage compound dictionaries is verified through their application to two kinds of natural language processing tasks. First, we describe a method of acquisition of terms based on initial terminological resources. Secondly, we propose an automatic spelling checking algorithm of simple and compound words in a finite-state automaton dictionary.
In this paper we introduce an example-based parser for Chinese. One strong point of the parsers is its high reliability. We propose a formal definition for reliability and derive from it K as a metric for the evaluation of parsers. In a row of experiments we try to identify some factors which support the reliability of the parser. It is suggested that these factors are independent of the parsing approach and can be realized in TAGs. 1. Introduction Example-based parsers adhere to the lazy learning algorithm while converting tree-bank entries into a parser. So-called treebank grammars, (Bod, 1992; Charniak, 1996) are eager learners, i.e. they abstract knowledge structures or statistical information from the treebank and reason on the basis of these abstractions. Explanation-based parsing is a different eager learning approach aiming at the extraction of specialized grammars out of a general-purpose grammars on the bases of parsing examples (Rayner &amp; Christer, 1994; Srivinas &amp; Joshi, 1...
The availability of semantically tagged corpora is becoming a very important and urgent need for training and evaluation within a large number of applications but also they are the natural application and accompaniment of semantic lexicons of which they constitute both a useful testbed to evaluate their adequacy and a repository of corpus examples for the attested senses. It is therefore essential that sound criteria are defined for their construction and a specific methodology is set up for the treatment of various semantic phenomena relevant to this level of description. In this paper we present some observations and results concerning an experiment of manual lexical-semantic tagging of a small Italian corpus performed within the framework of the ELSNET project. The ELSNET experimental project has to be considered as a feasibility study. It is part of a preparatory and training phase, started with the Romanseval/Senseval experiment (Calzolari et al., 1998), and ending up with the lexical-semantic annotation of larger quantities of semantically annotated texts such as the syntactic-semantic Treebank which is going to be annotated within an Italian National Project (SI-TAL). Indeed, the results of the ELSNET experiment have been of utmost importance for the definition of the technical guidelines for the lexical-semantic level of description of the Treebank.
The value of language resources is greatly enhanced if they share a common markup with an explicit minimal semantics. Achieving this goal for lexical databases is difficult, as large-scale resources can realistically only be obtained by up-translation from pre-existing dictionaries, each with its own proprietary structure. This paper describes the approach we have taken in the Concede project, which aims to develop compatible lexical databases for six Central and Eastern European languages. Starting with sample entries from original presentation-oriented electronic representations of dictionaries, we transformed the data into an intermediate TEI-compatible representation to provide a common baseline for evaluating and comparing the dictionaries. We then developed a more restrictive encoding, formalised as an XML DTD with a clearly-defined semantic interpretation. We present this DTD and discuss a sample conversion from TEI, together with an application which hyperlinks a HTML represent...
Structural ambiguity, particularly attachment of prepositional phrases, is a serious type of global ambiguity in Natural Language. The disambiguation becomes crucial when a syntactic analyzer must make the correct decision among at least two equally grammatical parse-trees for the same sentence. This paper attempts to find answers to the problem of how attachment ambiguity can be resolved by utilizing Machine Learning (ML) techniques. ML is founded on the assumption that the performance in cognitive tasks is based on the similarity of new situations (testing) to stored representations of earlier experiences (training). Therefore, a large amount of training data is an important prerequisite for providing a solution to the problem. A combination of unsupervised and restricted supervised acquisition of such data will be reported. Training is performed both on a subset of the content of the Gothenburg Lexical Database (GLDB), and on instances of large corpora annotated with coarse-grained semantic information. Testing is performed on corpora instances using a range of different algorithms and metrics. The application language is written Swedish.
`Linguistic annotation' covers any descriptive or analytic notations applied to raw language data. The basic data may be in the form of time functions - audio, video and/or physiological recordings - or it may be textual. The added notations may include transcriptions of all sorts (from phonetic features to discourse structures), part-of-speech and sense tagging, syntactic analysis, `named entity' identification, co-reference annotation, and so on. While there are several ongoing efforts to provide formats and tools for such annotations and to publish annotated linguistic databases, the lack of widely accepted standards is becoming a critical problem. Proposed standards, to the extent they exist, have focused on file formats. This paper focuses instead on the logical structure of linguistic annotations. We survey a wide variety of existing annotation formats and demonstrate a common conceptual core, the annotation graph. This provides a formal framework for constructing, maintaining and searching linguistic annotations, while remaining consistent with many alternative data structures and file formats.
Grammars are core elements of many NLP applications. Grammars can be developed in two ways: built by hand or extracted from corpora. In this paper, we compare a handcrafted grammar with a Treebank grammar. We contend that recognizing substructures of the grammars&apos; basic units is necessary not only because it allows grammars to be compared at a higher level, but also because it provides the building blocks for consistent and efficient integration of the grammars.
This paper presents a novel methodology of disambiguating prepositional phrase attachments. We create patterns of attachments by classifying a collection of prepositional relations derived from Treebank parses. As a by-product, the arguments of every prepositional relation are semantically disambiguated. Attachment decisions are generated as the result of a learning process, that builds upon some of the most popular current statistical and machine learning techniques. We have tested this methodology on (1) Wall Street Journal articles, (2) textual definitions of concepts from a dictionary and (3) an ad hoc corpus of Web documents, used for conceptual indexing and information extraction.
Computers are now widely used in the preparation of dictionaries. There are many advantages in maintaining and updating a dictionary in electronic form, most obviously that printed versions can be typeset directly from the electronic copy. But more than that, electronic dictionaries are beginning to be used by computers in retrieval systems. This chapter looks at electronic dictionaries and examines how lexical databases can help to refine and improve retrieval and analysis programs. It also traces the development of the uses of computers and dictionaries, and assesses various types of resources. Much research still needs to be done on the structure and contents of lexical and linguistic databases, especially for the semantic component, but the examples discussed in this chapter give some idea of the potential.
This paper describes the design criteria and annotation guidelines of Sinica Treebank. The three design criteria are: Maximal Resource Sharing, Minimal Structural Complexity, and Optimal Semantic Information. One of the important design decisions following these criteria is the encoding of thematic role information. An on-line interface facilitating empirical studies of Chinese phrase structure is also described.