Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
1.1 Notion of word....................................... 4 1.2 Tests of wordhood..................................... 5 1.3 Compatibility with other guidelines............................ 6
In this paper we introduce an example-based parser for Chinese. One strong point of the parsers is its high reliability. We propose a formal definition for reliability and derive from it K as a metric for the evaluation of parsers. In a row of experiments we try to identify some factors which support the reliability of the parser. It is suggested that these factors are independent of the parsing approach and can be realized in TAGs. 1. Introduction Example-based parsers adhere to the lazy learning algorithm while converting tree-bank entries into a parser. So-called treebank grammars, (Bod, 1992; Charniak, 1996) are eager learners, i.e. they abstract knowledge structures or statistical information from the treebank and reason on the basis of these abstractions. Explanation-based parsing is a different eager learning approach aiming at the extraction of specialized grammars out of a general-purpose grammars on the bases of parsing examples (Rayner & Christer, 1994; Srivinas & Joshi, 1...
This paper describes a hybrid proposal to combine n-grams and Stochastic Context-Free Grammars (SCFGs) for language modeling. A classical n-gram model is used to capture the local relations between words, while a stochastic grammatical model is considered to represent the long-term relations between syntactical structures. In order to define this grammatical model, which will be used on large-vocabulary complex tasks, a category-based SCFG and a probabilistic model of word distribution in the categories have been proposed. Methods for learning these stochastic models for complex tasks are described, and algorithms for computing the word transition probabilities are also presented. Finally, experiments using the Penn Treebank corpus improved by 30% the test set perplexity with regard to the classical n-gram models.
Data oriented parsing systems employ redundant stochastic tree substitution grammars (STSGs) to analyse natural language utterances on the basis of an annotated corpus (a treebank). An important component of such systems is the way in which the substitution probability of a parse tree fragment is estimated from its occurrences in the treebank. In the standard method for doing this, the probability of a fragment is directly correlated with its occurrence frequency in the collection of all fragments of all corpus trees. We show that this results in undesirable statistical biases. We therefore propose an alternative method, which estimates the substitution probability of a fragment as the probability that it has been involved in the derivation of a corpus tree. We show that this method has more plausible properties.
This dissertation describes a natural language processing research in the field of nominal compounds in general and technical English. The starting point for the studies presented was INTEX, a tool for automatic treatment of large corpora.<br />While analyzing the problem of large coverage listing and describing of compounds, we addressed the following issues:<br />1) Which methods of compound description should be used?<br />2) For what kind of applications is this description useful?<br />The first issue is treated in the context of electronic lexical databases such as they are admitted in the INTEX system. We analyze the inflectional morphology of compounds in French, English and Polish. We propose a method of automatic generation of their inflected forms. We describe the construction of two electronic dictionaries: one for general English compounds, and the other for simple and compound terms of the computer science technical English. We also present a library of finite-state automata and transducers for the recognition of English cardinal and ordinal numerals.<br />The utility of large coverage compound dictionaries is verified through their application to two kinds of natural language processing tasks. First, we describe a method of acquisition of terms based on initial terminological resources. Secondly, we propose an automatic spelling checking algorithm of simple and compound words in a finite-state automaton dictionary.
The value of language resources is greatly enhanced if they share a common markup with an explicit minimal semantics. Achieving this goal for lexical databases is difficult, as large-scale resources can realistically only be obtained by up-translation from pre-existing dictionaries, each with its own proprietary structure. This paper describes the approach we have taken in the Concede project, which aims to develop compatible lexical databases for six Central and Eastern European languages. Starting with sample entries from original presentation-oriented electronic representations of dictionaries, we transformed the data into an intermediate TEI-compatible representation to provide a common baseline for evaluating and comparing the dictionaries. We then developed a more restrictive encoding, formalised as an XML DTD with a clearly-defined semantic interpretation. We present this DTD and discuss a sample conversion from TEI, together with an application which hyperlinks a HTML represent...
We present some novel machine learning techniques for the identification of subcategorization information for verbs in Czech. We compare three different statistical techniques applied to this problem. We show how the learning algorithm can be used to discover previously unknown subcategorization frames from the Czech Prague Dependency Treebank. The algorithm can then be used to label dependents of a verb in the Czech treebank as either arguments or adjuncts. Using our techniques, we ar able to achieve 88% precision on unseen parsed text.
I have been developing a computer program named RebLin. RebLin has special functions other than searching electronic dictionaries. In this paper, I will show how multi-functional RebLin expands the possibility of electronic dictionaries. First, I will compare RebLin with other computer programs which handle electronic dictionaries, analyzing how each of the software deal with inflected forms and derivatives. Then, I will explain how RebLin retrieves useful information out of other electronic resources such as those on the internet, digitized movies and a lexical database. Finally, I will analyze a log file which records words which users input in the search field of RebLin, and show how they make use of the various electronic resources which RebLin provides them with.
The present study used the picture perception paradigm to examine the extent to which three well-documented psychophysiological measures demonstrate consistency across time in response to emotional stimuli. The three measures were the eye-blink startle response and the activation in two facial muscle regions (zygomatic and corrugator). Twenty-seven young women were assessed on two occasions, 2 weeks apart. Whereas activation in the corrugator and zygomatic muscle regions demonstrated the predicted patterns at both assessments (with some attenuation in the zygomatic muscle regions), the startle response had limited consistency across the two assessments. The startle response revealed the predicted linear pattern of valence modulation during the first assessment. During the second assessment, startle magnitude response was a quadratic function of valence ratings and a linear function of arousal ratings. The unexpected pattern of startle response during the second session appeared to be related to the content of the pleasant slides, with action slides generating quadratic valence modulation and erotic slides continuing to exhibit the expected linear valence modulation.
Computer-driven systems for constructing composite faces of suspects (E-fit; Mac-a-Mug) have largely replaced mechanical systems (Photofit; the Identikit) in police use, yet little is known of their comparative effectiveness in rendering an accurate likeness. Participants (N = 24) constructed 2 of 4 familiar or unfamiliar faces, for one of which they used Photofit and for the other, E-fit. A likeness of each face was made first under target-absent conditions and then with photographs of the target present. The accuracy of the resulting composites was assessed by familiarity ratings, names elicited, and matching accuracy. The computer-driven system showed consistent superiority only when a familiar face was constructed in the presence of photographs; when participants worked from memory, E-fit was no better than Photofit. The implications of these findings for theories of face retrieval and the operational use of composites are discussed.
This paper discusses research on the English of Mexican Americans, arguing that in focusing primarily on description of vernacular Chicano English, the literature describing English spoken by Mexican Americans presents an incomplete picture of the complexity of their linguistic situation. Researchers may lose sight of the range of linguistic behaviors found within the Mexican American community, or even within a single family, where it is not uncommon to find fluent Spanish speakers, speakers with limited Spanish proficiency, and speakers of both nonstandard and standard dialects of English. The paper examines the range of linguistic behaviors found within three generations of a primarily English-dominant, middle class Mexican American family. It finds that even within this closeknit group of speakers, there exist distinct linguistic norms, ranging from those associated more closely with Chicano English to those associated with standard English. Even those speakers who make use of few if any linguistic resources associated with Chicano English or Spanish distinguish themselves linguistically from non-Mexican Americans. The paper considers the conscious and unconscious linguistic choices made by the speakers to be acts of identity, suggesting that the linguistic behaviors described are important means of constructing aspects of their social identities. (Contains 20 references.) (SM) Reproductions supplied by EDRS are the best that can be made from the original document. 1 RE-EXAMINING THE ENGLISH OF MEXICAN AMERICANS MS. AMANDA R. DORAN UNIVERSITY OF TEXAS AUSTIN, TEXAS PERMISSION TO REPRODUCE AND DISSEMINATE THIS MATERIAL HAS BEEN GRANTED BY dml TO THE EDUCATIONAL RESOURCES INFORMATION CENTER (ERIC) U.S. DEPARTMENT OF EDUCATION Office of Educational Research and Improvement ED E CATIONAL RESOURCES INFORMATION CENTER (ERIC) This document has been reproduced as received from the person or organization originating it. Minor changes have been made to improve reproduction quality. Points of view or opinions stated in this document do not necessarily represent official OERI position or policy.
It is generally recognized that the common nonterminal labels for syntactic constituents (NP, VP, etc.) do not exhaust the syntactic and semantic information one would like about parts of a syntactic tree. For example, the Penn Treebank gives each constituent zero or more &apos;function tags&apos; indicating semantic roles and other related information not easily encapsulated in the simple constituent labels. We present a statistical algorithm for assigning these function tags that, on text already parsed to a simplelabel level, achieves an F-measure of 87%, which rises to 99% when considering &apos;no tag&apos; as a valid choice.
This paper describes the design criteria and annotation guidelines of Sinica Treebank. The three design criteria are: Maximal Resource Sharing, Minimal Structural Complexity, and Optimal Semantic Information. One of the important design decisions following these criteria is the encoding of thematic role information. An on-line interface facilitating empirical studies of Chinese phrase structure is also described.
Three state-of-the-art statistical parsers are combined to produce more accurate parses, as well as new bounds on achievable Treebank parsing accuracy. Two general approaches are presented and two combination techniques are described for each approach. Both parametric and non-parametric models are explored. The resulting parsers surpass the best previously published performance results for the Penn Treebank.
Abstract Traditional theories of finance posit that the pricing of securities in financial markets should be done according to the quality of their underlying technical fundamentals. However, research on financial markets has tended to indicate that factors other than technical fundamentals are often used by market participants to gauge the value of securities. This phenomenon may be quite prevalent in markets for initial public offerings (IPOs), where securities lack a financial history. The imagery and affect associated with securities can be a powerful basis upon which to judge their worth. Advanced business students in a securities analysis course were asked to evaluate a number of industry groups represented on the New York Stock Exchange in terms of a set of judgmental variables. After providing imagery and affective evaluations for each industry group, the participants judged the likelihood that they would invest in companies associated with each industry. Imagery and affective ratings were highly correlated with one another and with the likelihood of investing. Judgments of performance correlated poorly to moderately with actual market performance as measured by weighted average returns for the industry groups studied. The results suggest that imagery and affect are part of a coherent psychological framework for evaluating classes of securities, but that framework may have low validity for predicting performance.
We present some novel machine learning techniques for the identification of subcategorization information for verbs in Czech. We compare three different statistical techniques applied to this problem. We show how the learning algorithm can be used to discover previously unknown subcategorization frames from the Czech Prague Dependency Treebank. The algorithm can then be used to label dependents of a verb in the Czech treebank as either arguments or adjuncts. Using our techniques, we are able to achieve 88% precision on unseen parsed text.
This article considers approaches which rerank the output of an existing probabilistic parser. The base parser produces a set of candidate parses for each input sentence, with associated probabilities that define an initial ranking of these parses. A second model then attempts to improve upon this initial ranking, using additional features of the tree as evidence. The strength of our approach is that it allows a tree to be represented as an arbitrary set of features, without concerns about how these features interact or overlap and without the need to define a derivation or a generative model which takes these features into account. We introduce a new method for the reranking task, based on the boosting approach to ranking problems described in Freund et al. (1998). We apply the boosting method to parsing the Wall Street Journal treebank. The method combined the log-likelihood under a baseline model (that of Collins [1999]) with evidence from an additional 500,000 features over parse trees that were not included in the original model. The new model achieved 89.75 % F-measure, a 13 % relative decrease in F-measure error over the baseline model’s score of 88.2%. The article also introduces a new algorithm for the boosting approach which takes advantage of the sparsity of the feature space in the parsing data. Experiments show significant efficiency gains for the new algorithm over the obvious implementation of the boosting approach. We argue that the method is an appealing alternative—in terms of both simplicity and efficiency—to work on feature selection methods within log-linear (maximum-entropy) models. Although the experiments in this article are on natural language parsing (NLP), the approach should be applicable to many other NLP problems which are naturally framed as ranking tasks, for example, speech recognition, machine translation, or natural language generation.
BOOK NOTICES 209 Linguistic databases. Ed. by John Nerbonne. (CSLI lecture notes 77.) Stanford, CA: CSLI, 1998. Pp. xxi, 243. The papers in this collection were originally presented at the 'Linguistic Databases' conference, University of Groningen, 23-24 March, 1995. Because ofthe almostproverbial rapidity with which information technology develops, the collection as a whole is dated already, but there is still much of interest to be found. Not all papers read at the conference are in this volume, but the papers cover a wide range of subjects, mostly practical in nature, not theoretical. After a clear and readable introduction by Nerbonne, the papers are presented in no particularorder, though the editor groups the papers in five main areas: syntactic corpora and databases, phonetic databases, applications in linguistic theory, applications, and extending basic technologies. The papers themselves are not presented according to this grouping, however, and at first sight the book appears rather disorganized. The wide variety of subjects can be deduced from the titles of the papers presented: 'Test suites for natural language processing', 'From annotated corpora to databases: The SgmlQL language', 'Markup of a test suite with SGML', 'An open systems approach for an acoustic-phonetic continuous speech database: The S_tools database-management system ', "The reading database of syllable structure',? database application for the generation of phonetic atlas maps', 'Swiss French polyphone and polyvar: Telephone speech databases to model inter- and intra-speaker variability', 'Investigating argument structure: The Russian nominalization database', "The use of a psycholinguistic database in the simplification of text for aphasie readers', "The computer learner corpus: A testbed for electronic EFL tools', 'Linking WordNet to a corpus query system', 'Multilingual data processing in the CELLAR environment '. The issue whether to use open free systems or closed proprietary systems is addressed in several papers. Some papers present applications developed both in open and closed systems. This is one area where developments have been going very fast, and nowadays freely available databases are often as capable as their commercial counterparts. Some of the applications presented in this collection are available from the Internet, and url's are often given. The collection can serve as a good introduction to the field for relative outsiders as ample references and links are given. The papers themselves vary greatly in subject matter so not all will be of interest to every reader. My particular favorite was 'From annotated corpora to databases: the SgmlQL language '. [BOUDEWUN REMPT.j Understanding phonology. By Carlos Gussenhoven and Haike Jacobs. (Understanding language series.) London: Arnold, 1998. Pp. xii, 286. This textbook is intended as an introduction to phonology aimed at 'students with little or no prior knowledge of linguistics' (back cover). As in many other textbooks, it uses exercises as a learning tool. Two types ofexercises are proposed. The ones identified by a key, 'intended as an expository aid' (xi), are provided with a solution in an appendix (though it is not always so much a clear cut answer as a guide for reflection, which is, to my view, a lot better). The ones identified by a dot are intended as practice material, and no solution is offered. I thought the idea of having two types of exercises a good one since it gives the reader the opportunity both for individual work and for discussion with others. Also, whenever it may apply, an optimality theoretic analysis is offered to describe a phonological process. Ch. 1, "The production of speech', is a basic introduction to phonology, phonetics, and phonation. Ch. 2,'Some typology: Sameness and difference', cleverly covers the universal and language specific aspects of phonological structures and typology. Ch. 3,'Making the form fit', addresses phonological grammar and adaptation by presenting the nativization of loan words in both the rules and the constraints approaches. Ch. 4, 'Underlying and surface representations', Ch. 5, 'Distinctive features', and Ch. 6, 'Ordered rules', deal with the basic notions of generative phonology within the SPE type formalism and introduce the reader to the school of linear phonology. Ch. 7,? case study: The diminutive suffix in Dutch', shows how these notions are applied. In Ch. 8, 'Levels of representation', Gussenhoven and Jacobs present an intermediate level of representation between the underlying representation and...
This research was designed to contribute to the development of new training systems to convey shape to the visually impaired. This dissertation consisted of two experiments which addressed the similarities and differences of visually impaired persons' and sighted persons' impressions of outlines of real objects perceived under auditory and touch display conditions, as well as tactile impressions of actual real objects, wood cutouts of those objects, tactor-pin diagrams of those objects, and raised-line drawings of those objects. Experiment 1 examined auditory and tactile perceptual equivalence, as well as sighted and visually impaired perceptual equivalence. Within each sight level group, Experiment 1 demonstrated a lack of equivalence for auditory and tactile perceptual structures. Experiment 1 demonstrated both similarities and differences in the perceptual structure comparisons between the sighted and the visually impaired. Experiment 2 examined the impact of sight level and four stimulus types on tactile percent correct identifications, response latencies, and familiarity ratings. Experiment 2 demonstrated sighted and visually impaired participants performed equally on a percent correct identification task. This Experiment also revealed the best performances were exhibited when real objects were presented, followed by wood cutouts, followed by tactor-pins, and performance with raised-line drawing (RLD) presentation was poorest. Experiment 2 also examined the exploratory procedures utilized during tactile identification of the four stimulus types. Experiment 2 revealed wood cutout presentation was associated with the largest percentage of different exploratory procedures being utilized, followed by real objects and tactor-pins which were equivalent, but utilized a larger percentage than the RLDs. Wood cutouts were the only other objects beside the real objects to use all exploratory procedures. The results of Experiment 1 and Experiment 2 indicate the utility of changing current industry standards for displaying shape information.
BOOK NOTICES 221 Indo-European perfects' (117-34), by Bridget Drinka. This paper investigates issues like the unidirectionality hypothesis and universal paths, with evidence drawn from a number of early Indo-European languages, e.g. Avestan, Sanskrit, and Homeric Greek. Other grammaticalization papers include "The sequencing of grammaticization effects: A twist from North America' (291-314) by Marianne Mithun and 'Grammaticalization of complex verbal constructions in Finnish' (363-76) by Taru Salminen. Several papers deal with language contact. Salikoko S. Mufwene's 'What research on creóle genesis can contribute to historical linguistics' (315-38) discusses issues like the social nature of creolization and the role of language contact in the histories of French and English. Another language contact paper is 'Yiddish and Hebrew: Borrowing through oral language contact' (135-48), by Elaine Gold, who argues that the source of the Hebrew component in Yiddish was not Hebrew texts but rather oral language contact. Syntactic change is not neglected. Ellen F. Prince's contribution, "The bonowing of meaning as a cause of internal syntactic change' (339-62), argues that 'at least some cases of (language-internal) syntactic change may result from (language-external) pragmatic and semantic borrowing' (339). Prince discusses three phenomena in support of this claim: Yiddish dos sentences, Yinglish 'Yiddish movement ', and the Yiddish pluperfect. Another paper on syntactic change is 'On the conservatism of embedded clauses' (255-68) by Kenjiro Matsuda which looks at some possible reasons for the resistance of such clauses to change (e.g. processing difficulties, pragmatic factors, and so on). All historical linguists should find something of interest in this volume, especially given the broad range of topics covered. The editors are to be commended for ajob well-done, and we can look forward to the publication of papers from the next ICHL. [Marc Pierce, University of Michigan.] Sprache und bürgerliche Nation: Beitr äge zur deutschen und europäischen Sprachgeschichte des 19. Jahrhunderts. Ed. by Dieter Cherubim, Siegfried Grosse, and Klaus J. Mattheier. Berlin & New York: Walter de Gruyter, 1998. Pp. ix, 456. This volume consists mainly of revised versions ofpapers presented at the 2nd Bad Homburger Kolloquium zur Sprachgeschichte des 19. Jahrhunderts, held in November 1993. (Three of the papers presented at the Kolloquium had already been promised to other publications; they were consequently replaced by four new papers written by conference participants ). A briefdescription ofthe contents follows. The range of topics covered is impressively broad; a number of important issues from the period 1790-1914 (as the term '19 Jahrhundert' is defined here) are discussed. Topics discussed include the status of German in various foreign countries, contact between German and other languages, the question of 'nation', and the language ofvarious social groups as well as their influence on the development of German. Papers include Klaus J. Mattheier's 'Kommunikationsgeschichte des 19. Jahrhunderts. Überlegungen zum Forschungsstand und zu Perspektiven der Forschungsentwicklung' (1-45), 'Deutsch in Belgien im neunzehnten Jahrhundert' (71-86) by Roland Willemyns, and 'Vom Dienstmädchen zur Professoringattin. Probleme bei der Aneigung bürgerlichen Sprachverhaltens und Sprachbewußtseins' (259-81) by Isa Schikorsky. Mattheier's paper is largely bibliographical in intent; it sketches some relevant issues (e.g. language contact, orthography, and the problem of a corpus) and provides a wealth of further references. Willemyns examines the status of German in Belgium—a weighty issue, given historical events. Schikorsky discusses the case of Elise Egloff, a Kindermädchen who eventually married a prominent professor of anatomy and pathology in Heidelberg, focusing on the problems that Elise faced in conforming to the linguistic norms of her future in-laws. Other contnbutions include ' "An mein Volk". Sprachliche Mittel monarchischer Appelle' (16796 ) by Hartmut Schmidt, 'Zum Einfluß der proletarischen und der bürgerlichen Frauenbewegung auf den politischen Wortschatz (um 1900)' (341-59). by Elisabeth Berner, and 'Morphologische und syntaktisch -stilistische Eigentümlichkeiten in deutschen Texten aus dem letzten Drittel des 19. Jahrhunderts' (420-43), by Siegfried Grosse. Schmidt discusses various aspects of a number of royal proclamations, e.g. Wilhelm H's 'An das deutsche Volk' issued on the entry of Germany into World War I, concentrating mainly on stylistic considerations. Berner's paper examines the usage of various lexemes, e.g. Emanzipation...
Pastiche is central to the resistant politics of Kathy Acker's writing--yet she would appear to agree with Fredric Jameson's influential critique of pastiche as "the wearing of a linguistic mask, speech in a dead language" (17). Her 1986 novel Don Quixote is all about having to speak "in a dead language" in the absence of a more "healthy" norm. It begins with the death of the protagonist, a female version of Cervantes's knight, who then goes on to narrate much of the subsequent story. Acker explains, "BEING DEAD, DON QUIXOTE COULD NO LONGER SPEAK. BEING BORN INTO AND PART OF A MALE WORLD, SHE HAD NO SPEECH OF HER OWN. ALL SHE COULD DO WAS READ MALE TEXTS WHICH WEREN'T HERS" (39). The novel then proceeds by plagiarism and pastiche, as Quixote goes on a quest--for a heterosexual love unsullied by patriarchal power relations--through fragments of numerous existing texts. Quixote rereads and pieces together a whole range of textual scraps, from Machiavelli's The Prince to a Godzilla movie. What becomes clear in her eccentric survey of (primarily) Western culture is that the lost, healthy linguistic norm is more than unhealthy for female readers--indeed, it is deadly.
This article focuses on the user-friendliness of lexical information sources. Whereas our previous study on user-friendliness (Euralex 1998) emphasized the context-sensitive needs of dictionary users, our present study goes one step further and suggests that it would be possible to compile interactive lexical databases that would be both context- and user-sensitive. We approach the function of lexical databases from two perspectives: from their role as primary information sources and from their role as lexical interfaces to other knowledge bases. Our approach is generally based on frame-semantics.We apply semantic frames to capture the different ways of conceptualization used when searching a knowledge base for social and health care services.
Computers are now widely used in the preparation of dictionaries. There are many advantages in maintaining and updating a dictionary in electronic form, most obviously that printed versions can be typeset directly from the electronic copy. But more than that, electronic dictionaries are beginning to be used by computers in retrieval systems. This chapter looks at electronic dictionaries and examines how lexical databases can help to refine and improve retrieval and analysis programs. It also traces the development of the uses of computers and dictionaries, and assesses various types of resources. Much research still needs to be done on the structure and contents of lexical and linguistic databases, especially for the semantic component, but the examples discussed in this chapter give some idea of the potential.
This paper proposes a new error-driven HMM-based text chunk tagger with context-dependent lexicon. Compared with standard HMM-based tagger, this tagger uses a new Hidden Markov Modelling approach which incorporates more contextual information into a lexical entry. Moreover, an error-driven learning approach is adopted to decrease the memory requirement by keeping only positive lexical entries and makes it possible to further incorporate more context-dependent lexical entries. Experiments show that this technique achieves overall precision and recall rates of 93.40% and 93.95% for all chunk types, 93.60% and 94.64% for noun phrases, and 94.64% and 94.75% for verb phrases when trained on PENN WSJ TreeBank section 00-19 and tested on section 20-24, while 25-fold validation experiments of PENN WSJ TreeBank show overall precision and recall rates of 96.40% and 96.47% for all chunk types, 96.49% and 96.99% for noun phrases, and 97.13% and 97.36% for verb phrases.
In this contribution we discuss how a fuzzy querying interface can support the generation of linguistic database summaries - a special technique of data mining. Links between our approach to linguistic summaries and the well-known technique of association rules is shown. The implementation of linguistic summaries generation using the authors’ FQUERY for Access package is presented.
This paper presents results for a maximum-entropy-based part of speech tagger, which achieves superior performance principally by enriching the information sources used for tagging. In particular, we get improved results by incorporating these features: (i) more extensive treatment of capitalization for unknown words; (ii) features for the disambiguation of the tense forms of verbs; (iii) features for disambiguating particles from prepositions and adverbs. The best resulting accuracy for the tagger on the Penn Treebank is 96.86% overall, and 86.91% on previously unseen words.
In this paper, we propose a new ambiguity representation scheme; Structure Preference Relation (SPR), which consists of useful quantitative distribution information for ambiguous structures. Two automatic acquisition algorithms, the first acquired from a treebank, and the second acquired from raw texts, are introduced, and some experimental results which prove the availability of the algorithms are also given. Finally, we introduce some SPR applications in linguistics and natural language processing, such as preference-based parsing and the discovery of representative ambiguous structures, and propose some future research directions.
The Verbmobil treebanks of spoken German, English, and Japanese are part of the Verbmobil project, which has the overriding goal to develop a speaker-independent system for the translation of spontaneous speech. In the framework of this language technology project, the treebanks provide training data for a variety of language technology modules. The treebanks consist of annotated syntactic tree structures based on transcribed dialogs in the scenarios of appointment negotiations, travel arrangements, and personal computer maintenance. The annotation schemes of the treebanks have been developed taking into account the specific characteristics of spoken language dialogs: repetitions, hesitations, false starts'', etc. * The work reported here was funded by the German Ministry of Education and Research (BMBF) in the framework of the Verbmobil project under grant FKZ:01 IV 701 M0.
In this paper we present the results of a quantitative evaluation of the discrepancies between the Italian and English lexica in terms of lexical gaps. This evaluation has been carried out in the context of MultiWordNet, an ongoing project that aims at building a multilingual lexical database. The quantitative evaluation of the English-to-Italian lexical gaps shows that the English and Italian lexica are highly comparable and gives empirical support to the MultiWordNet model. 1.
No language in the world is homogeneous, or ever will be. Whereas earlier forms of English were characterised by extreme variation on all levels and Middle English is in fact best described as a loose conglomerate of unstable varieties, we usually lack any more detailed insight into what functions this variation had for the individual speaker. The social correlates so well known from modern sociolinguistics, such as age, sex, education, religion, can normally not be applied to the existing texts, nor can even the geographical range of recorded forms be determined with any degree of certainty. Finally, if modern dialect or other non-standard features are contrasted with (as the term non-standard implies) an accepted standard form of a language, this method would necessarily fail with Middle English even if we knew more about it than we do and, in view of the state of surviving documents, ever will. It is safe to assume that for its speakers the linguistic heterogeneity of Middle English was ordered in some way, but it was so only for continually shifting speech communities, whose number and individual geographical spread we know very little about. The scene changed dramatically in the fifteenth century: the emergence of a new standard language began to re-institute a linguistic norm for written supraregional English. This development was a natural consequence of the acceptance of English in public domains, and was speeded up by the change-over to English as the Chancery language in 1430.
International audience
This article focuses on ongoing work done for Portuguese concerning the phenomenon of lexical co-occurrence known as collocation (cf. Cruse, 1986, inter al.). Instances of the syntactic variety formed by noun plus adjective have been especially observed. Collocational instances are not lexical entries, and thus should not be stored in the lexicon as multiword lexical units. Their processing can be conceived through relations linking the lexical components. Mechanisms for dealing with the collocation-hood of the expressions are required to be included in the systems, topographically, in their lexical modules. Lexical databases like wordnets, with a general architecture typically structured on semantic relations, make room for the specification of this phenomenon. This can be handled through the definition of ad-hoc relations expressing the different semantic effects the adjectival modification bring to nominal phrases, collocationally. 1
This paper describes the methodology that is being used to augment the Penn Treebank annotation with sense tags and other types of semantic information. Inspired by the results of SENSEVAL, and the high inter-annotator agreement that was achieved there, similar methods were used for a pilot study of 5000 words of running text from the Penn Treebank. Using the same techniques of allowing the annotators to discuss difficult tagging cases and to revise WordNet entries if necessary, comparable inter-annotator rates have been achieved. The criteria for determining appropriate revisions and ensuring clear sense distinctions are described. We are also using hand correction of automatic predicate argument structure information to provide additional thematic role labeling. 1.
In this paper, we present a method for comparing Lexicalized Tree Adjoining Grammars extracted from annotated corpora for three languages: English, Chinese and Korean. This method makes it possible to do a quantitative comparison between the syntactic structures of each language, thereby providing a way of testing the Universal Grammar Hypothesis, the foundation of modern linguistic theories.
In this paper, we present a neural-networks-based knowledge discovery and data mining (KDDM) methodology based on granular computing, neural computing, fuzzy computing, linguistic computing, and pattern recognition. The major issues include 1) how to make neural networks process both numerical and linguistic data in a data base, 2) how to convert fuzzy linguistic data into related numerical features, 3) how to use neural networks to do numerical-linguistic data fusion, 4) how to use neural networks to discover granular knowledge from numerical-linguistic data bases, and 5) how to use discovered granular knowledge to predict missing data. In order to answer the above concerns, a granular neural network (GNN) is designed to deal with numerical-linguistic data fusion and granular knowledge discovery in numerical-linguistic databases. From a data granulation point of view, the GNN can process granular data in a database. From a data fusion point of view, the GNN makes decisions based on different kinds of granular data. From a KDDM point of view, the GNN is able to learn internal granular relations between numerical-linguistic inputs and outputs, and predict new relations in a database. The GNN is also capable of greatly compressing low-level granular data to high-level granular knowledge with some compression error and a data compression rate. To do KDDM in huge data bases, parallel GNN and distributed GNN will be investigated in the future.
Article choice can pose difficult problems in applications such as machine translation and automated summarization. In this paper, we investigate the use of corpus data to collect statistical generalizations about article use in English in order to be able to generate articles automatically to supplement a symbolic generator. We use data from the Penn Treebank as input to a memory-based learner (TiMBL 3.0; We discuss competitive results obtained using a variety of lexical, syntactic and semantic features that play an important role in automated article generation.
Bagging and boosting, two effective machine learning techniques, are applied to natural language parsing. Experiments using these techniques with a trainable statistical parser are described. The best resulting system provides roughly as large of a gain in F-measure as doubling the corpus size. Error analysis of the result of the boosting technique reveals some inconsistent annotations in the Penn Treebank, suggesting a semi-automatic method for finding inconsistent treebank annotations.
This paper presents the first-ever results of applying statistical parsing models to the newly-available Chinese Treebank. We have employed two models, one extracted and adapted from BBN's SIFT System (Miller et al., 1998) and a TAG-based parsing model, adapted from (Chiang, 2000). On sentences with ≤40 words, the former model performs at 69% precision, 75% recall, and the latter at 77% precision and 78% recall.
The accuracy of statistical parsing models can be improved with the use of lexical information. Statistical parsing using Lexicalized tree adjoining grammar (LTAG), a kind of lexicalized grammar, has remained relatively unexplored. We believe that is largely in part due to the absence of large corpora accurately bracketed in terms of a perspicuous yet broad coverage LTAG. Our work attempts to alleviate this difficulty. We extract different LTAGs from the Penn Treebank. We show that certain strategies yield an improved extracted LTAG in terms of compactness, broad coverage, and supertagging accuracy. Furthermore, we perform a preliminary investigation in smoothing these grammars by means of an external linguistic resource, namely, the tree families of an XTAG grammar, a hand built grammar of English.
BOOK NOTICES 209 Linguistic databases. Ed. by John Nerbonne. (CSLI lecture notes 77.) Stanford, CA: CSLI, 1998. Pp. xxi, 243. The papers in this collection were originally presented at the 'Linguistic Databases' conference, University of Groningen, 23-24 March, 1995. Because ofthe almostproverbial rapidity with which information technology develops, the collection as a whole is dated already, but there is still much of interest to be found. Not all papers read at the conference are in this volume, but the papers cover a wide range of subjects, mostly practical in nature, not theoretical. After a clear and readable introduction by Nerbonne, the papers are presented in no particularorder, though the editor groups the papers in five main areas: syntactic corpora and databases, phonetic databases, applications in linguistic theory, applications, and extending basic technologies. The papers themselves are not presented according to this grouping, however, and at first sight the book appears rather disorganized. The wide variety of subjects can be deduced from the titles of the papers presented: 'Test suites for natural language processing', 'From annotated corpora to databases: The SgmlQL language', 'Markup of a test suite with SGML', 'An open systems approach for an acoustic-phonetic continuous speech database: The S_tools database-management system ', "The reading database of syllable structure',? database application for the generation of phonetic atlas maps', 'Swiss French polyphone and polyvar: Telephone speech databases to model inter- and intra-speaker variability', 'Investigating argument structure: The Russian nominalization database', "The use of a psycholinguistic database in the simplification of text for aphasie readers', "The computer learner corpus: A testbed for electronic EFL tools', 'Linking WordNet to a corpus query system', 'Multilingual data processing in the CELLAR environment '. The issue whether to use open free systems or closed proprietary systems is addressed in several papers. Some papers present applications developed both in open and closed systems. This is one area where developments have been going very fast, and nowadays freely available databases are often as capable as their commercial counterparts. Some of the applications presented in this collection are available from the Internet, and url's are often given. The collection can serve as a good introduction to the field for relative outsiders as ample references and links are given. The papers themselves vary greatly in subject matter so not all will be of interest to every reader. My particular favorite was 'From annotated corpora to databases: the SgmlQL language '. [BOUDEWUN REMPT.J Understanding phonology. By Carlos Gussenhoven and Haike Jacobs. (Understanding language series.) London: Arnold, 1998. Pp. xii, 286. This textbook is intended as an introduction to phonology aimed at 'students with little or no prior knowledge of linguistics' (back cover). As in many other textbooks, it uses exercises as a learning tool. Two types ofexercises are proposed. The ones identified by a key, 'intended as an expository aid' (xi), are provided with a solution in an appendix (though it is not always so much a clear cut answer as a guide for reflection, which is, to my view, a lot better). The ones identified by a dot are intended as practice material, and no solution is offered. I thought the idea of having two types of exercises a good one since it gives the reader the opportunity both for individual work and for discussion with others. Also, whenever it may apply, an optimality theoretic analysis is offered to describe a phonological process. Ch. 1, "The production of speech', is a basic introduction to phonology, phonetics, and phonation. Ch. 2,'Some typology: Sameness and difference', cleverly covers the universal and language specific aspects of phonological structures and typology. Ch. 3,'Making the form fit', addresses phonological grammar and adaptation by presenting the nativization of loan words in both the rules and the constraints approaches. Ch. 4, 'Underlying and surface representations', Ch. 5, 'Distinctive features', and Ch. 6, 'Ordered rules', deal with the basic notions of generative phonology within the SPE type formalism and introduce the reader to the school of linear phonology. Ch. 7,? case study: The diminutive suffix in Dutch', shows how these notions are applied. In Ch. 8, 'Levels of representation', Gussenhoven and Jacobs present an intermediate level of representation between the underlying representation and...
1.1 Tagging criteria....................................... 4 1.2 POS tagset......................................... 5 1.3 Size of the POS tagset................................... 6
We aim at finding the minimal set of fragments which achieves maximal parse accuracy in Data Oriented Parsing. Experiments with the Penn Wall Street Journal treebank show that counts of almost arbitrary fragments within parse trees are important, leading to improved parse accuracy over previous models tested on this treebank. We isolate a number of dependency relations which previous models neglect but which contribute to higher parse accuracy.
This paper demonstrates that machine learning is a suitable approach for rapid parser development. From 1000 newly treebanked Korean sentences we generate a deterministic shift-reduce parser. The quality of the treebank, particularly crucial given its small size, is supported by a consistency checker. 1 Introduction Given the enormous complexity of natural language, parsing is hard enough as it is, but often unforeseen events like the crises in Bosnia or East-Timor create a sudden demand for parsers and machine translation systems for languages that have not benefited from major attention of the computational linguistics community up to that point. Good machine translation relies strongly on the context of the words to be translated, a context that often goes well beyond neighboring surface words. Often basic relationships, like that between a verb and its direct object, provide crucial support for translation. Such relationships are usually provided by parsers. The NLP resources f...