Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
An architecture for federating heterogeneousdictionary databases is described. It proposes acommon description language and query language toprovide for the exchange of information betweendatabases with different organizations, on differentplatforms and in different DBMSs. The common querylanguage has an SQL like structure. The first versionof the description language follows the TEI standardtag definitions for dictionaries with the expectationthat the description language will be expanded in thefuture. A practical implementation of the proposalsusing WWW technology for two multi-lingualdictionaries is described.
This article focuses on the user-friendliness of lexical information sources. Whereas our previous study on user-friendliness (Euralex 1998) emphasized the context-sensitive needs of dictionary users, our present study goes one step further and suggests that it would be possible to compile interactive lexical databases that would be both context- and user-sensitive. We approach the function of lexical databases from two perspectives: from their role as primary information sources and from their role as lexical interfaces to other knowledge bases. Our approach is generally based on frame-semantics.We apply semantic frames to capture the different ways of conceptualization used when searching a knowledge base for social and health care services.
BOOK NOTICES 487 seriously interested in Salish studies as well as by local Washington libraries concerned with promoting traditional Lushootseed language and culture. [Edward T. Vajda, Western Washington University.] An ethnographic grammar of the Eipo language spoken in the central mountains of Irian Jaya (West New Guinea), Indonesia. By Volker Heeschen. (Mensch, Kultur und Umwelt im zentralen Bergland von West Neuguinea 23.) Berlin: Dietrich Reimer Verlag, 1998. Pp. 412. Eipo belongs to the Mek family, a group ofclosely related languages spoken in several mountain valleys between areas occupied by speakers of Dani and Ok languages. The author's latest contribution to the ethnolinguistics of this remote area, this large-format paperbackjoins several previous volumes in the same series devoted to Eipo, including: Wörterbuch EipoDeutsch -English (Volker Heeschen, 1983, vol. 6,), with 5,682 main entries, and Kommunikation bei den Eipo (Volker Heeschen, 1989, vol. 19), a description of communicative styles and language change in a small speech community. The present work provides the first extensive description ofEipo phonology and grammar. Heeschen gathered his voluminous data during more than a dozen field trips made since 1977. The book is 'ethnographic' in the sense that H links his descriptions to specific cultural and pragmatic contexts. Rather than attempting to portray Eipo as conforming to a fixed norm, H describes the rules creating grammatical forms in Eipo, a language with about 400 speakers (22-23), as relatively fluid when compared to languages spokenby larger, more extensive populations. The book consists of three parts divided into several chapters each. Part 1 introduces Eipo culture and history (13-35) and discusses Eipo's position within the Mek language family (72-94). H identifies Mek as a low-level genetic grouping similar to the several dozen posited by William A. Foley (The Papuan languages ofNew Guinea, Cambridge: Cambridge University Press, 1986). However, H uses typological and lexical data to support Wurm's classification of Mek within the Trans-New Guinea phylum (Stephen A. Wurm, Papuan languages ofOceania, Tübingen: Gunter Narr, 1982), though the evidence suggests a very distant connection. In a somewhat rambling fashion, H discusses a medley of approaches used by previous scholars to describe 'exotic' languages (36-71) and selects elements from a variety of traditions for their relevance in describing Eipo. H justifies this eclecticism based on his own observations regarding linguistic selfawareness and language creation among the Eipo (95-114). Part 2 provides a meticulous, data- rather than theory-driven description of Eipo phonetics and phonology (115-40), word classes and morphosyntax (141-264), and syntax (265-356). Each sectioncontains numerous paradigms and other grammatical schemata, plus legions of example phrases and sentences interpreted from the vantage of the author's keen understanding of the ambient cultural context, a factor which if omitted would render the literal translations of many examples unintelligible. Finally, Part 3 (357-80) provides nine previously unpublished texts in Eipo and neighboring Mek languages. These texts deal with local myths and legends and are accompanied by interlinear glosses, a translation into idiomatic English, and copious ethnographic explanations. H's work represents a solid, multifaceted contribution to the study of New Guinea ethnography and linguistics and contains much that will be of interest to general typologists. Because Mek languages were until recently more poorly described than neighboring groups, this book contributes important data to the ongoing task ofestablishing genetic relationships between New Guinea's several hundred languages. Finally, H's conclusions regarding rates of vocabulary change in a language not known to have ever counted more than a few hundred speakers, with its typical absence of any fixed conservative norm, may have implications for broader studies of language contact and genetic linguistics. [Edward J. Vajda, Western Washington University.] People, countries, and the Rainbow Serpent: Systems of classification among the Lardil of Mornington Island. By David McKnight. (Oxford studies in anthropological linguistics 12.) New York & Oxford: Oxford University Press, 1999. Pp. x, 270. This book reflects over five years of field work conducted at intervals beginning in 1966 and contains a treasure trove of data on Lardil language and culture that would almost certainly have otherwise disappeared unrecorded. McKnight elicited information from his native speaker informants in a...
In this paper, we propose a new ambiguity representation scheme; Structure Preference Relation (SPR), which consists of useful quantitative distribution information for ambiguous structures. Two automatic acquisition algorithms, the first acquired from a treebank, and the second acquired from raw texts, are introduced, and some experimental results which prove the availability of the algorithms are also given. Finally, we introduce some SPR applications in linguistics and natural language processing, such as preference-based parsing and the discovery of representative ambiguous structures, and propose some future research directions.
The accuracy of statistical parsing models can be improved with the use of lexical information. Statistical parsing using Lexicalized tree adjoining grammar (LTAG), a kind of lexicalized grammar, has remained relatively unexplored. We believe that is largely in part due to the absence of large corpora accurately bracketed in terms of a perspicuous yet broad coverage LTAG. Our work attempts to alleviate this difficulty. We extract different LTAGs from the Penn Treebank. We show that certain strategies yield an improved extracted LTAG in terms of compactness, broad coverage, and supertagging accuracy. Furthermore, we perform a preliminary investigation in smoothing these grammars by means of an external linguistic resource, namely, the tree families of an XTAG grammar, a hand built grammar of English.
A BILINGUAL LEXICAL DATABASE FOR FRAME SEMANTICS Get access Thierry Fontenelle Thierry Fontenelle 19 Rue du Merschgrund (L-8373 Hobscheid, Luxembourg)University of Liège(B-4000 Liege, Belgium) (fontenel@pt_lu) Search for other works by this author on: Oxford Academic Google Scholar International Journal of Lexicography, Volume 13, Issue 4, December 2000, Pages 232–248, https://doi.org/10.1093/ijl/13.4.232 Published: 01 December 2000
A Classification Information Model is a pattern classification model.The model decides the proper class of an input instance by integrating individual decisions, each of which is made with each feature in the pattern.Each individual decision is weighted according to the distributional property of the feature deriving the decision. An individual decision and its weight are represented as classification information which is extracted from the training instances.In the word sense disambiguation based on the model, the proper sense of an input instance is determined by the weighted sum of whole individual decisions derived from the features contained in the instance.
This paper describes a specific part of the Prague Dependency Treebank annotation, the step from the surface dependency structure towards the underlying representation of the sentence. The first section explains the theoretical basis of the project. In Section 2 all the procedure of conversion to the tectogrammatical structure is summarized and Section 3 presents in detail the present stage of the automated part of the conversion procedure.
The ability to circumlocute successfully is of utmost importance in compensating for gaps in lexical knowledge. Although all studies indicate that one’s ability to circumlocute increases with increasing proficiency, it is interesting that little attention has been paid to those learners who have the greatest ability to circumlocute, native‐like speakers. This study addresses the norms of native and native‐like circumlocution. It expands the discussion of strategies involved in this skill to include the means by which speakers frame their message and thereby set the linguistic context for their listeners. Participants in this study, both native and native‐like speakers, were found to employ similar strategies while circumlocuting, including the use of synonyms, analogies, and descriptions. These participants also consistently framed their speech to facilitate listener comprehension, and they frequently included in their discourse some reference to their status as a nonexpert in the field. Similarities in native and native‐like circumlocution found in this study help to provide some empirical validation to the notion of “native‐like.”
Many English Corpus Linguistics projects reported in ICAME Journal and elsewhere involve grammatical analysis or tagging of English texts (eg Atwell 1983, Leech et al 1983, Booth 1985, Owen 1987, Souter 1989a, O’Donoghue 1991, Belmore 1991, Kytö and Voutilainen 1995, Aarts 1996, Qiao and Huang 1998). Each new project has to review existing tagging schemes, and decide which to adopt and/or adapt. The AMALGAM project can help in this decision, by providing descriptions and analyses of a range of tagging schemes, and an internet-based service for researchers to try out the range of tagging schemes on their own data. The project AMALGAM (Automatic Mapping Among Lexico-Grammatical Annotation Models) explored a range of Part-of-Speech tagsets and phrase structure parsing schemes used in modern English corpus-based research. The PoS-tagging schemes include: Brown (Greene and Rubin 1981), LOB (Atwell 1982, Johansson et al 1986), Parts (man 1986), SEC (Taylor and Knowles 1988), POW (Souter 1989b), UPenn (Santorini 1990), LLC (Eeg-Olofsson 1991), ICE (Greenbaum 1993), and BNC (Garside 1996). The parsing schemes include some which have been used for hand annotation of corpora or manual post-editing of automatic parsers, and others which are unedited output of a parsing program. Project deliverables include: – a detailed description of each PoS-tagging scheme, at a comparable level of detail. This includes a list of PoS-tags with descriptions and example uses from the source Corpus. The description of the use of PoS-tags is also illustrated in a multi-tagged corpus: a set of sample texts PoS-tagged in parallel with each PoS-tagset (and proofread by experts), for comparative studies – an analysis of the different lexical tokenization rules used in the source Corpora, to arrive at a ‘Corpus-neutral’ tokenization scheme (and consequent adjustments to the PoS-tagsets in our study to accept modified tokenization) – an implementation of each PoS-tagset in conjunction with our standardised tokenizer, as a family of PoS-taggers, one for each PoS-tagset – a method for ‘PoS-tagset conversion’, taking a text tagged according to one PoS-tagset and outputting the text annotated with another PoS-tagset – a sample of texts parsed according to a range a parsing schemes: a Multi-Treebank resource for comparative studies – an Internet service allowing researchers worldwide free access to the above resources, including a simple email-based method for PoS-tagging any English text with any or all PoS-tagset(s).
Language deviation is a kind of langUage form diverging from the language norm, In poetry, there are eight kinds of language deviation: lexical deviation.Phonological deviation. grammatical deviation. graphological deviation. semantic deviation.deviation of register. deviation of historical period and dialectal deviation. From the point of view of the aesthetic function, we analyze the language deviation of poetry in outer to appreciate the poems better.
The Translational English Corpus (TEC) held at the Centre for Translation Studies at UMIST is a full-text, synchronic, general, monolingual, written, single, translational corpus of English. It is also a direct, multi-source-language, mono-translation-mode (written mode), mono-translation-method (human translation), largely into-mother-tongue, professional, published corpus. At the time of writing TEC represents four text categories: newspapers, biography, fiction, and inflight magazines. Hatim (1999) has argued that so far in translation studies an important distinction has been ignored. This distinction is between what is "in" and what is "of" the text. "In" refers to the language itself, which can be ana-lyzed through text analysis, while "of" refers to the text in its entirety, its overall effect in terms of ideology. In this paper, I will argue that TEC can indeed be a valuable, self-contained, single resource for studying precisely the "of" of translational language, that is, its ideological impact in the target language and culture. To this purpose, I will discuss some methodological issues as well as the methods for carrying out critical linguistic analy-sis. I will also suggest ways of investigating norms of lexical use in TEC and their possible ideological implications through the lexico-grammatical and collocational analysis of a set of key words relating to Europe in translated newspaper articles.
This paper describes the design criteria and annotation guidelines of Sinica Treebank. The three design criteria are: Maximal Resource Sharing, Minimal Structural Complexity, and Optimal Semantic Information. One of the important design decisions following these criteria is the encoding of thematic role information. An on-line interface facilitating empirical studies of Chinese phrase structure is also described.
This paper presents a query tool for syntactically annotated corpora. The query tool is developed to search the Verbmobil treebanks annotated at the University of Tübingen. However, in principle it also can be adapted to other corpora such as the Negra Corpus, the Penn Treebank or the French treebank developed in Paris. The tool uses a query language that allows to search for tokens, syntactic categories, grammatical functions and binary relations of (immediate) dominance and linear precedence between nodes. The overall idea is to extract in an initializing phase the relevant information from the corpus and store it in a relational database. An incoming query is then translated into a corresponding SQL query that is evaluated on the database.
In developing the concepts of phonological and lexical subtypes of dyslexia, criteria have been proposed based on the projection of linear regression for raw test scores. Substantial discrepancies in subtype prevalence may arise from nonlinearities in the test scores. A norming process developed for a new nonword test, the Martin and Pratt Nonword Reading Test (Martin & Pratt, 2000), was applied to the Word Identification subtest of the Woodcock Reading Mastery Test (Woodcock, 1987) and the Regular and Irregular Word Tests published in Coltheart and Leahy (1996). These tests were administered to a representative sample of 863 children aged 6 to 15 years in the Southern Tasmanian State School population. An inverse normal transform results in a distribution which is approximately normal within age groups. On this scale the age effect was well approximated by a linear increase with the logarithm of (age −5 years). This process can be adapted to provide norms for word lists more economically and allows convenient spreadsheet formulae for norms. Substantial differences from other norms may be attributed to school district family income differences found in this sample. Male means are lower than female for all tests, but this reflects comparable high performance and disproportionate poor performance by males on reading tests.
Event-related brain potentials were recorded to study whether verbs and nouns activate topographically distinct cortical generators. Fifteen subjects performed a primed lexical decision task with verb/verb and noun/noun pairs. The relatedness between prime and target items was varied in three steps (unrelated, moderately, and strongly related) and the EEG was recorded from 124 scalp electrodes. The topography of cortical sources of the N400 effect was evaluated by standardized differences scores and by cortical current source estimates which were constrained by the individual MRI-determined cortex anatomy. A behavioral priming effect and a substantial N400 effect was found for both word categories. However, the topography of the grand average N400 effect of verbs and nouns did not differ, neither for raw nor for standardized amplitudes. Cortical current source estimates of the N400 effect revealed a very broad and scattered distribution of active locations with pronounced interindividual differences. Cortical current source estimates obtained with the L1-norm and L2-norm model, respectively, differed in the distribution of sources over the cortex but converged on the same "hot spots." The data give no indication that the N400 effect is generated by word category-specific networks which have a different topography. The marked individual differences are discussed with respect to the involved processes and the current source estimation procedures.
In this paper, we present a neural-networks-based knowledge discovery and data mining (KDDM) methodology based on granular computing, neural computing, fuzzy computing, linguistic computing, and pattern recognition. The major issues include 1) how to make neural networks process both numerical and linguistic data in a data base, 2) how to convert fuzzy linguistic data into related numerical features, 3) how to use neural networks to do numerical-linguistic data fusion, 4) how to use neural networks to discover granular knowledge from numerical-linguistic data bases, and 5) how to use discovered granular knowledge to predict missing data. In order to answer the above concerns, a granular neural network (GNN) is designed to deal with numerical-linguistic data fusion and granular knowledge discovery in numerical-linguistic databases. From a data granulation point of view, the GNN can process granular data in a database. From a data fusion point of view, the GNN makes decisions based on different kinds of granular data. From a KDDM point of view, the GNN is able to learn internal granular relations between numerical-linguistic inputs and outputs, and predict new relations in a database. The GNN is also capable of greatly compressing low-level granular data to high-level granular knowledge with some compression error and a data compression rate. To do KDDM in huge data bases, parallel GNN and distributed GNN will be investigated in the future.
1. Introduction Not every historical linguist embraces the idea of Chomsky's syntactocentrism with enthusiasm. It may be untimely to say unkind things about it, but there are syntactic problems which cannot be resolved satisfactorily only by formal operations. Under the current psycholinguistic views there seem to be some chances of recognizing the old conceptual world of the speaker and thus contributing to a more appropriate understanding of the writings he has left. Following chiefly Jackendoff's ideas expressed in The architecture of the language faculty (1997) -- yet with due respect for other linguistic and psycholinguistic orientations -- I will discuss grammatical relations which involve word order, thematic roles and word-formation (compounding) and which by structural standards prove so intractable. A common trait of them all is that they are structurally ambiguous and consequently differ in meanning, or that they are simply semantically opaque. 2. Word order An example of how weakly significant word order in Old English can be is the first part of the following sentence: 1) Storm oft holm gebringep, geofen in grimmum selum (Maxims I 112/50) which has been understood as either 'The sea often brings a storm, the ocean in stormy seasons' (Gordon 1954: 342) 'The sea often brings a storm' (Bosworth, entry gebringan) or 'The often brings forth a flood' (Reszkiewicz 1971:35) 'storm oft brings ocean into a furious condition' (Bosworth, entry soel) The interpretative difficulty lies in the fact that the functions of a grammatical subject and a grammatical object are not clearly transparent: the nouns and holm are both singular and each can agree with the finite form of the verb, gebringep, which as a two (or even three) argument requires a subject and an object. This brings up a question: which is which? Structurally speaking each can perform either function. They are both masculine, singular, of a-inflection of which nominative/accusative syncretism is a norm. Besides, there is no adjectival or pronominal modifier to help, neither can alliteration be helpful. Reszkiewicz searched for a clue to the functional identification in the position of the noun with regard to the and came to the conclusion that: Older Old English, especially poetry, lacked both the definite and the indefinite articles; the object often preceded the governing verb (Reszkiewicz 1971: 35). Although the grounds on which such a decision is reached are formally defens ible, empirically they are less so as they can be falsified by a sentence, also a gnomic verse, which reads: (2) Moegen mon sceal mid mete fedan (Maxims I 118/44) in which it is the subject man and not moegen which is closer to the finite form of the verb, sceal (moegen and man also show inflectional syncretism in this respect); this sententious saying means: 'One shall nourish strength with meat' (food) (Gordon 1954: 344) 'A man must feed strength with meat' (Bosworth, entry fedan) The proponents of either of the two meanings of the gnomic storm verse would probably try to persuade us that their views are compatible with the formal grammatical relations. But which of the meanings would satisfy the pragmatics of the discourse? Although the senses of particular lexical items are clear, a real cognitive image is still concealed. As a historical linguist I am more comfortable asking questions than answering them, so my glimpse into the Old English cognitive mind will be based on the possible, we now try to see, life as it would have been over a millenium of years ago. Since the conceptual structure of our example is not immediately predictable from the syntactic structure, nor is it found in the lexical structures, I will try to consider the language context first and then to search for similar uses of and holm. …
The effects of lexical difficulty and talker variability on word recognition were examined in four groups of listeners: native English/normal hearing; native English/hearing impaired; non-native English/normal hearing; and non-native English/hearing impaired (hearing level matched to the native hearing impaired). Lexical difficulty was measured by the difference in performance to 75 lexically ‘‘easy’’ and ‘‘hard’’ words based on word frequency and Neighborhood Activation Theory [Luce and Pisoni (1998)]. The effect of talker variability was measured by the difference in performance between single and multiple talker (nine talkers) conditions. The familiarity of the 150 words was rated on a seven-point scale. An up–down adaptive procedure was used to determine the sound pressure level for 50% performance. Non-native listeners in both normal and hearing-impaired groups required a greater intensity for equal intelligibility than for the comparative native normal and hearing-impaired listeners. Results, however, showed significant effects of lexical difficulty and talker variability in all four groups. Structural equation modeling demonstrated that an auditory factor estimated by pure tone average, etc., accounts for four times more variance to performance than does a linguistic fluency factor measured by word familiarity ratings and native versus non-native status, however, the linguistic fluency factor is also essential to the model fit.
1.1 Tagging criteria....................................... 4 1.2 POS tagset......................................... 5 1.3 Size of the POS tagset................................... 6
1.1 Notion of word....................................... 4 1.2 Tests of wordhood..................................... 5 1.3 Compatibility with other guidelines............................ 6
Computers are now widely used in the preparation of dictionaries. There are many advantages in maintaining and updating a dictionary in electronic form, most obviously that printed versions can be typeset directly from the electronic copy. But more than that, electronic dictionaries are beginning to be used by computers in retrieval systems. This chapter looks at electronic dictionaries and examines how lexical databases can help to refine and improve retrieval and analysis programs. It also traces the development of the uses of computers and dictionaries, and assesses various types of resources. Much research still needs to be done on the structure and contents of lexical and linguistic databases, especially for the semantic component, but the examples discussed in this chapter give some idea of the potential.
The purpose of this article is to analyze the presence of traces of what will call a sacerdotal dimension, which is related to the theme of hermaphroditism, in Savinio's novel, as well as to verify what cultural implications are suggested by this sacerdotal dimension. Its traces will be found in the development of the young protagonist's (Nivasio's) thoughts and discoveries, which make up the book's plot, and which are surrounded by a mythic halo. Infanzia di Nivasio Dolcemare (1941) (2) should not be considered as merely an autobiographical work, but as an autobiography made dreamlike and strange, exclusively dedicated to a lost age, that of and puberty, up to the moment of the discovery of sex, after which childhood is over, in fact, and the narrative closes. (3) would like to propose a preliminary reflection: although much has been said of the linguistic techniques employed by Savinio, such as pastiche, plurilinguism, and others, one must always keep in mind that the author's linguistic stratagems were all used to promote a peculiar autobiographical dimension. (4) Throughout his literary career, which cannot be analyzed here in its entirety, Savinio often seems to have been concerned with representing the formation and the expression of subjectivity in relation with the world. The textual I of this subjectivity never assumes traditional forms, however, but constantly seeks to deviate from linguistic norms, to transgress in its formation of images and characters. Yet its goal is always also to reconstruct the sense of a subjective progression, of the journey of an individual ego through the things of the world. The story of Infanzia di Nivasio Dolcemare is set in Greece, as is frequent in Savinio, but it is a Greece relived through the memory of the domestic situations of childhood. This re-evocation is founded on deformed obsessions and presences. The deformation already begins in the very names of the protagonists, which were created according to an anagrammatic logic. Nivasio's father's name is Visanio, for instance, but even the name Nivasio itself is an anagram of Savinio. The complexity of this is amplified even more if one considers that the name Savinio is, in its turn, not the anagram, but the pseudonym of Andrea De Chirico. This mechanism of transformation and concealment constantly reopens the problem oft he writer's autobiographical intentions. In the anagram, as in the pseudonym. Savinio tries to evade what we might call linear biography, so as to complicate or to metamorphose it by the use of linguistic means, or by the evocation of mythic figures such as the Hermaphrodite and the Argonaut. Yet one also constantly has the impression that all of Savinio's experiments are intended to communicate something other than mere deformation or surrealism; to communicate, rather, the peculiar progression of a subject that elaborates various self-representations on a number of different levels and according to diverse strategies. At the beginning of the novel. Nivasio conjures up the initial moment of every autobiography: that of his own birth. The complexity of the situation consists in the fact that Savinio immediately projects this event into a descriptive sphere in which all the external objects and circumstances surrounding the birth scene appear. The effect is that of estrangement: Il giorno in cui Nivasio usci dal grembo materno, ii sole picchiava a martello sulla citta della civetta. Cinque da una parte e cinque dall'altra, le lunghe steariche colorate sorgevano agli angoli del caminetto, si piegavano sui candelabri di bronzo, piangevano lunghi lacrimoni. La culla spumeggiava. in un angolo. Di minuto in minuto un rapido fruscio d'acqua rameggiava nei muri passava sulle finestre che opponevano le loro persiane chiuse all'assalto del caldo portentoso. (19) The very tears, rather than being attributed to the newborn baby, are produced by the steariche, or tallow candles. …
This paper proposes a new error-driven HMM-based text chunk tagger with context-dependent lexicon. Compared with standard HMM-based tagger, this tagger uses a new Hidden Markov Modelling approach which incorporates more contextual information into a lexical entry. Moreover, an error-driven learning approach is adopted to decrease the memory requirement by keeping only positive lexical entries and makes it possible to further incorporate more context-dependent lexical entries. Experiments show that this technique achieves overall precision and recall rates of 93.40% and 93.95% for all chunk types, 93.60% and 94.64% for noun phrases, and 94.64% and 94.75% for verb phrases when trained on PENN WSJ TreeBank section 00-19 and tested on section 20-24, while 25-fold validation experiments of PENN WSJ TreeBank show overall precision and recall rates of 96.40% and 96.47% for all chunk types, 96.49% and 96.99% for noun phrases, and 97.13% and 97.36% for verb phrases.
This paper describes the generation of iconic and categorical representations of word meaning, in propositional form, from the WordNet lexical database. These are derived from the list of synonyms, the descriptive gloss, and from the hypernym and meronym relations of each WordNet word sense. We demonstrate that these representations promote identification and discrimination, these being suggested qualities of representations of meaning, and finally suggest that these representations have further applications in language engineering. 1.
This paper presents our efforts to retrieve proper name information from a machine readable dictionary. We explain the problem in brief and then we discuss related work that other researchers have attempted in this area. We describe the embedded knowledge that one can find in the Collins English Dictionary (CED). Finally we present our ideas about how to extract this knowledge and how we plan to store it in a lexical database of proper names.
This paper describes the methodology that is being used to augment the Penn Treebank annotation with sense tags and other types of semantic information. Inspired by the results of SENSEVAL, and the high inter-annotator agreement that was achieved there, similar methods were used for a pilot study of 5000 words of running text from the Penn Treebank. Using the same techniques of allowing the annotators to discuss difficult tagging cases and to revise WordNet entries if necessary, comparable inter-annotator rates have been achieved. The criteria for determining appropriate revisions and ensuring clear sense distinctions are described. We are also using hand correction of automatic predicate argument structure information to provide additional thematic role labeling. 1.
We present some novel machine learning techniques for the identification of subcategorization information for verbs in Czech. We compare three different statistical techniques applied to this problem. We show how the learning algorithm can be used to discover previously unknown subcategorization frames from the Czech Prague Dependency Treebank. The algorithm can then be used to label dependents of a verb in the Czech treebank as either arguments or adjuncts. Using our techniques, we are able to achieve 88% precision on unseen parsed text.
The present study investigated the relationship between daily diary affect ratings and ambulatory cardiovascular activity in 117 male Vietnam combat veterans (61 with posttraumatic stress disorder [PTSD] and 56 without PTSD). Participants completed 12-14 hr of ambulatory monitoring and daily diary affect ratings. Compared with veterans without PTSD, veterans with PTSD reported higher negative affect and lower positive affect in daily diary ratings. No differences were detected for mean laboratory initial recordings or mean ambulatory heart rate (HR), systolic blood pressure (SBP), or diastolic blood pressure (DBP). However, compared with veterans without PTSD, veterans with PTSD demonstrated higher SBP and DBP variability and a higher proportion of HR activity (compared with initial recording values) during daily activity. There was a significant Time of Day x Group interaction for mean HR, with a trend for PTSD participants to maintain HR levels during evening hours.
The value of language resources is greatly enhanced if they share a common markup with an explicit minimal semantics. Achieving this goal for lexical databases is difficult, as large-scale resources can realistically only be obtained by up-translation from pre-existing dictionaries, each with its own proprietary structure. This paper describes the approach we have taken in the Concede project, which aims to develop compatible lexical databases for six Central and Eastern European languages. Starting with sample entries from original presentation-oriented electronic representations of dictionaries, we transformed the data into an intermediate TEI-compatible representation to provide a common baseline for evaluating and comparing the dictionaries. We then developed a more restrictive encoding, formalised as an XML DTD with a clearly-defined semantic interpretation. We present this DTD and discuss a sample conversion from TEI, together with an application which hyperlinks a HTML represent...
We present a method for automatically detecting errors in a manually marked corpus using anomaly detection. Anomaly detection is a method for determining which elements of a large data set do not conform to the whole. This method fits a probability distribution over the data and applies a statistical test to detect anomalous elements. In the corpus error detection problem, anomalous elements are typically marking errors. We present the results of applying this method to the tagged portion of the Penn Treebank corpus.
BOOK NOTICES 209 Linguistic databases. Ed. by John Nerbonne. (CSLI lecture notes 77.) Stanford, CA: CSLI, 1998. Pp. xxi, 243. The papers in this collection were originally presented at the 'Linguistic Databases' conference, University of Groningen, 23-24 March, 1995. Because ofthe almostproverbial rapidity with which information technology develops, the collection as a whole is dated already, but there is still much of interest to be found. Not all papers read at the conference are in this volume, but the papers cover a wide range of subjects, mostly practical in nature, not theoretical. After a clear and readable introduction by Nerbonne, the papers are presented in no particularorder, though the editor groups the papers in five main areas: syntactic corpora and databases, phonetic databases, applications in linguistic theory, applications, and extending basic technologies. The papers themselves are not presented according to this grouping, however, and at first sight the book appears rather disorganized. The wide variety of subjects can be deduced from the titles of the papers presented: 'Test suites for natural language processing', 'From annotated corpora to databases: The SgmlQL language', 'Markup of a test suite with SGML', 'An open systems approach for an acoustic-phonetic continuous speech database: The S_tools database-management system ', "The reading database of syllable structure',? database application for the generation of phonetic atlas maps', 'Swiss French polyphone and polyvar: Telephone speech databases to model inter- and intra-speaker variability', 'Investigating argument structure: The Russian nominalization database', "The use of a psycholinguistic database in the simplification of text for aphasie readers', "The computer learner corpus: A testbed for electronic EFL tools', 'Linking WordNet to a corpus query system', 'Multilingual data processing in the CELLAR environment '. The issue whether to use open free systems or closed proprietary systems is addressed in several papers. Some papers present applications developed both in open and closed systems. This is one area where developments have been going very fast, and nowadays freely available databases are often as capable as their commercial counterparts. Some of the applications presented in this collection are available from the Internet, and url's are often given. The collection can serve as a good introduction to the field for relative outsiders as ample references and links are given. The papers themselves vary greatly in subject matter so not all will be of interest to every reader. My particular favorite was 'From annotated corpora to databases: the SgmlQL language '. [BOUDEWUN REMPT.j Understanding phonology. By Carlos Gussenhoven and Haike Jacobs. (Understanding language series.) London: Arnold, 1998. Pp. xii, 286. This textbook is intended as an introduction to phonology aimed at 'students with little or no prior knowledge of linguistics' (back cover). As in many other textbooks, it uses exercises as a learning tool. Two types ofexercises are proposed. The ones identified by a key, 'intended as an expository aid' (xi), are provided with a solution in an appendix (though it is not always so much a clear cut answer as a guide for reflection, which is, to my view, a lot better). The ones identified by a dot are intended as practice material, and no solution is offered. I thought the idea of having two types of exercises a good one since it gives the reader the opportunity both for individual work and for discussion with others. Also, whenever it may apply, an optimality theoretic analysis is offered to describe a phonological process. Ch. 1, "The production of speech', is a basic introduction to phonology, phonetics, and phonation. Ch. 2,'Some typology: Sameness and difference', cleverly covers the universal and language specific aspects of phonological structures and typology. Ch. 3,'Making the form fit', addresses phonological grammar and adaptation by presenting the nativization of loan words in both the rules and the constraints approaches. Ch. 4, 'Underlying and surface representations', Ch. 5, 'Distinctive features', and Ch. 6, 'Ordered rules', deal with the basic notions of generative phonology within the SPE type formalism and introduce the reader to the school of linear phonology. Ch. 7,? case study: The diminutive suffix in Dutch', shows how these notions are applied. In Ch. 8, 'Levels of representation', Gussenhoven and Jacobs present an intermediate level of representation between the underlying representation and...
This paper discusses research on the English of Mexican Americans, arguing that in focusing primarily on description of vernacular Chicano English, the literature describing English spoken by Mexican Americans presents an incomplete picture of the complexity of their linguistic situation. Researchers may lose sight of the range of linguistic behaviors found within the Mexican American community, or even within a single family, where it is not uncommon to find fluent Spanish speakers, speakers with limited Spanish proficiency, and speakers of both nonstandard and standard dialects of English. The paper examines the range of linguistic behaviors found within three generations of a primarily English-dominant, middle class Mexican American family. It finds that even within this closeknit group of speakers, there exist distinct linguistic norms, ranging from those associated more closely with Chicano English to those associated with standard English. Even those speakers who make use of few if any linguistic resources associated with Chicano English or Spanish distinguish themselves linguistically from non-Mexican Americans. The paper considers the conscious and unconscious linguistic choices made by the speakers to be acts of identity, suggesting that the linguistic behaviors described are important means of constructing aspects of their social identities. (Contains 20 references.) (SM) Reproductions supplied by EDRS are the best that can be made from the original document. 1 RE-EXAMINING THE ENGLISH OF MEXICAN AMERICANS MS. AMANDA R. DORAN UNIVERSITY OF TEXAS AUSTIN, TEXAS PERMISSION TO REPRODUCE AND DISSEMINATE THIS MATERIAL HAS BEEN GRANTED BY dml TO THE EDUCATIONAL RESOURCES INFORMATION CENTER (ERIC) U.S. DEPARTMENT OF EDUCATION Office of Educational Research and Improvement ED E CATIONAL RESOURCES INFORMATION CENTER (ERIC) This document has been reproduced as received from the person or organization originating it. Minor changes have been made to improve reproduction quality. Points of view or opinions stated in this document do not necessarily represent official OERI position or policy.
In the sections 1-4 this paper presents a Norwegian lexical project trying to catch up with some shortcomings in relation to modern lexicography in adding a systematic semantic component to existing bases of morphology. The project also aims to reuse consisting lexicographical products and as a spin off to improve and systematise them. The project is still in its beginning, and have so far just started the definition of adjectives. Section 5 presents the adjectiv classification in WordNet and discuss some of the questions raised in WordNet presentations of adjectives. The section concludes with an alternative adjective classification built on Norwegian and German studies, which seems to answer some of these questions.
Computer-driven systems for constructing composite faces of suspects (E-fit; Mac-a-Mug) have largely replaced mechanical systems (Photofit; the Identikit) in police use, yet little is known of their comparative effectiveness in rendering an accurate likeness. Participants (N = 24) constructed 2 of 4 familiar or unfamiliar faces, for one of which they used Photofit and for the other, E-fit. A likeness of each face was made first under target-absent conditions and then with photographs of the target present. The accuracy of the resulting composites was assessed by familiarity ratings, names elicited, and matching accuracy. The computer-driven system showed consistent superiority only when a familiar face was constructed in the presence of photographs; when participants worked from memory, E-fit was no better than Photofit. The implications of these findings for theories of face retrieval and the operational use of composites are discussed.
It is often assumed that the emergence of a new concept can be traced to the appearance of a new word in a language. The aim of this article is to show that the process is more complex than this by examining the evolution of a number of words and contexts in the field of high altitude topography, when high mountains were being explored for the first time. This is a lexical field of a somewhat special nature, relating to a reality that had always been visible, but had never been described before. While there are undoubtedly examples where a concept crystallises around a newly adopted word, the norm is a slow process of uncertainty and hesitation.
The existence of a large body of literature in the Tyneside and Northumbrian dialects, dating from the late 18th century and continuing to the present day, testifies to a strong and enduring sense of regional identity closely associated with an acute sense of the differences between these dialects and Standard English/RP. Although much of this literature is conservative in nature and conservationist in intent, more recent examples in the local and popular press attempt to represent the salient features of the modern urban dialect (Geordie). This article examines extracts from a selection of texts, dating from George (Geordie) Ridley’s The Blaydon Races (1862) to cartoons in the Newcastle local newspaper, The Evening Chronicle(1996-7) and from Viz comic (1998). In the texts examined, semi-phonetic spelling is used to represent features of the Geordie accent. This article demonstrates that, whilst some features of the traditional dialect have been dropped by the more recent writers, others, such as the monophthongal/u:/in words such as ‘ town’, ‘brown’, spelt <toon, broon> are retained, notably in lexical items having referents which are closely bound up with local identity. Features found only in 20th-century texts often indicate very localized shibboleths, distinguishing Geordies from ‘Makkems’ (citizens of Sunderland, about 15 miles south-east of Newcastle). Recent sociolinguistic research points to a tendency for supra-local norms to replace the more traditional forms indicated by the spellings in dialect literature. This article argues that the prominence of local forms in dialect literature may represent an assertion of local identity in the face of the perceived threat of cultural and linguistic homogenization.
This research was designed to contribute to the development of new training systems to convey shape to the visually impaired. This dissertation consisted of two experiments which addressed the similarities and differences of visually impaired persons' and sighted persons' impressions of outlines of real objects perceived under auditory and touch display conditions, as well as tactile impressions of actual real objects, wood cutouts of those objects, tactor-pin diagrams of those objects, and raised-line drawings of those objects. Experiment 1 examined auditory and tactile perceptual equivalence, as well as sighted and visually impaired perceptual equivalence. Within each sight level group, Experiment 1 demonstrated a lack of equivalence for auditory and tactile perceptual structures. Experiment 1 demonstrated both similarities and differences in the perceptual structure comparisons between the sighted and the visually impaired. Experiment 2 examined the impact of sight level and four stimulus types on tactile percent correct identifications, response latencies, and familiarity ratings. Experiment 2 demonstrated sighted and visually impaired participants performed equally on a percent correct identification task. This Experiment also revealed the best performances were exhibited when real objects were presented, followed by wood cutouts, followed by tactor-pins, and performance with raised-line drawing (RLD) presentation was poorest. Experiment 2 also examined the exploratory procedures utilized during tactile identification of the four stimulus types. Experiment 2 revealed wood cutout presentation was associated with the largest percentage of different exploratory procedures being utilized, followed by real objects and tactor-pins which were equivalent, but utilized a larger percentage than the RLDs. Wood cutouts were the only other objects beside the real objects to use all exploratory procedures. The results of Experiment 1 and Experiment 2 indicate the utility of changing current industry standards for displaying shape information.
In this paper, we present an ongoing project of tagging early French dictionaries, dictionnaire Universel de Basnage de Bauval (1701) in order to build a lexical database which would enable fine-grained queries for linguistic and historical studies. The tagging model is SGML and as a tagset basis we chose to adopt the Text Encoding Initiative Guidelines for print dictionaries. We show that in spite of numerous inconsistencies, computerizing Le dictionnaire Universel is a feasible task and exhibits many regularities in the text which can be partially automated with finite state automata. Applying a systematic grammar on the text renews the metalexicographic studies of early dictionaries.
In this paper we present a whole Natural Language Processing (NLP) system for Spanish. The core of this system is the parser, which uses the grammatical formalism Lexical-Functional Grammars (LFG). Another important component of this system is the anaphora resolution module. To solve the anaphora, this module contains a method based on linguistic information (lexical, morphological, syntactic and semantic), structural information (anaphoric accessibility space in which the anaphor obtains the antecedent) and statistical information. This method is based on constraints and preferences and solves pronouns and definite descriptions. Moreover, this system fits dialogue and non-dialogue discourse features. The anaphora resolution module uses several resources, such as a lexical database (Spanish WordNet) to provide semantic information and a POS tagger providing the part of speech for each word and its root to make this resolution process easier. Keywords: Anaphora resolution, semantics, LFG grammar and parsing, EuroWordNet. 1.
Flogging a Dead Language: Identity Politics, Sex, and the Freak Reader in Acker’s Don Quixote Nicola Pitchford Pastiche is central to the resistant politics of Kathy Acker’s writing—yet she would appear to agree with Fredric Jameson’s influential critique of pastiche as “the wearing of a linguistic mask, speech in a dead language” (17). Her 1986 novel Don Quixote is all about having to speak “in a dead language” in the absence of a more “healthy” norm. It begins with the death of the protagonist, a female version of Cervantes’s knight, who then goes on to narrate much of the subsequent story. Acker explains, “BEING DEAD, DON QUIXOTE COULD NO LONGER SPEAK. BEING BORN INTO AND PART OF A MALE WORLD, SHE HAD NO SPEECH OF HER OWN. ALL SHE COULD DO WAS READ MALE TEXTS WHICH WEREN’T HERS” (39). The novel then proceeds by plagiarism and pastiche, as Quixote goes on a quest—for a heterosexual love unsullied by patriarchal power relations—through fragments of numerous existing texts. Quixote rereads and pieces together a whole range of textual scraps, from Machiavelli’s The Prince to a Godzilla movie. What becomes clear in her eccentric survey of (primarily) Western culture is that the lost, healthy linguistic norm is more than unhealthy for female readers—indeed, it is deadly. The novel is motivated by the idea of both reading and speaking “in a dead language”—but “flogging a dead language” seems a more apt description of Acker’s strategy, in more ways than one. For both the reader in and the reader of the novel, the act of rereading that pastiche entails can seem like flogging a dead horse, in the sense of merely covering once again the familiar ground of the already said. Of course, the same has been said of any reading in postmodernity where all language may well be dead, having belonged properly to a previous historical moment that gave it life and from which it has now been dissociated by forces of commercial appropriation and cultural amnesia. But this generic deadness that Jameson identifies as inherent in postmodern writing is not quite what I wish to explore. Rather, I want to attempt to account for what I see as a particular familiarity, and perhaps a particular tendency toward exhaustion and redundancy, that accompanies reading Acker’s texts from this period in her career, a period characterized by Acker’s extensive use of pastiche or what she frequently refers to as “plagiarism.” In what follows, I look at what happens to and through the act of reading, to ask how reading is connected to agency. Despite the considerable difficulty of Acker’s experimental novels, reading them can become an activity weighed down by a certain deadening obviousness. I want to suggest that this lifelessness derives from Acker’s attempts to construct, through pastiche, a community of readers defined by their opposition to traditional literary culture. I want also to argue that her deployment of pastiche in specific contexts—especially sexual contexts—in fact complicates and undermines the static and oversimplified role that she sometimes seems to offer her reader. In such moments, a complex interplay of various possible readerly identifications creates a contingent and particular version of agency. In a chapter on Acker in his recent book on literary celebrity, Joe Moran has suggested that in both her public persona and her work, Acker “puts forward two contrasting views of identity—one textual and one essentialist” (142). While he locates these competing versions in Acker’s characters and in her own public performance of the (death of the) author, I wish to extend his observations and apply them also to the modes of reading suggested by her texts, and by this novel in particular. The “textual” version of identity, generally celebrated by those critics friendly to Acker’s work, is readily apparent in the cut-and-paste technique of Don Quixote, in which borrowed textual fragments are reanimated by their juxtaposition. Here, language (my metaphorical dead horse), along with the social identities it produces, is like Quixote’s skinny nag Rocinante, who by all rights should be dead but who keeps lurching doggedly...