Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
In this article, I conduct a quantitative analysis of do absence in negative declaratives in the present tense in a dialect from the north-east of Scotland, Buckie. Analysis of nearly 800 contexts of use reveals that this variation is entirely conditioned by linguistic internal constraints. The most significant of these is person and number of the subject — 3rd person singular subjects and plural NPs have no do absence, while do is variable in the remaining pronouns. I argue that a syntactic explanation best accounts for this patterning of use. Where there is no overt - s inflection in the present tense (influenced by the “northern subject rule”), do is not obligatory in Buckie Scots. Frequency effects, lexical restrictions and processing constraints are called upon to account for the range of frequencies of do absence seen in the variable contexts. Lastly, there is no significant change in use of do across three generations of speakers, highlighting the community members’ relative immunity to prescriptive norms.
BOOK NOTICES 487 seriously interested in Salish studies as well as by local Washington libraries concerned with promoting traditional Lushootseed language and culture. [Edward T. Vajda, Western Washington University.] An ethnographic grammar of the Eipo language spoken in the central mountains of Irian Jaya (West New Guinea), Indonesia. By Volker Heeschen. (Mensch, Kultur und Umwelt im zentralen Bergland von West Neuguinea 23.) Berlin: Dietrich Reimer Verlag, 1998. Pp. 412. Eipo belongs to the Mek family, a group ofclosely related languages spoken in several mountain valleys between areas occupied by speakers of Dani and Ok languages. The author's latest contribution to the ethnolinguistics of this remote area, this large-format paperbackjoins several previous volumes in the same series devoted to Eipo, including: Wörterbuch EipoDeutsch -English (Volker Heeschen, 1983, vol. 6,), with 5,682 main entries, and Kommunikation bei den Eipo (Volker Heeschen, 1989, vol. 19), a description of communicative styles and language change in a small speech community. The present work provides the first extensive description ofEipo phonology and grammar. Heeschen gathered his voluminous data during more than a dozen field trips made since 1977. The book is 'ethnographic' in the sense that H links his descriptions to specific cultural and pragmatic contexts. Rather than attempting to portray Eipo as conforming to a fixed norm, H describes the rules creating grammatical forms in Eipo, a language with about 400 speakers (22-23), as relatively fluid when compared to languages spokenby larger, more extensive populations. The book consists of three parts divided into several chapters each. Part 1 introduces Eipo culture and history (13-35) and discusses Eipo's position within the Mek language family (72-94). H identifies Mek as a low-level genetic grouping similar to the several dozen posited by William A. Foley (The Papuan languages ofNew Guinea, Cambridge: Cambridge University Press, 1986). However, H uses typological and lexical data to support Wurm's classification of Mek within the Trans-New Guinea phylum (Stephen A. Wurm, Papuan languages ofOceania, Tübingen: Gunter Narr, 1982), though the evidence suggests a very distant connection. In a somewhat rambling fashion, H discusses a medley of approaches used by previous scholars to describe 'exotic' languages (36-71) and selects elements from a variety of traditions for their relevance in describing Eipo. H justifies this eclecticism based on his own observations regarding linguistic selfawareness and language creation among the Eipo (95-114). Part 2 provides a meticulous, data- rather than theory-driven description of Eipo phonetics and phonology (115-40), word classes and morphosyntax (141-264), and syntax (265-356). Each sectioncontains numerous paradigms and other grammatical schemata, plus legions of example phrases and sentences interpreted from the vantage of the author's keen understanding of the ambient cultural context, a factor which if omitted would render the literal translations of many examples unintelligible. Finally, Part 3 (357-80) provides nine previously unpublished texts in Eipo and neighboring Mek languages. These texts deal with local myths and legends and are accompanied by interlinear glosses, a translation into idiomatic English, and copious ethnographic explanations. H's work represents a solid, multifaceted contribution to the study of New Guinea ethnography and linguistics and contains much that will be of interest to general typologists. Because Mek languages were until recently more poorly described than neighboring groups, this book contributes important data to the ongoing task ofestablishing genetic relationships between New Guinea's several hundred languages. Finally, H's conclusions regarding rates of vocabulary change in a language not known to have ever counted more than a few hundred speakers, with its typical absence of any fixed conservative norm, may have implications for broader studies of language contact and genetic linguistics. [Edward J. Vajda, Western Washington University.] People, countries, and the Rainbow Serpent: Systems of classification among the Lardil of Mornington Island. By David McKnight. (Oxford studies in anthropological linguistics 12.) New York & Oxford: Oxford University Press, 1999. Pp. x, 270. This book reflects over five years of field work conducted at intervals beginning in 1966 and contains a treasure trove of data on Lardil language and culture that would almost certainly have otherwise disappeared unrecorded. McKnight elicited information from his native speaker informants in a...
We present a fuzzy logic based approach to the derivation of linguistic summaries of sets of data (databases), and show that it may be viewed as an example of a new paradigm shift from computing on numbers to computing on words that has been recently strongly advocated by Zadeh. We present an implementation of linguistic database summaries for sales data of a computer retailer that clearly shows that the new approach is viable and yields a new quality by providing human consistent results.
A Classification Information Model is a pattern classification model.The model decides the proper class of an input instance by integrating individual decisions, each of which is made with each feature in the pattern.Each individual decision is weighted according to the distributional property of the feature deriving the decision. An individual decision and its weight are represented as classification information which is extracted from the training instances.In the word sense disambiguation based on the model, the proper sense of an input instance is determined by the weighted sum of whole individual decisions derived from the features contained in the instance.
Familiarity for Nouns and Verbs: Not the Same as, and Better than, Frequency Natalie Kacinik (kacinn01@student.ucr.edu) Connie Shears (shearc01@student.ucr.edu) Christine Chiarello (christine.chiarello@ucr.edu) University of California, Riverside Life Sciences Psychology Building Riverside, CA 92521, U.S.A Much of the study of language processing has centered on the single word recognition paradigm. Many cognitive theories regarding semantic memory have emerged from this research, and several lexical dimensions (i.e., imageability, frequency) have been found to influence the processing and recognition of nouns (Balota, Ferraro, and Connor, 1991). Indeed, efforts are typically made to balance word stimuli on factors such as length and frequency. However, the importance of different orthographic and semantic dimensions in determining the speed and accuracy with which words are responded to has not been extensively investigated. Moreover, most of this research has either not considered different word types (nouns vs. verbs), or has focused on concrete, imageable nouns, largely because of the lack of word norming corpora available for other word types. Recently, new measures have been developed (Chiarello, Shears, & Lund, 1999) computing typicality of grammatical class (noun vs verb) and examining grammatical class differences in imageability and frequency, using established corpora such as Francis and Kucera (FK, 1982), as well as using the more contemporary Usenet corpus. While these semantic dimensions and word class comparisons have provided valuable tools for word recognition researchers, most studies have failed to consider word familiarity as an important determinant of speed and accuracy of responding (but see Gernsbacher, 1984, and Balota, Cortese, & Pilotti, We report a series of regression analyses using data obtained from 2 lexical decision experiments and other corpora. We investigated the influence of variables identified in Chiarello et al. (1999) [i.e., imageability, length, noun- verb distributional distance (NVDD), FK and Usenet frequency, and recently collected familiarity ratings] on the speed and accuracy of lexical decision responses to nouns and verbs. Familiarity, measured on a 7 pt. scale, was defined as ‘common in everyday experience’. Overall, familiarity was found to be highly correlated with RT ( r = -.70, p<.001), thereby accounting for nearly half of the variance. Although significantly correlated with imageability, NVDD, FK and Usenet frequency ( r =.22,.23,.39, and.40, all ps<.005), regression analyses indicated that much of the RT variance accounted for by familiarity was unique. The importance of these variables in predicting RT also varied by word class (nouns vs verbs). Specifically, familiarity, then frequency, and then imageability were found to be the most important predictors of noun RT, whereas familiarity, then imageability, then frequency, and finally NVDD were found to be the most important predictors of verb RT. In conclusion, our results support and extend Gernsbacher’s (1984) earlier demonstration of familiarity as a powerful contributor to word recognition, possibly because it is a contemporary metric of actual encounters, related to the variety of contexts a word has been experienced in, and the ease with which individuals can recall those contexts (Audet & Burgess, 1999). Our findings indicate the need for researchers to consider the importance of processing differences based on the familiarity of stimuli to the subject population (i.e., controlling only for frequency and imageability may not be enough). Finally, we also demonstrate the need for researchers to carefully consider the issue of word class, as different dimensions appear to be more or less important for the processing of nouns and verbs. References Audet, C., & Burgess, C. (1999). Using a high- dimensional memory model to evaluate the properties of abstract and concrete words. In M. Hahn, & S.C. Stoness (Eds.), Proceedings of the Cognitive Science Society. Mahwah, NJ: Erlbaum. Balota, D.A., Cortese, M.J., & Pilotti, M. (1999). Item- level analyses of lexical decision performance: Results from a mega-study. In Abstracts of the 40th Annual Meeting of the Psychonomics Society. Los Angeles, CA: Psychonomic Society. Balota, D., Ferraro, R., & Connor, L. (1991). On the early influence of meaning in word recognition: A review of the literature. In P.J. Schwanenflugel (E.d. ), The psychology of word meanings. Hillsdale, NJ: Erlbaum. Chiarello, C., Shears, C., & Lund, K. (1999). Imageability and distributional typicality measures of nouns and verbs in contemporary English. Behavior Research Methods, Instruments, & Computers, 31, 603-637. Francis, W., & Kucera, H. (1982). Frequency analysis of English usage: Lexicon and grammar. Boston: Houghton Mifflin. Gernsbacher, M.A. (1984). Resolving 20 years of inconsistent interactions between lexical familiarity and orthography, concreteness, and polysemy. Journal of Experimental Psychology: General, 113, 256-281.
A word sense disambiguation system which is going to be used aspart of a NLP system needs to be large scale, able to beoptimised towards a specific task and above all accurate. This paperdescribes the knowledge sources used in a disambiguation system able toachieve all three of these criteria. It is a hybrid system combining sub-symbolic, stochastic and rule-based learning. The paper reportsthe results achieved in Senseval and analyses them to show the system'sstrengths and weaknesses relative to other similar systems.
An architecture for federating heterogeneousdictionary databases is described. It proposes acommon description language and query language toprovide for the exchange of information betweendatabases with different organizations, on differentplatforms and in different DBMSs. The common querylanguage has an SQL like structure. The first versionof the description language follows the TEI standardtag definitions for dictionaries with the expectationthat the description language will be expanded in thefuture. A practical implementation of the proposalsusing WWW technology for two multi-lingualdictionaries is described.
This paper describes the evaluation of a WSD method withinSENSEVAL. This method is based on Semantic Classification Trees (SCTs)and short context dependencies between nouns and verbs. The trainingprocedure creates a binary tree for each word to be disambiguated. SCTsare easy to implement and yield some promising results. The integrationof linguistic knowledge could lead to substantial improvement.
This article reports the results of apreliminary analysis of translation equivalents infour languages from different language families,extracted from an on-line parallel corpus of GeorgeOrwell's Nineteen Eighty-Four. The goal ofthe study is to determine the degree to whichtranslation equivalents for different meanings of apolysemous word in English are lexicalized differentlyacross a variety of languages, and to determinewhether this information can be used to structure orcreate a set of sense distinctions useful in naturallanguage processing applications. A coherenceindex is computed that measures the tendency fordifferent senses of the same English word to belexicalized differently, and from this data aclustering algorithm is used to create sensehierarchies.
We describe a simple approach to word sensedisambiguation using information filtering andextraction. The method fully exploits and extends theinformation available in the Hector dictionary. Thealgorithm proceeds by the application of severalfilters to prune the candidate set of word sensesreturning the most frequent if more than one remains.The experimental methodology and its implication arealso discussed.
This work combines a set of available techniques – whichcould be further extended – to perform noun sense disambiguation. We use several unsupervised techniques (Rigau et al., 1997) that draw knowledge from a variety of sources. In addition, we also apply a supervised technique in order to show that supervised and unsupervised methods can be combined to obtain better results. This paper tries to prove that using an appropriate method to combine those heuristics we can disambiguate words in free running text with reasonable precision.
SENSEVAL set itself the task of evaluating automaticword sense disambiguation programs (see Kilgarriff andRosenzweig, this volume, for an overview of theframework and results). In order to do this, it wasnecessary to provide a `gold standard' dataset of `correct' answers. This paper will describe thelexicographic part of the process involved in creatingthat dataset. The primary objective was for a group oflexicographers to manually examine keywords in a largenumber of corpus contexts, and assign to each contexta sense-tag for the keyword, taken from the Hectordictionary. Corpus contexts also had to be manuallypart-of-speech (POS) tagged. Various observationsmade and insights gained by the lexicographers duringthis process will be presented, including a critiqueof the resources and the methodology.
Wisdom is a system for performing word sense disambiguation (WSD)using a limited number of linguistic features and a simplesupervised learning algorithm. The most likely sense tag for aword is determined by calculating co-occurrence statistics forwords appearing within a small window. This paper gives abrief description of the components in the Wisdom system and thealgorithm used to predict the correct sense tag. Some results forWisdom from the Senseval competition are presented, and directionsfor future work are also explored.
In this paper we present some observations concerning an experiment of (manual/automatic) semantic tagging of a small Italian corpus performed within the framework of the SENSEVAL/ROMANSEVAL initiative. Themain goal of the initiative was to set up a framework for evaluation of Word Sense Disambiguation systems (WSDS) through the comparative analysis of their performance on the same type of data. In this experiment there are two aspects which are of relevance: first, the preparation of the reference annotated corpus, and, second, the evaluation of the systems against it. In both aspects we are mainly interested here in the analysis of the linguistic side which can lead to a better understanding of the problem of semantic annotation of a corpus, be itmanual or automatic annotation. In particular, we will investigate, firstly, the reasons for disagreement between human annotators, secondly, some linguistically relevant aspects of the performance of the Italian WSDS and, finally, the lessons learned from the present experiment.
Age of acquisition (AoA) has been reported to be a predictor of the speed of reading words aloud (word naming) and lexical decision, with early-acquired words being responded to faster than later-acquired words in both tasks. All previous studies of AoA effects have, however, relied upon adult estimates of word learning age the validity of which it is easy to cast doubt upon. Using objective age of acquisition norms derived from children's naming data, this study shows that AoA effects do not depend upon the use of adult ratings. In addition to effects of real AoA, influences of word frequency and orthographic neighbourhood size were obtained in both word naming and lexical decision. Imageability affected lexical decision but not word naming, while the characteristics of the word's initial phoneme affected word naming but not lexical decision.
In developing the concepts of phonological and lexical subtypes of dyslexia, criteria have been proposed based on the projection of linear regression for raw test scores. Substantial discrepancies in subtype prevalence may arise from nonlinearities in the test scores. A norming process developed for a new nonword test, the Martin and Pratt Nonword Reading Test (Martin & Pratt, 2000), was applied to the Word Identification subtest of the Woodcock Reading Mastery Test (Woodcock, 1987) and the Regular and Irregular Word Tests published in Coltheart and Leahy (1996). These tests were administered to a representative sample of 863 children aged 6 to 15 years in the Southern Tasmanian State School population. An inverse normal transform results in a distribution which is approximately normal within age groups. On this scale the age effect was well approximated by a linear increase with the logarithm of (age −5 years). This process can be adapted to provide norms for word lists more economically and allows convenient spreadsheet formulae for norms. Substantial differences from other norms may be attributed to school district family income differences found in this sample. Male means are lower than female for all tests, but this reflects comparable high performance and disproportionate poor performance by males on reading tests.
Four experiments were conducted to determine whether the Hyperspace Analogue to Language (HAL) model of semantic memory could differentiate between two different populations. An analysis of the differences in densities (or average distances between word neighbors in semantic space) in HAL matrices—generated from text corpora derived from younger and older adults—confirmed that HAL was able to distinguish between the two age groups. This difference was again detected when structured interview data were used to build the corpora. A third experiment, designed to test the specificity of HAL in detecting differences between groups, did not detect any difference in the densities of the memory representations when older adults generated both the test corpora. The final experiment, conducted on the language of adults with Alzheimer’s and normal adults, again demonstrated that HAL could discriminate between the two populations. These results suggest that HAL is capable of modeling, on the basis of changes in mean density, some of the differences between populations without modifying the model itself but, rather, by changing the text corpus from which the model creates its representations in semantic space.
Language deviation is a kind of langUage form diverging from the language norm, In poetry, there are eight kinds of language deviation: lexical deviation.Phonological deviation. grammatical deviation. graphological deviation. semantic deviation.deviation of register. deviation of historical period and dialectal deviation. From the point of view of the aesthetic function, we analyze the language deviation of poetry in outer to appreciate the poems better.
This paper describes the grling-sdm system, which is asupervised probabilistic classifier that participated in the 1998SENSEVAL competition for word-sense disambiguation. This systemuses model search to select decomposable probability models describingthe dependencies among the feature variables.These types of models have been found to be advantageous in terms ofefficiency and representational power. Performance on the SENSEVALevaluation data is discussed.
The ability to circumlocute successfully is of utmost importance in compensating for gaps in lexical knowledge. Although all studies indicate that one’s ability to circumlocute increases with increasing proficiency, it is interesting that little attention has been paid to those learners who have the greatest ability to circumlocute, native‐like speakers. This study addresses the norms of native and native‐like circumlocution. It expands the discussion of strategies involved in this skill to include the means by which speakers frame their message and thereby set the linguistic context for their listeners. Participants in this study, both native and native‐like speakers, were found to employ similar strategies while circumlocuting, including the use of synonyms, analogies, and descriptions. These participants also consistently framed their speech to facilitate listener comprehension, and they frequently included in their discourse some reference to their status as a nonexpert in the field. Similarities in native and native‐like circumlocution found in this study help to provide some empirical validation to the notion of “native‐like.”
We propose the use of the bootstrap resampling technique as a tool to assess the within-subject reliability of experimental modulation effects on event-related potentials (ERPs). The assessment of the within-subject reliability is relevant in all those cases when the subject score is obtained by some estimation procedure, such as averaging. In these cases, possible deviations from the assumptions on which the estimation procedure relies may lead to severely biased results and, consequently, to incorrect functional inferences. In this study, we applied bootstrap analysis to data from an experiment aimed at investigating the relationship between ERPs and memory processes. ERPs were recorded from two groups of subjects engaged in a recognition memory task. During the study phase, subjects in Group A were required to make an orthographic judgment on 160 visually presented words, whereas subjects in Group B were only required to pay attention to the words. During the test phase all subjects were presented with the 160 previously studied words along with 160 new words and were required to decide whether the current word was “old” or “new.” To assess the effect of word imagery value, half of the words had a high imagery value and half a low imagery value. Analyses of variance performed on ERPs showed that an imagery-induced modulation of the old/new effect was evident only for subjects who were not engaged in the orthographic task during the study phase. This result supports the hypothesis that this modulation is due to some aspect of the recognition memory process and not to the stimulus encoding operations that occur during the recognition memory task. However, bootstrap analysis on the same data showed that the old/new effect on ERPs was not reliable for all the subjects. This result suggests that only a cautious inference can be made from these data.
This paper specifically addresses the question of polysemy with respect toverbs, and whether or not the sense distinctions that are made in on-linelexical resources such as WordNet are appropriate for computational lexicons.The use of sets of related syntactic frames and verb classes are examined as ameans of simplifying the task of defining different senses, and the importanceof concrete criteria such as different predicate argument structures, semanticclass constraints and lexical co-occurrences is emphasized.
Web-accessible conferencing softwareand ``conversational ethics'' drawn from Habermas andRawls have successfully brought together on-lineparticipants separated by geography and viewpoint, andoccasionally resulted in consensus regarding otherwisedivisive issues such as abortion. The author describessuccesses, limitations, and costs of incorporatingthese technologies and discourse ethics in a religiousstudies class. Results are striking, but thepedagogical benefits involve technical risks and highlabor and time costs. This experience, coupled withrecent research, suggests that electronic pedagogies,like other teaching strategies, work for some, but notall students: this argues that we take up electronicteaching as one approach among many.
Senseval was the first open, community-based evaluation exercise for WordSense Disambiguation programs. It took place in the summer of 1998,with tasks for English, French and Italian. There were participating systems from 23 researchgroups. This special issueis an account of the exercise. In addition to describing the contentsof the volume, this introduction considers how the exercise has shedlight on some general questions about wordsenses and evaluation.
The paper examines the task of Word Sense Disambiguation (WSD) criticallyand compares it with Part of Speech (POS) tagging, arguing that the abilityof a writer to create new senses distinguishes the tasks and makes it moreproblematic to test WSD by the mark-up-and-model paradigm, because newsenses cannot be marked up against dictionaries. This serves to set WSDapart and puts limits on its effectiveness as an independent NLP task.Moreover, it is argued that current WSD methods based on very small wordsamples are also potentially misleading because they may or may not scaleup. Since all-word WSD methods are now available and are producing figurescomparable to the smaller scale tasks, it is argued that we shouldconcentrate on the former and find ways of bootstrapping test materialsfor such tests in the future.
An algorithm for analyzing ordinal scaling results is described. Frequency data on ordinal categories are modeled for unidimensional psychological attributes according to Thurstone’s judgment scaling model. The algorithm applies maximum likelihood estimation of model parameters. The Cramér-Rao bounds of the standard errors of the estimated parameters are calculated, and a stress measure and a goodness-of-fit measure are supplied.
Two studies examined the relationship between self-monitoring and factors influencing romantic attraction to others. In Study 1, participants completed an Internet-mediated version of the Self-Monitoring Scale (Gangestad & Snyder, 1985) and indicated which of two people (one physically attractive, one with a more desirable personality) they found most attractive. Results matched previous findings (Snyder, Berscheid, & Glick, 1985), but the effect was smaller. Study 2, a paper-and-pencil replication of Study 1, examined whether the weaker effect was due to Internet mediation and found no differences in the choices made by high and low self-monitors. Results suggested that while determinants of attraction may vary for different populations, Internet research methods can tap the same phenomena as traditional laboratory studies.
Three studies focused on the development and enhancement of narrative skills within a preschool classroom. The purpose of Study 1 was to collect local norms on narrative development. Fifty-two preschool African American English speakers representing 3-, 4-, and 5-year-old age groups, narrated a familiar storybook. Some children in each age group evidenced use of nine story element types. Developmental changes were characterized by growth in types as well as tokens of story elements. Study 2 demonstrated that preschoolers’ narratives can be influenced by the narratives of their peers. Paired children narrated a familiar storybook to each other. The stories of paired children were significantly more similar in form (shared story element types) and content (shared lexical types) than those of unpaired children. Study 3 provided a preliminary test of an intervention designed to exploit the effect of peer models for long-term gain in narrative abilities. Two tutees practiced book narration following the clinician-prompted models of their peer tutors. As a result, the tutees demonstrated an expanded repertoire of story elements and an increased frequency of use of story element types in both trained and untrained stories. Their rate of growth in story element use was superior to that of their classmates who had not participated in the intervention. The benefit of peers for achieving instructional congruence in cases of clinicianclient mismatch is emphasized.
QUAID (question-understanding aid) is a software tool that assists survey methodologists, social scientists, and designers of questionnaires in improving the wording, syntax, and semantics of questions. The tool identifies potential problems that respondents might have in comprehending the meaning of questions on questionnaires. These problems can be scrutinized by researchers when they revise questions to improve question comprehension and, thereby, enhance the reliability and validity of answers. QUAID was designed to identify nine classes of problems, but only five of these problems are addressed in this article: unfamiliar technical term, vague or imprecise relative term, vague or ambiguous noun phrase, complex syntax, and working memory overload. We compared the output of QUAID with ratings of language experts who evaluated a corpus of questions on the five classes of problems. The corpus consisted of 505 questions on 11 surveys developed by the U.S. Census Bureau. Analyses of hit rates, false alarm rates,d′ scores, recall scores, and precision scores revealed that QUAID was able to identify these five problems with questions, although improvements in QUAID’s performance are anticipated in future research and development.
Bagging and boosting, two effective machine learning techniques, are applied to natural language parsing. Experiments using these techniques with a trainable statistical parser are described. The best resulting system provides roughly as large of a gain in F-measure as doubling the corpus size. Error analysis of the result of the boosting technique reveals some inconsistent annotations in the Penn Treebank, suggesting a semi-automatic method for finding inconsistent treebank annotations.
The Verbmobil treebanks of spoken German, English, and Japanese are part of the Verbmobil project, which has the overriding goal to develop a speaker-independent system for the translation of spontaneous speech. In the framework of this language technology project, the treebanks provide training data for a variety of language technology modules. The treebanks consist of annotated syntactic tree structures based on transcribed dialogs in the scenarios of appointment negotiations, travel arrangements, and personal computer maintenance. The annotation schemes of the treebanks have been developed taking into account the specific characteristics of spoken language dialogs: repetitions, hesitations, false starts'', etc. * The work reported here was funded by the German Ministry of Education and Research (BMBF) in the framework of the Verbmobil project under grant FKZ:01 IV 701 M0.
1.1 Notion of word....................................... 4 1.2 Tests of wordhood..................................... 5 1.3 Compatibility with other guidelines............................ 6
Computers are now widely used in the preparation of dictionaries. There are many advantages in maintaining and updating a dictionary in electronic form, most obviously that printed versions can be typeset directly from the electronic copy. But more than that, electronic dictionaries are beginning to be used by computers in retrieval systems. This chapter looks at electronic dictionaries and examines how lexical databases can help to refine and improve retrieval and analysis programs. It also traces the development of the uses of computers and dictionaries, and assesses various types of resources. Much research still needs to be done on the structure and contents of lexical and linguistic databases, especially for the semantic component, but the examples discussed in this chapter give some idea of the potential.
This paper describes a specific part of the Prague Dependency Treebank annotation, the step from the surface dependency structure towards the underlying representation of the sentence. The first section explains the theoretical basis of the project. In Section 2 all the procedure of conversion to the tectogrammatical structure is summarized and Section 3 presents in detail the present stage of the automated part of the conversion procedure.
This paper describes the design criteria and annotation guidelines of Sinica Treebank. The three design criteria are: Maximal Resource Sharing, Minimal Structural Complexity, and Optimal Semantic Information. One of the important design decisions following these criteria is the encoding of thematic role information. An on-line interface facilitating empirical studies of Chinese phrase structure is also described.
In this paper, we propose a new ambiguity representation scheme; Structure Preference Relation (SPR), which consists of useful quantitative distribution information for ambiguous structures. Two automatic acquisition algorithms, the first acquired from a treebank, and the second acquired from raw texts, are introduced, and some experimental results which prove the availability of the algorithms are also given. Finally, we introduce some SPR applications in linguistics and natural language processing, such as preference-based parsing and the discovery of representative ambiguous structures, and propose some future research directions.
Thesauri have been widely used in bibliographic databases for 30 years. Recently, CD-ROMs of a variety of dictionaries with their GUIs are spreaded to current users. On the other hands, End users dose not use thesauri as for their bulky printed matter. The kind of software browsing graphically and managing thesauri does not appear in PC environment. The browsing tool for lexical database with hierarchical structure has been developed using Java.. This paper describes the functions, the components, and examples of its usage. The problems of the browser and the functions to be extended are discussed.
We aim at finding the minimal set of fragments which achieves maximal parse accuracy in Data Oriented Parsing. Experiments with the Penn Wall Street Journal treebank show that counts of almost arbitrary fragments within parse trees are important, leading to improved parse accuracy over previous models tested on this treebank. We isolate a number of dependency relations which previous models neglect but which contribute to higher parse accuracy.
1.1 Tagging criteria....................................... 4 1.2 POS tagset......................................... 5 1.3 Size of the POS tagset................................... 6
The accuracy of statistical parsing models can be improved with the use of lexical information. Statistical parsing using Lexicalized tree adjoining grammar (LTAG), a kind of lexicalized grammar, has remained relatively unexplored. We believe that is largely in part due to the absence of large corpora accurately bracketed in terms of a perspicuous yet broad coverage LTAG. Our work attempts to alleviate this difficulty. We extract different LTAGs from the Penn Treebank. We show that certain strategies yield an improved extracted LTAG in terms of compactness, broad coverage, and supertagging accuracy. Furthermore, we perform a preliminary investigation in smoothing these grammars by means of an external linguistic resource, namely, the tree families of an XTAG grammar, a hand built grammar of English.