Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Web-accessible conferencing softwareand ``conversational ethics'' drawn from Habermas andRawls have successfully brought together on-lineparticipants separated by geography and viewpoint, andoccasionally resulted in consensus regarding otherwisedivisive issues such as abortion. The author describessuccesses, limitations, and costs of incorporatingthese technologies and discourse ethics in a religiousstudies class. Results are striking, but thepedagogical benefits involve technical risks and highlabor and time costs. This experience, coupled withrecent research, suggests that electronic pedagogies,like other teaching strategies, work for some, but notall students: this argues that we take up electronicteaching as one approach among many.
Senseval was the first open, community-based evaluation exercise for WordSense Disambiguation programs. It took place in the summer of 1998,with tasks for English, French and Italian. There were participating systems from 23 researchgroups. This special issueis an account of the exercise. In addition to describing the contentsof the volume, this introduction considers how the exercise has shedlight on some general questions about wordsenses and evaluation.
The paper examines the task of Word Sense Disambiguation (WSD) criticallyand compares it with Part of Speech (POS) tagging, arguing that the abilityof a writer to create new senses distinguishes the tasks and makes it moreproblematic to test WSD by the mark-up-and-model paradigm, because newsenses cannot be marked up against dictionaries. This serves to set WSDapart and puts limits on its effectiveness as an independent NLP task.Moreover, it is argued that current WSD methods based on very small wordsamples are also potentially misleading because they may or may not scaleup. Since all-word WSD methods are now available and are producing figurescomparable to the smaller scale tasks, it is argued that we shouldconcentrate on the former and find ways of bootstrapping test materialsfor such tests in the future.
Two studies examined the relationship between self-monitoring and factors influencing romantic attraction to others. In Study 1, participants completed an Internet-mediated version of the Self-Monitoring Scale (Gangestad & Snyder, 1985) and indicated which of two people (one physically attractive, one with a more desirable personality) they found most attractive. Results matched previous findings (Snyder, Berscheid, & Glick, 1985), but the effect was smaller. Study 2, a paper-and-pencil replication of Study 1, examined whether the weaker effect was due to Internet mediation and found no differences in the choices made by high and low self-monitors. Results suggested that while determinants of attraction may vary for different populations, Internet research methods can tap the same phenomena as traditional laboratory studies.
International audience
Three state-of-the-art statistical parsers are combined to produce more accurate parses, as well as new bounds on achievable Treebank parsing accuracy. Two general approaches are presented and two combination techniques are described for each approach. Both parametric and non-parametric models are explored. The resulting parsers surpass the best previously published performance results for the Penn Treebank.
The Verbmobil treebanks of spoken German, English, and Japanese are part of the Verbmobil project, which has the overriding goal to develop a speaker-independent system for the translation of spontaneous speech. In the framework of this language technology project, the treebanks provide training data for a variety of language technology modules. The treebanks consist of annotated syntactic tree structures based on transcribed dialogs in the scenarios of appointment negotiations, travel arrangements, and personal computer maintenance. The annotation schemes of the treebanks have been developed taking into account the specific characteristics of spoken language dialogs: repetitions, hesitations, false starts'', etc. * The work reported here was funded by the German Ministry of Education and Research (BMBF) in the framework of the Verbmobil project under grant FKZ:01 IV 701 M0.
No language in the world is homogeneous, or ever will be. Whereas earlier forms of English were characterised by extreme variation on all levels and Middle English is in fact best described as a loose conglomerate of unstable varieties, we usually lack any more detailed insight into what functions this variation had for the individual speaker. The social correlates so well known from modern sociolinguistics, such as age, sex, education, religion, can normally not be applied to the existing texts, nor can even the geographical range of recorded forms be determined with any degree of certainty. Finally, if modern dialect or other non-standard features are contrasted with (as the term non-standard implies) an accepted standard form of a language, this method would necessarily fail with Middle English even if we knew more about it than we do and, in view of the state of surviving documents, ever will. It is safe to assume that for its speakers the linguistic heterogeneity of Middle English was ordered in some way, but it was so only for continually shifting speech communities, whose number and individual geographical spread we know very little about. The scene changed dramatically in the fifteenth century: the emergence of a new standard language began to re-institute a linguistic norm for written supraregional English. This development was a natural consequence of the acceptance of English in public domains, and was speeded up by the change-over to English as the Chancery language in 1430.
We present a fuzzy logic based approach to the derivation of linguistic summaries of sets of data (databases), and show that it may be viewed as an example of a new paradigm shift from computing on numbers to computing on words that has been recently strongly advocated by Zadeh. We present an implementation of linguistic database summaries for sales data of a computer retailer that clearly shows that the new approach is viable and yields a new quality by providing human consistent results.
We discuss the advantages of lexicalized tree-adjoining grammar as an alternative to lexicalized PCFG for statistical parsing, describing the induction of a probabilistic LTAG model from the Penn Treebank and evaluating its parsing performance. We find that this induction method is an improvement over the EM-based method of (Hwa, 1998), and that the induced model yields results comparable to lexicalized PCFG.
Part 1 The lexical database: nouns in WordNet, George A. Miller modifiers in WordNet, Katherine J. Miller a semantic network of English verbs, Christiane Fellbaum design and implementation of the WordNet lexical database and searching software, Randee I. Tengi. Part 2: automated discovery of WordNet relations, Marti A. Hearst representing verb alterations in WordNet, Karen T. Kohl et al the formalization of WordNet by methods of relational concept analysis, Uta E. Priss. Part 3 Applications of WordNet: building semantic concordances, Shari Landes et al performance and confidence in a semantic annotation task, Christiane Fellbaum et al WordNet and class-based probabilities, Philip Resnik combining local context and WordNet similarity for word sense identification, Claudia Leacock and Martin Chodorow using WordNet for text retrieval, Ellen M. Voorhees lexical chains as representations of context for the detection and correction of malapropisms, Graeme Hirst and David St-Onge temporal indexing through lexical chaining, Reem Al-Halimi and Rick Kazman COLOR-X - using knowledge from WordNet for conceptual modelling, J.F.M. Burg and R.P. van de Riet knowledge processing on an extended WordNet, Sanda M. Harabagiu and Dan I Moldovan appendix - obtaining and using WordNet.
In this contribution we discuss how a fuzzy querying interface can support the generation of linguistic database summaries - a special technique of data mining. Links between our approach to linguistic summaries and the well-known technique of association rules is shown. The implementation of linguistic summaries generation using the authors’ FQUERY for Access package is presented.
Shoebox Template Materials for Gwich'in Dictionary. Materials compiled to aid in the creation of a future Gwich'in Dictionary. Includes Gwich'in Dictionary Database; Templates (Nouns, verbs, cross-references); Gwich'in Listening Exercise Dipthongs; Lexical Database (SIL's Shoebox) Template; Holikachuk examples, various other materials. ca.84pp Comments: 2000 is an estimated date.
The ambiguity related to the use of movement and localisation declension cases in Basque is a serious problem in the morphological generation phase in Machine Translation. We present the approach we have developed to solve this ambiguity. Information about the [±animate] semantic feature of the lemma we want to decline is necessary to choose the appropriate suffix. The lexical database used does not contain such information. Besides, it would be very hard to add it manually to the 28.000 substantives contained in it. For this reason, we made two experiments to obtain the [±animate] feature from other resources. First, our aim was to get automatically the required knowledge from corpora, but the results were not good. Secondly, after a minimal manual tagging and using semantic relations between words extracted from definitions of a monolingual dictionary, we tagged more than half of the words in real texts with the [±animate] feature.
In this paper, we present a method for comparing Lexicalized Tree Adjoining Grammars extracted from annotated corpora for three languages: English, Chinese and Korean. This method makes it possible to do a quantitative comparison between the syntactic structures of each language, thereby providing a way of testing the Universal Grammar Hypothesis, the foundation of modern linguistic theories.
Since 1995, a few statistical parsing algorithms have demonstrated a breakthrough in parsing accuracy, as measured against the UPenn TREEBANK as a gold standard. In this paper we report adapting a lexicalized, probabilistic context-free parser to information extraction and evaluate this new technique on MUC-7 template elements and template relations.
This paper presents results for a maximum-entropy-based part of speech tagger, which achieves superior performance principally by enriching the information sources used for tagging. In particular, we get improved results by incorporating these features: (i) more extensive treatment of capitalization for unknown words; (ii) features for the disambiguation of the tense forms of verbs; (iii) features for disambiguating particles from prepositions and adverbs. The best resulting accuracy for the tagger on the Penn Treebank is 96.86% overall, and 86.91% on previously unseen words.
Article choice can pose difficult problems in applications such as machine translation and automated summarization. In this paper, we investigate the use of corpus data to collect statistical generalizations about article use in English in order to be able to generate articles automatically to supplement a symbolic generator. We use data from the Penn Treebank as input to a memory-based learner (TiMBL 3.0; We discuss competitive results obtained using a variety of lexical, syntactic and semantic features that play an important role in automated article generation.
The characterization of the Aragonese used in the Heredia's works has met with surprise almost always on account of its heterogeneity, accountable by the different models used and the diverse people that have played a role in his scriptorium. To this should be added the different peculiarities of the language itself used by his patron, a language that has never been taken into account up to now. In order to fill this gap this paper presents an edition and a study of a letter from Castellan de Amposta, kept in the Archivo de la Corona de Aragon. This analysis reveals the use of a variety of linguistic norms: forms habitually found in Aragonese are to be found side by side in the letter with forms usual in Catalan. There are also forms that coincide with Castilian usage.
This paper demonstrates that machine learning is a suitable approach for rapid parser development. From 1000 newly treebanked Korean sentences we generate a deterministic shift-reduce parser. The quality of the treebank, particularly crucial given its small size, is supported by a consistency checker. 1 Introduction Given the enormous complexity of natural language, parsing is hard enough as it is, but often unforeseen events like the crises in Bosnia or East-Timor create a sudden demand for parsers and machine translation systems for languages that have not benefited from major attention of the computational linguistics community up to that point. Good machine translation relies strongly on the context of the words to be translated, a context that often goes well beyond neighboring surface words. Often basic relationships, like that between a verb and its direct object, provide crucial support for translation. Such relationships are usually provided by parsers. The NLP resources f...
The paper reports on a multi-layered corpus of Italian, annotated at the syntactic and lexicosemantic levels, whose development is supported by a dedicated software augmented with an intelligent interface. The issue of evaluating this type of resource is also addressed.
Abstract Traditional theories of finance posit that the pricing of securities in financial markets should be done according to the quality of their underlying technical fundamentals. However, research on financial markets has tended to indicate that factors other than technical fundamentals are often used by market participants to gauge the value of securities. This phenomenon may be quite prevalent in markets for initial public offerings (IPOs), where securities lack a financial history. The imagery and affect associated with securities can be a powerful basis upon which to judge their worth. Advanced business students in a securities analysis course were asked to evaluate a number of industry groups represented on the New York Stock Exchange in terms of a set of judgmental variables. After providing imagery and affective evaluations for each industry group, the participants judged the likelihood that they would invest in companies associated with each industry. Imagery and affective ratings were highly correlated with one another and with the likelihood of investing. Judgments of performance correlated poorly to moderately with actual market performance as measured by weighted average returns for the industry groups studied. The results suggest that imagery and affect are part of a coherent psychological framework for evaluating classes of securities, but that framework may have low validity for predicting performance.
This paper presents joint research between a Spanish team and an American one on the development and exploitation of a Spanish treebank. Such treebanks for other languages have proven valuable for the development of high-quality parsers and for a wide variety of language studies. However, when the project started, at the end of 1997, there was no syntactically annotated corpus for Spanish. This paper describes the design of such a treebank and its initial application to parser construction. 1. Constructing a Spanish treebank 1.1. Preliminary considerations This paper presents joint research between a Spanish team and an American one on the development and ex-ploitation of a Spanish treebank. Such treebanks for other languages have proven valuable for the development of high-quality parsers and for a wide variety of language studies. As there was no previous experience in building a syntactically annotated corpus for Spanish, the first effort consisted necessarily in writing a set of annotation guide-lines. The starting point was the existing documentation at that time, especially the Penn Treebank project (Marcus,
In this paper we present the results of a quantitative evaluation of the discrepancies between the Italian and English lexica in terms of lexical gaps. This evaluation has been carried out in the context of MultiWordNet, an ongoing project that aims at building a multilingual lexical database. The quantitative evaluation of the English-to-Italian lexical gaps shows that the English and Italian lexica are highly comparable and gives empirical support to the MultiWordNet model. 1.
The purpose of this article is to briefly introduce an interactive POS tagging system developed as a project at the Institute for Humanities and Cultural Studies in Tehran, Iran. The system is designed as part of the annotation procedure for a Persian corpus called The Farsi Linguistic Database (FLDB) (a project at the Institute for Humanities and Cultural Studies in Tehran which comprises a selection of contemporary Modern Persian literature, formal and informal spoken varieties of the language, and a series of dictionary entries and word lists [Assi 1997: 5]) and is the first attempt ever to tag a Persian corpus. In Section 1, the project itself will be introduced; Section 2 presents an evaluation of the project, and Section 3 is allocated to some suggestions for future work.
This paper reports the results of a vignette- and questionnaire-based research project over the World-Wide Web investigating the influence of moral intensity (MI) on decision making in a business context. A qualitative analysis of the feedback in terms of e-mail communications was used to provide insights into the reactions and responses of participants to both the research method and the topic of research. Implications are discussed, and some methodological recommendations are derived. Second, analysis of the quantitative results of the Web-based questionnaire administration indicated that three of the six MI components were particularly important determinants of several outcome variables. This pattern of results essentially replicated findings yielded by a previous mail administration of the survey, even though a smaller amount of variation in the outcome variables was accounted for. Neither occupational background nor the region of origin of participants measurably influenced the results.
We present a treebank project for French. We have annotated a newspaper corpus of 1 Million words with part of speech, inflection, compounds, lemmas and constituency. We describe the tagging and parsing phases of the project, and for each, the automatic tools, the guidelines and the validation process. We then present some uses of the corpus as well as some directions for future work.
This document describes the bracketing guidelines for the Penn Chinese Treebank Project. The goal of the project is the creation of a 100-thousand-word corpus of Mandarin Chinese text with syntactic bracketing. The Chinese Treebank has been released via the Linguistic Data Consortium (LDC) and is available to the public.\nThis document can be divided into six parts. Section I discusses six fundamental grammatical relations that are represented in the Treebank. Section II introduces the bracketing tagset, which includes 23 syntactic labels, 26 functional tags, and 7 tags for null elements. Section III, IV and V specify our annotation schemata for noun phrases, verbs phrases, and other minor categories, respectively. Section VI describes our treatment for empty categories, such as trace for syntactic movement, PRO for control, and pro for argument drop. Section VII and VIII cover the coordinated clauses and subordinating clauses. Section IX, X and XI specify the way we handle punctuation, ambiguity, and some problematic cases.
This volume contains the papers presented at the workshop on Recent Advances in Natural Language Processing and Information Retrieval, held on 8 October 2000 in conjunction with the 38 th Annual Meeting of the Association for Computational Linguistics (ACL).The aim of the workshop was to foster the interaction between researchers in the areas of Natural Language Processing (NLP) and Information Retrieval (IR), and furthermore, to promote discussion on the current and potential benefits of common approaches to related research challenges. In putting the workshop together, our goal was to bring these two communities together to establish opportunity for communication. We hope there will be similar workshops at IR related conferences so that a gradual closing of the existing gaps will occur. This workshop is not the first on the topic, but one which reflects what we believe are increasing trends as each field reaches its limits.The growing research and application possibilities provided by the increased amount of networked information have motivated new attempts to explore the relationship between NLP and IR. For researchers in IR, a compelling challenge is to move from (monolingual) document retrieval within controlled text collections, to actually retrieving information, rather than individual documents, from multilingual, heterogeneous and dynamic webs of interlinked documents and online services. The reciprocal challenge for NLP research is to scale up, adapt and possibly reshape techniques and resources to help bridge the gap between document and information retrieval in practical applications.The central topic of the workshop was the application of language technologies to information retrieval, including•the role of lexical-syntactic information in mono- and multilingual IR, including morphology, phrase detection and treatment, word sense disambiguation adapted to IR needs, acquisition and use of lexical resources, etc.•empirical evidence regarding the use of NL techniques in different retrieval scenarios, typification of such scenarios, and the discussion of evaluation measures beyond precision/recall variants.•interaction between NLP and IR techniques in topics related to both areas such as Cross-Language and Interactive Text Retrieval, Question Answering, Information Extraction, Text Summarization, Text Data Mining, etc.The original call for papers resulted in 25 submissions, from which 10 papers were selected on the basis of a thorough reviewing process. A majority of the papers contained in this volume study how to improve retrieval processes using syntactic and semantic information, including lexical expansions for web querying, semantic indexing, phrasal indexing in different languages, acquisition and use of lexical databases, etc. The remaining papers combine NL techniques from related areas - summarization, information extraction - to extend search capabilities and presentation of results, including summarization of search engine hit lists and summarization for text categorization, and a search interface that accepts template-like general constraints and is able to return specific information items such as locations, people or companies that satisfy user's constraints. Although none of these papers deal directly with cross-language issues (well covered in the Cross-Language Evaluation Forum held at Lisbon two weeks earlier), the monolingual systems presented here cover five different languages (English, Japanese, Korean, Italian and Czech).
This paper presents our efforts to retrieve proper name information from a machine readable dictionary. We explain the problem in brief and then we discuss related work that other researchers have attempted in this area. We describe the embedded knowledge that one can find in the Collins English Dictionary (CED). Finally we present our ideas about how to extract this knowledge and how we plan to store it in a lexical database of proper names.
Age of acquisition (AoA) has been reported to be a predictor of the speed of reading words aloud (word naming) and lexical decision, with early-acquired words being responded to faster than later-acquired words in both tasks. All previous studies of AoA effects have, however, relied upon adult estimates of word learning age the validity of which it is easy to cast doubt upon. Using objective age of acquisition norms derived from children's naming data, this study shows that AoA effects do not depend upon the use of adult ratings. In addition to effects of real AoA, influences of word frequency and orthographic neighbourhood size were obtained in both word naming and lexical decision. Imageability affected lexical decision but not word naming, while the characteristics of the word's initial phoneme affected word naming but not lexical decision.
Three studies focused on the development and enhancement of narrative skills within a preschool classroom. The purpose of Study 1 was to collect local norms on narrative development. Fifty-two preschool African American English speakers representing 3-, 4-, and 5-year-old age groups, narrated a familiar storybook. Some children in each age group evidenced use of nine story element types. Developmental changes were characterized by growth in types as well as tokens of story elements. Study 2 demonstrated that preschoolers’ narratives can be influenced by the narratives of their peers. Paired children narrated a familiar storybook to each other. The stories of paired children were significantly more similar in form (shared story element types) and content (shared lexical types) than those of unpaired children. Study 3 provided a preliminary test of an intervention designed to exploit the effect of peer models for long-term gain in narrative abilities. Two tutees practiced book narration following the clinician-prompted models of their peer tutors. As a result, the tutees demonstrated an expanded repertoire of story elements and an increased frequency of use of story element types in both trained and untrained stories. Their rate of growth in story element use was superior to that of their classmates who had not participated in the intervention. The benefit of peers for achieving instructional congruence in cases of clinicianclient mismatch is emphasized.
This paper describes the evaluation of a WSD method withinSENSEVAL. This method is based on Semantic Classification Trees (SCTs)and short context dependencies between nouns and verbs. The trainingprocedure creates a binary tree for each word to be disambiguated. SCTsare easy to implement and yield some promising results. The integrationof linguistic knowledge could lead to substantial improvement.
SENSEVAL set itself the task of evaluating automaticword sense disambiguation programs (see Kilgarriff andRosenzweig, this volume, for an overview of theframework and results). In order to do this, it wasnecessary to provide a `gold standard' dataset of `correct' answers. This paper will describe thelexicographic part of the process involved in creatingthat dataset. The primary objective was for a group oflexicographers to manually examine keywords in a largenumber of corpus contexts, and assign to each contexta sense-tag for the keyword, taken from the Hectordictionary. Corpus contexts also had to be manuallypart-of-speech (POS) tagged. Various observationsmade and insights gained by the lexicographers duringthis process will be presented, including a critiqueof the resources and the methodology.
In this paper we present some observations concerning an experiment of (manual/automatic) semantic tagging of a small Italian corpus performed within the framework of the SENSEVAL/ROMANSEVAL initiative. Themain goal of the initiative was to set up a framework for evaluation of Word Sense Disambiguation systems (WSDS) through the comparative analysis of their performance on the same type of data. In this experiment there are two aspects which are of relevance: first, the preparation of the reference annotated corpus, and, second, the evaluation of the systems against it. In both aspects we are mainly interested here in the analysis of the linguistic side which can lead to a better understanding of the problem of semantic annotation of a corpus, be itmanual or automatic annotation. In particular, we will investigate, firstly, the reasons for disagreement between human annotators, secondly, some linguistically relevant aspects of the performance of the Italian WSDS and, finally, the lessons learned from the present experiment.
This paper describes the generation of iconic and categorical representations of word meaning, in propositional form, from the WordNet lexical database. These are derived from the list of synonyms, the descriptive gloss, and from the hypernym and meronym relations of each WordNet word sense. We demonstrate that these representations promote identification and discrimination, these being suggested qualities of representations of meaning, and finally suggest that these representations have further applications in language engineering. 1.
This work combines a set of available techniques – whichcould be further extended – to perform noun sense disambiguation. We use several unsupervised techniques (Rigau et al., 1997) that draw knowledge from a variety of sources. In addition, we also apply a supervised technique in order to show that supervised and unsupervised methods can be combined to obtain better results. This paper tries to prove that using an appropriate method to combine those heuristics we can disambiguate words in free running text with reasonable precision.
In developing the concepts of phonological and lexical subtypes of dyslexia, criteria have been proposed based on the projection of linear regression for raw test scores. Substantial discrepancies in subtype prevalence may arise from nonlinearities in the test scores. A norming process developed for a new nonword test, the Martin and Pratt Nonword Reading Test (Martin & Pratt, 2000), was applied to the Word Identification subtest of the Woodcock Reading Mastery Test (Woodcock, 1987) and the Regular and Irregular Word Tests published in Coltheart and Leahy (1996). These tests were administered to a representative sample of 863 children aged 6 to 15 years in the Southern Tasmanian State School population. An inverse normal transform results in a distribution which is approximately normal within age groups. On this scale the age effect was well approximated by a linear increase with the logarithm of (age −5 years). This process can be adapted to provide norms for word lists more economically and allows convenient spreadsheet formulae for norms. Substantial differences from other norms may be attributed to school district family income differences found in this sample. Male means are lower than female for all tests, but this reflects comparable high performance and disproportionate poor performance by males on reading tests.