Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Abstract: Ontologies are becoming extremely useful tools for sophisticated software engineering. Designing applications, databases, and knowledge bases with reference to a common ontology can mean shorter development cycles, easier and faster integration with other software and content, and a more scalable product. Although ontologies are a very promising solution to some of the most pressing problems that confront software engineering, they also raise some issues and difficulties of their own. Consider, for example, the questions below: • How can a formal ontology be used effectively by those who lack extensive training in logic and mathematics? • How can an ontology be used automatically by applications (e.g. Information Retrieval and Natural Language Processing applications) that process free text? • How can we know when an ontology is complete? In this paper we will begin by describing the upperlevel ontology SUMO (Suggested Upper Merged Ontology), which has been proposed as the initial version of an eventual Standard Upper Ontology (SUO). We will then describe the popular, free, and structured WordNet lexical database. After this preliminary discussion, we will describe the methodology that we are using to align WordNet with the SUMO. We close this paper by discussing how this alignment of WordNet with SUMO will provide answers to the questions posed above. Ontologies are becoming extremely useful tools for sophisticated software engineering. Designing applications, databases, and knowledge bases with reference to a common ontology can mean shorter development cycles, easier and faster integration with other software and content, and a more scalable product. Although ontologies are a very promising solution to some of the most pressing problems that confront software engineering, they also raise some issues and difficulties of their own. Consider, for example, the questions below: • How can a formal ontology be used effectively by those who lack extensive training in logic and mathematics? • How can an ontology be used automatically by applications (e.g. Information Retrieval and Natural Language Processing applications) that process free text? • How can we know when an ontology is complete? In this paper we will begin by describing the upperlevel ontology SUMO (Suggested Upper Merged Ontology), which has been proposed as the initial version of an eventual Standard Upper Ontology (SUO). We will then describe the popular, free, and structured WordNet lexical database. After this preliminary discussion, we will describe the methodology that we are using to align WordNet with the SUMO. We close this paper by discussing how this alignment of WordNet with SUMO will provide answers to the questions posed above. keywords: natural language, ontology 1. SUMO The SUMO (Suggested Upper Merged Ontology) is an ontology that was created at Teknowledge Corporation with extensive input from the SUO mailing list, and it has been proposed as a starter document for the IEEE-sanctioned SUO Working Group [1]. The SUMO was created by merging publicly available ontological content into a single, comprehensive, and cohesive structure [2,3]. As of February 2003, the ontology contains 1000 terms and 4000 assertions. The ontology can be browsed online (http://ontology.teknowledge.com), and source files for all of the versions of the ontology can be freely downloaded (http://ontology.teknowledge.com/cgibin/cvsweb.cgi/SUO/).
We present a system for automatically identifying PropBank-style semantic roles based on the output of a statistical parser for Combinatory Categorial Grammar. This system performs at least as well as a system based on a traditional Treebank parser, and outperforms it on core argument roles.
We present an extension of the classic A* search procedure to tabular PCFG parsing. The use of A* search can dramatically reduce the time required to find a best parse by conservatively estimating the probabilities of parse completions. We discuss various estimates and give efficient algorithms for computing them. On average-length Penn treebank sentences, our most detailed estimate reduces the total number of edges processed to less than 3% of that required by exhaustive parsing, and a simpler estimate, which requires less than a minute of pre-computation, reduces the work to less than 5%. Un-like best-first and finite-beam methods for achieving this kind of speed-up, an A* method is guaranteed to find the most likely parse, not just an approximation. Our parser, which is simpler to implement than an upward-propagating best-first parser, is correct for a wide range of parser control strategies and maintains worst-case cubic time.
▪ Abstract Using Australian languages as examples, cultural selection is shown to shape linguistic structure through invisible hand processes that pattern the unintended outcomes (structures in the system of shared linguistic norms) of intentional actions (particular utterances by individual agents). Examples of the emergence of culturally patterned structure through use are drawn from various levels: the semantics of the lexicon, grammaticalized kin-related categories, and culture-specific organizations of sociolinguistic diversity, such as moiety lects, “mother-in-law” registers, and triangular kin terms. These phenomena result from a complex of diachronic processes that adapt linguistic structures to culture-specific concepts and practices, such as ritualization and phonetic reduction of frequently used sequences, the input of shared cultural knowledge into pragmatic interpretation, semanticization of originally context-dependent inferences, and the input of linguistic ideologies into the systematization of lectal variants. Some of these processes, such as the emergence of subsection terminology and moiety lects, operate over speech communities that transcend any single language and can only be explained if the relevant processes take the multilingual speech community as their domain of operation. Taken together, the cases considered here provide strong evidence against nativist assumptions that see linguistic structures simply as instantiations of biologically given “mentalese” concepts already present in the mind of every child and give evidence in favor of a view that sees individual language structures as also conditioned by historical processes, of which functional adaptation of various kinds is most important. They also illustrate how, in the domain of language, stable socially shared structures can emerge from the summed effects of many communicative micro-events by individual agents.
BACKGROUND: Human affective responses appear to be regulated by limbic and paralimbic circuits. However, much less is known about the neurochemical systems engaged in this regulation. The mu-opioid neurotransmitter system is distributed in, and thought to regulate the function of, brain regions centrally implicated in affective processing. OBJECTIVE: To examine the involvement of mu-opioid neurotransmission in the regulation of affective states in healthy human volunteers. DESIGN: Measures of mu-opioid receptor availability in vivo were obtained with positron emission tomography and the mu-opioid receptor selective radiotracer [11C]carfentanil during a neutral state and during a sustained sadness state. Subtraction analyses of the binding potential maps were then performed within subjects, between conditions, on a voxel-by-voxel basis. SETTING: Imaging center at a university medical center. PARTICIPANTS: Fourteen healthy female volunteers. Intervention Sustained neutral and sadness states, randomized and counterbalanced in order, elicited by the cued recall of an autobiographical event associated with that emotion. MAIN OUTCOME MEASURES: Changes in mu-opioid receptor availability and negative and positive affect ratings between conditions. Increases or reductions in the in vivo receptor measure reflect deactivation or activation of neurotransmitter release, respectively. RESULTS: The sustained sadness condition was associated with a statistically significant deactivation in mu-opioid neurotransmission in the rostral anterior cingulate, ventral pallidum, amygdala, and inferior temporal cortex. This deactivation was reflected by increases in mu-opioid receptor availability in vivo. The deactivation of mu-opioid neurotransmission in the rostral anterior cingulate, ventral pallidum, and amygdala was correlated with the increases in negative affect ratings and the reductions in positive affect ratings during the sustained sadness state. CONCLUSIONS: These data demonstrate dynamic changes in mu-opioid neurotransmission in response to an experimentally induced negative affective state. The direction and localization of these responses confirms the role of the mu-opioid receptor system in the physiological regulation of affective experiences in humans.
Abstract Two experiments examined the effects of word familiarity on word recognition and text comprehension during silent reading. Readers' eye movements were monitored as they read sentences containing words that varied in familiarity as assessed by printed estimates of word frequency, subjective ratings of familiarity, and a multiple‐choice test of meaning knowledge. Effects of word frequency were unaffected by differences in subjective familiarity rating for high frequency words. Differential effects of familiarity rating were observed in low frequency conditions. In addition, processing time on high and low frequency words did not differ when familiarity was held constant for moderately familiar words. Readers spent more initial processing time on novel words than familiar words. Performance on a vocabulary test administered after the reading session demonstrated that readers successfully acquired and retained new word meanings. Finally, reanalysis of word processing time as a function of vocabulary test performance demonstrated a systematic relationship between online processing patterns and memory for novel word meaning.
Pleasant stimuli typically elicit greater electromyographic (EMG) activity over zygomaticus major and less activity over corrugator supercilii than do unpleasant stimuli. To provide a systematic comparison of these 2 measures, the authors examined the relative form and strength of affective influences on activity over zygomaticus major and corrugator supercilii. Self-reported positive and negative affective reactions and facial EMG were collected as women (n = 68) were exposed to series of affective pictures, sounds, and words. Consistent with speculations based on known properties of the neurophysiology of the facial musculature, results revealed a stronger linear effect of valence on activity over corrugator supercilii versus zygomaticus major. In addition, positive and negative affect ratings indicated that positive and negative affect have reciprocal effects on activity over corrugator supercilii, but not zygomaticus major.
Since Eloise Jelinek has been interested in the issues of negation, focus and information structure, to the research of which she has contributed substantially, we want to use this nice occasion and present here partial results of an analysis of the Topic-Focus articulation (TFA) of Czech and of the impact of these results on inquiries into coreferrence in coherent discourse. In Czech linguistics, TFA has been systematically explored thanks to the classical Prague School of functional and structural linguistics. As reflecting the ‘given – new ’ strategy in discourse, TFA has been considered to belong to the main objects of linguistic study. Continuing the results gained by V. Mathesius, J. Firbas and others since the 1920s, the explicit linguistic descriptive framework characterized in Sgall et al. (1986), Hajičová (1993), Hajičová E., Partee B. and P. Sgall (1998) includes a possibility to describe TFA not only as concerning the intrinsic dynamics of the process of communication, patterned in the utterance (sentence occurrence), but also as constituting the structure of the sentence itself, i.e. grammar. Within this framework, TFA is understood as one of the basic aspects of (underlying) sentence structure, which characterizes the sentence as a unit of the interactive system of language; TFA thus is seen as a manifestation of the sentence being anchored in the context.
This paper deals with the problem of how to interrelate theory-specific treebanks and how to transform one treebank format to another. Currently, two approaches to achieve these goals can be differentiated. The first creates a mapping algorithm between treebank formats. Categories of a source format are transformed into a target format via a given set of general or language-specific mapping rules. The second relates treebanks via a transformation to a general model of linguistic categories, for example based on the EAGLES recommendations for syntactic annotations of corpora, or relying on the HPSG framework. This paper proposes a new methodology as a solution for these desiderata.
One of the major problems in the implementation of natural language processing (NLP) or machine translation (MT) is a complete lexicon: the place where the system's information about words is stored. There are difficulties in deciding what information should be stored in a lexicon and even greater difficulties in acquiring this information in proper form. The OriNet system was designed to incorporate a multiple lexical database and tools under one consistent functional interface in order to facilitate systems requiring syntactic, semantic and lexical information of the Oriya language. We divide the work into two independent tasks. One is to write the source file that contains basic lexical data, the content of these files being the lexical substance of OriNet. The second is to create a set of programs that would accept the source files and process them ultimately to display for the user. This paper describes ongoing work on designing an object oriented model for the OriNet system. It uses object oriented programming, particularly the rich library of classes and programming principles which Java offers. It also provides a convenient tool to conceptualise the process of the OriNet system. This technique also allows flexibility and extensibility of the system with more robustness.
Broad-coverage, deep unification grammar development is time-consuming and costly. This problem can be exacerbated\nin multilingual grammar development scenarios. Recently (Cahill et al., 2002) presented a treebank-based methodology\nto semi-automatically create broadcoverage, deep, unification grammar resources for English. In this paper we\npresent a project which adapts this model to a multilingual grammar development scenario to obtain robust, wide-coverage, probabilistic Lexical-Functional Grammars\n(LFGs) for English and German via automatic f-structure annotation algorithms based on the Penn-II and TIGER\ntreebanks. We outline our method used to extract a probabilistic LFG from the TIGER treebank and report on the quality of the f-structures produced. We achieve an f-score of 66.23 on the evaluation of 100 random sentences against a manually constructed gold standard.
We present a new part-of-speech tagger that demonstrates the following ideas: (i) explicit use of both preceding and following tag contexts via a dependency network representation, (ii) broad use of lexical features, including jointly conditioning on multiple consecutive words, (iii) effective use of priors in conditional loglinear models, and (iv) fine-grained modeling of unknown word features. Using these ideas together, the resulting tagger gives a 97.24% accuracy on the Penn Treebank WSJ, an error reduction of 4.4% on the best previous single automatically learned tagging result.
This article describes three statistical models for natural language parsing. The models extend methods from probabilistic context-free grammars to lexicalized grammars, leading to approaches in which a parse tree is represented as the sequence of decisions corresponding to a head-centered, top-down derivation of the tree. Independence assumptions then lead to parameters that encode the X-bar schema, subcategorization, ordering of complements, placement of adjuncts, bigram lexical dependencies, wh-movement, and preferences for close attachment. All of these preferences are expressed by probabilities conditioned on lexical heads. The models are evaluated on the Penn Wall Street Journal Treebank, showing that their accuracy is competitive with other models in the literature. To gain a better understanding of the models, we also give results on different constituent types, as well as a breakdown of precision/recall results in recovering various types of dependencies. We analyze various characteristics of the models through experiments on parsing accuracy, by collecting frequencies of various structures in the treebank, and through linguistically motivated examples. Finally, we compare the models to others that have been applied to parsing the treebank, aiming to give some explanation of the difference in performance of the various models.
The paper presents an HPSG-based annotation scheme for constructing a Bulgarian treebank: BulTreeBank. It differs from other grammar-based annotation schemes in having a hybrid status with respect to the partial parsing component and the full parsing module. As the
The question of how treebank annotation schemes should be related to linguistic theories has been debated as long as treebanks have existed. Historically speaking, it is probably true to say that there has been a development from mostly theoryneutral annotation schemes to more theoretically oriented frameworks or even annotation
Intensified aquaculture has strong impact on fish health by stress and infectious diseases and has stimulated the interest in the orchestration of cytokines and growth factors, particularly their influence by environmental factors, however, only scarce data are available on the GH/IGF-system, central physiological system for development and tissue shaping. Most recently, the capability of the host to cope with tissue damage has been postulated as critical for survival. Thus, the present study assessed the combined impacts of estrogens and bacterial infection on the insulin-like growth factors (IGF) and tumor-necrosis factor (TNF)-α. Juvenile rainbow trout were exposed to 2 different concentrations of 17β-estradiol (E2) and infected with Yersinia ruckeri. Gene expressions of IGF-I, IGF-II and TNF-α were measured in liver, head kidney and spleen and all 4 estrogen receptors (ERα1, ERα2, ERβ1 and ERβ2) known in rainbow trout were measured in liver. After 5 weeks of E2 treatment, hepatic up-regulation of ERα1 and ERα2, but down-regulation of ERß1 and ERß2 were observed in those groups receiving E2-enriched food. In liver, the results further indicate a suppressive effect of Yersinia-infection regardless of E2-treatment on day 3, but not of E2-treatment on IGF-I whilst TNF-α gene expression was not influenced by Yersinia-infection but was reduced after 5 weeks of E2-treatment. In spleen, the results show a stimulatory effect of Yersinia-infection, but not of E2-treatment on both, IGF-I and TNF-α gene expressions. In head kidney, E2 strongly suppressed both, IGF-I and TNF-α. To summarise, the treatment effects were tissue- and treatment-specific and point to a relevant role of IGF-I in infection.
This paper reports on experiments in classifying the semantic role annotations assigned to prepositional phrases in both the Penn Treebank and FrameNet. In both cases, experiments are done to see how the prepositions can be classified given the dataset's role inventory, using standard word-sense disambiguation features. In addition to using traditional word collocations, the experiments incorporate class-based collocations in the form of WordNet hypernyms. For Treebank, the word collocations achieve slightly better performance: 78.5% versus 77.4% when separate classifiers are used per preposition. When using a single classifier for all of the prepositions together, the combined approach yields a significant gain at 85.8% accuracy versus 81.3% for word-only collocations. For FrameNet, the combined use of both collocation types achieves better performance for the individual classifiers: 70.3% versus 68.5%. However, classification using a single classifier is not effective due to confusion among the fine-grained roles.
In this paper we show how the trees in the Penn treebank can\nbe associated automatically with simple quasi-logical forms. Our approach is based on combining two independent strands of work: the first is the observation that there is a close correspondence between quasi-logical forms and LFG f-structures [van Genabith and Crouch, 1996]; the second is the development of an automatic f-structure annotation algorithm for the Penn treebank [Cahill et al, 2002a; Cahill\net al, 2002b]. We compare our approach with that of [Liakata and Pulman, 2002].
ABSTRACT. A recent development in Chinese renders many occurrences of the construction [Adjunct PP + Verb + N[P.sub.2]] into [Verb + N2 + N[P.sub.1]]. A crucial difference between the two constructions is that in the former the verb and its N[P.sub.2] object can be separated. In the latter no separation is permitted, and N[P.sub.2] consists of only its head [N.sub.2]; that is, the verb and [N.sub.2], have formed a V-N COMPOUND. This paper attempts to account for such V-N compounding in Chinese. We suppose that the lexical structure representation of verbs in the Chinese V-N compound is similar to that of denominal verbs (Hale & Keyser 1993b). Thus, the formation of the Chinese V-N compound can be derived most simply by head movement, which is both morphologically driven and constrained by the Minimal Link Condition (Chomsky 1994). * INTRODUCTION. A recent development in Chinese renders many cases of [Adjunct PP (i.e. P + N[P.sub.1]) + Verb + N[P.sub.2]] into [Verb + [N.sub.2] + N[P.sub.1]], especially when the adjunct PP is locative (Hua 1997, Wang 1997, Xing 1997, Liu 1998a,b, Wang 1998), as shown by 1a-b, 2a-b and 3a-b. (1) (1) a. women xiang Niuyue qian ju we to New York move home [right arrow] b. women qian ju Niuyue (2) we move home New York 'We move to New York' (2) a. ta zai Hafu zhi jiao he at Harvard University engage teaching [right arrow] b. ta zhi jiao Hafu he engage teaching Harvard University He teaches at Harvard University' (3) a. Mali zai Jianada liu xue Mary in Canada engage study [right arrow] b. Mali liu xue Jianada Mary engage study Canada 'Mary studies in Canada' The crucial difference between the two constructions is that in [Adjunct PP + Verb + N[P.sub.2]] the verb and its N[P.sub.2] object can be separated by an aspect marker, a measure phrase, or a modifier of N[P.sub.2] (Li & Thompson 1981), as in 4a, 5a, and 6a. On the other hand, in [Verb + [N.sub.2] + N[P.sub.1]] no such separation is permitted, and N[P.sub.2] consists of its head noun [N.sub.2] only, as shown in 4b, 5b, and 6b. (4) a. women xiang Niuyue qian-le ju we to New York move-Asp home 'We have moved to New York' b. *women qian-le ju Niuyue we move-ASP home New York (5) a. ta zai Hafu zhi-le yi nian jiao he at Harvard University engage-ASP one year teaching 'He has taught at Harvard University for one year' b. *ta zhi-le yi nian jiao Hafu he engage-ASP one year teaching Harvard University (6) a. Mali zai Jianada liu-guo liang ci xue Mary in Canada engage-ASP two CL study 'Mary has studied in Canada twice' b. *Mali liu-guo liang ci xue Jianada Mary engage-ASP two CL study Canada In sum, the verb and [N.sub.2] in 1b, 2b, and 3b have formed a V-N COMPOUND. The V-N compound has, in fact, long existed in Chinese (Chao 1968), but until recently very few of them could take an NP object without being treated as ungrammatical or unacceptable (Li & Thompson 1981). By contrast, [[P + N[P.sub.1]] + Verb + N[P.sub.2]] has been regarded as the norm and has been strongly required or preferred (Hua 1997, Chow 2000). Since the 1970s, however, many cases of [[P + N[P.sub.1]] + Verb + N[P.sub.2]] have been transformed into [V-[N.sub.2] + N[P.sub.1]] (Xing 1997, Diao 1998). (3) Now the [V-[N.sub.2] + N[P.sub.1]] construction has become so common that it is no longer treated as ungrammatical or unacceptable but as a legitimate type of predicate structure (Gao 1998, Liu 1998a,b, Wang 1998). …
1006 Reviews 'national' or Parisian counterpart, and to give a clear idea of its distinctive place in the French media landscape. Martin manages to give this overview in a very readable manner and without being superficial. He acknowledges and draws on the excellent work that has been done on individual titles, periods, and geographical areas. A particularly welcome aspect is the significant space devoted to considering newspapers as eco? nomic and social entities, highlighting not just the Citizen Hersants and the starjour? nalists but also the networks of correspondents in the smallest ofvillages, the typographers, the delivery drivers, the sellers, and the readers. Especially fascinating fromthe perspective of social history is the analysis of the evolution, content, and role of the 'avis de deces' rubric: starting as simple quasi-administrative announcements, often appearing after the funeral, these came to be used as a substitute for the individual 'faire-part', then as a signifierof social status. They were also a major source of income fornewspapers. Martin is sensitive throughout to the impact of new technologies, up to and including the Internet, and notes that the regional press has often pioneered their use in France. The volume is impressively useable: there is an accurate general index and a separate index of newspaper titles, running to over seven pages; a useful chronology; a detailed table of contents; and an annotated summary bibliography to complement the abundant and detailed notes. These tools will help a range of readers make the most of a volume that achieves its purpose and invites furtherstudy. University of Leeds Paul Rowe La Neologie en francais contemporain: examen du concept et analyse de productions neologiquesrecentes. By Jean-Francois Sablayrolles. Paris: Champion. 2000. 588 pp.?86.90. ISBN 2-7453-0275-2. This is a scholarly and thought-provoking contribution to a field which has in the past suffered from either too narrow an academic approach, or from being subject to merely anecdotal treatment in amusing collections of neologisms. Jean-Francois Sablayrolles is ambitious, and largely successful, in his attempt to link broad theoret? ical discussion to a significant body of data. After a concise history of the notion of 'neologism' in Greek, Latin, and French, he reviews the differentapproaches to the subject by French linguists and then summarizes how twentieth-century theoretical models, from the structuralists to generativists and the most recent work of Melcu'k, have dealt with the processes of lexical creativity. He notes that, generally speaking, they have been assigned a very secondary and marginal role. In the second part of the book Sablayrolles proposes his own definitions of neologisme and neologie, and examines the types of unit and process that these involve. Perennial issues such as the role of dictionaries, upon which linguists have to rely, albeit often grudgingly, for their data, and the problem of differentiating between polysemy and homonymy, are given a fresh airing. More original are the brief dis? cussion of links between politico-cultural ideology and attitudes to neologisms, and speculation on the possibility of calculating the lifespan of a neologism. In the third part of the book the author analyses and compares the data that he has gathered from his six corpora, and ends with a discussion of the differentfunctions of neologisms. These range from their attention-catching use in newspaper headlines to their role in political polemics and their largely ludic function in the work of the writers R. Jorif and Ph. Meyer. Somewhat problematic is his inclusion of a corpus scolaire, drawn from the written work of secondary-school students. Many of these examples are non-standard verb forms such as ils croivent and il a acqueri. One can argue that these are not lexical, and possibly not new. Unlike the rest of his data, these forms are probably, as he concedes, neither 'voulus' nor 'conscients' (p. 317). Surely they MLRy 98.4, 2003 1007 are either part of a system which just happens to be differentfrom the norm or an attempt to conjugate a lexical item which is simply alien to the system? Theoretically at least, Sablayrolles appears to give the status of neologism to all new forms, what? ever their source or motivation. (Perhaps intentional, conscious creation...
The article studies the particular features of the dynamics of lexical norms in Ukrainian language on the \nmaterials of dictionaries and mass-media. \nIt analyses the process of vocabulary enrichment by new lexical units, the phenomena of semantic \ntransformation and stylistic transposition.
This paper describes log-linear parsing models for Combinatory Categorial Grammar (CCG). Log-linear models can easily encode the long-range dependencies inherent in coordination and extraction phenomena, which CCG was designed to handle. Log-linear models have previously been applied to statistical parsing, under the assumption that all possible parses for a sentence can be enumerated. Enumerating all parses is infeasible for large grammars; however, dynamic programming over a packed chart can be used to efficiently estimate the model parameters. We describe a parellelised implementation which runs on a Beowulf cluster and allows the complete WSJ Penn Treebank to be used for estimation.
Many extensions to text-based, data-intensive knowledge management approaches, such as Information Retrieval or Data Mining, focus on integrating the impressive recent advances in language technology. For this, they need fast, robust parsers that deliver linguistic data which is meaningful for the subsequent processing stages. This paper introduces such a parsing system. Its output is a hierarchical structure of syntactic relations, functional dependency structures.
Recent work in machine translation and information extraction has demonstrated the utility of a level that represents the predicate-argument structure. It would be especially useful for machine translation to have two such Proposition Banks, one for each language under consideration. A Proposition Bank for English has been developed over the last few years, and we describe here our development of a tool for facilitating the development of a Chinese Proposition Bank. We also discuss some issues specific to the Chinese Treebank that complicate the matter of mapping syntactic representation to a predicate-argument level, and report on some preliminary evaluation of the accuracy of the semantic tagging tool. 1
Digitizing and annotating texts and field recordings Given that several initiatives worldwide currently explore the new field of documentation of endangered languages, the E-MELD project proposes to survey and unite procedures, techniques and results in order to achieve its main goal, ''the formulation and promulgation of best practice in linguistic markup of texts and lexicons''. In this context, this year's workshop deals with the processing of recorded texts. I assume the most valuable contribution I could make to the workshop is to show the procedures and methods used in the Awetí Language Documentation Project. The procedures applied in the Awetí Project are not necessarily representative of all the projects in the DOBES program, and they may very well fall short in several respects of being best practice, but I hope they might provide a good and concrete starting point for comparison, criticism and further discussion. The procedures to be exposed include: * taping with digital devices, * digitizing (preliminarily in the field, later definitely by the TIDEL-team at the Max Planck Institute in Nijmegen), * segmenting and transcribing, using the transcriber computer program, * translating (on paper, or while transcribing), * adding more specific annotation, using the Shoebox program, * converting the annotation to the ELAN-format developed by the TIDEL-team, and doing annotation with ELAN. Focus will be on the different types of annotation. Especially, I will present, justify and discuss Advanced Glossing, a text annotation format developed by H.-H. Lieb and myself designed for language documentation. It will be shown how Advanced Glossing can be applied using the Shoebox program. The Shoebox setup used in the Awetí Project will be shown in greater detail, including lexical databases and semi-automatic interaction between different database types (jumping, interlinearization). ( Freie Universität Berlin and Museu Paraense Emílio Goeldi, with funding from the Volkswagen Foundation.)
This paper describes a fast algorithm that selects features for conditional maximum entropy modeling. Berger et al. (1996) presents an incremental feature selection (IFS) algorithm, which computes the approximate gains for all candidate features at each selection stage, and is very time-consuming for any problems with large feature spaces. In this new algorithm, instead, we only compute the approximate gains for the top-ranked features based on the models obtained from previous stages. Experiments on WSJ data in Penn Treebank are conducted to show that the new algorithm greatly speeds up the feature selection process while maintaining the same quality of selected features. One variant of this new algorithm with look-ahead functionality is also tested to further confirm the good quality of the selected features. The new algorithm is easy to implement, and given a feature space of size F, it only uses O(F) more space than the original IFS algorithm.
This paper presents Advanced Glossing, a proposal for a general glossing format designed for language documentation, and a specific setup for the Shoebox-program that implements Advanced Glossing to a large extent. Advanced Glossing (AG) goes beyond the traditional Interlinear Morphemic Translation, keeping syntactic and morphological information apart from each other in separate glossing tables. AG provides specific lines for different kinds of annotation – phonetic, phonological, orthographical, prosodic, categorial, structural, relational, and semantic, and it allows for gradual and successive, incomplete, and partial filling in case that some information may be irrelevant, unknown or uncertain. The implementation of AG in Shoebox sets up several databases. Each documented text is represented as a file of syntactic glossings. The morphological glossings are kept in a separate database. As an additional feature interaction with lexical databases is possible. The implementation makes use of the interlinearizing automatism provided by Shoebox, thus obtaining the table format for the alignment of lines in cells, and for semi-automatic filling-in of information in glossing tables which has been extracted from databases
A semantic database has been extended with visual information to enable video annotation. This paper describes a lexical database, WordNet. We show its limitations with respect to describing visual characteristics, and describe an extension to WordNet that contains specific visual information. Having such a semantic database makes video annotation possible for broadcast news: a domain that can cover any topic and involve a wide variety of events, objects and scenes. Combining basic visual analysis techniques and a semantic database containing visual descriptions avoids the problem developing large numbers of specific object and event detectors. Such a semantic database can be of great value for the analysis of multi-modal information. As far as we know, such a database has not been developed before.
Among the most consistent findings in the warnings literature is the so-called 'familiarity effect.' Research has shown that the more familiar an individual is with a product or situation the less likely he or she is to notice, read, recall, or comply with hazard communications. The effect has been found across numerous product types and situations using various operational definitions of familiarity and measures of warning effectiveness. However, research has also shown that subjective familiarity ratings are not highly correlated with actual product experience. Thus, individuals must be capable of developing a false or exaggerated sense of familiarity. One possible source of this exaggerated familiarity is exposure to product advertising.\n\nThree experiments were conducted to investigate whether the familiarity effect can be produced from exposure to product advertising. The relationships between advertising exposure and perceived familiarity and between perceived familiarity, perceived safety and warning effectiveness were examined. Experiment 1 explored participants' attitudes and beliefs about well-known and obscure brands of household, consumer products and sought to determine how past, direct product experience influences those attitudes and beliefs. Experiments 2 and 3 examined how the number of advertising exposures and the safety-related content of advertisements influence attitudes and beliefs about the advertised products and the effectiveness of on product warnings.\n\nResults of Experiment 1 revealed that past experience can not fully explain consumers' attitudes and beliefs about household, consumer products. Experiments 2 and 3 showed that advertising influences perceived product familiarity and knowledge. While there was a trend of greater perceived safety with increased ad exposures, the effect was not significant. No effects of advertising on warning recall were found. Implications for the design of product advertisements and product packaging as well as directions for future research are discussed.
This paper describes the use of clustering at three stages within a larger research effort to identify semantic frames used in English automatically. The first of two tasks within this effort has been the identification of sets of semantically related verb senses that invoke a common semantic frame. Within this task, clustering has been used both to build sets of verb senses with the potential of invoking a common semantic frame and then to merge sets with a high degree of overlap. The paper is organized as follows: Section 2 introduces frame semantics. Section 3 outlines the methodology used to identify sets of semantically related verb senses that invoke a common semantic frame, while section 4 presents the specific clustering algorithm used within that process. Section 5 discusses the use of this clustering algorithm for the identification of semantically related verbs in two machine-readable lexical resources: the machine-readable version of the Longman Dictionary of Contemporary English (LDOCE, 1978 edition) and WordNet, an online lexical database (http://www.cogsci.princeton.edu/-wn; version 1.7.1 has been used for the work reported here). Section 6 presents the use of clustering to merge overlapping sets of verb senses formed in previous steps. Section 7 discusses the results of these clusterings, paying particular attention to the effect of LDOCE's restricted defining vocabulary on the clustering process.