Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Schema matching is prerequisite to an automated transformation of XML documents. Because previous works about schema matching compute all semantically-possible matchings, they produce many-to-many matching relationships. Such imprecise matchings are inappropriate for an automated transformation of XML documents. This paper presents an efficient schema matching algorithm that computes precise one-to-one matchings between two schemas. The proposed algorithm consists of two steps: preliminary matching relationships between leaf nodes in the two schemas are computed and one-to-one matchings are finally extracted based on a proposed path similarity. Specifically, for a sophisticated schema matching, the proposed algorithm is based on a domain ontology as well as a lexical database that includes abbreviations and synonyms. Experimental results with real schemas from an e-commerce field show that the proposed method is superior to previous works, resulting in an accuracy of 97% in average.
OBJECTIVES: To compare the effect of a specialized care facility (SCF) on quality of life (QoL) for residents with middle- to late-stage dementia over a 1-year period with residence in traditional institutional facilities. DESIGN: A prospective, matched-group design with assessments of QoL every 3 months for 1 year. SETTING: Twenty-four long-term care centers and four designated assisted living environments in an urban center in western Canada. PARTICIPANTS: One hundred eighty-five residents with Global Deterioration Scores of 5 or greater were enrolled: 62 in the intervention SCF group and 123 in the traditional institutional facilities groups. INTERVENTION: The SCF is a 60-bed purpose-built facility with 10 people living in six bungalows. The facility followed an ecologic model of care that is responsive to the unique interplay of each person and the environment. This model encompasses a vision of long-term care that is more comfortable, more like home, and offers more choice, meaningful activity, and privacy than traditional settings. MEASUREMENTS: QoL outcomes were assessed using the Brief Cognitive Rating Scale, Functional Assessment Staging, Cohen-Mansfield Agitation Inventory, Pleasant Events Scale-Alzheimer's disease, Multidimensional Observation Scale of Elderly Subjects, and Apparent Affect Rating Scale. RESULTS: The intervening SCF group demonstrated less decline in activities of daily living, more sustained interest in the environment, and less negative affect than residents in the traditional institutional facilities. There were no differences between groups in concentration, memory, orientation, depression, or social withdrawal. CONCLUSION: The present study suggests that QoL for adults with middle- to late-stage dementia is the same or better in a purpose-built and staffed SCF than in traditional institutional settings.
The rhetorical forms in Umbuso KaShaka (The Realm of Shaka), the Zulu translation of Nada the Lily, are analysed within the framework of Descriptive Translation Studies. The rhetorical forms investigated are individualisation, stereotyping, validation and structuring. A range of translation strategies is employed by the translator to establish the rhetorical forms. For the rhetorical form, individualisation, the strategies re-lexification, catalysis and idiosyncracy are used. Strategies like repetition, filial address, participatory response and fixed expressions are applied to set up the rhetorical form stereotyping. Lexical and semantic transfer, functional and cultural equivalents, cultural substitution and loanshift establish validation. Structuring is dealt with on grammatical (anastrophe and anacolutha) and textual level. Structuring on textual level includes components like punctuation, paragraphing, addition and omission. It is shown that the early Zulu translations are characterised by slight shifts in rhetorical form to restore their original (authentic) form, function and significance, so that they truly reflect Zulu language and culture. On the one hand, the translator was guided by a set of translator's norms relating to this period, namely to obey certain prescribed rules in order to be regarded as a good translator, i.e.to be faithful to the source text. On the other hand, however, an attempt was made to meet the expectations of the target system.
Existing treebanks of written language, as e.g., TIGER [2], TuBa-D/Z [11], Penn Treebank [1] etc., usually consist of sentences that can be considered as grammatically well-formed. The SINBAD treebank we present here covers a completely new domain, namely suboptimal syntactic structures, i.e., sentences which are neither fully grammatical nor completely ungrammatical, but merely suboptimal.1 The treebank consists of a collection of German sentences that are rated suboptimal or ungrammatical in the literature, as well as of sentences drawn from our own experimental work on graded grammaticality judgments. In the literature, these structures are usually compared with grammatical structures which express the same meaning, and for ease of comparison these were sometimes included in the treebank as well. With this data collection we provide access to negative evidence which does not occur in ordinary corpora of written or spoken language. It is characteristic for suboptimal structures that these data are judged incoherently varying between different speakers and in different contexts. It is therefore important to provide a systematic collection of these judgments in order to allow researchers better access to past judgements on the phenomena they are interested in and thus contribute towards greater consistency, even in tricky cases. Since most work in syntactic theory is based on suboptimal or ungrammatical structures, the treebank aims at providing linguists with a data basis for their research. This requires a rich syntactic annotation with linguistically relevant concepts. The linguistic framework of the annotation is that of generative grammar in the sense that the trees are strictly binary branching and contain traces and empty categories. The
Concept mapping is a knowledge elicitation technique which stimulates learners to articulate and synthesize their actual states of knowledge during the learning process. Several approaches have been proposed for automating the assessment procedure of learners' concept maps based on an expert's map as a reference point. However, these approaches do not handle cases where learners have misspelled a concept or they have used a synonym or a concept related to the appropriate one. In this paper we present an alternative approach in which the process of the error identification is performed through the use of an expert map and of WordNet which is an electronic lexical database containing semantic relationships between words. This way we handle cases such as misspelled concepts, synonyms and related concepts. After error detection WordNet is also employed for providing the learner with appropriate feedback based on the identified errors, with the intention of helping the learner to correct them.
Here, we study the design of domain lexicons and address the problems of lexical knowledge representation and linking between different knowledge sources. A domain lexicon contains substantial domain specific vocabularies with associated phonological, morphological, syntactic and semantic/pragmatic information as well as links to general knowledge bases which are rich enough to support knowledge-intensive models for practical NLP systems. We take financial domain as an example to illustrate the representation structures for syntactic and semantic knowledge. In order to suit for both maximal reusability and deep analysis, our domain lexicon is designed with uniform knowledge representation and fine-grain feature encoding. We also address the issues of how to bridge the gaps between coarse-grain general lexicons and fine-grain domain lexicons.
. This paper discusses some pitfalls in corpus research and suggests solutions on the basis of examples and computer simulations. We first address reliability problems in language transcriptions, agreement between transcribers, and how disagreements can be dealt with. We then show that the frequencies of occurrence obtained from a corpus cannot always be analyzed with the traditional χ2 test, as corpus data are often not sequentially independent and unit independent. Next, we stress the relevance of the power of statistical tests, and the sizes of statistically significant effects. Finally, we point out that a t-test based on log odds often provides a better alternative to a χ2 analysis based on frequency counts.
This paper presents an overview of a project to acquire wide-coverage, probabilistic Lexical-Functional Grammar (LFG) resources from treebanks. Our approach is based on an automatic annotation algorithm that annotates “raw” treebank trees with LFG f-structure information approximating to basic predicate-argument/dependency structure. From the f-structure-annotated treebank we extract probabilistic unification grammar resources. We present the annotation algorithm, the extraction of lexical information and the acquisition of wide-coverage and robust PCFGbased LFG approximations including long-distance dependency resolution. We show how the methodology can be applied to multilingual, treebank-based unification grammar acquisition. Finally we show how simple (quasi-)logical forms can be derived automatically from the f-structures generated for the treebank trees.
A statistical estimator attempts to guess an unknown probability distribution by analyzing a sample from this distribution. One desirable property of an estimator is that its guess is increasingly likely to get arbitrarily close to the actual distribution as the sample size increases. This property is called consistency. Data Oriented Parsing (DOP) employs all fragments of the trees in a training treebank, including the full parse-trees themselves, as the rewrite rules of a probabilistic tree-substitution grammar. Since the most popular DOP-estimator (DOP1) was shown to be inconsistent, there is an outstanding theoretical question concerning the possibility of DOP-estimators with reasonable statistical properties. This question constitutes the topic of the current paper. First, we show that, contrary to common wisdom, any unbiased estimator for DOP is futile because it will not generalize over the training treebank. Subsequently, we show that a consistent estimator that generalizes over the treebank should involve a local smoothing technique. This exposes the relation between DOP and existing memory-based models that work with full memory and an analogical function such as k-nearest neighbor, which is known to implement backoff smoothing. Finally, we present a new consistent backoff-based estimator for DOP and discuss how it combines the memory-based preference for the longest match with the probabilistic preference for the most frequent match.
OBJECTIVE: To validate a culturally relevant body image instrument among urban African Americans through three distinct studies. RESEARCH METHODS AND PROCEDURES: In Study 1, 38 medical practitioners performed content validity tests on the instrument. In Study 2, three research staff rated the body image of 283 African-American public housing residents (75% women, mean age = 44 years), with the residents completing body image, BMI, and percentage body fat measures. In Study 3, 35 African Americans (57% men, mean age = 42) completed body image measures and evaluated their cultural relevance. RESULTS: In Study 1, 97% to 100% of practitioners sorted the jumbled figures into the correct ascending order. The correlation between the body image figures and the practitioners' weight classifications of the figures was high (r = 0.91). In Study 2, observers arrived at similar ratings of body size with excellent consistency (alpha = 0.95). Ratings of body image were strongly correlated with participant BMI (r = 0.89 to 0.93 across observers and 0.81 for all participants) and percentage of body fat (r = 0.77 to 0.89 across observers and 0.76 for all participants). In Study 3, body image ratings with the new scale were positively correlated with other validated figural scales. The majority of participants reported that figures in the new body image scale looked most like themselves and other African Americans and were easiest to identify themselves with. DISCUSSION: The instrument displayed strong psychometric performance and cultural relevance, suggesting that the scale is a promising tool for examining body image and obesity among African Americans.
A major methodological challenge in environmental sound research is to select appropriate stimuli. When an experiment involves a large number of sound sources, making custom recordings or producing sounds live is frequently impractical or, for certain sounds, impossible. Existing databases of environmental sound recordings provide a researcher with a useful alternative. However, finding and selecting suitable sounds in such databases can be difficult because of the great variety of sounds present, poor documentation, questionable recording quality, and required purchasing costs. This article describes a number of practical issues to consider during the stimulus selection process, offers a preliminary compilation of existing resources for obtaining environmental sound recordings, provides some normative perceptual data that can be used as a reference for selecting stimuli and evaluating performance, and lists required characteristics and structural aspects of a research-oriented environmental sound database.
Introduction The CorpusEye project (http://corp.hum.sdu.dk ) at the University of Denmark aims at designing and programming an internet based corpus search interface that (1) offers standardised search tools and a unified descriptive formalism across different corpus types and different languages, and (2) allows users to exploit grammatical information in annotated corpora in a user-friendly and menubased way. All corpora in CorpusEye have been annotated with VISL's Constraint Grammar based parsers, in the case of treebanks using an additional PSG module or equivalent (Bick 2003). At the time of writing, the material covers 8 languages and ca. 600 million words.
This paper investigates the usefulness of sentence-internal prosodic cues in syntactic parsing of transcribed speech. Intuitively, prosodic cues would seem to provide much the same information in speech as punctuation does in text, so we tried to incorporate them into our parser in much the same way as punctuation is. We compared the accuracy of a statistical parser on the LDC Switchboard treebank corpus of transcribed sentence-segmented speech using various combinations of punctuation and sentence-internal prosodic information (duration, pausing, and f0 cues).
A statistical corpus-based approach for acquiring selectional preferences of verbs is proposed. By parsing through text corpora, we obtain examples of context nouns that are considered to be the selectional preferences of a given verb. The approach is to generalize initial noun classes to the most appropriate levels on a semantic hierarchy. We present an iterative algorithm for generalization by combining an agglomerative merging and a model selection technique called the Bayesian Information Criterion (BIC). In our experiments, we consider the Web as the large corpora. We also propose approaches for extracting examples from the Web. Preliminarily experimental results are given to show the feasibility and effectiveness of our approach. 1.
This paper explores the interaction between conceptual structure and morpho-syntax. In particular, we show that ontology-based conceptual classification can be used to predict internal relations in compounds. We propose an ontology-based approach to predict the semantic relation between the two component words in Mandarin VV compounds. A Mandarin VV compound is classified according to the eventive relation between the two simplex verbs. These relations specify how the eventive meanings of the two simplex verbs combine to form the meaning of the compound. The three types of eventive relations that we deal with in this paper are: coordinate, modificational, and resultative. Since the way in which two events combine with each other depends upon their event types, we hypothesize that the eventive relations can be predicted by the conceptual classified event types of the two simplex verbs. An approach of ontology-based prediction is proposed based on this hypothesis. The assignment of ontology classification for each simplex verb is based on SUMO and Sinica BOW. The correlation between the ontology class of each verb position and each eventive type is trained and scored based on a manually tagged lexical database. We encode the ontology information of each VV compound in a 3-tuple based on these correlation scores. This 3-tuple is represented as a three-dimensional vector and used to predict the eventive type of new VV compounds. Our classification experiment on unknown VV compounds yields good recall and precision. 1.
The present work falls in the line of activities promoted by the European Languguage Resource Association (ELRA) Production Committee (PCom) and raises issues in methods, procedures and tools for the reusability, creation, and management of Language Resources. A two-fold purpose lies behind this experiment. The first aim is to investigate the feasibility, define methods and procedures for combining two Italian lexical resources that have incompatible formats and complementary information into a Unified Lexicon (UL). The adopted strategy and the procedures appointed are described together with the driving criterion of the merging task, where a balance between human and computational efforts is pursued. The coverage of the UL has been maximized, by making use of simple and fast matching procedures. The second aim is to exploit this newly obtained resource for implementing the phonological and morphological layers of the CLIPS lexical database. Implementing these new layers and linking them with the already exisitng syntactic and semantic layers is not a trivial task. The constraints imposed by the model, the impact at the architectural level and the solution adopted in order to make the whole database ‘speak ’ efficiently are presented. Advantages vs. disadvantages are discussed. 1. Background and Motivations The work described here raises issues in methods, procedures and tools for the reusability, creation, and management of Language Resources (LRs) and has been
In the design of a Multilingual Lexical Database, one of the biggest problems is constituted by conceptual mismatches between languages, and the resulting matter of lexical gaps. Lexical gaps concern words for which there is no direct translation in a target language, but which nonetheless need to receive a translation within the system. In this article, it will be shown that the various possible ways of dealing with these lexical gaps can be classified in four basic groups. Using the SIM<it>u</it>LLDA system as an example (Janssen 2002), the advantages of the structured interlingua approach over the other possibilities will be explained. With the SIM<it>u</it>LLDA set-up, it is possible to derive correct lexical definitions for lexical gaps from the lexical database. How this process of “lexical gap filling” works will be shown using a concrete example of a lexical gap: the treatment of the English words <it>river</it> and <it>stream </it>in contrast with the French words <it>fleuve</it> and <it>rivière</it>.
This paper presents a prosodic phrasing model for Korean to be used in a text-to-speech synthesis (TTS) system. Read text corpora were morpho-syntactically parsed and prosodically labeled following the Penn Korean Treebank (Han, Chunghye, Ko, Eon-Suk, Yi, Heejong, Palmer, M., 2002. Penn Korean Treebank: development and evaluation. In: Proceedings of the 16th Pacific Asian Conference on Language and Computation. Korean Society for Language and Information.) and K-ToBI prosodic labeling conventions (Sun-Ah, J., 2000. K-ToBI (Korean ToBI) labelling conventions. Version 3.1. Available from: URL.), respectively. Decision trees were trained with morpho-syntactic and textual distance features to predict locations of accentual and intonational phrase breaks. Our phrasing model cross-validated on a 300-sentence corpus (6936 words or 21,436 syllables, with an average of 72 syllables or 23 words per sentence) predicted non-breaks with F=92.4% and breaks with F=88.0% (F=72.8% for accentual phrase breaks and F=71.3% for intonational phrase breaks).
The purpose of this paper is to describe the TuBa-D/Z treebank of written German and to compare it to the independently developed TIGER treebank (Brants et al., 2002). Both treebanks, TIGER and TuBa-D/Z, use an annotation framework that is based on phrase structure grammar and that is enhanced by a level of predicate-argument structure. The comparison between the annotation schemes of the two treebanks focuses on the different treatments of free word order and discontinuous constituents in German as well as on differences in phrase-internal annotation.
this paper I discuss a corpus-based approach to the analysis of some phenomena of lexical semantics using empirical data drawn from a Russian-German paraUel corpus ofDostoevskij's Idiot together with its German translations. This parallel corpus is part of the Austrian Academy Corpus (AAC) at the Austrian Academy of Sciences in Vienna. The subject of investigation is lexical co-occurrences which determine the combinatorial profile ofa word. A corpus-based analysis oflexical co-occurrences contributes to both monolingual and buingual lexicography by providing new and more detailed insights into the contextual behaviour of a word. From the diachronic perspective the semantic change comes about at the periphery of the combinatorial profile of a given word. Some of the peripheral co-occurrences can become so frequent that they drift from periphery to centre while others fall out of use and start to be perceived as norm violations. The comparison ofcombinatorial profiles ofthe same word in the 1860s and in present day Russian proves to be an efficient instrument for defining the combinatorial norms of a given word against the background of its near-synonyms. The comparison of a given word with all possible translation equivalents has a similar function.
This paper is a contribution to the issue -- which has, in the course of the last decade, become critical -- of the basic requirements and validation criteria for lexical language resources in Standard Arabic. The work is based on a critical analysis of the architecture of the DIINAR.1 lexical database, the entries of which are associated with grammar-lexis relations operating at word-form level (i.e. in morphological analysis). Investigation shows a crucial difference, in the concept of 'lexical database', between source program and generated lexica. The source program underlying DIINAR.1 is analysed, and some figures and ratios are presented. The original categorisations are, in the course of scrutiny, partly revisited. Results and ratios given here for basic entries on the one hand, and for generated lexica of inflected word-forms on the other. They aim at giving a first answer to the question of the ratios between the number of lemma-entries and inflected word-forms that can be expected to be included in, or generated by, a Standard Arabic lexical dB. These ratios can be considered as one overall language-specific criterion for the analysis, evaluation and validation of lexical dB-s in Arabic.
Abstract This paper surveys work on applying the insights of lexicalized grammars to low‐level discourse, to show the value of positing an autonomous grammar for low‐level discourse in which words (or idiomatic phrases) are associated with discourse‐level predicate–argument structures or modification structures that convey their syntactic‐semantic meaning and scope. It starts by describing a lexicalized Tree Adjoining Grammar for discourse (D‐LTAG). It then reviews an initial experiment in parsing text automatically, using both a lexicalized TAG and D‐LTAG, and then touches upon issues involved in how lexico‐syntactic elements contribute to discourse semantics. The paper concludes with a brief description of the Penn Discourse TreeBank, a resource being developed for the study of discourse structure and semantics.
The Human Use Regulatory Affairs Advisor (HURAA) is a Web-based facility that provides help and training on the ethical use of human subjects in research, based on documents and regulations in United States federal agencies. HURAA has a number of standard features of conventional Web facilities and computer-based training, such as hypertext, multimedia, help modules, glossaries, archives, links to other sites, and page-turning didactic instruction. HURAA also has these intelligent features: (1) an animated conversational agent that serves as a navigational guide for the Web facility, (2) lessons with case-based and explanation-based reasoning, (3) document retrieval through natural language queries, and (4) a context-sensitive Frequently Asked Questions segment, calledPoint & Query. This article describes the functional learning components of HURAA, specifies its computational architecture, and summarizes empirical tests of the facility on learners.
The modern dominant approach to translation is that which aims at intelligibility of the meaning and normality of the style in the translated text. Many translation theorists maintain that a translator should attempt to produce a target text which is clear and understandable in meaning, normal in style, and natural in language. To attain this goal, one is allowed and sometimes obliged, to make some adjustments such as expansion and reduction in the process of transfer according to the linguistic norms of the receptor language. However, in translating a highly sensitive religious text like the Qur'an, the accuracy of the meaning conveyed in the translated text is of primary importance, since it is the source text which is considered as the main criterion in evaluating the translation. Therefore, no translator is allowed to sacrifice accuracy of the meaning for the sake of intelligibility and naturalness in the translation, nor is it legitimate to make the meaning unintelligible or distorted by producing a target text which is awkward andunnatural. What on should do is to produce a translation natural in language, intelligible and accurate in meaning
The architecture of a lexical database in which multilingual semantic networks would be stored requires the incorporation of e xible mechanisms and services, which would enable the efcient navigation within and across lexical data. We report on WordNet Management System (WMS), a system that functions as the interconnection and communication link between a user and a number of interlinked WordNets. Semantic information is being accessed through a distributed network of servers, forming a large-scale multilingual semantic network.
In this paper we address the following questions from our experience of the last two and a half years in developing a large-scale corpus of Arabic text annotated for morphological information, part-of-speech, English gloss, and syntactic structure: (a) How did we 'leapfrog' through the stumbling blocks of both methodology and training in setting up the Penn Arabic Treebank (ATB) annotation? (b) How did we reconcile the Penn Treebank annotation principles and practices with the Modern Standard Arabic (MSA) traditional and more recent grammatical concepts? (c) What are the current issues and nagging problems? (d) What has been achieved and what are our future expectations?
This paper reports on an ongoing project that uses varied language resources and advanced NLP tools for a linguistic classification task in discourse semantics. The system we present is designed to assign a “situation entity ” class label to each predicator in English text. The project goal is to achieve the best-possible identification of situation entities in naturally-occurring written texts by implementing a robust system that will deal with real corpus material, rather than just with constructed textbook examples of discourse. In this paper we focus on the combination of multiple information sources, which we see as being vital for a robust classification system. We use a deep syntactic grammar of English to identify morphological, syntactic, and discourse clues, and we use various lexical databases for fine-grained semantic properties of the predicators. Experiments performed to date show that enhancing the output of the grammar with information from lexical resources improves recall but lowers precision in the situation entity classification task. 1.
The identification of dysfunctional thoughts is a central effort in cognitive therapy. This paper describes the first version of a computer module that classifies dysfunctional thoughts automatically. It is part of COGNO, a system we are developing to give automatic feedback on dysfunctional thoughts. The system uses rules that were developed from language markers identified in a sample of 149 dysfunctional thoughts. The system was tested with an independent set of 112 example thoughts. The system detects the majority of dysfunctional thoughts, but works reliably only for some thought categories. Automatic thought classification may be a first step toward developing natural dialogue systems in cognitive therapy.
Schwarz (2001, 2002) proposed the ex-Wald distribution, obtained from the convolution of Wald and exponential random variables, as a model of simple and go/no-go response time. This article provides functions for the S-PLUS package that produce maximum likelihood estimates of the parameters for the ex-Wald, as well as for the shifted Wald and ex-Gaussian, distributions. In a Monte Carlo study, the efficiency and bias of parameter estimates were examined. Results indicated that samples of at least 400 are necessary to obtain adequate estimates of the ex-Wald and that, for some parameter ranges, much larger samples may be required. For shifted Wald estimation, smaller samples of around 100 were adequate, at least when fits identified by the software as having ill-conditioned maximums were excluded. The use of all functions is illustrated using data from Schwarz (2001). The S-PLUS functions and Schwarz’s data may be downloaded from the Psychonomic Society’s Web archive, www. psychonomic.org/archive/.
The current research explored the processes that predominate during the anticipation of an emotionally salient event. Experiment 1 (N536), employed three different conditional stimuli followed by pictorial pleasant, unpleasant or neutral unconditioned stimuli. Half the participants were trained with visual CSs, the other half with tactile CSs. In the group trained with visual CSs, startle eyeblinks were larger and faster during CSs that were paired with unpleasant pictures than CSs paired with neutral or pleasant pictures respectively, indicating an affect startle pattern. This linear trend was not found in the group trained with tactile CSs. Experiment 2 (N564) aimed to investigate whether the affective pattern found in the startle data in Experiment 1 could also be found using a behavioural measure of emotion. This time participants’ reaction time during a post-experimental affective priming taskwas used as dependantmeasure to assess the presence of emotional learning. Instead of a simple differential conditioning task, an occasion setting paradigm was employed and participants were trained using either a feature positive or feature negative design with pleasant or unpleasant picture USs. For participants trained with unpleasant USs, valence ratings collected before and after conditioning training suggested the presence of emotional learning, whereas no such pattern was found for participants trained with pleasant USs. These findings were not confirmed in the priming data.
Reviewed by: Prague Linguistic Circle papers ed. by Eva Hajičová et al. Zdenek Salzmann Prague Linguistic Circle papers. Travaux du cercle linguistique de Prague, new series, vol. 4. Ed. by Eva Hajičová, Petr Sgall, Jiříhana, and Tomáš Hoskovec. Amsterdam: John Benjamins, 2002. Pp. vii, 376. ISBN 158811175X. $125 (Hb). This is the fourth volume in the TCLP series begun in 1995. It consists of thirteen papers assigned to five parts. The majority of contributors are from the Czech Republic; also represented are scholars from Germany (two) and from the United States, Russia, and Israel (one from each country). Part 1, ‘The Prague tradition in retrospect’ (1–108), contains three papers by now-deceased members of the Circle. The first is by Josef Vachek (1909–1996), the Circle’s most courageous defender during the bleak post-World War II years. His long paper, ‘Prolegomena to the history of the Prague School of Linguistics’, was first published in 1999, three years after his death (I reviewed the original Czech version in Language 76.471–72, 2000). The second paper, by Oldřich Leška (1927–1997), is titled ‘Anton Marty’s philosophy of language’ (83–99). Swiss-born Marty taught at Prague’s German University until his death in 1914, and some of his lectures were attended by Vilém Mathesius, the Circle’s founder. According to Leška, ‘it was easy for Mathesius [to transform Marty’s] general ideas [about language] into an effective system of descriptive functional grammar’ (92). The last paper in this part is Vladimír Skalička’s ‘Die Typologie des Ungarischen’ (101–8). Skalička (1909–1991), a Finno-Ugricist, taught general linguistics, and many of his publications dealt with linguistic typology; this chapter is a translation of a paper published in 1967 in Hungarian. The first paper of Part 2, ‘Grammar’ (109–81), is by Eva Hajičová. In her article ‘Theoretical description of language as a basis of corpus annotation: The case of Prague Dependency Treebank’, the author argues that annotation (tagging) scenarios should include an underlying level of syntactic annotation, and she then attempts to demonstrate that the design of such a scenario is feasible if it is based on a sound, explicit linguistic theory. Next, Yishai Tobin asks if ‘conditionals’ in Hebrew and English are the same or different, and concludes that the Hebrew system of conditionals classifies alternative perceptions of possibilities according to a totally different set of semantic criteria. Indo-Europeanists would be interested in Tomáš Hoskovec’s article ‘Sur la paradigmatisation du verbe indo-européen’ (begun in TCLP 3), in which he describes the structure of verbal paradigms in Greek and Latin, and in Baltic and Slavic languages. Part 3, ‘Topic-focus articulation’ (183–305), begins with a paper by Vladimir Borschev and Barbara H. Partee. While a full account of the problem of the Russian genitive of negation in existential sentences is yet to be attained, a picture may slowly be emerging: aspects of both functional and formal (syntactic) approaches may have to be employed for an adequate description. The aim of Libuše Dušková’s paper is to show that the functional sentence perspective (FSP) structure of the different syntactic realizations of the second participant in verbal action with the FSP function of the theme is to some extent differentiated rather than interchangeable. In a paper originally published in 1995, Jaroslav Peregrin claims that the theory of generalized quantifiers can provide for an adequate framework for the analysis of the topic-focus articulation. Klaus von Heusinger presents a different approach to dealing with information structure. He argues for a two-level semantics, with one level corresponding to the meaning of a sentence, and the other to the background meaning. In Part 4, ‘General views’ (307–62), Petr Sgall writes about the nature, sources, and consequences of freedom in natural languages. Philip A. Luelsdorff proposes that English has four tenses (future, present, past, and generic) and four aspects (to future, progressive, perfect, and passive); the tenses express...
Idiolects are person-dependent similarities in language use. They imply that texts by one author show more similarities in language use than texts between authors. Sociolects, on the other hand, are group-dependent similarities in language use. They imply that texts by a group of authors, for instance in terms of gender or time period, share more similarities within a group than between groups. Although idiolects and sociolects are commonly used terms in the humanities, they have not been investigated a great deal from corpus and computational linguistic points of view. To test several idiolect and sociolect hypotheses a factorial combination was used of time period (Modernism, Realism), gender of author (male, female) and author (Eliot, Dickens, Woolf, Joyce) totaling 16 corresponding literary texts. In a series of corpus linguistic studies using Boolean and vector models, no conclusive evidence was found for the selected idiolect and sociolect hypotheses. In final analyses testing the semantics within each literary text, this lack of evidence was explained by the low homogeneity within a literary text.
Recent performance appraisal research has highlighted the important role played by contextual and individual factors in shaping rating behavior. This article reviews cumulated empirical data supporting the proposition that in factors and constraints present in the organization, contexts in which appraisal systems reside and rater attributes, such as personality factors or beliefs, systematically affect rating behavior. The effects of these context and rater factors are reflected in ratings accuracy, ratings discrimination among raters/dimensions, and rating elevation.
Abstract. The use of lexicons and corpora advances both linguistic re-search and performance of current natural language processing (NLP) systems. We present a tool that exploits such resources, specifically En-glish and German lexical databases and the World Wide Web to recognise English inclusions in German newspaper articles. The output of the tool can assist lexical resource developers in monitoring changing patterns of English inclusion usage. The corpus used for the classification covers three different domains. We report the classification results and illustrate their value to linguistic and NLP research. 1
Abstract. An important aspect of discourse understanding and generation involves the recognition and processing of discourse relations. These are conveyed by discourse connectives, i.e., lexical items like because and as a result or implicit connectives expressing an inferred discourse relation. The Penn Discourse TreeBank (PDTB) provides annotations of the argument structure, attribution and semantics of discourse connectives. In this paper, we provide the rationale of the tagset, detailed descriptions of the senses with corpus examples, simple semantic definitions of each type of sense tags as well as informal descriptions of the inferences allowed at each level. 1
We present a method to approximate a LTAG grammar by a CFG. A key process in the approximation method is finite enumeration of partial parse results that can be generated during parsing. We applied our method to the XTAG English grammar and LTAG grammars which are extracted from the Penn Treebank, and investigated characteristics of the obtained CFGs. We perform CFG filtering for LTAG by the obtained CFG. In the experiments, we describe that the obtained CFG is useful for CFG filtering for LTAG parser. 1
The method of organization of word meanings is a crucial issue with lexical databases. Our purpose in this research is to extract word hierarchies from corpora automatically. Our initial task to this end is to determine adjective hyperonyms. In order to find adjective hyperonyms, we utilize abstract nouns. We constructed linguistic data by extracting semantic relations between abstract nouns and adjectives from corpus data and classifying abstract nouns based on adjective similarity using a self-organizing semantic map, which is a neural network model (Kohonen 1995). In this paper we describe how to hierarchically organize abstract nouns (adjective hyperonyms) in a semantic map mainly using CSM. We compare three hierarchical organizations of abstract nouns, according to CSM, frequency (Tf.CSM) and an alternative similarity measure based on coefficient overlap, to estimate hyperonym relations between words.
This paper presents a deterministic dependency parser based on memory-based learning, which parses English text in linear time. When trained and evaluated on the Wall Street Journal section of the Penn Treebank, the parser achieves a maximum attachment score of 87.1%. Unlike most previous systems, the parser produces labeled dependency graphs, using as arc labels a combination of bracket labels and grammatical role labels taken from the Penn Treebank II annotation scheme. The best overall accuracy obtained for identifying both the correct head and the correct arc label is 86.0%, when restricted to grammatical role labels (7 labels), and 84.4% for the maximum set (50 labels).
In information retrieval and text mining, information on word senses is usually taken from dictionaries or lexical databases that have been prepared by lexicographers. We propose an automatic method for word sense induction, i.e. for the discovery of a set of sense descriptors to a given ambiguous word. The approach is based on the statistics of word co-occurrence as derived from Web pages. The underlying assumption is that the senses of an ambiguous word are best described by terms that, although bearing a strong association to this word, are mutually exclusive, i.e. whose association strength within the retrieved Web pages is as weak as possible. Measuring association strength is based upon a novel confidence gain approach that relates the observed co-occurrence frequency for two sense descriptor candidates to an average co-occurrence frequency for pairs of arbitrary words. The proposed approach is fully unsupervised and takes into account the contemporary meanings of words, as reflected in texts from the Internet. Our results are evaluated using a list of ambiguous words commonly referred to in the literature.
Abstract. WordNet is an electronic lexical database structured around psychological and linguistic principles. As such, it should play a part in any e ort to integrate cognitive factors into knowledge based systems. Yet, some of its basic assumptions have been attacked, and the suggestion made that a major restructuring would make it more cognitively transparent. We investigate these allegations from a psycholinguistic perspective and conclude that WordNet is in fact rigorous in terms of the cognitive principles it embodies. What is lacking is a methodology for translating the explicit and implicit knowledge in WordNet into a usable, formal ontologies. We show some ways in which WordNet should be extended to facilitate this process. We agree that WordNet is not in itself ready for use as a formal ontology, but we argue that it is an invaluable tool for describing the conceptualized structure of our world, and should be used as a fundamental resource. 1