Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Word recognition and generation is a fundamental part of the processing of natural language and it requires computationally effective morphological processors, especially for languages with rich morphology such as Modern Greek. Various models have been proposed for developing computerized systems to accomplish the task of recognition of morphosyntactic features of words In the work presented here, the lazy tagging approach was examined, in which taggers are expected to work in the simplest possible way. The model of functional decomposition was extended and adapted for Modern Greek as a target language, following the lazy word-parsing approach, in order to cover a number of morphological phenomena that are encountered in Modern Greek, namely inflection, affixation, and longdistance dependencies. To achieve a more efficient word recognition, several automata of different levels of computing power, based on the original model, were introduced and evaluated according to the criteria of complexity, recognition speed, and accuracy of the results. The proposed system was used for processing a large-scale corpus, and the results are presented and discussed. To accomplish their task, taggers can rely upon large lexical databases, which are expected to be organized in such a way as to provide rapid access to the stored data and efficient memory management. Directed graphs can be used to describe and organize a lexical database of large magnitude in a compact manner. These data structures are named here matrix lexica, where the letters are described as nodes of directed graphs and the lemmata as paths (set of edges). It is expected that matrix lexica will support a tagger efficiently by providing a high speed of resolution, sound mathematical foundation, low memory requirements, and ability to handle distorted input in future developments.
Accurate linguistic annotation is a core requirement of natural language processing systems. The demand for accuracy in the face of rapid prototyping constraints and numerous target languages has led to the employment of machine learning methods for developing linguistic annotation systems. The popularity of applying machine learning methods to computational linguistics problems has given rise to a large supply of trainable natural language processing systems. Most problems of interest have an array of off-the-shelf products or downloadable code implementing solutions using various techniques. In situations where these solutions are developed independently, it is observed that their errors tend to be independently distributed. In this thesis we discuss approaches for capitalizing on this situation in a sample problem domain, Penn Treebank-style parsing. The machine learning community provides us with techniques for combining outputs of classifiers, but parser output is more structured and interdependent than classifications. To overcome this, two novel strategies for combining parsers are used: learning to control a switch between parsers and constructing a hybrid parse from multiple parsers' outputs. In this thesis we give supervised and unsupervised techniques for each of these strategies as well as performance and robustness results from evaluation of the techniques. One shortcoming of combining off-the-shelf parsers is that the parsers are not developed with the intention to perform well on complementary data or to compensate for each others' weaknesses. The individual parsers are globally optimized. We present two techniques for producing an ensemble of parsers in such a way that their outputs can be constructively combined. All of the ensemble members will be created using the same underlying parser induction algorithm, and the method for producing complementary parsers is only loosely coupled to that algorithm.
This paper proposes a new inference approach for Chinese probabilisticcontext-free grammar, which implements the EM algorithm based on the bracketmatching schemes. Two characteristics of the algorithm are as follows: 1) To pre-process the training texts with automatic constituent boundary prediction tools,which can provide stronger syntactic restriction upon training texts in lower compu-tational costs; 2) To develop an initial rule set by integrating different knowledgeresources, including a set of basic syntactic rules generated by an automatic gram-mar construction t00l and a set of special rules summarized by linguists or extractedfrom treebanks, and provide a better initialization for the learning process. There-fore, a linguistically-motivated and broad-coverage Chinese PCFG rule set can beeasily generated through this algorithm. Current experimental results prove goodlearning efficiency of this algorithm and high reliability of the generated rule set.
We argue that the current dominant paradigm in parser evaluation work, which combines use of the Penn Treebank reference corpus and of the Parseval scoring metrics, is not well-suited to the task of general comparative evaluation of diverse parsing systems. In (Gaizauskas et al., 1998), we propose an alternative approach which has two key components. Firstly, we propose parsed corpora for testing that are much flatter than those currently used, whose “gold standard ” parses encode only those grammatical constituents upon which there is broad agreement across a range of grammatical theories. Secondly, we propose modified evaluation metrics that require parser outputs to be ‘faithful to’, rather than mimic, the broadly agreed structure encoded in the flatter gold standard analyses. This paper addresses a crucial issue for the (Gaizauskas et al., 1998) approach, namely, the creation of the evaluation resources that the approach requires, i.e. annotated corpora recording the flatter parse analyses. We argue that, due to the nature of the resources required, they can be derived in a comparatively inexpensive fashion from existing parse annotated resources, where available. 1.
En se situant dans le cadre d'un projet appele Semantic dictionary viewed as a lexical database dont l'objectif est d'identifier et de denombrer les informations linguistiquement pertinentes pouvant etre derivees des definitions du sens, l'A. examine ici les processus productifs de derivation semantique des verbes en russe, et en particulier les processus associes a un changement metonymique qui se refletent dans le cadre du cas profond d'un verbe. Il tente ainsi de montrer que les roles semantiques permettent de decouvrir l'invariant semantique d'une grande partie des derivations productives
This paper examines y deletion in Seoul Korean on a large socio-linguistic database. It first shows that Seoul Korean has two distinct processes y deletion: categorical y deletion, which occurs after a palatal consonant, and variable y deletion occurring only before the vowel e. The first part of this paper examines variable y deletion. The examina- tion of the process reveals that it is conditioned by such internal factors as the existence of a preceding consonant, the nature of the preceding consonant, and word- internal position where y occurs, i.e., whether y appears in the word-initial or non- ward- initial syllable. This process is also affected by external factors such as the speech style, social class, and age of the speaker. The difference in the deletion rate of y among the three age groups (and among the social class groups, additionally) \nis taken to suggest that y deletion before e, i.e., monophthongization of ye to e, is a change in progress. This paper then attempts a phonological account of two different process of y deletion, proposing two OCP (coronal) constraints. The two constraints are proposed as having rather different strengths: the stronger one triggers categorical deletion and the weaker, variable deletion. A phonetic explanation of the ongoing y deletion change is also attempted. It is shown that ye has the shortest perceptual distance between its component segments among all y diphthongs of Seoul Korean. It is suggested that a relative lack of perceptual distinction between the two segments of the diphthong ye is responsible for the ongoing loss of y in this diphthongal sequence.
Inscriptions, a type of source which has hitherto hardly been taken into account in historical linguistics, can provide data for a whole series of questions thrown up by research into the linguistic geography of Early New High German and the Late Middle Low German of the same period. In respect of the pressure towards standardization exerted by M. Luther's bible, various degrees can be differentiated. The greatest willingness to adopt this linguistic norm is found in the case of New High German monophthongization and diphthongization, as well as in the lowering of u to o before nasals and the use of the verbal prefix ver-. In these cases, the symbols on the appropriate maps show agreement with Luther 's usage. In some other cases, namely, the use of for MHG /ei/, the occurrence of vowel forms without unrounding, the use of the uncontracted form of the word nicht and the use of the subjunctive in the predicates of concessive clauses, we find that the form favoured by Luther is the one which appears most frequently in the inscriptions. Nevetheless, especially in sixteenth-century inscriptions, we find several examples of the replacement of standard forms by deviant variants valid at the same time. In contrast, the number of deviations from Luther's language visibly decreases in the seventeenth century. In some cases, the treatment of the norms of the Luther's bible in the inscriptions indicates that the linguistic usage of the reformer had no or only indirect significance for the course of development here. Thus, in the case of the feminine inflectional system, while it is true that the epigraphic sources largely follows Luther's practice, but it is also the case that from about 1600 on a tendency towards deviation from his usage emerges. Here, restructuring is a process which was clearly not influenced by Luther. Again, in the case of abstract suffix -nis, the course of development seems to have been determined by factors inherent in the language rath
This document describes adaptation of the basic Sarnoff JND Vision Model created in response to the NASA/ARPA need for a general-purpose model to predict the perceived image quality attained by flat-panel displays. The JND model predicts the perceptual ratings that humans will assign to a degraded color-image sequence relative to its nondegraded counterpart. Substantial flexibility is incorporated into this version of the model so it may be used to model displays at the sub-pixel and sub-frame level. To model a display (e.g., an LCD), the input-image data can be sampled at many times the pixel resolution and at many times the digital frame rate. The first stage of the model downsamples each sequence in time and in space to physiologically reasonable rates, but with minimum interpolative artifacts and aliasing. Luma and chroma parts of the model generate (through multi-resolution pyramid representation) a map of differences-between test and reference called the JND map, from which a summary rating predictor is derived. The latest model extensions have done well in calibration against psychophysical data and against image-rating data given a CRT-based front-end. THe software was delivered to NASA Ames and is being integrated with LCD display models at that facility,
Recent approaches to statistical parsing include those that estimate an approximation of a stochastic, lexicalized grammar directly from a treebank and others that rebuild trees with a number of tree-constructing operators, which are applied in order according to a stochastic model when parsing a sentence. In this paper we take an entirely different approach to statistical parsing, as we propose a method for parsing using a Hidden Markov Model. We describe the stochastic model and the tree construction procedure, and we report results on the Wall Street Journal Corpus.
This paper describes the combination compound unit (CU) recognizer with syntactic verifier using partial parsing mechanism. The recognizer finds all the CUs, combined concept including collocations, idioms, and compound nouns, in input sentence. CU information reduces the search space of syntactic analysis and a portion of Part-Of-Speech (POS) ambiguities. Syntactic verification is to obtain precise CU recognition results by means of pruning wrongly recognized units that are caused by improper variable hypotheses. The experimental results show the precision of CU recognition is increased to 99.69% with 31 CFG rules on cyclic trie structure for 1,268 WSJ articles in the Penn Treebank. They also show CU recognition increases the understandability of translation for Web documents.
Abstract. This article sets out some of the results of wider research on the linguistic databases of a natural language generation system. One of the necessary steps in the building of such databases is to determine the linguistic means the generator must have in order to produce a linguistic form that corresponds to the semantic representation given as an input. We wish to focus here on the theoretical choices and issues rather than on the application itself. We assume that texts have a syntactic structure, whose characteristics are partly comparable to the syntactic structure of a sentence, and that a connective can be considered as a textual predicate which has arguments that are constrained in the same way as the arguments of a verb. This article will concentrate more specifically on one particular semantic relation — the simultaneity of two events — and will show how the taxonomy of the associated connectives can be elaborated. Finally, we will set out some of the major developments of this research, which concern the interface between conceptual and linguistic knowledge.
Children's judgements about pain at age 8-10 years were examined comparing two groups of children who had experienced different exposure to nociceptive procedures in the neonatal period: extremely low birthweight (ELBW) <or = 1000 g (N = 47) and full birthweight (FBW) > or = 2500 g (N = 37). The 24 pictures that comprise the Pediatric Pain Inventory, depicting events in four settings: medical, recreational, daily living, and psychosocial, were used as the pain stimuli. The subjects rated pain intensity using the Color Analog Scale and pain affect using the Facial Affective Scale. Child IQ and maternal education were statistically adjusted in group comparisons. Pain intensity and pain affect related to activities of daily living and recreation were significantly higher than psychosocial and medically related pain on both scales in both groups of children. Although the two groups of children did not differ overall in their perceptions of pain intensity or affect, the ELBW children rated medical pain intensity significantly higher than psychosocial pain, unlike the FBW group. Also, duration of neonatal intensive care unit stay for the ELBW children was related to increased pain affect ratings in recreational and daily living settings. Despite altered response to pain in the early years reported by parents, on the whole at 8-10 years of age ELBW children judged pain in pictures similarly to their term peers. However, differences were evident, which suggests that studies are needed of biobehavioural reactivity to pain beyond infancy, as well as research into beliefs, attitudes, and perceptions about pain during the course of childhood in formerly ELBW children.
198 LANGUAGE, VOLUME 74, NUMBER 1 (1998) cussion of specific examples reveals some characteristic aspects of agrammatic output. Ch. 4, 'The grammar of connected agrammatic speech', completes the summary of grammatical characteristics of nonfluent aphasia. It argues against the label 'telegraphic speech' and against the theory of 'economy of articulatory effort' in explaining agrammatic behavior. Ch. 5, 'Speech, writing, and oral reading', discusses differences and similarities of impairment among these different output modes. Examples of spontaneous writing illustrate the possibility of dramatic dissociation between written and spoken language. Examples from oral reading show that a patient may tend to make the same types of errors in reading as in speech. Ch. 6, 'Bilingual and polyglot aphasia', by Loraeme K. Obler, José Centeno, and Nancy Eng, discusses the similarities and differences between languages in aphasies who are multilingual. They point out that deficits and recovery are usually parallel but that variance can occur (resulting, for example, from structural differences between languages). The authors then discuss special behaviors and brain organization of the bilingual aphasie. They end the chapter with advice to speech-language pathologists on diagnosis of patients who speak a language or dialect unfamiliar to the clinician. Ch. 7, 'Inventing therapy for aphasia', by Audrey L. Holland and Claire Penn, is a fascinating discussion of how to devise therapy when the patient and therapist do not share the patient's primary language. They address such issues as how to choose which language to work with and the importance of cross-cultural influences on therapy design. In sum, this book is primarily designed to inform clinicians in their efforts to provide therapy. However, the liberal exemplification of agrammatic output should be of interest to even nonclinical researchers. [Sherri K. Shaw, University of Texas, Austin. ] Xenismen: Die Nachahmung fremder Sprachen. By Wolfgang Moser (Europ äische Hochschulschriften, Reihe 21, 159.) Frankfurt am Main Peter Lang, 1996. Pp. 284. Xenisms are phenomena that characterize somebody or something as foreign. They can be nonlinguistic (a kimono, national anthems, schnitzel) or linguistic (Chinese characters, a foreign accent, loan words). Most importantly, xenisms must be considered as typical of a foreign people or language, even if this doesn't coincide with reality: most Germans don't wear leather breeches, and Chinese people do have an ItI phoneme distinct from IV. In his PhD thesis (University of Graz, Austria), Moser gives a detailed linguistic and semiotic analysis ofxenisms that imitate foreign languages, limiting the scope of his study to intentionally used interlingual xenisms in written texts. He then discusses the role of foreignness and points out that on a scale 'total ignorance—intimate knowledge', foreignness is more or less near to ignorance but not identical to it. You have to know at least something (even something wrong) about other people to see them as foreign and different from yourself. You must know that Fritz is a German name to evoke 'germanness', but you need not know more (e.g. that hardly anybody in Germany would call their son Fritz nowadays, but rather Michael or Kevin). This explains why linguistic xenisms do not appear at random but are associated with the languages and cultures in contact with one's own. The introductory chapter (1 3-29) is completed by a series ofprivative dichotomies: Xenisms can be cotextual or implanted, reduced or fully decodable, interlingual or intralingual, spontaneous or conventional, foreign or pseudo-foreign, foreign elements (mostly words) or foreign usages (e.g. word order rules). Finally, the relation of xenisms to a certain language can be more or less vague or precise. Apart from the introductory chapter, the book is divided into two main parts: The first part (31-129) analyzes the linguistic structure of xenisms; the second part (131-254) deals with xenisms as semiotic phenomena. In the structural analysis, M explains xenisms as deviations from linguistic norms (following Eugenio Coseriu's definition of norm). Their forms range from the imitation of hieroglyphs to 'typical' name endings and are broadly illustrated with examples from Asterix and its translations. The second part of this chapter analyzes the various languages that are used to characterize different people in Jaroslav Hasek's The good soldier Svejk and how the xenistic effects...
412 LANGUAGE, VOLUME 74, NUMBER 2 (1998) machine translation. Its substantial bibliography and its author and subject indexes are useful research tools in themselves. However, it cannot be considered a pedagogically-oriented introduction to the theory and practice of translation for the earlier stages of translator training. The book results from a lecture series given by the author in Finland in 1993 and is pitched at quite a demanding level although the claim that the 'individual chapters are relatively self-contained [and] can be read largely independently' (xiii) is justified. For this reason, it will be a very useful source of supplementary readings in advanced translator training. Also, teachers of translation, translation critics, and translators looking for some time out to reflect on the nature of their craft will benefit from it. They may end up agreeing, however, that at times W is not completely innocent of the 'pretentious, glutinous, heavily metaphorical or extremely abstract prose' (3) that he chides others who write about translation for using. Also, readers who know German will benefit more than those who don't from the fairly large sections of German text that are sometimes included to exemplify a point. While we should be surprised at a book on translation that doesn't include at least some other-language material, the growing internationalization of TS means that authors can no longer assume familiarity with particular languages by their readers. In such cases, interlinear glosses and/or literal back translations should be provided. While rejecting the impossibility of translation, W' s consideration of the text-related and translatorrelated problems inherent in translation in no way diminishes his appreciation of the complexity of the task. However, his balanced perspective always encourages a positive and hopeful outlook that interlingual communication is indeed possible, e.g. translation 'contains both culture-specific and cultureuniversal components' (90). Compensatory linguistic behaviors and adaptive skills and strategies can be acquired to enable the interlingual/intercultural gap to be bridged where necessary. W has little use for (sentence-based) generative theory which allows for a linguistic creativity that is both inapplicable and uninteresting to TS. On the other hand, modern cognitive linguistics is seen as having a most useful input into (text-based) translation theory and practice which should seek to operate in 'an interdisciplinary, cognitively embedded framework' (xiii). Such an approach will allow for the creativity displayed in linguistic performance to be given the central significance it deserves. W reminds us that the modern phase of 'TS... is still a fairly young and methodologically unstable field of research' (2), but this volume is indeed a most useful advance. [Robert Early, University of the South Pacific, Vanuatu.] Dictionary of Caribbean English usage. Ed. by Richard Allsopp. (French and Spanish supplement edited by J. E. Allsopp.) Oxford: Oxford University Press, 1996. Pp. lxxviii, 697. This dictionary (hereafter referred to as DCE) is a groundbreaking publication derived from fieldwork and approximately one thousand bibliographic sources. It presents extensive data on English spoken in the Anglophone West Indies (including the Bahamas, Belize, and Guyana) and is certain to become a valuable resource in the fields of both Creole and English studies. Allsopp should be congratulated for finishing what must have seemed a daunting project when begun more than 25 years ago. The aims of DCE are different from an earlier landmark in lexicography, the Dictionary ofJamaican English (ed. by F. G. Cassidy and R. B. Le Page, Cambridge University Press, 1967). That dictionary established the goal ofhistorically describing the lexicon of Jamaican English, including both creóle varieties as well as more standard forms. DCE, however, is less devoted to historical principles than to language planning, seeking to establish a norm for Caribbean English while identifying some regional variation. What is identified as 'Caribbean English' is in fact a narrowly-defined representation of lexical entries and idioms associated primarily with 'acrolectal ' and 'mesolectal' varieties. DCE purposely excludes entries associated with 'deeper' creóle forms, and consequently, it assumes an uncomfortably (and admittedly) prescriptive tone (see xxvi). A should keep in mind that the bundle of features which creolists often consider as constituting the so-called basilect, mesolect, and ACROLECT are ambiguous at best and artificial at...
In two studies, the alternate-form reliability of the Snodgrass picture fragment completion test of implicit memory (Snodgrass, Smith, Feenan, & Corwin, 1987) was examined. In this test, identification thresholds are established for fragmented pictures. The same fragmented pictures are then shown again, intermixed with new fragmented pictures. Implicit memory is indicated by a decrease in identification threshold from the first to second presentation. Alternate-form reliability was low to moderate, depending on the measure used, regardless of the length of the test. A third study showed that explicit memory instructions did not increase the reliability. Recommendations for use of the test in correlational and experimental research are presented.
Psychologists have used artificial neural networks for a few decades to simulate perception, language acquisition, and other cognitive processes. This paper discusses the use of artificial neural networks in research on semantics—in particular, in the investigation of abstract noun meanings. It is widely acknowledged that a word’s meaning varies with its contexts of use, but it is a complex task to identify which context elements are relevant to a word’s meaning. The present study illustrates how connectionist networks can be used to examine this problem. A simple feedforward network learned to distinguish among six abstract nouns, on the basis of characteristics of their contexts, in a corpus of randomly selected naturalistic sentences.
The article reports on one of the more sophisticated critical editions ever to be published in electronic format. The Wife of Bath is richly encoded, provides access to literally thousands of manuscript images, and enables users to assess the relationships between the numerous extant manuscript editions. The authors assess the methods used in the edition's development and the lessons learned through its production.
The authors documented a linguistic norm account of direction of comparison asymmetry effects in relational judgments (e.g., seeing hyenas as more similar to dogs than dogs are similar to hyenas). The asymmetry effect is magnified by discrepancies in prominence between subject and reference, and has previously been explained using A. Tversky's (1977) feature-matching model. Given a linguistic norm to place more prominent objects in the referent position, violation of this norm might reduce sentence clarity, which then weakens the magnitude of subsequent relational judgments. 53 undergraduates completed a study that tested the degree to which both clarity and feature-matching predicted the magnitude of relational judgments by construction in 3 regression models, on for each of the 3 types of relation statements: similarity, difference, and spatial relation. Clarity perceptions predict the magnitude of relational judgments independently of the cognitive manipulation of the features of the compared objects. The pattern of findings suggests that a linguistic norm interpretation may account for variance in relational judgments independently of Tversky's feature-matching model. (PsycINFO Database Record (c) 2016 APA, all rights reserved)
Four studies bridged the areas of personality-mood and mood-cognition relations by investigating the effects of Extraversion and Neuroticism on the evaluation of affectively pleasant, unpleasant, and neutral word pairs.Specifically measured were affectivity ratings, categorization according to affect, judgments of associative strength, and response latencies.A strong, consistent cognitive bias toward affective as opposed to neutral stimuli was found across participants.Although some biases were systematically related to personality and mood, effects of individual differences were present only under specific conditions.The results are discussed in terms of a personality-mood framework and its implications for cognitive functioning.
We investigated the feasibility of a computer-graphics-based method of assessing stereomotion thresholds (Silicon Graphics Stereoview stereoscopic system). Stereomotion thresholds for a rectangle oscillating in depth were determined with the use of a dual randomly interleaved staircase design. In a group of 31 naive observers, the average thresholds of 5.97′ of arc forcrossed stereomotion and 6.00′ of arc foruncrossed stereomotion were comparable to those assessed in earlier work done with optics-based techniques. By assessing the thresholds for a rectangle that was defined either by lateral motion or by changing size, in a group of experienced observers, we were able to show that any potential residual translational motion present in the display would not have influenced the stereomotion thresholds. Our findings suggest that this computer-graphics-based technique may be a reasonable alternative to optics-based methods of assessing stereomotion thresholds.
Chadwyck-Healey has a long tradition of electronic publishing. Beginning with production of CD-based literary corpora, it has recently moved many of its products to a web-accessible online environment. The article reflects on experiences with both CD and web-based publications.
This paper compares different methods of generating intonation for an American English Text-to-Speech synthesis system. We look at a primarily rule-based approach and two data-driven approaches. For data-driven modeling we used two separate data sets, each representing a somewhat different prosodic style. One database was recordings of a portion of 1989 Wall Street Journal text from the Penn Treebank Project. The second database was recordings of interactive prompts used in telephone network services. Both were read by the same female speaker. Approximately two and one-half hours of speech was phonetically and prosodically segmented and labeled (first automatically, and subsequently verified manually). The prosodic labeling used ToBI [7] tones and breaks. Three different intonation models were compared: (1) a predominantly rule-based model based on ToBI labels [3]; (2) a parametric model using the Tilt approach [8]; and (3) a Vector Quantized model based on an underlying parametric re...
Abstract. We conducted a statistical analysis of several subsets of words from the Hoosier Mental Lexicon in order to examine some factors underlying the subjective familiarity ratings collected by Nusbaum, Pisoni, and Davis (1984). In this analysis, we grouped words into High-FAM (average familiarity rating greater than 6 on a 7-point scale), Mid-FAM (between 4.5 and 3.5), or Low-FAM (less than 2) sets, with primary interest being placed on the small set of Low-familiarity words (502 total), which were unknown to the majority of native listeners. This Low-FAM set consisted mainly of non-English borrowings, technical terms, and archaic words. These items received much higher ratings from Speech Research Laboratory staff, who we assume have greater language experience, showing that the familiarity scale employed is not continuous for very low ratings, especially in terms of the properties of Low-FAM words. As expected, High-FAM words were the majority and had greater frequency of occurrence, greater neighborhood density and frequency, and were shorter on average than Mid or Low words. However, when these sets of words were controlled for length, we found that High- and Mid-FAM words did not differ significantly in neighborhood density and
This paper discusses the relation for soaps between sensory attributes and both liking and image attributes. A clear relation emerges between sensory attribute level and liking, but no clear relation emerges between sensory attributes and image attributes. There are two possible conclusions to be drawn from these results. One conclusion is that consumers can validly assign ratings to the image attribute of a soap, but that there is no way to trace this image rating back to sensory inputs. This conclusion suggests that more research is needed to understand the meaning of image attributes. The second conclusion is that consumers cannot validly rate the image attributes of a soap, even though they can complete the questionnaire. This second conclusion implies that consumers can validly rate some attributes (e.g., sensory, liking), but not others (e.g., image), and that it may be misleading to collect and attempt to analyze image ratings for health and beauty aids products.
Part-of-speech tagging methodology has succeeded, but on problems that may lack real-world application. Redirection of the field is indicated, toward potentially more useful, but harder and more sophisticated tagging tasks: (1) using much more detailed tagsets (semantically and syntactically); (2) testing performance on treebanks reflecting the huge gamut of domains, etc., characterizing real-world applications; (3) understanding the magnitude of the unknown-word and unknown-tag problems, then overcoming them. Tagging results are presented on two versions of a new, highly variegated treebank, featuring tagsets of 2720 and 443 tags, respectively, and utilizing a dictionaryless, decision-tree tagger.
Finding simple, non-recursive, base noun phrases is an important subtask for many natural language processing applications. While previous empirical methods for base NP identification have been rather complex, this paper instead proposes a very simple algorithm that is tailored to the relative simplicity of the task. In particular, we present a corpus-based approach for finding base NPs by matching part-of-speech tag sequences. The training phase of the algorithm is based on two successful techniques: first the base NP grammar is read from a "treebank" corpus; then the grammar is improved by selecting rules with high "benefit" scores. Using this simple algorithm with a naive heuristic for matching rules, we achieve surprising accuracy in an evaluation on the Penn Treebank Wall Street Journal.
with a preface by George Miller WordNet, an electronic lexical database, is considered to be the most important resource available to researchers in computational linguistics, text analysis, and many related areas. Its design is inspired by current psycholinguistic and computational theories of human lexical memory. English nouns, verbs, adjectives, and adverbs are organized into synonym sets, each representing one underlying lexicalized concept. Different relations link the synonym sets.The purpose of this volume is twofold. First, it discusses the design of WordNet and the theoretical motivations behind it. Second, it provides a survey of representative applications, including word sense identification, information retrieval, selectional preferences of verbs, and lexical chains.Contributors: Reem Al-Halimi, Robert C. Berwick, J. F. M. Burg, Martin Chodorow, Christiane Fellbaum, Joachim Grabowski, Sanda Harabagiu, Marti A. Hearst, Graeme Hirst, Douglas A. Jones, Rick Kazman, Karen T. Kohl, Shari Landes, Claudia Leacock, George A. Miller, Katherine J. Miller, Dan Moldovan, Naoyuki Nomura, Uta Priss, Philip Resnik, David St-Onge, Randee Tengi, Reind P. van de Riet, Ellen Voorhees.
Punctuation has usually been ignored by researchers in computational linguistics over the years. Recently, it has been realized that a true understanding of written language will be impossible if punctuation marks are not taken into account. This paper contains the details of a computer-aided exercise to investigate English punctuation practice for the special case of comma (the most significant punctuation mark) in a parsed corpus. The study classifies the various "structural" uses of the comma according to the syntax-patterns in which a comma occurs. The corpus (Penn Treebank) consists of syntactically annotated sentences with no part-of-speech tag information about the individual words.
Three experiments examined contributions of study phase awareness of word identity to subsequent word-identification priming by manipulating visual attention to words at study. In Experiment 1, word-identification priming was reduced for ignored relative to attended words, even though ignored words were identified sufficiently to produce negative priming in the study phase. Word-identification priming was also reduced after color naming relative to emotional valence rating (Experiment 2) or word reading (Experiment 3), even though an effect of emotional valence upon color naming (Experiment 2) indicated that words were identified at study. Thus, word-identification priming was reduced even when word identification occurred at study. Word-identification priming may depend on awareness of word identity at the time of study.
Abstract The relationship between multiple stigmas and others' perceptions of the stigmas was investigated. The participants received 6 stigma targets; controllability of the onset of the stigma, race of the target, and gender of the target were manipulated. The participants' affective ratings differed on the basis of perceived controllability of the onset of the stigma. After receiving additional information about the targets, the participants assigned more blame to the Black stigma targets than to the White stigma targets.
In many Oceanic languages in northwest Melanesia the default attribute construction ('a big house') is one whose morphosyntax looks like that of a possession construction: the attribute occupies the (possessed) head slot, the noun the (possessor) modifier slot ('a big one of a house'), that is, the opposite of the cross-linguistic norm and a rare phenomenon worldwide. I briefly describe these constructions, which are morphosyntactically varied, then examine their history, proposing that a major factor in their genesis was the presence in Proto-Oceanic of a small class of adjectival nouns whose reflexes in languages scattered across Oceania either may still behave as noun phrase heads or retain features reflecting this earlier status. The adjectival noun class had a small membership but high token frequency, and provided the template for a pattern extension that in a number of northwest Melanesian languages drew in the much larger adjectival verb class. I also address the question of why this change occurred in northwest Melanesia but not elsewhere in Oceania
Although the influence of emotional arousal on declarative memory has been documented behaviorally, the mechanisms underlying arousal-memory interactions and their representation in the human brain remain uncertain. One route through which arousal achieves its effects on memory performance is by regulating consolidation processes. Animal research has revealed that the amygdala strengthens hippocampal-dependent memory consolidation in a limited time window following participation in an arousing task. To examine whether this integrative function of amygdalo-hippocampal structures extends to the human brain, we tested unilateral-temporallobectomy patients on an adaptation of a classic paradigm in which levels of physiological arousal at encoding modulate retention over time. Subjects rated emotionally arousing (taboo) and neutral words on an arousal scale while their skin conductance responses (SCRs) were monitored. Recall for the words was assessed immediately and after a 1-hr delay. Both temporal-lobectomy patients and control subjects generated enhanced SCRs and arousal ratings for the arousing words at the time of encoding. However, only control subjects exhibited an increase in memory for the arousing words over time. This group difference in the effect of arousal on the rate of forgetting suggests that the role of medial temporal lobe structures in memory consolidation for arousing events is conserved across species.
This paper describes a method for determining syntactic structure in coordinate constructions. It is based on the information taken from semantic similarities, selectional restrictions, and some other linguistic cues. We discuss the role the information plays in resolving ambiguities that appear in coordinate constructions, describe the means of acquiring the necessary information automatically from two on-line corpora and a lexical database, and devise two algorithms for disambiguating coordinate constructions. An experiment that follows shows effectiveness of our method and its applicability to resolving ambiguities in some other syntactic structures. 1
BOOK NOTICES 207 sure's concept ofmotivation should not be associated with words but rather with cotext and context. Everything is relative in language, and the systematic character of language does not lie in separate phonemic, lexical, grammatical, and textual systems. In fact the systematic and universal features of language are made up by the human capabilities of thinking and experiencing. Speakers are able to use a limited number of signs to express highly complicated ideas, and they can decode expressions in very complex situations. To do this requires applying the basic principle of language and language description—the idea of economy. This elementary principle of human behavior can be seen when a speaker tries to avoid unnecessary redundancy and when simpler ways of pronunciation are preferred to more difficult ones. In addition language economy can be seen on the deeper level of reinterpreting linguistic units new to a certain speaker. This gives sense to an utterance since under the principle ofeconomy, a speaker must assume that no text is uttered without meaning. As a result of new expressions and new interpretations produced by the principle ofeconomic use, language might change over time. Considering language change again, D emphasizes how the individual reflects about language. For the individual the main goal of language and speaking is to impart and to decode sense or meaning. D rejects the idea of the 'invisible hand phenomenon' as well as the concept of teleology in language change. Again he links the systematic character of language to the speaker's purposeful acts rather than to the whole speech community or to language as an abstract system. In short, D questions the traditional structuralist conception of system in language. He emphasizes the systematic cognitive behavior of each individual that uses the relative means of language. In order to support his opinion D illustrates his ideas with many detailed examples mainly taken from German, English, and French. [Dieter Aichele, Fachhochschule Neubrandenburg.] The Oxford English-Hebrew Dictionary. Ed. by N. S. Doniach and A. Kahane. Oxford: Oxford University Press, 1996. Pp. xxiii, 1091. The late N. S. Doniach (d. 16 April 1994), the chief editor ofthis dictionary, is well known to Semitologists for editing the excellent Oxford EnglishArabic dictionary ofcurrent usage (1972). The introduction by Professor A. Kahane explains D's goal of using various styles of modern Hebrew in this volume, including colloquial language and slang. From abacus to Zulu, this dictionary, happily, has it all! It even has the f-word with many of its most common idiomatic usages, such as '__ up' and '__ off!' (351). The tome's particularly noteworthy features include up-to-date terminology ofall sorts, such as that dealing with computers. However, some inconsistencies can be found. 'Software' is written as toxna with a vav (883), but it is spelled without a vav (using a kamats katan) under 'hardware' (400). And curiously, the name of the vowel kamats is written kamatz, yet the vowel chataf kamats is spelled differently on the very next line (ix). AU Hebrew words are given in their fully vocalized or pointed forms, including dagesh, which is said to have the phonetic value 'stress mark', a puzzling statement (ix). Phonological matters on the whole, however, have been handled well. It was a wise decision for the editors to give preference to a word's modern pronunciation if it differs from its traditional pointing (xxiii). Along these lines, I wanted to check the vocalization and pronunciation of the irregular Classical Hebrew plural for bayit 'house'—battiim or (bottiim); to my surprise, however, it was not given (426). British English has been chosen as the norm of this dictionary, a reasonable and expected decision by Oxford University Press. An American has little difficulty getting used to British spellings such as 'programme'; however, 'program' is also listed, but with the stipulation, 'US and Comput.' (724). However, it may be difficult at first for an American to appreciate 'farther' and 'father' transcribed exactly the same. There are a few discrepancies to report. American English 'buggy' is given as eglat-tinok (114), but 'pram' lists only eglat-tinokot, its plural (708), whereas the more formal 'perambulator' (called 'formal ' by the editors) lists...
Four studies bridged the areas of personality-mood and mood-cognition relations by investigating the effects of Extraversion and Neuroticism on the evaluation of affectively pleasant, unpleasant, and neutral word pairs. Specifically measured were affectivity ratings, categorization according to affect, judgments of associative strength, and response latencies. A strong, consistent cognitive bias toward affective as opposed to neutral stimuli was found across participants. Although some biases were systematically related to personality and mood, effects of individual differences were present only under specific conditions. The results are discussed in terms of a personality-mood framework and its implications for cognitive functioning. Personality traits influence the frequency and intensity of experienced
Two barriers to the use of the Thematic Apperception Test (TAT) in motivation research were addressed: its low internal consistency and its time-consuming coding system. Sixty males and 60 females wrote five stories to TAT pictures either on the computer or by hand. Half of each group were timed and half untimed. The writing of stories was guided by four sets of questions, and stories were coded for need for power (n Pow) by the corresponding four paragraphs. Cronbach’s alpha for the five stories was .46; for the 20 paragraphs, Cronbach’s alpha was .65. We conclude that, to the extent that measuring internal consistency is appropriate for a thought-sampling instrument like the TAT, internal consistency should be calculated by paragraphs. Significantly more words were produced in the untimed condition, but n Pow did not differ by gender, hand-written versus computer-written, or timed versus untimed conditions. The five pictures elicited significantly different amounts of n Pow. It is recommended that researchers who give the TAT on the computer use the untimed condition. Suggestions are made for increasing the scoring validity and for using the computer to decrease the time required for human coders.
This paper examines the feasibility of incremental annotation, i.e. using existing annotation on a text as the basis for further annotation rather than starting the new annotation from scratch. It contains a theoretical component, describing basic methodology and potential obstacles, as well as a practical component, describing an experiment which tests the efficiency of incremental annotation. Apart from guidelines for the execution of such pilot experiments, the experiment demonstrates that incremental annotation is most effective when supported by thorough pre-planning and documentation. Unplanned, opportunistic use of existing annotation is much less effective in its reduction of annotation time and furthermore increases the development time of the annotation software, so that this type of incremental annotation appears only practical for large amounts of heritage data.
We tested the validity of the paddle method for measuring both the kinesthetic and visual-kinesthetic perception of inclination. In three conditions, subjects performed three different tasks: (1) rotating a manual paddle to a set of verbally given inclinations (blindfolded subjects), (2) rotating a manual paddle to the same set of verbally given inclinations after specific kinesthetic training (blindfolded subjects), and (3) rotating the paddle to a set of fixed visual inclinations after the kinesthetic training. The results showed a high degree of accuracy and precision in the second and third task but not in the first one. When subjects were asked to rotate a manual paddle to a set of verbally given inclinations, they used three main anchors (0°, 45°, 90°). Furthermore, the paddle method is biased by a kinesthetic deficiency, namely a rotational problem of the wrist that can be corrected by means of specific training.
This paper is based on the following set of assumptions: (1) Ways of speaking characteristic of a given speech-community constitute a manifestation of a tacit system of `cultural rules', or, as the author calls them, `cultural scripts'; (2) to understand a society's ways of speaking, we have to identify and articulate its implicit `cultural scripts'; (3) to be able to do this without ethnocentric bias we need a universal, language-independent perspective; and (4) this can be attained if the `rules' in question are stated in terms of lexical universals, that is, universal human concepts lexicalized in all languages of the world. In various earlier publications Anna Wierzbicka tried to show how cultural scripts can be set out and justified with reference, in particular, to Japanese, Chinese, Polish, (White) Anglo-American, Black American and Anglo-Australian cultural norms. In this paper she applies the `cultural script' approach to German and compares German norms with `Anglo' norms (that is, norms prevailing in English-specking societies). Wierzbicka notes that in recent decades great changes have undoubtedly occurred in German ways of speaking and, it can be presumed, in underlying cultural values. For example, the dramatic spread of the use of the `familiar' form of address ( du, as opposed to Sie), and the decline in the use of titles (e.g. Herr Müller instead of Prof. Müller) point to significant changes in interpersonal relations, in the direction of more egalitarian informality. At the same time, she shows that evidence of contemporary public signs which are discussed here suggests that some traditional German values, like the value of social discipline and of Ordnung (order) based on legitimate authority, are far from obsolete. She also shows that in studying such values we can rely on concepts more precise and more illuminating than `authoritarianism' or `authoritarian personality', often used in the past in analyses of German culture and society, and that the `cultural scripts' approach offers a rigorous and efficient tool for studying change and variation, as well as continuity, in social attitudes and cultural values. Above all, the author argues that rather than perpetuating stereotypes based on prejudice and lack of understanding `cultural scripts' help outsiders grasp the `cultural logic' underlying unfamiliar ways of speaking which may otherwise look like a strange collection of idiosyncracies—or worse.
This experiment examined individual differences in emotional responsivity by recording the startle eyeblink reflex while 57 college students viewed affect-laden pictures and then rated their pleasantness. All participants first completed measures of affect intensity, alexithymia, and depression. Startleprobes were some-times presented at 120, 300, 800, or 4,500 ms after slide onset. By 300 ns, blinks elicited during negative slides were larger than those elicited during positive ones. Negative slides were also rated as more unpleasant. Moreover, all three personality variables moderated either the valence ratings, startle modification, or both. High-affect intensity was associated with diminished modulation of startle, but more extreme ratings. Alexithymia had no effect on the startle measure, but high-alexithymia participants did show more moderate ratings. Depressed participants exhibited accelerated (120 ms) modulation of startle. The results suggest the importance of measuring both physiological responses and subjective feelings in the study of individual differences in emotion.
In this paper, we consider the different methods that have been developed to quantify random generation behavior and incorporate these measurement scales into a Windows95 computer program called RgCalc. RgCalc analyzes the quality of human attempts at random generation and can provide computer-generated, pseudorandom sequences for comparison. The program is designed to be appropriate for the analysis of various types of random generation situations employed in the psychological literature. The different algorithms for the evaluation of a dataset are detailed and an outline of the program is described. Performance measures are available for assessing various aspects of the response distribution, the sequencing of pairs, the ordinal relationships between sets of items, and the tendency to repeat alternatives over different lengths. A factor analysis is used to illustrate the multiple dimensions underlying human randomization processes.
Previous research on the effects of age of acquisition on lexical processing has relied on adult estimates of the age at which children learn words. The authors report 2 experiments in which effects of age of acquisition on lexical retrieval are demonstrated using real age-of-acquisition norms. In Experiment 1, real age of acquisition emerged as a powerful predictor of adult object-naming speed. There were also significant effects of visual complexity, word frequency, and name agreement. Similar results were obtained in reanalyses of data from 2 other studies of object naming. In Experiment 2, real age of acquisition affected immediate but not delayed object-naming speed. The authors conclude that age-of-acquisition effects are real and suggest that age of acquisition influences the speed with which spoken word forms can be retrieved from the phonological lexicon.
This paper describes applications of stochastic and symbolic NLP methods to treebank annotation. In particular we focus on (1) the automation of treebank annotation, (2) the comparison of conflicting annotations for the same sentence and (3) the automatic detection of inconsistencies. These techniques are currently employed for building a German treebank.