Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
This paper examines the use of an unsupervised statistical model for determining the attachment of ambiguous coordinate phrases (CP) of the form nl p n2 cc n3. The model presented here is based on JAR98], an unsupervised model for determining prepositional phrase attachment. After training on unannotated 1988 Wall Street Journal text, the model performs at 72% accuracy on a development set from sections 14 through 19 of the WSJ TreeBank [MSM93].
This article examines the reliance of U.S. campuses on international teaching assistants (ITA) for staffing undergraduate course and the strategies that may affect ratings of their speaking competence. This increasing reliance has led to student complaints about incomprehensibility of ITA. This problem has been examined by looking through the eyes of the students, administrators and taxpayers. Therefore, the responsibility had been placed on the ITA, whose burden it was to learn the language and culture more fully. The goal of having ITA learn the language and culture better was eclipsing another important issue that needs consideration, the issue of teaching assistant's feelings of loss of control over the students' perception about themselves.
Research has shown that seeing another person (i.e., a model) perform at a certain level can influence the goal choice and performance of an observer. This study extended these findings by examining the model's effects on two potential explanatory mediators: expectations of reaching different performance levels and valence at those levels. The moderating effects of observer task experience and self‐esteem were also examined. In a repeated measures design, results showed that model performance influenced observers' goals, task performance, expectancies, and valence ratings. Path analyses indicated that expectancies mediated observational effects on goal choice. Results also indicated significant Model x Trial interactions, with a diminishing effect of model influence as personal experience developed. Results are discussed in terms of mediating effects of expectancies and valences on model influences on goal choice as well as the different weights given to social and personal information.
This paper explores the role of lexicalization and prun-ing of grammars for base noun phrase identification. We modify our original framework (Cardie & Pierce 1998) to extract lexicalized treebank grammars that assign a score to each potential noun phrase based upon both the part-of-speech tag sequence and the word sequence of the phrase. We evaluate the mod-ified framework on the “simple ” and “complex ” base NP corpora of the original study. As expected, we find that lexicalization dramatically improves the perfor-mance of the unpruned treebank grammars; however, for the simple base noun phrase data set, the lexical-ized grammar performs below the corresponding unlex-icalized but pruned grammar, suggesting that lexical-ization is not critical for recognizing very simple, rel-atively unambiguous constituents. Somewhat surpris-ingly, we also find that error-driven pruning improves the performance of the probabilistic, lexicalized base noun phrase grammars by up to 1.0 % recall and 0.4% precision, and does so even using the original pruning strategy that fails to distinguish the effects of lexical-ization. This result may have implications for many probabilistic grammar-based approaches to problems in natural language processing: error-driven pruning is a remarkably robust method for improving the perfor-mance of probabilistic and non-probabilistic grammars alike.
This paper explores the automatic construction of a multilingual Lexical Knowledge Base from pre-existing lexical resources. We present a new and robust approach for linking already existing lexical/semantic hierarchies. We used a constraint satisfaction algorithm (relaxation labeling) to select-among all the candidate translations proposed by a bilingual dictionary- the right English WordNet synset for each sense in a. taxonomy automatically derived from a Spanish monolingual dictionary. Although on average, there are 15 possible WordNet connections for each sense in the taxonomy, the method achieves an accuracy over 80%. Finally, we also propose several ways in which this technique could be applied to enrich and improve existing lexical databases.
In recent work we have presented a formal framework for linguistic annotation based on labeled acyclic digraphs. These `annotation graphs' offer a simple yet powerful method for representing complex annotation structures incorporating hierarchy and overlap. Here, we motivate and illustrate our approach using discourse-level annotations of text and speech data drawn from the CALLHOME, COCONUT, MUC-7, DAMSL and TRAINS annotation schemes. With the help of domain specialists, we have constructed a hybrid multi-level annotation for a fragment of the Boston University Radio Speech Corpus which includes the following levels: segment, word, breath, ToBI, Tilt, Treebank, coreference and named entity. We show how annotation graphs can represent hybrid multi-level structures which derive from a diverse set of file formats. We also show how the approach facilitates substantive comparison of multiple annotations of a single signal based on different theoretical models. The discussion shows how annotation graphs open the door to wide-ranging integration of tools, formats and corpora.
In order to preserve the basic unity of a language like Spanish, which is spoken in 20 countries, all the speakers must adopt a respectful and careful attitude when they use it orally. In the area of phonetics, educated Spanish speakers, even those who study the language, express themselves in a somewhat careless way and what is most surprising is that such negligence is accepted in Spain by the educated linguistic norm.
We present a new approach to partial parsing of natural language texts that relies on machine learning methods. The approach combines corpus-based grammar induction with a very simple pattern-matching algorithm and an optional constituent verification step. The grammar induction algorithm acquires a set of rules for each level of linguistic analysis using a new technique for errordriven pruning of treebank grammars. The constituent verification step employs standard inductive learning techniques as an additional precision-enhancing device. We evaluate the approach on four partial parsing data sets and find that performance is very good (over 93% precision and recall) for applications that require or prefer fairly simple constituent bracketing. As the complexity of the partial parsing task increases, however, our approach lags the performance of competing approaches. We explain these differences in terms of the knowledge sources employed by each method and describe a number of features...
Hebrew Studies 40 (1999) 269 Reviews grams that serve as excellent visual aids. The work is written in a highly technical language that assumes familiarity with the terminology. shorthand, and conceptual framework of Generative Grammar. and this limits its accessibility and appeal to a wide readership. The inclusion of a glossary andlor a brief overview of this method of linguistic study would help address this problem. The book's title is somewhat inaccurate since it does not offer a truly comparative study of Hebrew and Arabic syntax. It is principally an analysis of Hebrew which makes reference to other languages to explicate and illustrate Shlonsky's ideas on Hebrew. While Arabic is cited more frequently than any other language besides Hebrew. there are lengthy sections of the book in which it never figures in the discussion. For example. there is not a single reference to Arabic in the thirty-page chapter treating subject-verb inversion. A further problem concerns the inconsistent manner in which Shlonsky appeals to the Arabic evidence. Throughout the book, he argues his case by referring to several different forms of the language. including Standard Arabic and the dialects from Palestine. Southern Palestine. Morocco, and Egypt. Each of these is a unique linguistic system that is distinct from the others but Shlonsky does not pay sufficient attention to the differences among them. This method hinders the purported comparative focus of his work since the reader is not sure which Arabic is meant to be the primary point of comparison. Consistent reliance upon one form of the language would have been a more beneficial approach to adopt. Such relatively minor problems do not significantly detract from the many strengths of this volume and the important contribution it makes. Shlonsky's book is required reading for anyone interested in serious study of Hebrew and will be a major work in the field. John Kaltner Rhodes College Memphis, TN 38112 kaltner@rhodes.edu THE SEMANTICS OF ASPECT AND MODALITY: EVIDENCE FROM ENGLISH AND BIBLICAL HEBREW. By Galia Hatav. Studies in Language Companion Series 34. Pp. x + 224. Philadelphia. PA: John Benjamins, 1997. Cloth, $85.00. Originally a dissertation at Tel-Aviv University. this work "aims to provide a general (semantic) theory for temporality...but it also systemati- Hebrew Studies 40 (l999) 270 Reviews cally examines the verb system in Biblical-Hebrew...which lacks tenses, as will be demonstrated, and thus enables us to see the nature of aspect and modality more clearly" (p. I). The author's theoretical assumption is "that TAM, i.e., the Tense-Aspect-Modal system in language, should be defmed within truth conditional semantics, in terms of temporality, rather than within a pragmatic approach which deals with it in terms of perspective, attitude, and the like" (although pragmatics is not ignored altogether, p. 195). Moreover, she seeks to analyze the data within the framework of a threefold distinction. Here she builds on the work of Hans Reichenbach, who argued that the contrast between the time of speech (S-time) and the time in which the event actually took place (E-time) is not sufficient to account for verbal uses. A third category is needed, namely, the time of reference (R-time), a somewhat fuzzy concept that Hatav defmes as a time unit that contains (or is contained in, or is ordered with) the E-time (pp. 3-5). In the introductory chapter the author, after explaining these and other assumptions, surveys previous attempts to account for the verbal system of biblical Hebrew; unfortunately, she seems unaware of Bruce K. Waltke and M. O'Connor, Biblical Hebrew Syntax, which gives considerable attention to verbal aspect She further informs us that she examined sixty-two chapters taken from the Pentateuch and the Former Prophets (excluding poetic material, since "poetry often violates otherwise valid linguistic norms," p. 24) and gives us some details regarding her method. Following the lead of discourse-analysis scholars, such as R. Longacre, she argues that "the biblical Hebrew system organizes the text into sequential and non-sequential material," but that in addition to sequentiality, three other temporal parameters are needed. These four parameters are individually considered in the following chapters. Chapter 2, accordingly, deals...
OBJECTIVE: To evaluate the efficacy of Simulated Presence, a personalized approach to enhance well-being among nursing home residents with Alzheimer's disease and related dementia's (ADRD). DESIGN: Latin-Square, double blinded, 3-factor design with restrictive randomization of three treatments (the study intervention, a placebo audio tape of a person reading the newspaper, and usual care). The three factors were treatment, time, and facility type. SETTING: Nine nursing homes in Eastern Massachusetts and Southern New Hampshire. PARTICIPANTS: Fifty-four subjects with documented ADRD who were aged 50 years or older, medically stable, had resided in their current nursing home for at least 3 months, and who had no planned discharge. All subjects had a history of agitated or withdrawn behaviors. INTERVENTION: The purpose of Simulated Presence is to provide a personalized intervention for persons with moderate to severe cognitive impairment. Through a unique testing process, some of the best loved memories of the ADRD person's lifetime are identified and then those memories are introduced to the patient in the format of a telephone conversation using a continuous play audio tape system. The intervention may be used for extended periods of time because each repetition is viewed as a fresh, live telephone call as a result of the short-term memory deficit of the person with ADRD. MEASUREMENTS: Direct observations of outcomes included using a newly developed scale, the Scale for the Observation of Agitation in Persons with Dementia, an agitation visual analog scale, the Positive Affect Rating Scale (mood and "interest"), a withdrawal visual analog scale, and facial diagrams of mood. Reported measures included daily staff observation logs of responses to interventions offered, and weekly staff surveys using the short-form Cohen-Mansfield Agitation Inventory and the Multidimensional Observation Scale for Elderly Subjects (mood and "interest"). Severity of dementia was assessed by the Mini-Mental State Exam, the Test for Severe Impairment, the Bedford Alzheimer's Nursing Scale, and the ADL Self-Performance Scale. RESULTS: Chi-square analysis of direct observations, using facial diagrams, revealed that Simulated Presence was equivalent to usual care (P =.141) and superior to placebo for producing a happy facial expression (P =.001). A positive effect was also documented in nursing staff observation logs using Analysis of Variance techniques (ANOVA) for subjects during Simulated Presence phases compared with the placebo phases (P <.001) and usual care phases (P <.001). According to ANOVA analyses of "interest" from weekly surveys, Simulated Presence was superior to both usual care (P =.001) and placebo (P =.008). We were unable to find evidence of significant differences (P <.05) among interventions for other direct observations and weekly reports of overall agitation or mood aspects of withdrawal. Subjects accepted the intervention most of the time, except for five subjects who refused it more than 50% of the time. CONCLUSION: This study provided evidence that Simulated Presence can be effective in enhancing well-being and decreasing problem behaviors in the nursing home setting as a substitute for or complement to usual care.
This pilot study examined African-American psychiatric patients' reactions to the Cultural Mistrust Inventory, a measure of blacks' mistrust of white society. Twenty-two black psychiatric patients were screened for the Culturally-Sensitive Diagnostic Interview Research Project. All patients were debriefed after the screening interview including queries about their reactions to the experience, whether they would be willing to participate in the next interview, and their reasons for participating or not. Patients' responses were recorded verbatim and were categorized in terms of their valence (positive, neutral, or negative) and affectivity (yes or no) by independent raters. Agreement between raters in terms of the valence of patients' reactions was very good, but it was poor to fair in terms of affectivity ratings. The majority of these black patients' responses were positive and nonaffective. Administration of the Cultural Mistrust Inventory to black psychiatric patients does not cause negative emotional reactions.
Most bipolar models of affective processing in social psychology assume that positive and negative valent processes are represented along a single continuum that rangesfrom very positive to very negative. Recent research has raised the possibility, however, that the motivational systems for positive/approach and negative/defensive valent processing (positivity and negativity, respectively) are separable. In this article, the authors use unipolar positivity, negativity, and ambivalence ratings and bipolar valence, dominance, and arousal ratings of 472 slides from the International Affective Picture System to examine several aspects of the bivariate model of evaluative space. Analysis confirmed a positivity offset and negativity bias in the activation functions of the valent systems as wel as multiple modes of evaluative activation (e.g., reciprocal, uncoupled positivity, uncoupled negativity). Together, these data suggest that the bipolar structure of affective processes should be tested rather than assumed.
Word recognition and generation is a fundamental part of the processing of natural language and it requires computationally effective morphological processors, especially for languages with rich morphology such as Modern Greek. Various models have been proposed for developing computerized systems to accomplish the task of recognition of morphosyntactic features of words In the work presented here, the lazy tagging approach was examined, in which taggers are expected to work in the simplest possible way. The model of functional decomposition was extended and adapted for Modern Greek as a target language, following the lazy word-parsing approach, in order to cover a number of morphological phenomena that are encountered in Modern Greek, namely inflection, affixation, and longdistance dependencies. To achieve a more efficient word recognition, several automata of different levels of computing power, based on the original model, were introduced and evaluated according to the criteria of complexity, recognition speed, and accuracy of the results. The proposed system was used for processing a large-scale corpus, and the results are presented and discussed. To accomplish their task, taggers can rely upon large lexical databases, which are expected to be organized in such a way as to provide rapid access to the stored data and efficient memory management. Directed graphs can be used to describe and organize a lexical database of large magnitude in a compact manner. These data structures are named here matrix lexica, where the letters are described as nodes of directed graphs and the lemmata as paths (set of edges). It is expected that matrix lexica will support a tagger efficiently by providing a high speed of resolution, sound mathematical foundation, low memory requirements, and ability to handle distorted input in future developments.
Finding simple, non-recursive, base noun phrases is an important subtask for many natural language processing applications. While previous empirical methods for base NP identification have been rather complex, this paper instead proposes a very simple algorithm that is tailored to the relative simplicity of the task. In particular, we present a corpus-based approach for finding base NPs by matching part-of-speech tag sequences. The training phase of the algorithm is based on two successful techniques: first the base NP grammar is read from a "treebank" corpus; then the grammar is improved by selecting rules with high "benefit" scores. Using this simple algorithm with a naive heuristic for matching rules, we achieve surprising accuracy in an evaluation on the Penn Treebank Wall Street Journal.
En se situant dans le cadre d'un projet appele Semantic dictionary viewed as a lexical database dont l'objectif est d'identifier et de denombrer les informations linguistiquement pertinentes pouvant etre derivees des definitions du sens, l'A. examine ici les processus productifs de derivation semantique des verbes en russe, et en particulier les processus associes a un changement metonymique qui se refletent dans le cadre du cas profond d'un verbe. Il tente ainsi de montrer que les roles semantiques permettent de decouvrir l'invariant semantique d'une grande partie des derivations productives
Determining the attachments of prepositions and subordinate conjunctions is a key problem in parsing natural language. This paper presents a trainable approach to making these attachments through transformation sequences and error-driven learning. Our approach is broad coverage, and accounts for roughly three times the attachment cases that have previously been handled by corpus-based techniques. In addition, our approach is based on a simplified model of syntax that is more consistent with the practice in current state-of-the-art language processing systems. This paper sketches syntactic and algorithmic details, and presents experimental results on data sets derived from the Penn Treebank. We obtain an attachment accuracy of 75.4% for the general case, the first such corpus-based result to be reported. For the restricted cases previously studied with corpusbased methods, our approach yields an accuracy comparable to current work (83.1%).
This paper compares different methods of generating intonation for an American English Text-to-Speech synthesis system. We look at a primarily rule-based approach and two data-driven approaches. For data-driven modeling we used two separate data sets, each representing a somewhat different prosodic style. One database was recordings of a portion of 1989 Wall Street Journal text from the Penn Treebank Project. The second database was recordings of interactive prompts used in telephone network services. Both were read by the same female speaker. Approximately two and one-half hours of speech was phonetically and prosodically segmented and labeled (first automatically, and subsequently verified manually). The prosodic labeling used ToBI [7] tones and breaks. Three different intonation models were compared: (1) a predominantly rule-based model based on ToBI labels [3]; (2) a parametric model using the Tilt approach [8]; and (3) a Vector Quantized model based on an underlying parametric re...
Forty-nine young adults (M age = 22 years) and 30 elderly adults (M age = 69 years) rated the 60 pictorial stimuli from the Boston Naming Test (BNT) on familiarity, providing the first such normative data for these stimuli along this dimension. Participants also made speeded lexical decisions about the word item representations of each BNT picture. B NT word frequency values were also examined in relation to BNT familiarity and speeded lexical decision performance. For both young and elderly adults, lexical decision reaction times to the word representations of BNT stimuli were negatively related to word frequency and familiarity of the BNT pictures. These patterns suggest that increases in word frequency and picture familiarity facilitate (i. e., speed up) the processing of BNT word representations. Furthermore, speed of processing appears to be a relevant dimension of BNT performance, at least when young and elderly adults free from clinical aphasia are involved.
Inscriptions, a type of source which has hitherto hardly been taken into account in historical linguistics, can provide data for a whole series of questions thrown up by research into the linguistic geography of Early New High German and the Late Middle Low German of the same period. In respect of the pressure towards standardization exerted by M. Luther's bible, various degrees can be differentiated. The greatest willingness to adopt this linguistic norm is found in the case of New High German monophthongization and diphthongization, as well as in the lowering of u to o before nasals and the use of the verbal prefix ver-. In these cases, the symbols on the appropriate maps show agreement with Luther 's usage. In some other cases, namely, the use of for MHG /ei/, the occurrence of vowel forms without unrounding, the use of the uncontracted form of the word nicht and the use of the subjunctive in the predicates of concessive clauses, we find that the form favoured by Luther is the one which appears most frequently in the inscriptions. Nevetheless, especially in sixteenth-century inscriptions, we find several examples of the replacement of standard forms by deviant variants valid at the same time. In contrast, the number of deviations from Luther's language visibly decreases in the seventeenth century. In some cases, the treatment of the norms of the Luther's bible in the inscriptions indicates that the linguistic usage of the reformer had no or only indirect significance for the course of development here. Thus, in the case of the feminine inflectional system, while it is true that the epigraphic sources largely follows Luther's practice, but it is also the case that from about 1600 on a tendency towards deviation from his usage emerges. Here, restructuring is a process which was clearly not influenced by Luther. Again, in the case of abstract suffix -nis, the course of development seems to have been determined by factors inherent in the language rath
Finding simple, non-recursive, base noun phrases is an important subtask for many natural language processing applications. While previous empirical methods for base NP identification have been rather complex, this paper instead proposes a very simple algorithm that is tailored to the relative simplicity of the task. In particular, we present a corpus-based approach for finding base NPs by matching part-of-speech tag sequences. The training phase of the algorithm is based on two successful techniques: first the base NP grammar is read from a ``treebank'' corpus; then the grammar is improved by selecting rules with high ``benefit'' scores. Using this simple algorithm with a naive heuristic for matching rules, we achieve surprising accuracy in an evaluation on the Penn Treebank Wall Street Journal.
Treebanks, such as the Penn Treebank (PTB), offer a simple approach to obtaining a broad coverage grammar: one can simply read the grammar off the parse trees in the treebank. While such a grammar is easy to obtain, a square-root rate of growth of the rule set with corpus size suggests that the derived grammar is far from complete and that much more treebanked text would be required to obtain a complete grammar, if one exists at some limit. However, we offer an alternative explanation in terms of the underspecification of structures within the treebank. This hypothesis is explored by applying an algorithm to compact the derived grammar by eliminating redundant rules - rules whose right hand sides can be parsed by other rules. The size of the resulting compacted grammar, which is significantly less than that of the full treebank grammar, is shown to approach a limit. However, such a compacted grammar does not yield very good performance figures. A version of the compaction algorithm taking rule probabilities into account is proposed, which is argued to be more linguistically motivated. Combined with simple thresholding, this method can be used to give a 58% reduction in grammar size without significant change in parsing performance, and can produce a 69% reduction with some gain in recall, but a loss in precision.
This paper discusses the relation for soaps between sensory attributes and both liking and image attributes. A clear relation emerges between sensory attribute level and liking, but no clear relation emerges between sensory attributes and image attributes. There are two possible conclusions to be drawn from these results. One conclusion is that consumers can validly assign ratings to the image attribute of a soap, but that there is no way to trace this image rating back to sensory inputs. This conclusion suggests that more research is needed to understand the meaning of image attributes. The second conclusion is that consumers cannot validly rate the image attributes of a soap, even though they can complete the questionnaire. This second conclusion implies that consumers can validly rate some attributes (e.g., sensory, liking), but not others (e.g., image), and that it may be misleading to collect and attempt to analyze image ratings for health and beauty aids products.
The authors documented a linguistic norm account of direction of comparison asymmetry effects in relational judgments (e.g., seeing hyenas as more similar to dogs than dogs are similar to hyenas). The asymmetry effect is magnified by discrepancies in prominence between subject and reference, and has previously been explained using A. Tversky's (1977) feature-matching model. Given a linguistic norm to place more prominent objects in the referent position, violation of this norm might reduce sentence clarity, which then weakens the magnitude of subsequent relational judgments. 53 undergraduates completed a study that tested the degree to which both clarity and feature-matching predicted the magnitude of relational judgments by construction in 3 regression models, on for each of the 3 types of relation statements: similarity, difference, and spatial relation. Clarity perceptions predict the magnitude of relational judgments independently of the cognitive manipulation of the features of the compared objects. The pattern of findings suggests that a linguistic norm interpretation may account for variance in relational judgments independently of Tversky's feature-matching model. (PsycINFO Database Record (c) 2016 APA, all rights reserved)
In this paper we describe a word clustering method for class-based n-gram model. The measurement for clus-tering is the entropy on a corpus different from the cor-pus for n-gram model estimation. The search method is based on the greedy algorithm. We applied this method to a Japanese EDR corpus and English Penn Treebank corpus. The perplexities of word-based n-gram model on EDR corpus and Penn Treebank are 153.1 and 203.5 re-spectively. And Those of class-based n-gram model, esti-mated through our method, are 146.4 and 136.0 respec-tively. The result tells us that our clustering methods is better than the Brown’s method and the Ney’s method called leaving-one-out. 1.
Historically, research on the construct of body image has focused on its stability. Many researchers are beginning to reexamine whether the body image construct is stable, and they have shown that the construct is subject to change after experimentally induced situations and after major life events. This study attempted to determine whether minor life events and mood had a significant relationship to body image ratings and whether a change in minor life events and mood over the course of one month would predict body image ratings. For men, it was found that minor life events were not significantly related to body image ratings, though higher mood scores were significantly related to lower ratings of physical appearance. For women, a greater number of positive minor life events was significantly related to engaging in more behaviors to keep oneself physically attractive, and higher mood scores were significantly related to lower ratings of physical appearance. For men, changes in minor life events or mood over the course of one month did not predict change in body image ratings. For women, an increase in positive minor life events predicted an increase in behaviors associated with keeping oneself physically attractive. A post-hoc analysis was conducted to determine whether individuals whose mood worsened over the course of one month would show greater changes in body image ratings. However, this post-hoc hypothesis was not supported. The main hypotheses were reanalyzed with the subsample stratified into younger and older adult men and women. Though the sample size was small, there appeared to be differences between older and younger adults, with younger adults more susceptible to body image fluctuations than older adults. In the overall sample, body image ratings changed little over the course of one month, though this discovery fits well within an overall personality contruct model proposed by Mischel (1968). Effect sizes for this study were small, and the sample size was too small to confidently find significant relationships or make predictions. Other limitations of this study as well as future directions in research are discussed.
Problems in retrieval of names form large data bases and in nominal record linkage are discussed with respect to computational solutions. The quest for robust methods that can handle the typical variability of historical nominal information is discussed, with some emphasis on probabilistic methods. It is argued that comparison and assessment of different systems used on the same data could enhance our understanding of methodological issues.
This paper explores the relationship between WordNet and other conventional linguistically-based lexical resources. We introduce an algorithm for aligning word senses from different resources, and use it in our experiment to sketch the role played by WordNet, as far as sense discrimination is concerned, when put in the context of other lexical databases. The results show how and where the resources systematically differ from one another with respect to the degree of polysemy, and suggest how we can (i) overcome the inadequacy of individual resources to achieve an overall balanced degree of sense discrimination, and (ii) use a combination of semantic classification schemes to enrich lexical information for NLP. 1 Introduction Lexical resources used in natural language processing (NLP) have evolved from handcrafted lexical entries to machine readable lexical databases and large corpora which allow statistical manipulation. The availability of electronic versions of linguistic resources...
This paper describes a method in determining syntactic structure for coordinate constructions. It is based on the information taken from semantic similarities, selectional restrictions, and some other linguistic cues. We discuss the role the information plays in resolving ambiguities that appear in coordinate constructions, describe the means of acquiring the necessary information automatically from two on-line corpora and a lexical database, and devise two algorithms for disambiguating coordinate constructions. An experiment that follows shows effectiveness of our method and its applicability to resolving ambiguities in some other syntactic structures.
In comparison to conventional displays, 3D stereoscopic displays convey additional information about the 3D structure of a scene by providing information that can be used to extract depth. In the present study we evaluated the psychovisual impact of stereoscopic images on viewers. Thirty-three non-expert viewers rated sensation of depth, perceived sharpness, subjective image quality, and relative preference for stereoscopic over non-stereoscopic images. Rating methods were based on procedures described in ITU- Rec. 500. Viewers also rated sequences in which the left- and right-eye images were processed independently, using a generic MPEG-2 codec, at bit-rates of 6, 3, and 1 Mbits/s. The main finding was that viewers preferred the stereoscopic version over the non-stereoscopic version of the sequences, provided that the sequence did not contain noticeable stereo artifacts, such as exaggerated disparity. Perceived depth was rated greater for stereoscopic than for non-stereoscopic sequences, and perceived sharpness of stereoscopic sequences was rated the same or lower compared to non-stereoscopic sequences. Subjective image quality was influenced primarily by apparent sharpness of the video sequences, and less so by perceived depth.
We describe a system for extracting concepts from unstructured text. We do this by clustering document words and then assembling a structure which relates these words semantically. The clustering process identifies words which co--occur across a set of documents and creates groups of words which suggest a semantic context common across the document set. This context is formalized by identifying semantic relationships between the cluster words using a lexical database to build a Semantic Relationship Graph (SRG). This SRG is a directed graph which conveys a robust representation of the sub-- and super--class relationships between the correct word senses. We show how this process can be applied to a user--selected set of HTML documents; the SRGs can subsequently aid in searching the World Wide Web for documents which are similar. 1 Introduction Mining textual information presents challenges over data mining of relational or transaction databases because there are no predefined fields...
Children's judgements about pain at age 8-10 years were examined comparing two groups of children who had experienced different exposure to nociceptive procedures in the neonatal period: extremely low birthweight (ELBW) <or = 1000 g (N = 47) and full birthweight (FBW) > or = 2500 g (N = 37). The 24 pictures that comprise the Pediatric Pain Inventory, depicting events in four settings: medical, recreational, daily living, and psychosocial, were used as the pain stimuli. The subjects rated pain intensity using the Color Analog Scale and pain affect using the Facial Affective Scale. Child IQ and maternal education were statistically adjusted in group comparisons. Pain intensity and pain affect related to activities of daily living and recreation were significantly higher than psychosocial and medically related pain on both scales in both groups of children. Although the two groups of children did not differ overall in their perceptions of pain intensity or affect, the ELBW children rated medical pain intensity significantly higher than psychosocial pain, unlike the FBW group. Also, duration of neonatal intensive care unit stay for the ELBW children was related to increased pain affect ratings in recreational and daily living settings. Despite altered response to pain in the early years reported by parents, on the whole at 8-10 years of age ELBW children judged pain in pictures similarly to their term peers. However, differences were evident, which suggests that studies are needed of biobehavioural reactivity to pain beyond infancy, as well as research into beliefs, attitudes, and perceptions about pain during the course of childhood in formerly ELBW children.
Children's judgements about pain at age 8–10 years were examined comparing two groups of children who had experienced different exposure to nociceptive procedures in the neonatal period: extremely low birthweight (ELBW) ≤ 1000 g ( N = 47) and full birthweight (FBW) ≤ 2500 g ( N = 37). The 24 pictures that comprise the Pediatric Pain Inventory, depicting events in four settings: medical, recreational, daily living, and psychosocial, were used as the pain stimuli. The subjects rated pain intensity using the Color Analog Scale and pain affect using the Facial Affective Scale. Child IQ and maternal education were statistically adjusted in group comparisons. Pain intensity and pain affect related to activities of daily living and recreation were significantly higher than psychosocial and medically related pain on both scales in both groups of children. Although the two groups of children did not differ overall in their perceptions of pain intensity or affect, the ELBW children rated medical pain intensity significantly higher than psychosocial pain, unlike the FBW group. Also, duration of neonatal intensive care unit stay for the ELBW children was related to increased pain affect ratings in recreational and daily living settings. Despite altered response to pain in the early years reported by parents, on the whole at 8–10 years of age ELBW children judged pain in pictures similarly to their term peers. However, differences were evident, which suggests that studies are needed of biobehavioural reactivity to pain beyond infancy, as well as research into beliefs, attitudes, and perceptions about pain during the course of childhood in formerly ELBW children.
198 LANGUAGE, VOLUME 74, NUMBER 1 (1998) cussion of specific examples reveals some characteristic aspects of agrammatic output. Ch. 4, 'The grammar of connected agrammatic speech', completes the summary of grammatical characteristics of nonfluent aphasia. It argues against the label 'telegraphic speech' and against the theory of 'economy of articulatory effort' in explaining agrammatic behavior. Ch. 5, 'Speech, writing, and oral reading', discusses differences and similarities of impairment among these different output modes. Examples of spontaneous writing illustrate the possibility of dramatic dissociation between written and spoken language. Examples from oral reading show that a patient may tend to make the same types of errors in reading as in speech. Ch. 6, 'Bilingual and polyglot aphasia', by Loraeme K. Obler, José Centeno, and Nancy Eng, discusses the similarities and differences between languages in aphasies who are multilingual. They point out that deficits and recovery are usually parallel but that variance can occur (resulting, for example, from structural differences between languages). The authors then discuss special behaviors and brain organization of the bilingual aphasie. They end the chapter with advice to speech-language pathologists on diagnosis of patients who speak a language or dialect unfamiliar to the clinician. Ch. 7, 'Inventing therapy for aphasia', by Audrey L. Holland and Claire Penn, is a fascinating discussion of how to devise therapy when the patient and therapist do not share the patient's primary language. They address such issues as how to choose which language to work with and the importance of cross-cultural influences on therapy design. In sum, this book is primarily designed to inform clinicians in their efforts to provide therapy. However, the liberal exemplification of agrammatic output should be of interest to even nonclinical researchers. [Sherri K. Shaw, University of Texas, Austin. ] Xenismen: Die Nachahmung fremder Sprachen. By Wolfgang Moser (Europ äische Hochschulschriften, Reihe 21, 159.) Frankfurt am Main Peter Lang, 1996. Pp. 284. Xenisms are phenomena that characterize somebody or something as foreign. They can be nonlinguistic (a kimono, national anthems, schnitzel) or linguistic (Chinese characters, a foreign accent, loan words). Most importantly, xenisms must be considered as typical of a foreign people or language, even if this doesn't coincide with reality: most Germans don't wear leather breeches, and Chinese people do have an ItI phoneme distinct from IV. In his PhD thesis (University of Graz, Austria), Moser gives a detailed linguistic and semiotic analysis ofxenisms that imitate foreign languages, limiting the scope of his study to intentionally used interlingual xenisms in written texts. He then discusses the role of foreignness and points out that on a scale 'total ignorance—intimate knowledge', foreignness is more or less near to ignorance but not identical to it. You have to know at least something (even something wrong) about other people to see them as foreign and different from yourself. You must know that Fritz is a German name to evoke 'germanness', but you need not know more (e.g. that hardly anybody in Germany would call their son Fritz nowadays, but rather Michael or Kevin). This explains why linguistic xenisms do not appear at random but are associated with the languages and cultures in contact with one's own. The introductory chapter (1 3-29) is completed by a series ofprivative dichotomies: Xenisms can be cotextual or implanted, reduced or fully decodable, interlingual or intralingual, spontaneous or conventional, foreign or pseudo-foreign, foreign elements (mostly words) or foreign usages (e.g. word order rules). Finally, the relation of xenisms to a certain language can be more or less vague or precise. Apart from the introductory chapter, the book is divided into two main parts: The first part (31-129) analyzes the linguistic structure of xenisms; the second part (131-254) deals with xenisms as semiotic phenomena. In the structural analysis, M explains xenisms as deviations from linguistic norms (following Eugenio Coseriu's definition of norm). Their forms range from the imitation of hieroglyphs to 'typical' name endings and are broadly illustrated with examples from Asterix and its translations. The second part of this chapter analyzes the various languages that are used to characterize different people in Jaroslav Hasek's The good soldier Svejk and how the xenistic effects...
Using the cold pressor test, three experiments were conducted to investigate the effects of water temperature and labeling on three dependent measures in college women: behavioral pain tolerance (BPT), a sensory rating of the pain experience (SR) and a parallel affective rating of the experience (AR). Temperature of the cold pressor was varied as the physical factor; labels (discomfort, pain, vasoconstriction pain) were varied as the psychological factor. Experiment I varied only water temperature; colder temperatures led to significantly lower BPT scores and significantly higher SR and AR scores. Experiment 2 varied only labeling and demonstrated that BPT decreased and AR increased as labels became more painful-sounding; in contrast, SR was unaffected by labeling. In Experiment 3 both the psychological and physical factors were varied simultaneously. Results indicated significantly higher BPT scores as the water temperature increased and the pain label became more benign. In addition, both SR and AR were sensitive to changes in temperature, whereas only AR was affected by changes in labeling.
We present a method for the extraction of stochastic lexicalized tree grammars (SLTG) of different complexities from existing treebanks, which allows us to analyze the relationship of a grammar automatically induced from a treebank wrt. its size, its complexity, and its predictive power on unseen data. Processing of different S-LTG is performed by a stochastic version of the two-step Early-based parsing strategy introduced in (Schabes and Joshi, 1991). 1 Introduction In this paper we present a method for the extraction of stochastic lexicalized tree grammars (S-LTG) of different complexities from existing treebanks, which allows us to analyze the relationship of a grammar automatically induced from a treebank wrt. its size, its complexity, and its predictive power on unseen data. The use of S-LTGs is motivated for two reasons. First, it is assumed that S-LTG better capture distributional and hierarchical information than stochastic CFG (cf. (Schabes, 1992; Schabes and Waters, 1996)),...
This research documented a linguistic norm account of direction of comparison asymmetry effects in relational judgments (e.g., seeing hyenas as more similar to dogs than dogs are similar to hyenas). The asymmetry effect is magnified by discrepancies in prominence between subject and referent, and has previously been explained using Tversky's (1977) feature-matching model. Given a linguistic norm to place more prominent objects in the referent position, violation of this norm might reduce sentence clarity, which then weakens the magnitude of subsequent relational judgments. This research showed that clarity perceptions predict the magnitude of relational judgments independently of the cognitive manipulation of the features of the compared objects. The pattern of findings suggests that a linguistic norm interpretation may account for variance in relational judgments independently of Tversky's (1977) feature-matching model.
Although the influence of emotional arousal on declarative memory has been documented behaviorally, the mechanisms underlying arousal-memory interactions and their representation in the human brain remain uncertain. One route through which arousal achieves its effects on memory performance is by regulating consolidation processes. Animal research has revealed that the amygdala strengthens hippocampal-dependent memory consolidation in a limited time window following participation in an arousing task. To examine whether this integrative function of amygdalo-hippocampal structures extends to the human brain, we tested unilateral-temporallobectomy patients on an adaptation of a classic paradigm in which levels of physiological arousal at encoding modulate retention over time. Subjects rated emotionally arousing (taboo) and neutral words on an arousal scale while their skin conductance responses (SCRs) were monitored. Recall for the words was assessed immediately and after a 1-hr delay. Both temporal-lobectomy patients and control subjects generated enhanced SCRs and arousal ratings for the arousing words at the time of encoding. However, only control subjects exhibited an increase in memory for the arousing words over time. This group difference in the effect of arousal on the rate of forgetting suggests that the role of medial temporal lobe structures in memory consolidation for arousing events is conserved across species.
This paper describes a method for determining syntactic structure in coordinate constructions. It is based on the information taken from semantic similarities, selectional restrictions, and some other linguistic cues. We discuss the role the information plays in resolving ambiguities that appear in coordinate constructions, describe the means of acquiring the necessary information automatically from two on-line corpora and a lexical database, and devise two algorithms for disambiguating coordinate constructions. An experiment that follows shows effectiveness of our method and its applicability to resolving ambiguities in some other syntactic structures. 1
Machine learning techniques can be used to make lexicons adaptive. The main problems in adaptation are the addition of lexical material to an existing lexical database, and the recomputation of sublanguage-dependent lexical information when porting the lexicon to a new domain or application. Inductive lexicons combine available lexical information and corpus data to alleviate these tasks. In this paper, we introduce the general methodology for the construction of inductive lexicons, and discuss empirical results on a case study using the approach: prediction of the gender of nouns in Dutch. 1. Introduction In computational lexicography, lexicons of language engineering applications should come with acceptable lexical coverage, and with the information necessary for the intended applications. They should also come equipped with methods for the automatic extension and adaptation of the lexicon with new or modified lexical entries. Computational lexicology should therefore try to solve t...
Zahlreiche neuere Arbeiten für das Englische zeigen, daß statistische Analysen großer Korpora und Treebanks gute Heuristiken für die Zuordnung von Präpositionalphrasen liefern können. Entsprechende Untersuchungen für das Deutsche scheitern bisher an den fehlenden Daten. Wir zeigen jedoch, daß durch Einbeziehung weiterer Faktoren auch für das Deutsche mit guten Ergebnissen zu rechnen ist. Betrachtet werden der Einfluß unterschiedlicher Gewichte für Verben und Nomina, die Auswirkungen einer vorgeschalteten lexikalischen Disambiguierung sowie die Kopplung lexikalischer und grammatischer Präferenzen. Recent proposals have shown that statistical analyses of large English corpora and treebanks provide good heuristics for the attachment of prepositional phrases. Similar proposals for German have failed since such resources have not been available. We show that by using some additional factors we can achieve similar results for German. We demonstrate the influence of different weights for verbs and nouns, the influence of lexical disambiguation and the combination of lexical and grammatical preferences.
Recent approaches to statistical parsing include those that estimate an approximation of a stochastic, lexicalized grammar directly from a treebank and others that rebuild trees with a number of tree-constructing operators, which are applied in order according to a stochastic model when parsing a sentence. In this paper we take an entirely different approach to statistical parsing, as we propose a method for parsing using a Hidden Markov Model. We describe the stochastic model and the tree construction procedure, and we report results on the Wall Street Journal Corpus.
AIM: To compare alcohol abusers' and non-abusers' distraction for alcohol-related and emotional words, controlling for emotional valence of those words. DESIGN AND METHOD: The experiment compared 20 alcohol abusers and 20 non-abusers in terms of performance on a computerized Stroop colour-naming test using alcohol-related and non-alcohol-related words. FINDINGS: Abusers rated the alcohol stimuli greater in emotional valence than the emotional stimuli. Therefore, differences in emotional-valence ratings between the two groups were statistically controlled. Against expectation, both alcohol abusers and non-abusers were more distracted by alcohol stimuli than by positive or negative emotional stimuli. CONCLUSIONS: The results indicate that alcohol words are distracting for drinkers in general, and this may indicate a high level of salience for these kinds of stimuli.
Finding simple, non-recursive, base noun phrases is an important subtask for many natural language processing applications. While previous empirical methods for base NP identification have been rather complex, this paper instead proposes a very simple algorithm that is tailored to the relative simplicity of the task. In particular, we present a corpus-based approach for finding base NPs by matching part-of-speech tag sequences. The training phase of the algorithm is based on two successful techniques: first the base NP grammar is read from a "treebank" corpus; then the grammar is improved by selecting rules with high "benefit" scores. Using this simple algorithm with a naive heuristic for matching rules, we achieve surprising accuracy in an evaluation on the Penn Treebank Wall Street Journal.
Part-of-speech tagging methodology has succeeded, but on problems that may lack real-world application. Redirection of the field is indicated, toward potentially more useful, but harder and more sophisticated tagging tasks: (1) using much more detailed tagsets (semantically and syntactically); (2) testing performance on treebanks reflecting the huge gamut of domains, etc., characterizing real-world applications; (3) understanding the magnitude of the unknown-word and unknown-tag problems, then overcoming them. Tagging results are presented on two versions of a new, highly variegated treebank, featuring tagsets of 2720 and 443 tags, respectively, and utilizing a dictionaryless, decision-tree tagger.
This paper examines y deletion in Seoul Korean on a large socio-linguistic database. It first shows that Seoul Korean has two distinct processes y deletion: categorical y deletion, which occurs after a palatal consonant, and variable y deletion occurring only before the vowel e. The first part of this paper examines variable y deletion. The examina- tion of the process reveals that it is conditioned by such internal factors as the existence of a preceding consonant, the nature of the preceding consonant, and word- internal position where y occurs, i.e., whether y appears in the word-initial or non- ward- initial syllable. This process is also affected by external factors such as the speech style, social class, and age of the speaker. The difference in the deletion rate of y among the three age groups (and among the social class groups, additionally) \nis taken to suggest that y deletion before e, i.e., monophthongization of ye to e, is a change in progress. This paper then attempts a phonological account of two different process of y deletion, proposing two OCP (coronal) constraints. The two constraints are proposed as having rather different strengths: the stronger one triggers categorical deletion and the weaker, variable deletion. A phonetic explanation of the ongoing y deletion change is also attempted. It is shown that ye has the shortest perceptual distance between its component segments among all y diphthongs of Seoul Korean. It is suggested that a relative lack of perceptual distinction between the two segments of the diphthong ye is responsible for the ongoing loss of y in this diphthongal sequence.
This paper describes the combination compound unit (CU) recognizer with syntactic verifier using partial parsing mechanism. The recognizer finds all the CUs, combined concept including collocations, idioms, and compound nouns, in input sentence. CU information reduces the search space of syntactic analysis and a portion of Part-Of-Speech (POS) ambiguities. Syntactic verification is to obtain precise CU recognition results by means of pruning wrongly recognized units that are caused by improper variable hypotheses. The experimental results show the precision of CU recognition is increased to 99.69% with 31 CFG rules on cyclic trie structure for 1,268 WSJ articles in the Penn Treebank. They also show CU recognition increases the understandability of translation for Web documents.