Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
En esta tesis se estudian las Gramaticas Incontextuales Probabilisticas y su aplicacion en problemas de Modelizacion del Lenguaje, Dos son los grandes problemas que se va a considerar en este tipo de modelos: el aprendizaje de las funciones de probabilidad asociadas a las reglas, y su integracion como modelo de interpretacion en tareas complejas de Modelizacion del Lenguaje. En primero de los problemas que se estudia es la estimacion de las funciones de probabilidad asociadas a las reglas.Se presentan y estudian dos de los algoritmos clasicos de estimacion de las GIP, el algoritmo Inside-Outside y el algoritmo basado en las cuentas de Viterbi. Despues se proponen nuevos algoritmos de estimacion en los cuales se utiliza un subconjunto especifico de derivaciones e cada cadena. Finalmente,los algoritmos propuestos se aplican al conjunto de datos del Penn Treebank para ilustrar su comportamiento en la practica. Por utlimo se aborda el problema de la interpretacion e integracion de las Gramaticas Incontextuales Probabilisticas en problemas de Modelizacion del Lenguaje. A continuacion se hace una propuesta de integracion que combina modelos de $n$-gramas a nivel de palabras con una Gramatica Incontextual Probabilistica a nivel de categorias lexicas.
In this paper we develop a formalization of semantic relations that facilitates efficient implementations of relations in lexical databases or knowledge representation systems using bases. The formalization of relations is based on a modeling of hierarchical relations in Formal Concept Analysis. Further, relations are analyzed according to Relational Concept Analysis, which allows a representation of semantic relations consisting of relational components and quantificational tags. This representation utilizes mathematical properties of semantic relations. The quantificational tags imply inheritance rules among semantic relations that can be used to check the consistency of relations and to reduce the redundancy in implementations by storing only the basis elements of semantic relations. The research presented in this paper is an example of an application of Relational Concept Analysis to lexical databases and knowledge representation systems (cf. Priss 1996) which is part of a larger framework of research on natural language analysis and formalization.
Imaging work has begun to elucidate the spatial organization of emotions; the temporal organization, however, remains unclear. Adaptive behavior relies on rapid monitoring of potentially salient cues (typically with high emotional value) in the environment. To clarify the timing and speed of emotional processing in the two human brain hemispheres, event-related potentials (ERPs) were recorded during hemifield presentation of face images. ERPs were separately computed for disliked and liked faces, as individually assessed by postrecording affective ratings. After stimulation of either hemisphere, personal affective judgements of face images significantly modulated ERP responses at early stages, 80-116 ms after right hemisphere and 104-160 ms after left hemisphere stimulation. This is the first electrophysiological evidence for valence-dependent, automatic, i.e. pre-attentive emotional processing in humans.
The digitization of library documents and archives increasingly extends to audiovisual (AV) document repositories. As a consequence, new computer-aided techniques are being devised, providing opportunities for new uses of AV documents. As scholars work mainly by reading, annotating, reusing, and producing documents they are directly concerned by these changes. The first part of this article describes AV document use in the humanities, as well as the current and future influence computers might have on evolving practices. After establishing that “full-indexing” (indexing of the content for random access to any segment of an AV document) is a necessary condition if scholars are to develop new practices in using AV material, we will focus on the specific problems raised by AV indexing as opposed to text indexing, followed by a discussion of related AV indexing projects as well as standardization issues. The third part will propose a representation model for the description of AV material (AI-Strata) and an exchange format of AV annotations (AEDI), based on a free segmentation approach. An example of annotation is also provided. The last part is devoted to a discussion regarding potential long-term influences of digital AV indexing techniques on scholarly uses of AV documents.
EuroWordNet builds a multilingual database with wordnets for several European languages.Each language specific wordnet is structured along the same lines as WordNet (Miller et al. 1990): i.e. synonyms are grouped in synsets, which in their turn are related by means of basic semantic relations, such as hyponymy, meronymy, semantic roles.Each wordnet uniquely describes the lexicalization pattern of a language.Multilinguality is achieved by storing the languagespecific wordnets in a central lexical database in which equivalent word meanings across the languages are linked to a so-called Inter-Lingual-Index (ILI).In this paper, we will address the way multilingual relations are expressed in EuroWordNet and how they are (semi-)automatically extracted from the monolingual wordnets and bilingual dictionaries.Different kinds of equivalence mappings are described to deal with fuzzy mappings and gaps.The fact that these equivalence relations are established at a more global synset level and that it is possible to browse to closely related synsets, makes it possible to get a comprehensive conceptual match of concepts across languages, even when lexicalizations differ.Other evidence, such as co-occurrence probabilities and morpho-syntactic constraints, can then be used to find a translation or a correct combination of words in a target language that is appropriate in a context.
OBJECTIVE: To evaluate the efficacy of Simulated Presence, a personalized approach to enhance well-being among nursing home residents with Alzheimer's disease and related dementia's (ADRD). DESIGN: Latin-Square, double blinded, 3-factor design with restrictive randomization of three treatments (the study intervention, a placebo audio tape of a person reading the newspaper, and usual care). The three factors were treatment, time, and facility type. SETTING: Nine nursing homes in Eastern Massachusetts and Southern New Hampshire. PARTICIPANTS: Fifty-four subjects with documented ADRD who were aged 50 years or older, medically stable, had resided in their current nursing home for at least 3 months, and who had no planned discharge. All subjects had a history of agitated or withdrawn behaviors. INTERVENTION: The purpose of Simulated Presence is to provide a personalized intervention for persons with moderate to severe cognitive impairment. Through a unique testing process, some of the best loved memories of the ADRD person's lifetime are identified and then those memories are introduced to the patient in the format of a telephone conversation using a continuous play audio tape system. The intervention may be used for extended periods of time because each repetition is viewed as a fresh, live telephone call as a result of the short-term memory deficit of the person with ADRD. MEASUREMENTS: Direct observations of outcomes included using a newly developed scale, the Scale for the Observation of Agitation in Persons with Dementia, an agitation visual analog scale, the Positive Affect Rating Scale (mood and "interest"), a withdrawal visual analog scale, and facial diagrams of mood. Reported measures included daily staff observation logs of responses to interventions offered, and weekly staff surveys using the short-form Cohen-Mansfield Agitation Inventory and the Multidimensional Observation Scale for Elderly Subjects (mood and "interest"). Severity of dementia was assessed by the Mini-Mental State Exam, the Test for Severe Impairment, the Bedford Alzheimer's Nursing Scale, and the ADL Self-Performance Scale. RESULTS: Chi-square analysis of direct observations, using facial diagrams, revealed that Simulated Presence was equivalent to usual care (P =.141) and superior to placebo for producing a happy facial expression (P =.001). A positive effect was also documented in nursing staff observation logs using Analysis of Variance techniques (ANOVA) for subjects during Simulated Presence phases compared with the placebo phases (P <.001) and usual care phases (P <.001). According to ANOVA analyses of "interest" from weekly surveys, Simulated Presence was superior to both usual care (P =.001) and placebo (P =.008). We were unable to find evidence of significant differences (P <.05) among interventions for other direct observations and weekly reports of overall agitation or mood aspects of withdrawal. Subjects accepted the intervention most of the time, except for five subjects who refused it more than 50% of the time. CONCLUSION: This study provided evidence that Simulated Presence can be effective in enhancing well-being and decreasing problem behaviors in the nursing home setting as a substitute for or complement to usual care.
Electronic texts are claimed to exhibit features distinct from their more tangible cousins. The Snapshot project aims to observe and capture language usage in an electronic medium by creating an open corpus of World Wide Web documents. These documents are re-encoded using the TEI guidelines to create a flexible, persistent and portable data repository. This report gives an overview of the decisions made with respect to the re-encoding of HTML documents, and with the structuring the overall corpus.
We examined male and female abservers' reactions to the use of touch by a nurse towards a patient in a hospital situation. The results suggest that an overrall pattern for observers to react more favorably to the nurse touch compared to no touch in interacting with patients may be judged in part by the attitudes of males and females about the use of touch. The generally favorable reaction to the use of touch by a nurse is consistent with Lewis et al's (1995) study. However, in contrast to their study, there were no differences in the nurse's supportiveness, competence, affect ratings and nurse's confidence in the low touch and high touch conditions.
The traditions of grammatical terminology are quite different in the German and the French educational system. While the French ministry of education issues official regulations which are binding on a national level (they last did so in 1997), the federal system in Germany prevents the standardisation of the grammatical nomenclature for all the Länder of the FRG. Moreover, the historico-cultural background is different in both countries: the French public has been used to an active and centralist language policy for several centuries; in Germany governmental interference with the linguistic norm is often met with resistance, and the regulation of many details is in fact left to the important publishing houses. For this very reason it was left largely to the school-book publishing houses to decide how to put the recommendations of the German Secretaries of cultural affairs issued in 1982 into practice. The present article is based on a number of selected examples and provides a critical analysis of the usage of grammatical terminology in German and French school-books, and it pleads for a re-orientation – particularly in Germany: the inconsistencies and ad hoc solutions which can be observed quite frequently can only be overcome if the terminology is based on a well reflected linguistic theory. This is the only way to achieve an effect of synergy between grammar lessons in different school languages and ultimately a standardisation of the grammatical terminology on a European level.
The use of the Inside-Outside (IO) algorithm for the estimation of the probability distributions of Stochastic ContextFree Grammars is characterized by the use of all the derivations in the learning process. However, its application in real tasks for Language Modeling is restricted due to the that it needs to converge. Alternatively, several estimations algorithms which consider a certain subset of derivations in the estimation process have been proposed elsewhere. This set of derivations can be chosen according to structural criteria, or by selecting the k-best derivations. These alternatives are studied in this paper, and they are tested on the corpus of the Wall Street Journal processed in the Penn Treebank project.
Thesauri have been widely used in bibliographic databases for 30 years. Recently, CD-ROMs of a variety of dictionaries with their GUIs are spreaded to current users. On the other hand, End users dose not use thesauri as for their bulky printed matter. The king of software browsing graphically and managing thesauri does not appear in PC environment.. The browsing tool for lexical database with hierarchical structure has been developed using Java.. This paper describes the functions, the components, and examples of its usage. The problems of the browser and the functions to be extended are discussed.
Consciousness of the arbitrary nature of language, its essential conventionality (in the sense not of conformity, but of operating in accordance with tacitly agreed codes which have no naturally inherent laws) has become a major feature of twentieth-century poetry. It is almost, one is tempted to say, the defining feature of the modern, were it not that post-modern poetry is still more marked by it than classical modernism of the Pound-Eliot-Williams upheaval. By freeing poetry from the demands of consecutive syntax and claiming as its own the modern cinematic technique of juxtaposing images for primarily emotional/dramatic effect, modernism opened the way to further experiments in the breaking down of accepted assumptions governing verbal expression. Once the standard of 'correct', transparent English, whether spoken or, still more significantly, printed, was breached, it became possible to question the conventions of presentation which educated writers and readers had come to take as inviolable rules, and which are still treated as such in the language of scientific, journalistic and critical discourse. It became possible for transgression, or non-observance, of the rules to function on a positively sophisticated, rather than vulgarly negative, level — though this, of course, also presupposes general familiarity and conformity with them, since the abnormal effects of dispensing with them, or operating them in unfamiliar ways, depends on the existence of a strong feeling for them as the linguistic norm.
The rapid political changes that have taken place in Eastern Europe over the past five years have brought about a sudden disintegration of many well-established social, economic and cultural patterns. This process is clearly visible in the languages of the former 'socialist camp'. Standard Russian (the Russian literary language) has in the past few years undergone such far-reaching changes that it is already possible to speak of its having acquired a new functional status: by escaping the tight boundaries of a rigorous purism and by a renewal of its lexical and phraseological resources, it has become much more democratic, cosmopolitan and dynamic. Where there was formerly a set of rules drawn up by 'the guardians of the purity of the Russian language', who used as their models either the literature of the classics and of the 'socialist realist' period or else the clichés of bureaucratic 'journalese', a preference is now shown for the language of the 'de-sovietised' mass media and for a spontaneous living language (including sub-standard elements and various forms of slang) which is no longer subject to the control of censors or the army of in-house editors. It is significant that under the pressure of the present language situation even those linguists who until recently were not prepared to contemplate any departure from the strict principles of 'language purity' or 'language culture' (культура речи) are now forced either to rely on the fluid and in many respects subjective criterion of 'language taste' (языковой вкус) or else to recognise the existence of a 'vulgarisation of the present-day literary norm' (вульгаризация современной литературной нормы).1KeywordsShock TherapyRussian LanguageSlavonic LanguageLanguage SituationLanguage CultureThese keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves.
The act of healing essentially includes a spiritual or religious component, and language is often a prime medium for enacting it. Personal narratives are sites for the negotiation and construction of cultural and linguistic norms; healing stories recontextualize bodily struggles as social and spiritual conflicts. This paper examines a personal narrative of spiritual healing told in the language of Rastafari, the Jamaican religious movement. A discourse analysis of the narrative focuses on elements of Rasta Talk in order to discover how Rastafarian beliefs underlie and shape the telling, which is itself an act of faith and a profession of commitment. The healing itself, however, draws primarily on a variety of nonRasta spiritual and occult traditions of Jamaican folk culture; their relation to Rastafari, and the reasons for employing Rasta religious rhetoric in the narrative, are also explored. Rasta Talk, a register of Jamaican Creole (JC) undergoing functional expansion, is characteristically (though by no means exclusively) used by Jamaicans who follow the Rastafari religion. Rastafari is a syncretic Afro-Christian faith which invokes and reinterprets Old Testament Biblical imagery in the service of particular religious, cultural and political themes. This narrative of supernatural illness and cure applies a historical critique of colonialism and racism to the healthcare system, allows the teller to reposition himself discursively to alleviate suffering and stigma, and claims the moral high ground for AfroJamaican ethnomedical practices and traditional values through the enactment of Rastafarian principles. Rastafari is briefly introduced first in relation to other Jamaican faith traditions, and Rasta Talk is described. The narrative is outlined chronologically. Subsequent analysis links linguistic features to key elements of Rasta beliefs, which – together with elements of Jamaican folk medicine and culture – provide the necessary context to understand the healing narrative as an act of identity.
This paper explores the role of lexicalization and prun-ing of grammars for base noun phrase identification. We modify our original framework (Cardie &amp; Pierce 1998) to extract lexicalized treebank grammars that assign a score to each potential noun phrase based upon both the part-of-speech tag sequence and the word sequence of the phrase. We evaluate the mod-ified framework on the “simple ” and “complex ” base NP corpora of the original study. As expected, we find that lexicalization dramatically improves the perfor-mance of the unpruned treebank grammars; however, for the simple base noun phrase data set, the lexical-ized grammar performs below the corresponding unlex-icalized but pruned grammar, suggesting that lexical-ization is not critical for recognizing very simple, rel-atively unambiguous constituents. Somewhat surpris-ingly, we also find that error-driven pruning improves the performance of the probabilistic, lexicalized base noun phrase grammars by up to 1.0 % recall and 0.4% precision, and does so even using the original pruning strategy that fails to distinguish the effects of lexical-ization. This result may have implications for many probabilistic grammar-based approaches to problems in natural language processing: error-driven pruning is a remarkably robust method for improving the perfor-mance of probabilistic and non-probabilistic grammars alike.
This study examined the effects of performance information from the ratee's coworker both on rater's perceptions of the information and ratings of the ratee's performance. As part of a managerial role-play exercise, subjects were required to rate an employee's performance. Audiotaped interactions of the ratee and customers were manipulated to reflect either good or poor performance. In addition, the ratee's coworker provided performance information which varied on (a) who was the initiator of the report and (b) the favorability of the information. Results revealed that while performance information from the ratee's coworker did not significantly affect ratings of performance when it was consistent with the rater's direct observations, ratings were affected when the information from a coworker was inconsistent with the rater's direct observations even though the information from a coworker was perceived as less accurate and of less use. Copyright © 1999 John Wiley & Sons, Ltd.
This article discusses the principal problems of delimiting the scope of adverbs and their classification in terms of lexicology, lexicography and natural language processing, with a view to making the exhaustive lexical database for adverbs as well as to developing the description models which lead to the construction of the lexicon. Our discussion is not limited to the adverbs in traditional sense, but extended to the mono-/poly lexical words which can be considered as adverbials morphologically, syntactically and lexicologically. For instance, N + postposition(distributionally very limited), limited and/or productive inflectional forms of verbs and adjectives, and fixed forms of phrases and sentences which we consider as compound adverbs. To support our explanation, the morphological and syntactic properties of adverbs are discussed with typological and language universal point of view. Semantic classification is not included in this article because we think it can not be made rigorously without delimiting and analyzing possible various forms of adverbs(and adverbials) and their syntactic functions.
This study describes the typical course and variability in major areas of communicative development for 228 Swedish-speaking children between 8 and 16 months of age. The assessments were made by parental reports with the Swedish Early Communicative Development Inventories (SECDI) using a semi-longitudinal design. Age-based norms for understanding of phrases, vocabulary comprehension, vocabulary production and use of gestures are described at the 10th, 25th, 50th, 75th and 90th percentile levels. More lexical verbs were found among the first words in comprehension than in production. An extensive variability within individuals in onset and development was found for the assessed skills. The individual differences proved to be stable over 4–6 months. No gender differences were found for comprehension of phrases, total gestures, vocabulary compre-hension, or for vocabulary production. Strong, unique associations were found between total gestures and vocabulary comprehension and between vocabulary comprehension and vocabulary production. In contrast, no unique association was found between gestures and vocabulary production. The results generally concur with those reported for English-speaking American children by Fenson et al. (1993, 1994).
We tested new analytic procedures for combining an observer's image-ratings of lesion-likelihood with localization reports that are incomplete (unavailable on images rated as 'normal') and/or imprecise (possibly scored as 'correct' by chance), and for fitting a constrained ROC formulation to the rating data alone. Eight radiologist readers in a previous study had rated the likelihood of nodular lesions on each of 250 chest-film cases (39 with subtle nodules, 36 with 'typical' nodules and 175 normal cases) that were presented in two display modes (original films or on video workstation). Ratings in the four positive categories (2 to 5) were accompanied by reports that grossly localized the suspected nodules into one of 7 film- regions (upper, middle or lower portions of left or right lung field, or retrocardiac), but there was no localization for the cases rated as 'normal' (category 1). In each of 29 sets of data, we estimated the area below the ROC curve (A<SUB>z</SUB>) and its standard error using three different fits: (1) the usual ROC formulation, (2) the constrained ROC formulation and (3) the new procedure that included incomplete and imprecise localization data (I&I). Estimates of A<SUB>z</SUB> from the usual and constrained ROC fits were quite similar unless the standard ROC exhibited an upward 'hook,' but standard errors of A<SUB>z</SUB> were always the same or smaller for the constrained ROC fit. The I&I fit that included localization data often estimated A<SUB>z</SUB> to be either larger or smaller than the usual or constrained ROC fits that considered only the rating data, but its A<SUB>z</SUB> had substantially smaller standard errors in 28 of the 29 sets of observer data.
In this paper, we propose an error correction method using text corpora. In this method, recognition errors are corrected using phonetically similar examples in the text corpora. The reliability of the correction hypotheses are judged according to their semantic consistency and their phonetic similarity to the original input. We previously proposed an error correction method that uses a treebank [1]. However, the previous method was not flexible in its use of examples, because structural mismatches occurred between the input and examples due to recognition errors. In our new proposal, examples are treated as morpheme sequences. This enables us to use examples partially when there are no useful full-sentence-examples. We built our proposed method into a speech translation system and compared the translation quality for simple translation and translation with error correction. The rate of acceptable translation increased about 10% with our proposed method compared to simple translation.
We present a new approach to partial parsing of natural language texts that relies on machine learning methods. The approach combines corpus-based grammar induction with a very simple pattern-matching algorithm and an optional constituent verification step. The grammar induction algorithm acquires a set of rules for each level of linguistic analysis using a new technique for errordriven pruning of treebank grammars. The constituent verification step employs standard inductive learning techniques as an additional precision-enhancing device. We evaluate the approach on four partial parsing data sets and find that performance is very good (over 93% precision and recall) for applications that require or prefer fairly simple constituent bracketing. As the complexity of the partial parsing task increases, however, our approach lags the performance of competing approaches. We explain these differences in terms of the knowledge sources employed by each method and describe a number of features...
This paper presents a pronoun resolution algorithm that adheres to the constraints and rules of Centering Theory (Grosz et al., 1995) and is an alternative to Brennan et al.'s 1987 algorithm. The advantages of this new model, the Left-Right Centering Algorithm (LRC), lie in its incremental processing of utterances and in its low computational overhead. The algorithm is compared with three other pronoun resolution methods: Hobbs' syntax-based algorithm, Strube's S-list approach, and the BFP Centering algorithm. All four methods were implemented in a system and tested on an annotated subset of the Treebank corpus consisting of 2026 pronouns. The noteworthy results were that Hobbs and LRC performed the best.
Using the cold pressor test, three experiments were conducted to investigate the effects of water temperature and labeling on three dependent measures in college women: behavioral pain tolerance (BPT), a sensory rating of the pain experience (SR) and a parallel affective rating of the experience (AR). Temperature of the cold pressor was varied as the physical factor; labels (discomfort, pain, vasoconstriction pain) were varied as the psychological factor. Experiment I varied only water temperature; colder temperatures led to significantly lower BPT scores and significantly higher SR and AR scores. Experiment 2 varied only labeling and demonstrated that BPT decreased and AR increased as labels became more painful-sounding; in contrast, SR was unaffected by labeling. In Experiment 3 both the psychological and physical factors were varied simultaneously. Results indicated significantly higher BPT scores as the water temperature increased and the pain label became more benign. In addition, both SR and AR were sensitive to changes in temperature, whereas only AR was affected by changes in labeling.
We present a method for the extraction of stochastic lexicalized tree grammars (SLTG) of different complexities from existing treebanks, which allows us to analyze the relationship of a grammar automatically induced from a treebank wrt. its size, its complexity, and its predictive power on unseen data. Processing of different S-LTG is performed by a stochastic version of the two-step Early-based parsing strategy introduced in (Schabes and Joshi, 1991). 1 Introduction In this paper we present a method for the extraction of stochastic lexicalized tree grammars (S-LTG) of different complexities from existing treebanks, which allows us to analyze the relationship of a grammar automatically induced from a treebank wrt. its size, its complexity, and its predictive power on unseen data. The use of S-LTGs is motivated for two reasons. First, it is assumed that S-LTG better capture distributional and hierarchical information than stochastic CFG (cf. (Schabes, 1992; Schabes and Waters, 1996)),...
This research documented a linguistic norm account of direction of comparison asymmetry effects in relational judgments (e.g., seeing hyenas as more similar to dogs than dogs are similar to hyenas). The asymmetry effect is magnified by discrepancies in prominence between subject and referent, and has previously been explained using Tversky's (1977) feature-matching model. Given a linguistic norm to place more prominent objects in the referent position, violation of this norm might reduce sentence clarity, which then weakens the magnitude of subsequent relational judgments. This research showed that clarity perceptions predict the magnitude of relational judgments independently of the cognitive manipulation of the features of the compared objects. The pattern of findings suggests that a linguistic norm interpretation may account for variance in relational judgments independently of Tversky's (1977) feature-matching model.
Treebanks, such as the Penn Treebank (PTB), offer a simple approach to obtaining a broad coverage grammar: one can simply read the grammar off the parse trees in the treebank. While such a grammar is easy to obtain, a square-root rate of growth of the rule set with corpus size suggests that the derived grammar is far from complete and that much more treebanked text would be required to obtain a complete grammar, if one exists at some limit. However, we offer an alternative explanation in terms of the underspecification of structures within the treebank. This hypothesis is explored by applying an algorithm to compact the derived grammar by eliminating redundant rules - rules whose right hand sides can be parsed by other rules. The size of the resulting compacted grammar, which is significantly less than that of the full treebank grammar, is shown to approach a limit. However, such a compacted grammar does not yield very good performance figures. A version of the compaction algorithm taking rule probabilities into account is proposed, which is argued to be more linguistically motivated. Combined with simple thresholding, this method can be used to give a 58% reduction in grammar size without significant change in parsing performance, and can produce a 69% reduction with some gain in recall, but a loss in precision.
Most bipolar models of affective processing in social psychology assume that positive and negative valent processes are represented along a single continuum that rangesfrom very positive to very negative. Recent research has raised the possibility, however, that the motivational systems for positive/approach and negative/defensive valent processing (positivity and negativity, respectively) are separable. In this article, the authors use unipolar positivity, negativity, and ambivalence ratings and bipolar valence, dominance, and arousal ratings of 472 slides from the International Affective Picture System to examine several aspects of the bivariate model of evaluative space. Analysis confirmed a positivity offset and negativity bias in the activation functions of the valent systems as wel as multiple modes of evaluative activation (e.g., reciprocal, uncoupled positivity, uncoupled negativity). Together, these data suggest that the bipolar structure of affective processes should be tested rather than assumed.
We argue that the current dominant paradigm in parser evaluation work, which combines use of the Penn Treebank reference corpus and of the Parseval scoring metrics, is not well-suited to the task of general comparative evaluation of diverse parsing systems. We propose an alternative approach which has two key components. Firstly, we propose parsed corpora for testing that are much flatter than those currently used, whose &quot;gold standard&quot; parses encode only those grammatical constituents upon which there is broad agreement across a range of grammatical theories. Secondly, we propose modified evaluation metrics that require parser outputs to be `faithful to&apos;, rather than mimic, the broadly agreed structure encoded in the flatter gold standard analyses. 1. Introduction Interest in the evaluation of language technology has grown immensely in the past few years. This interest varies depending on the perspective one has on the technology: users and suppliers want to know how accurate, usabl...
AIM: To compare alcohol abusers' and non-abusers' distraction for alcohol-related and emotional words, controlling for emotional valence of those words. DESIGN AND METHOD: The experiment compared 20 alcohol abusers and 20 non-abusers in terms of performance on a computerized Stroop colour-naming test using alcohol-related and non-alcohol-related words. FINDINGS: Abusers rated the alcohol stimuli greater in emotional valence than the emotional stimuli. Therefore, differences in emotional-valence ratings between the two groups were statistically controlled. Against expectation, both alcohol abusers and non-abusers were more distracted by alcohol stimuli than by positive or negative emotional stimuli. CONCLUSIONS: The results indicate that alcohol words are distracting for drinkers in general, and this may indicate a high level of salience for these kinds of stimuli.
A series of four experiments were conducted to examine viewer perceptions of three sets of five nonrepresentational paintings. Increased complexity was embedded in the hierarchical structure of each set by carefully selecting colors and ordering them in each successive painting according to certain rules of transformation which created hierarchies. Experiment 1 supported the hypothesis that subjects would discern the hierarchical complexity underlying the sets of paintings. In Experiment 2 viewers rated the paintings on collative (complexity, disorder) and affective (pleasing, interesting, tension, and power) scales, and a factor analysis revealed that affective ratings were tied to complexity (Factor 1) but not to disorder (Factor 2). In Experiment 3, a measure of exploratory activity (free looking time) was correlated with complexity (Factor 1) but not with disorder (Factor 2). Multidimensional scaling was used in Experiment 4 to examine perceptions of the paintings seen in pairs. Dimension 1 contrasted Soft with Hard-Edged paintings, while Dimension 2 reflected the relative separation of figure from ground in these paintings. Together these results show that untrained viewers can discern hierarchical complexity in paintings and that this quality stimulates affective responses and exploratory activity.
BOOK NOTICES 207 sure's concept ofmotivation should not be associated with words but rather with cotext and context. Everything is relative in language, and the systematic character of language does not lie in separate phonemic, lexical, grammatical, and textual systems. In fact the systematic and universal features of language are made up by the human capabilities of thinking and experiencing. Speakers are able to use a limited number of signs to express highly complicated ideas, and they can decode expressions in very complex situations. To do this requires applying the basic principle of language and language description—the idea of economy. This elementary principle of human behavior can be seen when a speaker tries to avoid unnecessary redundancy and when simpler ways of pronunciation are preferred to more difficult ones. In addition language economy can be seen on the deeper level of reinterpreting linguistic units new to a certain speaker. This gives sense to an utterance since under the principle ofeconomy, a speaker must assume that no text is uttered without meaning. As a result of new expressions and new interpretations produced by the principle ofeconomic use, language might change over time. Considering language change again, D emphasizes how the individual reflects about language. For the individual the main goal of language and speaking is to impart and to decode sense or meaning. D rejects the idea of the 'invisible hand phenomenon' as well as the concept of teleology in language change. Again he links the systematic character of language to the speaker's purposeful acts rather than to the whole speech community or to language as an abstract system. In short, D questions the traditional structuralist conception of system in language. He emphasizes the systematic cognitive behavior of each individual that uses the relative means of language. In order to support his opinion D illustrates his ideas with many detailed examples mainly taken from German, English, and French. [Dieter Aichele, Fachhochschule Neubrandenburg.] The Oxford English-Hebrew Dictionary. Ed. by N. S. Doniach and A. Kahane. Oxford: Oxford University Press, 1996. Pp. xxiii, 1091. The late N. S. Doniach (d. 16 April 1994), the chief editor ofthis dictionary, is well known to Semitologists for editing the excellent Oxford EnglishArabic dictionary ofcurrent usage (1972). The introduction by Professor A. Kahane explains D's goal of using various styles of modern Hebrew in this volume, including colloquial language and slang. From abacus to Zulu, this dictionary, happily, has it all! It even has the f-word with many of its most common idiomatic usages, such as '__ up' and '__ off!' (351). The tome's particularly noteworthy features include up-to-date terminology ofall sorts, such as that dealing with computers. However, some inconsistencies can be found. 'Software' is written as toxna with a vav (883), but it is spelled without a vav (using a kamats katan) under 'hardware' (400). And curiously, the name of the vowel kamats is written kamatz, yet the vowel chataf kamats is spelled differently on the very next line (ix). AU Hebrew words are given in their fully vocalized or pointed forms, including dagesh, which is said to have the phonetic value 'stress mark', a puzzling statement (ix). Phonological matters on the whole, however, have been handled well. It was a wise decision for the editors to give preference to a word's modern pronunciation if it differs from its traditional pointing (xxiii). Along these lines, I wanted to check the vocalization and pronunciation of the irregular Classical Hebrew plural for bayit 'house'—battiim or (bottiim); to my surprise, however, it was not given (426). British English has been chosen as the norm of this dictionary, a reasonable and expected decision by Oxford University Press. An American has little difficulty getting used to British spellings such as 'programme'; however, 'program' is also listed, but with the stipulation, 'US and Comput.' (724). However, it may be difficult at first for an American to appreciate 'farther' and 'father' transcribed exactly the same. There are a few discrepancies to report. American English 'buggy' is given as eglat-tinok (114), but 'pram' lists only eglat-tinokot, its plural (708), whereas the more formal 'perambulator' (called 'formal ' by the editors) lists...
A well-known problem in the domain of quantitative linguistics and stylistics concerns the evaluation of the lexical richness of texts. Since the most obvious measure of lexical richness, the vocabulary size (the number of different word types), depends heavily on the text length (measured in word tokens), a variety of alternative measures has been proposed which are claimed to be independent of the text length. This paper has a threefold aim. Firstly, we have investigated to what extent these alternative measures are truly textual constants. We have observed that in practice all measures vary substantially and systematically with the text length. We also show that in theory, only three of these measures are truly constant or nearly constant. Secondly, we have studied the extent to which these measures tap into different aspects of lexical structure. We have found that there are two main families of constants, one measuring lexical richness and one measuring lexical repetition. Thirdly, we have considered to what extent these measures can be used to investigate questions of textual similarity between and within authors. We propose to carry out such comparisons by means of the empirical trajectories of texts in the plane spanned by the dimensions of lexical richness and lexical repetition, and we provide a statistical technique for constructing confidence intervals around the empirical trajectories of texts. Our results suggest that the trajectories tap into a considerable amount of authorial structure without, however, guaranteeing that spatial separation implies a difference in authorship.
This paper presents a theoretical approach of how simple, episodic associations are transduced into semantic and grammatical categorical knowledge. The approach is implemented in the hyperspace analogue to language (HAL) model of memory, which uses a simple global co-occurrence learning algorithm to encode the context in which words occur. This encoding is the basis for the formation of meaning representations in a high-dimensional context space. Results are presented, and the argument is made that this simple process can ultimately provide the language-comprehension system with semantic and grammatical information required in the comprehension process.
The lateralized readiness potential (LRP) is an electrophysiological indicator of the central activation of motor responses. Procedures for deriving the LRP on the basis of event-related brain potential (ERP) waveforms obtained over the left and right motor cortices are described, and some findings are summarized that show that the LRP is likely to reflect activation processes within the motor cortex. Two experiments investigating spatial S-R compatibility effects are reported that demonstrate that, because of systematic overlaps of motor and nonmotor asymmetries, LRP waveforms derived by the double subtraction method cannot always be interpreted unequivocally in terms of response activation. Such confounds can be detected when LRP waveforms are compared with difference waveforms obtained by the double subtraction method from ERPs elicited at other lateral scalp sites.
Participants judged the believability of simulated accounts varying in perceptual and emotional detail. In Experiment 1, both younger and older adults' tendency to believe that an account described an actually experienced event increased as either type of detail was added. In Experiment 2, while younger adults in a low suspicion condition judged the accounts with added details as more likely to be of actually-experienced events than the impoverished accounts, those given instructions designed to induce suspicion about the speakers' honesty found the more detailed accounts less believable. In Experiment 3, for both younger and older adults, both types of detail again increased believability ratings under low-suspicion conditions but did not affect ratings under high-suspicion conditions. In addition, there were systematic differences in the types of details the high- and low-suspicion participants reported using to make their judgments. Results are discussed in relation to the source monitoring framework (Johnson, Hashtroudi, & Lindsay, 1993) and reality monitoring and credibility judgments.
Children's judgements about pain at age 8–10 years were examined comparing two groups of children who had experienced different exposure to nociceptive procedures in the neonatal period: extremely low birthweight (ELBW) ≤ 1000 g ( N = 47) and full birthweight (FBW) ≤ 2500 g ( N = 37). The 24 pictures that comprise the Pediatric Pain Inventory, depicting events in four settings: medical, recreational, daily living, and psychosocial, were used as the pain stimuli. The subjects rated pain intensity using the Color Analog Scale and pain affect using the Facial Affective Scale. Child IQ and maternal education were statistically adjusted in group comparisons. Pain intensity and pain affect related to activities of daily living and recreation were significantly higher than psychosocial and medically related pain on both scales in both groups of children. Although the two groups of children did not differ overall in their perceptions of pain intensity or affect, the ELBW children rated medical pain intensity significantly higher than psychosocial pain, unlike the FBW group. Also, duration of neonatal intensive care unit stay for the ELBW children was related to increased pain affect ratings in recreational and daily living settings. Despite altered response to pain in the early years reported by parents, on the whole at 8–10 years of age ELBW children judged pain in pictures similarly to their term peers. However, differences were evident, which suggests that studies are needed of biobehavioural reactivity to pain beyond infancy, as well as research into beliefs, attitudes, and perceptions about pain during the course of childhood in formerly ELBW children.
This paper describes the combination compound unit (CU) recognizer with syntactic verifier using partial parsing mechanism. The recognizer finds all the CUs, combined concept including collocations, idioms, and compound nouns, in input sentence. CU information reduces the search space of syntactic analysis and a portion of Part-Of-Speech (POS) ambiguities. Syntactic verification is to obtain precise CU recognition results by means of pruning wrongly recognized units that are caused by improper variable hypotheses. The experimental results show the precision of CU recognition is increased to 99.69% with 31 CFG rules on cyclic trie structure for 1,268 WSJ articles in the Penn Treebank. They also show CU recognition increases the understandability of translation for Web documents.
Abstract. This article sets out some of the results of wider research on the linguistic databases of a natural language generation system. One of the necessary steps in the building of such databases is to determine the linguistic means the generator must have in order to produce a linguistic form that corresponds to the semantic representation given as an input. We wish to focus here on the theoretical choices and issues rather than on the application itself. We assume that texts have a syntactic structure, whose characteristics are partly comparable to the syntactic structure of a sentence, and that a connective can be considered as a textual predicate which has arguments that are constrained in the same way as the arguments of a verb. This article will concentrate more specifically on one particular semantic relation — the simultaneity of two events — and will show how the taxonomy of the associated connectives can be elaborated. Finally, we will set out some of the major developments of this research, which concern the interface between conceptual and linguistic knowledge.
We show how a treebank can be used to cluster words on the basis of their syntactic behavior. By extracting statistics on the structures in which words appear it is possible to discover similarities and differences in usage between words with the same part-of-speech. This clustering is compared to the conventional clustering based on co-occurrences. While conventional clustering can discover semantical similarities or the tendency to appear together, the method we present ignores these factors and places the focus on syntactical usage, in other words the sort of structures it appears in. We present a case study on prepositions, showing how they can be automatically subdivided by their syntactic behavior and we discuss the appropriateness of such a subdivision. We have also carried out experiments to compare the quality of clusters quantitatively. For this goal we used clusters based on syntactic behavior for improving the estimation of the distribution of the dependency relation between words. Since such a distribution is necessarily estimated with sparse data, an entropy test can show how informative the classes are about syntactic usage. Finally, we discuss a number of ways in which a classification of words can contribute to applications of natural language processing.
In comparison to conventional displays, 3D stereoscopic displays convey additional information about the 3D structure of a scene by providing information that can be used to extract depth. In the present study we evaluated the psychovisual impact of stereoscopic images on viewers. Thirty-three non-expert viewers rated sensation of depth, perceived sharpness, subjective image quality, and relative preference for stereoscopic over non-stereoscopic images. Rating methods were based on procedures described in ITU- Rec. 500. Viewers also rated sequences in which the left- and right-eye images were processed independently, using a generic MPEG-2 codec, at bit-rates of 6, 3, and 1 Mbits/s. The main finding was that viewers preferred the stereoscopic version over the non-stereoscopic version of the sequences, provided that the sequence did not contain noticeable stereo artifacts, such as exaggerated disparity. Perceived depth was rated greater for stereoscopic than for non-stereoscopic sequences, and perceived sharpness of stereoscopic sequences was rated the same or lower compared to non-stereoscopic sequences. Subjective image quality was influenced primarily by apparent sharpness of the video sequences, and less so by perceived depth.
This paper describes a method for determining syntactic structure in coordinate constructions. It is based on the information taken from semantic similarities, selectional restrictions, and some other linguistic cues. We discuss the role the information plays in resolving ambiguities that appear in coordinate constructions, describe the means of acquiring the necessary information automatically from two on-line corpora and a lexical database, and devise two algorithms for disambiguating coordinate constructions. An experiment that follows shows effectiveness of our method and its applicability to resolving ambiguities in some other syntactic structures. 1
Zahlreiche neuere Arbeiten für das Englische zeigen, daß statistische Analysen großer Korpora und Treebanks gute Heuristiken für die Zuordnung von Präpositionalphrasen liefern können. Entsprechende Untersuchungen für das Deutsche scheitern bisher an den fehlenden Daten. Wir zeigen jedoch, daß durch Einbeziehung weiterer Faktoren auch für das Deutsche mit guten Ergebnissen zu rechnen ist. Betrachtet werden der Einfluß unterschiedlicher Gewichte für Verben und Nomina, die Auswirkungen einer vorgeschalteten lexikalischen Disambiguierung sowie die Kopplung lexikalischer und grammatischer Präferenzen. Recent proposals have shown that statistical analyses of large English corpora and treebanks provide good heuristics for the attachment of prepositional phrases. Similar proposals for German have failed since such resources have not been available. We show that by using some additional factors we can achieve similar results for German. We demonstrate the influence of different weights for verbs and nouns, the influence of lexical disambiguation and the combination of lexical and grammatical preferences.
Word recognition and generation is a fundamental part of the processing of natural language and it requires computationally effective morphological processors, especially for languages with rich morphology such as Modern Greek. Various models have been proposed for developing computerized systems to accomplish the task of recognition of morphosyntactic features of words In the work presented here, the lazy tagging approach was examined, in which taggers are expected to work in the simplest possible way. The model of functional decomposition was extended and adapted for Modern Greek as a target language, following the lazy word-parsing approach, in order to cover a number of morphological phenomena that are encountered in Modern Greek, namely inflection, affixation, and longdistance dependencies. To achieve a more efficient word recognition, several automata of different levels of computing power, based on the original model, were introduced and evaluated according to the criteria of complexity, recognition speed, and accuracy of the results. The proposed system was used for processing a large-scale corpus, and the results are presented and discussed. To accomplish their task, taggers can rely upon large lexical databases, which are expected to be organized in such a way as to provide rapid access to the stored data and efficient memory management. Directed graphs can be used to describe and organize a lexical database of large magnitude in a compact manner. These data structures are named here matrix lexica, where the letters are described as nodes of directed graphs and the lemmata as paths (set of edges). It is expected that matrix lexica will support a tagger efficiently by providing a high speed of resolution, sound mathematical foundation, low memory requirements, and ability to handle distorted input in future developments.
Accurate linguistic annotation is a core requirement of natural language processing systems. The demand for accuracy in the face of rapid prototyping constraints and numerous target languages has led to the employment of machine learning methods for developing linguistic annotation systems. The popularity of applying machine learning methods to computational linguistics problems has given rise to a large supply of trainable natural language processing systems. Most problems of interest have an array of off-the-shelf products or downloadable code implementing solutions using various techniques. In situations where these solutions are developed independently, it is observed that their errors tend to be independently distributed. In this thesis we discuss approaches for capitalizing on this situation in a sample problem domain, Penn Treebank-style parsing. The machine learning community provides us with techniques for combining outputs of classifiers, but parser output is more structured and interdependent than classifications. To overcome this, two novel strategies for combining parsers are used: learning to control a switch between parsers and constructing a hybrid parse from multiple parsers' outputs. In this thesis we give supervised and unsupervised techniques for each of these strategies as well as performance and robustness results from evaluation of the techniques. One shortcoming of combining off-the-shelf parsers is that the parsers are not developed with the intention to perform well on complementary data or to compensate for each others' weaknesses. The individual parsers are globally optimized. We present two techniques for producing an ensemble of parsers in such a way that their outputs can be constructively combined. All of the ensemble members will be created using the same underlying parser induction algorithm, and the method for producing complementary parsers is only loosely coupled to that algorithm.
This paper proposes a new inference approach for Chinese probabilisticcontext-free grammar, which implements the EM algorithm based on the bracketmatching schemes. Two characteristics of the algorithm are as follows: 1) To pre-process the training texts with automatic constituent boundary prediction tools,which can provide stronger syntactic restriction upon training texts in lower compu-tational costs; 2) To develop an initial rule set by integrating different knowledgeresources, including a set of basic syntactic rules generated by an automatic gram-mar construction t00l and a set of special rules summarized by linguists or extractedfrom treebanks, and provide a better initialization for the learning process. There-fore, a linguistically-motivated and broad-coverage Chinese PCFG rule set can beeasily generated through this algorithm. Current experimental results prove goodlearning efficiency of this algorithm and high reliability of the generated rule set.
We argue that the current dominant paradigm in parser evaluation work, which combines use of the Penn Treebank reference corpus and of the Parseval scoring metrics, is not well-suited to the task of general comparative evaluation of diverse parsing systems. In (Gaizauskas et al., 1998), we propose an alternative approach which has two key components. Firstly, we propose parsed corpora for testing that are much flatter than those currently used, whose “gold standard ” parses encode only those grammatical constituents upon which there is broad agreement across a range of grammatical theories. Secondly, we propose modified evaluation metrics that require parser outputs to be ‘faithful to’, rather than mimic, the broadly agreed structure encoded in the flatter gold standard analyses. This paper addresses a crucial issue for the (Gaizauskas et al., 1998) approach, namely, the creation of the evaluation resources that the approach requires, i.e. annotated corpora recording the flatter parse analyses. We argue that, due to the nature of the resources required, they can be derived in a comparatively inexpensive fashion from existing parse annotated resources, where available. 1.
En se situant dans le cadre d'un projet appele Semantic dictionary viewed as a lexical database dont l'objectif est d'identifier et de denombrer les informations linguistiquement pertinentes pouvant etre derivees des definitions du sens, l'A. examine ici les processus productifs de derivation semantique des verbes en russe, et en particulier les processus associes a un changement metonymique qui se refletent dans le cadre du cas profond d'un verbe. Il tente ainsi de montrer que les roles semantiques permettent de decouvrir l'invariant semantique d'une grande partie des derivations productives