Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
This paper describes how traditional andnon-traditional methods were used to identifyseventeen previously unknown articles that webelieve to be by Stephen Crane, published inthe New-York Tribune between 1889 and1892. The articles, printed without byline inwhat was at the time New York City's mostprestigious newspaper, report on activities ina string of summer resort towns on New Jersey'snorthern shore. Scholars had previouslyidentified fourteen shore reports as Crane's;these possible attributions more than doublethat corpus. The seventeen articles confirmhow remarkably early Stephen Crane set hisdistinctive writing style and artistic agenda. In addition, the sheer quantity of the articlesfrom the summer of 1892 reveals how vigorouslythe twenty-year-old Crane sought to establishhimself in the role of professional writer. Finally, our discovery of an article about theNew Jersey National Guard's summer encampmentreveals another way in which Crane immersedhimself in nineteenth-century military cultureand help to explain how a young man who hadnever seen a battle could write so convincinglyof war in his soon-to-come masterpiece,The Red Badge of Courage. We argue that thejoint interdisciplinary approach employed inthis paper should be the way in whichattributional research is conducted.
Se presentan los valores normativos de la adapta The norms of the Spanish adaptation for shows 9 ci6n espanola de los conjuntos 9-14 (segunda parte) del 14 (second part) of the International Affective Picture Intemational Affective Picture System (lAPS). Los resul System (lAPS) are presented. The results are highly tados muestran una alta consistencia con los obtenidos consistent with those obtained in the first part of the en la primera parte de la adaptaci6n espanola y con los Spanish adaptation and in the original USA version. The valores originales norteamericanos. La distrlbuclon de picture distribution in the bi-dimensional space, defined las diapositivas en el espacio bidimensional valencia by the ratings of valence and arousal, displays the typical arousal adopta la tipica forma de boomerang, obssrvan boomerang form. In addition, a more pronounced slope dose una menor inclinaci6n, junto con una mayor disper -and a smaller dispersion- is observed in the si6n, en el brazo que se extiende hacia el polo agrada unpleasantextreme of the boomerang than in the pleasant ble que en el brazo que se extiende hacia el polo des one. The correlations between the Noth-American and agradable. Las correlaciones entre las evaluaciones the Spanish values are all highly significant. Nevertheless, norteamericanas y espanolas son todas altamente sig the differences found in the first part of the study are nificativas. No obstante, se confirman las diferencias confirmed: Spanish people perceive the affective pictures, encontradas en la primera parte del trabajo, en el sen as a whole, as more arousing and less dominant than tido de que los aspafioles perciben las imagenes North-Americans. Similarly, our results confirm the gender afectivas, en su conjunto, con mayor nivel de activaci6n differences previously found: women rate the pictures y menor nivel de control que los norteamericanos. Asi as more arousing and less dominant than men. Although mismo, sa confirman las diferencias de genero encon no significant differences are observed in valence ratings, tradas anteriormente: las mujeres otorgan a las image as a whole, the pictures evaluated as more pleasant are nes un mayor nivel de acnvaclon y un nivel menor de clearly different for men and women. The implications of control que los varones. Aunque no exlsten diferenclas the results regarding the theoretical model underlying the significativas en las estimaciones globales de la valencia lAPS are highlighted. afectiva, las imagenes evaluadas como mas agradables por varones y mujeres son claramente diferentes. Las Key words: emotion, affective valence, arousal, implicaciones de los resultados con respecto al modelo dominance, cross-cultural differences, gender te6rico que subyace al lAPS son resaltadas. differences.
In I 997 the Department of Nordic languages of Ghent University started up a Dutch-Danish dictionary project. The aim is to create a bilingual database of about 40.000 lemmas. The commissioner of the project is the CL VV (Committee for lexicographical interlingual resources), a Dutch-Belgian official agency that coordinates bilingual dictionary projects with Dutch as one of the involved languages. A Dutch-Danish bilingual lexical database is produced by means of the editor OMBI with a reverse function as one of its key features. RBN (Referentiebestand Nederlands), a corpus based predictionary lexical database, forms our material for the source language. In this article we intend to present this project by describing our way of working with the OMBI-program (network, advantages, problems) and with the RBN-material (macro, micro). A few editorial matters will be highlighted and illustrated with examples.
ABSTRACT This paper presents a study of the attributes of margarine, showing the depth of information about consumer perceptions and drivers of liking that emerges from a detailed analysis of relations among attributes. The paper develops three sets of analyses to understand relations among attributes: principal components analysis in order to identify basic dimensions of perception, linear functions relating overall liking to attribute liking or to image ratings in order to identify drivers of liking, and quadratic functions that relate overall liking or image ratings to sensory attribute levels in order to identify optimal sensory levels and to create sensory preference segments. The analyses show how consumer data can generate learning about the consumer perceptions on the one hand, and guidance for product development.
We study the impact of richer syntactic dependencies on the performance of the structured language model (SLM) along three dimensions: parsing accuracy (LP/LR), perplexity (PPL) and word-error-rate (WER, N-best re-scoring). We show that our models achieve an improvement in LP/LR, PPL and/or WER over the reported baseline results using the SLM on the UPenn Treebank and Wall Street Journal (WSJ) corpora, respectively. Analysis of parsing performance shows correlation between the quality of the parser (as measured by precision/recall) and the language model performance (PPL and WER). A remarkable fact is that the enriched SLM outperforms the baseline 3-gram model in terms of WER by 10% when used in isolation as a second pass (N-best re-scoring) language model.
All spiritual cultures and material cultures, together with values and the way of life resulted from the two categories of cultures, are sure to play a vital role in conditioning people's statements and actions.So certain linguistic expressions and fixed forms which reflect values and life-style come to be the linguistic norm and rhetoric norm that ordinary people observe.
The endogenous opioid system is involved in stress responses, in the regulation of the experience of pain, and in the action of analgesic opiate drugs. We examined the function of the opioid system and mu-opioid receptors in the brains of healthy human subjects undergoing sustained pain. Sustained pain induced the regional release of endogenous opioids interacting with mu-opioid receptors in a number of cortical and subcortical brain regions. The activation of the mu-opioid receptor system was associated with reductions in the sensory and affective ratings of the pain experience, with distinct neuroanatomical involvements. These data demonstrate the central role of the mu-opioid receptors and their endogenous ligands in the regulation of sensory and affective components of the pain experience.
Leech (1969), considering English poetry, treats poems as linguistically deviant forms of language in which the form and content are “foregrounded” against a background of nondeviant language. The deviant language that is used may be noticeably irregular or noticeably regular. The content of the poem may also be deviant as the poet creates meanings that are not expected to be taken literally. I have applied these ideas to the British Sign Language (BSL) poem “Trio” by the late Deaf poet Dorothy Miles. Analysis of this poem shows the same features in BSL poetry that Leech found in English poetry. The poetry of Dorothy (Dot) Miles is widely considered in Britain to be some of the best BSL poetry in the public domain. Her work is powerful, thoroughly crafted, and richly significant at many levels, and it easily justifies—and rewards—careful linguistic analysis. One well-known feature of Dot’s poetry is that it frequently “works” both in English and in BSL. This feature, however, will not be the focus of this chapter. Instead, I will consider the features of BSL that create the richness of the three-part poem, “Trio,” which is made up of three stanzas: Morning, Afternoon, and Evening.1 The starting point for this analysis is the fact that poetry is a deviation from ordinary language. Poetry is not only allowed to deviate from normal patterns of ordinary language, but also it is expected to do so. Careful choice of linguistic forms allows the poet to produce language that carries significance far greater than that of ordinary language. Leech (1969) has defined foregrounding as the deviation from linguistic norms for the sake of art. The foregrounding in “Trio” occurs in two different types of deviation that involve noticeably irregular and noticeably regular use of language. In the first type, poetic deviation creates a
This study investigated the effect of a weak magnetic field (50 microT, 20 Hz sinusoidal, 5 s duration) on concurrent perceptions of visual stimuli. Subjects were seated between Helmholtz coils and gave post-exposure ratings for the affective content and arousing nature of presented images. They were blind as to the presence or absence of a simultaneously presented field. Skin conductance and arousal ratings did not show significant differences between experimental and control conditions, but the affective content rating did (P = 0.041), with the images viewed under field exposure being rated as having a more positive affect. Such measures might thus be useful as additional indicators of magnetic field detection. A post-hoc analysis of skin conductance profiles showed that 48% of subjects exhibited a lowering of skin conductance during field exposure, 34% exhibited no apparent reaction, and 17% exhibited an increase. Overall ratings given by each of the groups appeared to relate to these physiological profiles.
Imagine discourse between the arts in which the conventions of what we might call ordinary cognition do not apply, on site of intense lobbying neither tethered by history or cultural integrity, nor, frequently, concerned with social cohesion or communicative norms. It will be discourse in which the categories of an imperial culture are abrogated (however temporarily) by an indigenous one, yet it will undoubtedly also be site of intense colonization. On it, likewise, there will be an appropriation of language on an unprecedented scale. Past experience will play little part. Memory short and episodic, rather than semantic. It primal discourse. Primal in that it the site of first contact. Primal also in that it most often considered be the meeting of primitive culture and an advanced. Primal, likewise, in behavioral sense: in it, the satisfaction of physiological needs tantamount. Indeed, body and mind here are in state of kinetic unrest. This scene of prolonged immat urity, yet ontological and epistemological questions held in private language are encouraged be made public. Here the verbal arts have no canon. Literature has no prevailing cultural standard of merit. Questions of the popular and the high cultural are not naturalized and the fictional and nonfictional carry the same degree of verisimilitude as works of propaganda, rhetoric, and didacticism. In modem times, the West has become the site of this tenacious yet frequently unacknowledged imperial discourse, the discourse between multifarious forms of artistic representation win the attention of children. It in such discourse that the picture book located. 1 Arguing the need for critical language for the discussion of children's picture books, Peter Hunt suggests that to pictures into the same mould as words seems be potentially unproductive, except in terms of establishing conventions, when, of course, it is, by definition, necessary (181). It impossible, however, conventionalize pictorial representation the same degree as linguistic representation. Linguistic systems are mastered painstakingly, piece by piece, referent by referent, word by word. Pictorial systems, by contrast, are mastered all at once; they involve what Flint Schier has called natural generativity and are therefore much less conventional than linguistic systems. Each system, nevertheless, relies on general agreement and on willingness engage in communicative activity: the pictorial system on deep recognitional capacities that link object and its picture, the linguistic system on lexical and syntactical regularities and rules. In the media-saturated culture of the contemporary West, the commonalities and differences of our separate but shared experiences are frequently offered up in televisual or hypertextual format in which what Hunt describes as force set of discursive practices that address and interpellate both adults and children as potential viewers or listeners. The linguistic and the pictorial are frequently experienced as synergistic or polylogic systems bound up in this mass media, media whose intention, according Jean Baudrillard, transcribe the complexity of contemporary life into an ongoing procession of meaningless simulacra, hyperreal, a real without origin or reality (2). Baudrillard's disenchanted vision of postmodernity, articulated most profoundly in the late 1970s and 1980s, produced an interesting ontological metaphor. Disneyland, he claimed, is there conceal the fact that it the 'real' country, all of 'real' America, which Disneyland (just as prisons are there conceal the fact that it the social, in its entirety, in its banal omnipresence, which carceral). Disneyland presented as imaginary in order make us believe that the rest real, when in fact all of Los Angeles and the America surrounding it are no longer real, but of the order of the hyperreal and of simulation (25). …
Results of the noun–verb pair comprehension and production tests from the Test Battery for Auslan Morphology and Syntax (A. Schembri et al., 2000) are presented, reanalyzed, and compared to data from 2 other cases dealing with noun–verb pairs: the Auslan lexical database and a comparison of Auslan and American Sign Language (ASL) signs. The data confirm the existence of formationally related noun–verb pairs in Auslan in which the verb displays a single movement and the noun displays a repeated movement. The data also suggest that the best exemplars of noun–verb pairs of this type in Auslan form a distinct set of iconic (mimetic) signs archetypically based on inherently reversible actions (such as opening and shutting). This strong iconic link perhaps explains why the derivational process appears to be of limited productivity, though it does appear to have 'spread' to a number of signs that appear to have no such iconicity. There appears to be considerable variability in the use of the derivational markings, particularly in connected discourse, even for signs of the 'open and shut' variety. Overall, the derivational process is apparently still closely linked to an iconic base, is incipient in the grammar of Auslan, and is best described as only partially grammaticalized. (PsycINFO Database Record (c) 2016 APA, all rights reserved)
We present three applications which share some of their linguistic processor. The first application "FILES" - Fully Integrated Linguistic Environment for Syntactic and Functional Annotation - is a fully integrated linguistic environment for syntactic and functional annotation of corpora currently being used for the Italian Treebank. The second application is a shallow parser - the same used in FILES - which has been endowed with a feedback module in order to inform students about their grammatical mistakes, if any, in German. Finally an LFG-based multilingual parser simulating parsing strategies with ambiguous sentences. We shall present the three applications in that sequence.
The present study examined the relationships between quantitative volume estimates of mesial temporal lobe structures based on structural magnetic resonance imaging (MRI) and the memory and emotional functioning of individuals with temporal lobe epilepsy (TLE). Twenty individuals identified as having TLE and 24 control participants were administered a test battery that included an experimental recognition memory test incorporating both verbal and nonverbal stimuli, an experimental test of emotional functioning that measured both subjective report and skin conductance response (SCR) to emotionally salient stimuli, and a battery of standardized tests and questionnaires assessing attention, personality, and emotion perception. Patients also completed standardized measures assessing intellectual function, memory, and quality of life. The patient group demonstrated deficits on tests of memory, attention, and emotion perception. Patients also demonstrated reduced SCR, however this result was found in response to both emotional and nonemotional stimuli and so is not necessarily indicative of deficits in emotional arousal. Inconsistent with expectations, patients reported normal experiential states of arousal in response to emotionally salient stimuli. MRI data were used to measure left- and right-hemisphere volumes of the hippocampus and amygdala in the patient group, and these volumes were used as predictors of performance on behavioral measures in multiple regression analyses. Consistent with predictions, reduced amygdala volume predicted lower arousal ratings to positive emotional stimuli. However, a similar relationship was not found for arousal ratings of negative stimuli. Other predicted relationships were not demonstrated. Amygdala volume did not show a relationship with SCR, and hippocampal volume did not show a relationship with memory performance. Additional hypotheses regarding the lateralization of hippocampal and amygdala function were not supported. Results of standardized tests suggested some potential relationships between hippocampal volume and attention, and between amygdala volume and psychological characteristics, although further research would be needed to establish the degree to which these results could be generalized to a larger population. Study findings support the continued development of NM morphometric techniques to predict patterns of strengths and weaknesses demonstrated by individuals with TLE.
The situation with regard to Russian language resources is fragmented and disorganized. For this reason, it is important to promote for Russian the development of its basic resources in one package that could be used for development of speech products. The paper presents a design of the Russian lexical databases, corpora and supporting tools (system for construction and support of lexical databases, system for transcription, morphological analyzer and normalyzer) developed for wide usage in speech engineering.
This paper presents a statistical parser for a wide-coverage Combinatory Categorial Grammar (CCG) derived from the Penn Treebank. The Treebank is translated to a corpus of canonical CCG derivations. We de ne a generative statistical model over CCG derivations and train it on the transformed Treebank.
This paper describes the idea behind a multilingual database (Muldi) designed to incorporate five constituent parts: monolingual and multilingual corpora, monolingual lexicons, lists of translation equivalents, and terminological records. The emphasis in Muldi is on the presentation, analysis, and description of syntagmatic information contained in lexical items. Types of translation equivalents as well as the problem of relationship between dictionary and corpus translation equivalents is also considered.
Results of the noun-verb pair comprehension and production tests from the Test Battery for Auslan Morphology and Syntax (Schembri et al., 2000) are re-presented, re-analyzed, and compared to data from two other cases also dealing with noun-verb pairs: the Auslan lexical database and a comparison of Auslan and American Sign Language (ASL) signs. The data elicited through the test battery and presented in this article confirm the existence of formationally related noun-verb pairs in Auslan in which the verb displays a single movement and the noun displays a repeated movement. The data also suggest that the best exemplars of noun-verb pairs of this type in Auslan form a distinct set of iconic (mimetic) signs archetypically based on inherently reversible actions (such as opening and shutting). This strong iconic link perhaps explains why the derivational process appears to be of limited productivity, though it does appear to have "spread" to a number of signs that appear to have no such iconicity. There appears to be considerable variability in the use of the derivational markings, particularly in connected discourse, even for signs of the "open and shut" variety. Overall, the derivational process is apparently still closely linked to an iconic base, is incipient in the grammar of Auslan, and is thus best described as only partially grammaticalized.
Interest in large-scale, grammar-based parsing has recently seen a large increase, in response to the complexities of language-based application tasks such as speech-to-speech translation, and enabled by efforts in large-scale, collaborative grammar engineering and in the induction of statistical grammars/parsers from treebanks. Parser throughput is an important consideration in real-world applications, since throughput tends to degrade as grammar size increases. Investigation of efficient approaches to parsing is therefore an important topic of research.
This paper presents a bottom-up generator that makes use of Information Retrieval techniques to rank potential generation candidates by comparing them to a data base of stored instances. We introduce two general techniques to address the search problem, expectation-driven search and dynamic grammar rule selection, and present the architecture of an implemented generation system called IGEN. Our approach uses a domain-specific generation grammar that is automatically derived from a semantically tagged treebank. We then evaluate the efficiency of our system.
AIM: In this paper the balance of affective and instrumental communication employed by nurses during the admission interview with recently diagnosed cancer patients was investigated. RATIONALE: The balance of affective and instrumental communication employed by nurses appears to be important, especially during the admission interview with cancer patients. METHODS: For this purpose, admission interviews between 53 ward nurses and simulated cancer patients were videotaped and analysed using the Roter Interaction Analysis system, in which a distinction is made between instrumental and affective communication. RESULTS: The results reveal that more than 60% of nurses' utterances were of an instrumental nature. Affective communication occurred, but was more related to global affect ratings like giving agreements and paraphrases than to discussing and exploring actively patients feelings by showing empathy, showing concern and optimism. CONCLUSION: In future, nurses should be systematically provided with (continuing) training programmes, in which they learn how to communicate effectively in relation to patients' emotions and feelings, and how to integrate emotional care with practical and medical tasks.
The present study explored psychological predictors of response to a series of three 25 second inhalations of 20% carbon dioxide-enriched air in 60 nonclinical participants. Multiple regression analyses indicated that only anxiety sensitivity physical concerns predicted self-reported fear, whereas both physical anxiety sensitivity concerns and behavioural inhibition sensitivity independently predicted affective ratings of emotional arousal. In contrast, the psychological concerns anxiety sensitivity dimension predicted ratings of emotional displeasure (valence), and both psychological anxiety sensitivity concerns and behavioural inhibition sensitivity independently predicted emotional dyscontrol. No variables significantly predicted heart rate. These data are in accord with current models of emotional reactivity that highlight the role of cognitive variables in the production of anxious and fearful responding to somatic perturbation, and help further clarify the particular predictors of anxiety-related responding to biological challenge.
This paper reports on the design of a lexical database for English which is currently under construction ("FrameNet-2 "), and describes the kinds of linguistic facts that the database is intended to make available, for both human and computer consumers. Building on a recently completed pilot study ("FrameNet-1 "), it is centered on the nature of the relation between lexical meanings and the conceptual structures which underlie them (semantic frames). The database will show the semantic and syntactic combinatorial possibilities (based on frame membership) of the lexical items it includes, as these are documented through grammatical and semantic annotations of sentences extracted from a large corpus of contemporary written English. The notions of pro-filing within a frame, frame inheritance including multiple inheritance, frame blending, and frame composition will be explained and illustrated, and the manner of storing information about them in the database will be outlined. The building of the database, with its necessary labor-intensive manual component, will be explained. 1
The characterization of the Aragonese used in the Heredia's works has met with surprise almost always on account of its heterogeneity, accountable by the different models used and the diverse people that have played a role in his scriptorium. To this should be added the different peculiarities of the language itself used by his patron, a language that has never been taken into account up to now. In order to fill this gap this paper presents an edition and a study of a letter from Castellan de Amposta, kept in the Archivo de la Corona de Aragon. This analysis reveals the use of a variety of linguistic norms: forms habitually found in Aragonese are to be found side by side in the letter with forms usual in Catalan. There are also forms that coincide with Castilian usage.
In this contribution we discuss how a fuzzy querying interface can support the generation of linguistic database summaries - a special technique of data mining. Links between our approach to linguistic summaries and the well-known technique of association rules is shown. The implementation of linguistic summaries generation using the authors’ FQUERY for Access package is presented.
Flogging a Dead Language: Identity Politics, Sex, and the Freak Reader in Acker’s Don Quixote Nicola Pitchford Pastiche is central to the resistant politics of Kathy Acker’s writing—yet she would appear to agree with Fredric Jameson’s influential critique of pastiche as “the wearing of a linguistic mask, speech in a dead language” (17). Her 1986 novel Don Quixote is all about having to speak “in a dead language” in the absence of a more “healthy” norm. It begins with the death of the protagonist, a female version of Cervantes’s knight, who then goes on to narrate much of the subsequent story. Acker explains, “BEING DEAD, DON QUIXOTE COULD NO LONGER SPEAK. BEING BORN INTO AND PART OF A MALE WORLD, SHE HAD NO SPEECH OF HER OWN. ALL SHE COULD DO WAS READ MALE TEXTS WHICH WEREN’T HERS” (39). The novel then proceeds by plagiarism and pastiche, as Quixote goes on a quest—for a heterosexual love unsullied by patriarchal power relations—through fragments of numerous existing texts. Quixote rereads and pieces together a whole range of textual scraps, from Machiavelli’s The Prince to a Godzilla movie. What becomes clear in her eccentric survey of (primarily) Western culture is that the lost, healthy linguistic norm is more than unhealthy for female readers—indeed, it is deadly. The novel is motivated by the idea of both reading and speaking “in a dead language”—but “flogging a dead language” seems a more apt description of Acker’s strategy, in more ways than one. For both the reader in and the reader of the novel, the act of rereading that pastiche entails can seem like flogging a dead horse, in the sense of merely covering once again the familiar ground of the already said. Of course, the same has been said of any reading in postmodernity where all language may well be dead, having belonged properly to a previous historical moment that gave it life and from which it has now been dissociated by forces of commercial appropriation and cultural amnesia. But this generic deadness that Jameson identifies as inherent in postmodern writing is not quite what I wish to explore. Rather, I want to attempt to account for what I see as a particular familiarity, and perhaps a particular tendency toward exhaustion and redundancy, that accompanies reading Acker’s texts from this period in her career, a period characterized by Acker’s extensive use of pastiche or what she frequently refers to as “plagiarism.” In what follows, I look at what happens to and through the act of reading, to ask how reading is connected to agency. Despite the considerable difficulty of Acker’s experimental novels, reading them can become an activity weighed down by a certain deadening obviousness. I want to suggest that this lifelessness derives from Acker’s attempts to construct, through pastiche, a community of readers defined by their opposition to traditional literary culture. I want also to argue that her deployment of pastiche in specific contexts—especially sexual contexts—in fact complicates and undermines the static and oversimplified role that she sometimes seems to offer her reader. In such moments, a complex interplay of various possible readerly identifications creates a contingent and particular version of agency. In a chapter on Acker in his recent book on literary celebrity, Joe Moran has suggested that in both her public persona and her work, Acker “puts forward two contrasting views of identity—one textual and one essentialist” (142). While he locates these competing versions in Acker’s characters and in her own public performance of the (death of the) author, I wish to extend his observations and apply them also to the modes of reading suggested by her texts, and by this novel in particular. The “textual” version of identity, generally celebrated by those critics friendly to Acker’s work, is readily apparent in the cut-and-paste technique of Don Quixote, in which borrowed textual fragments are reanimated by their juxtaposition. Here, language (my metaphorical dead horse), along with the social identities it produces, is like Quixote’s skinny nag Rocinante, who by all rights should be dead but who keeps lurching doggedly...
BOOK NOTICES 209 Linguistic databases. Ed. by John Nerbonne. (CSLI lecture notes 77.) Stanford, CA: CSLI, 1998. Pp. xxi, 243. The papers in this collection were originally presented at the 'Linguistic Databases' conference, University of Groningen, 23-24 March, 1995. Because ofthe almostproverbial rapidity with which information technology develops, the collection as a whole is dated already, but there is still much of interest to be found. Not all papers read at the conference are in this volume, but the papers cover a wide range of subjects, mostly practical in nature, not theoretical. After a clear and readable introduction by Nerbonne, the papers are presented in no particularorder, though the editor groups the papers in five main areas: syntactic corpora and databases, phonetic databases, applications in linguistic theory, applications, and extending basic technologies. The papers themselves are not presented according to this grouping, however, and at first sight the book appears rather disorganized. The wide variety of subjects can be deduced from the titles of the papers presented: 'Test suites for natural language processing', 'From annotated corpora to databases: The SgmlQL language', 'Markup of a test suite with SGML', 'An open systems approach for an acoustic-phonetic continuous speech database: The S_tools database-management system ', "The reading database of syllable structure',? database application for the generation of phonetic atlas maps', 'Swiss French polyphone and polyvar: Telephone speech databases to model inter- and intra-speaker variability', 'Investigating argument structure: The Russian nominalization database', "The use of a psycholinguistic database in the simplification of text for aphasie readers', "The computer learner corpus: A testbed for electronic EFL tools', 'Linking WordNet to a corpus query system', 'Multilingual data processing in the CELLAR environment '. The issue whether to use open free systems or closed proprietary systems is addressed in several papers. Some papers present applications developed both in open and closed systems. This is one area where developments have been going very fast, and nowadays freely available databases are often as capable as their commercial counterparts. Some of the applications presented in this collection are available from the Internet, and url's are often given. The collection can serve as a good introduction to the field for relative outsiders as ample references and links are given. The papers themselves vary greatly in subject matter so not all will be of interest to every reader. My particular favorite was 'From annotated corpora to databases: the SgmlQL language '. [BOUDEWUN REMPT.J Understanding phonology. By Carlos Gussenhoven and Haike Jacobs. (Understanding language series.) London: Arnold, 1998. Pp. xii, 286. This textbook is intended as an introduction to phonology aimed at 'students with little or no prior knowledge of linguistics' (back cover). As in many other textbooks, it uses exercises as a learning tool. Two types ofexercises are proposed. The ones identified by a key, 'intended as an expository aid' (xi), are provided with a solution in an appendix (though it is not always so much a clear cut answer as a guide for reflection, which is, to my view, a lot better). The ones identified by a dot are intended as practice material, and no solution is offered. I thought the idea of having two types of exercises a good one since it gives the reader the opportunity both for individual work and for discussion with others. Also, whenever it may apply, an optimality theoretic analysis is offered to describe a phonological process. Ch. 1, "The production of speech', is a basic introduction to phonology, phonetics, and phonation. Ch. 2,'Some typology: Sameness and difference', cleverly covers the universal and language specific aspects of phonological structures and typology. Ch. 3,'Making the form fit', addresses phonological grammar and adaptation by presenting the nativization of loan words in both the rules and the constraints approaches. Ch. 4, 'Underlying and surface representations', Ch. 5, 'Distinctive features', and Ch. 6, 'Ordered rules', deal with the basic notions of generative phonology within the SPE type formalism and introduce the reader to the school of linear phonology. Ch. 7,? case study: The diminutive suffix in Dutch', shows how these notions are applied. In Ch. 8, 'Levels of representation', Gussenhoven and Jacobs present an intermediate level of representation between the underlying representation and...
Bagging and boosting, two effective machine learning techniques, are applied to natural language parsing. Experiments using these techniques with a trainable statistical parser are described. The best resulting system provides roughly as large of a gain in F-measure as doubling the corpus size. Error analysis of the result of the boosting technique reveals some inconsistent annotations in the Penn Treebank, suggesting a semi-automatic method for finding inconsistent treebank annotations.
This paper demonstrates that machine learning is a suitable approach for rapid parser development. From 1000 newly treebanked Korean sentences we generate a deterministic shift-reduce parser. The quality of the treebank, particularly crucial given its small size, is supported by a consistency checker. 1 Introduction Given the enormous complexity of natural language, parsing is hard enough as it is, but often unforeseen events like the crises in Bosnia or East-Timor create a sudden demand for parsers and machine translation systems for languages that have not benefited from major attention of the computational linguistics community up to that point. Good machine translation relies strongly on the context of the words to be translated, a context that often goes well beyond neighboring surface words. Often basic relationships, like that between a verb and its direct object, provide crucial support for translation. Such relationships are usually provided by parsers. The NLP resources f...
BOOK NOTICES 221 Indo-European perfects' (117-34), by Bridget Drinka. This paper investigates issues like the unidirectionality hypothesis and universal paths, with evidence drawn from a number of early Indo-European languages, e.g. Avestan, Sanskrit, and Homeric Greek. Other grammaticalization papers include "The sequencing of grammaticization effects: A twist from North America' (291-314) by Marianne Mithun and 'Grammaticalization of complex verbal constructions in Finnish' (363-76) by Taru Salminen. Several papers deal with language contact. Salikoko S. Mufwene's 'What research on creóle genesis can contribute to historical linguistics' (315-38) discusses issues like the social nature of creolization and the role of language contact in the histories of French and English. Another language contact paper is 'Yiddish and Hebrew: Borrowing through oral language contact' (135-48), by Elaine Gold, who argues that the source of the Hebrew component in Yiddish was not Hebrew texts but rather oral language contact. Syntactic change is not neglected. Ellen F. Prince's contribution, "The bonowing of meaning as a cause of internal syntactic change' (339-62), argues that 'at least some cases of (language-internal) syntactic change may result from (language-external) pragmatic and semantic borrowing' (339). Prince discusses three phenomena in support of this claim: Yiddish dos sentences, Yinglish 'Yiddish movement ', and the Yiddish pluperfect. Another paper on syntactic change is 'On the conservatism of embedded clauses' (255-68) by Kenjiro Matsuda which looks at some possible reasons for the resistance of such clauses to change (e.g. processing difficulties, pragmatic factors, and so on). All historical linguists should find something of interest in this volume, especially given the broad range of topics covered. The editors are to be commended for ajob well-done, and we can look forward to the publication of papers from the next ICHL. [Marc Pierce, University of Michigan.] Sprache und bürgerliche Nation: Beitr äge zur deutschen und europäischen Sprachgeschichte des 19. Jahrhunderts. Ed. by Dieter Cherubim, Siegfried Grosse, and Klaus J. Mattheier. Berlin & New York: Walter de Gruyter, 1998. Pp. ix, 456. This volume consists mainly of revised versions ofpapers presented at the 2nd Bad Homburger Kolloquium zur Sprachgeschichte des 19. Jahrhunderts, held in November 1993. (Three of the papers presented at the Kolloquium had already been promised to other publications; they were consequently replaced by four new papers written by conference participants ). A briefdescription ofthe contents follows. The range of topics covered is impressively broad; a number of important issues from the period 1790-1914 (as the term '19 Jahrhundert' is defined here) are discussed. Topics discussed include the status of German in various foreign countries, contact between German and other languages, the question of 'nation', and the language ofvarious social groups as well as their influence on the development of German. Papers include Klaus J. Mattheier's 'Kommunikationsgeschichte des 19. Jahrhunderts. Überlegungen zum Forschungsstand und zu Perspektiven der Forschungsentwicklung' (1-45), 'Deutsch in Belgien im neunzehnten Jahrhundert' (71-86) by Roland Willemyns, and 'Vom Dienstmädchen zur Professoringattin. Probleme bei der Aneigung bürgerlichen Sprachverhaltens und Sprachbewußtseins' (259-81) by Isa Schikorsky. Mattheier's paper is largely bibliographical in intent; it sketches some relevant issues (e.g. language contact, orthography, and the problem of a corpus) and provides a wealth of further references. Willemyns examines the status of German in Belgium—a weighty issue, given historical events. Schikorsky discusses the case of Elise Egloff, a Kindermädchen who eventually married a prominent professor of anatomy and pathology in Heidelberg, focusing on the problems that Elise faced in conforming to the linguistic norms of her future in-laws. Other contnbutions include ' "An mein Volk". Sprachliche Mittel monarchischer Appelle' (16796 ) by Hartmut Schmidt, 'Zum Einfluß der proletarischen und der bürgerlichen Frauenbewegung auf den politischen Wortschatz (um 1900)' (341-59). by Elisabeth Berner, and 'Morphologische und syntaktisch -stilistische Eigentümlichkeiten in deutschen Texten aus dem letzten Drittel des 19. Jahrhunderts' (420-43), by Siegfried Grosse. Schmidt discusses various aspects of a number of royal proclamations, e.g. Wilhelm H's 'An das deutsche Volk' issued on the entry of Germany into World War I, concentrating mainly on stylistic considerations. Berner's paper examines the usage of various lexemes, e.g. Emanzipation...
Abstract Traditional theories of finance posit that the pricing of securities in financial markets should be done according to the quality of their underlying technical fundamentals. However, research on financial markets has tended to indicate that factors other than technical fundamentals are often used by market participants to gauge the value of securities. This phenomenon may be quite prevalent in markets for initial public offerings (IPOs), where securities lack a financial history. The imagery and affect associated with securities can be a powerful basis upon which to judge their worth. Advanced business students in a securities analysis course were asked to evaluate a number of industry groups represented on the New York Stock Exchange in terms of a set of judgmental variables. After providing imagery and affective evaluations for each industry group, the participants judged the likelihood that they would invest in companies associated with each industry. Imagery and affective ratings were highly correlated with one another and with the likelihood of investing. Judgments of performance correlated poorly to moderately with actual market performance as measured by weighted average returns for the industry groups studied. The results suggest that imagery and affect are part of a coherent psychological framework for evaluating classes of securities, but that framework may have low validity for predicting performance.
The present paper proposes to measure the ‘‘functional load of pitch-accent’’ for Japanese words. It is calculated from (1) the ratio of the number of accentually contrastive homophone pairs to the number of noncontrastive pairs; and (2) the difference between word-familiarity ratings for each pair. The functional load is larger when there are more contrastive homophones to noncontrastive homophones for the target word, and when the difference of word familiarity scores is smaller between each homophone pair. A large scale word-familiarity database is used for the calculation of functional load and other statistical properties of pitch accent. Distribution of accent and opposition types at each word length are investigated with respect to familiarity scores. Results show that oppositions which include an unaccented form dominate in short simplex words, from low to high-familiarity range. This suggests that the role of pitch accent in distinguishing homophones is biased to the presence–absence contrast, and not to the location per se of pitch accent. Preliminary results from a perception experiment suggest that top-down information in the lexicon, such as functional load and distribution of accent and opposition types, interacts with the bottom-up process in lexical access.
In this paper we present a whole Natural Language Processing (NLP) system for Spanish. The core of this system is the parser, which uses the grammatical formalism Lexical-Functional Grammars (LFG). Another important component of this system is the anaphora resolution module. To solve the anaphora, this module contains a method based on linguistic information (lexical, morphological, syntactic and semantic), structural information (anaphoric accessibility space in which the anaphor obtains the antecedent) and statistical information. This method is based on constraints and preferences and solves pronouns and definite descriptions. Moreover, this system fits dialogue and non-dialogue discourse features. The anaphora resolution module uses several resources, such as a lexical database (Spanish WordNet) to provide semantic information and a POS tagger providing the part of speech for each word and its root to make this resolution process easier. Keywords: Anaphora resolution, semantics, LFG grammar and parsing, EuroWordNet. 1.
In this paper, we present an ongoing project of tagging early French dictionaries, dictionnaire Universel de Basnage de Bauval (1701) in order to build a lexical database which would enable fine-grained queries for linguistic and historical studies. The tagging model is SGML and as a tagset basis we chose to adopt the Text Encoding Initiative Guidelines for print dictionaries. We show that in spite of numerous inconsistencies, computerizing Le dictionnaire Universel is a feasible task and exhibits many regularities in the text which can be partially automated with finite state automata. Applying a systematic grammar on the text renews the metalexicographic studies of early dictionaries.
Three state-of-the-art statistical parsers are combined to produce more accurate parses, as well as new bounds on achievable Treebank parsing accuracy. Two general approaches are presented and two combination techniques are described for each approach. Both parametric and non-parametric models are explored. The resulting parsers surpass the best previously published performance results for the Penn Treebank.
In this paper we present the results of a quantitative evaluation of the discrepancies between the Italian and English lexica in terms of lexical gaps. This evaluation has been carried out in the context of MultiWordNet, an ongoing project that aims at building a multilingual lexical database. The quantitative evaluation of the English-to-Italian lexical gaps shows that the English and Italian lexica are highly comparable and gives empirical support to the MultiWordNet model. 1.
No language in the world is homogeneous, or ever will be. Whereas earlier forms of English were characterised by extreme variation on all levels and Middle English is in fact best described as a loose conglomerate of unstable varieties, we usually lack any more detailed insight into what functions this variation had for the individual speaker. The social correlates so well known from modern sociolinguistics, such as age, sex, education, religion, can normally not be applied to the existing texts, nor can even the geographical range of recorded forms be determined with any degree of certainty. Finally, if modern dialect or other non-standard features are contrasted with (as the term non-standard implies) an accepted standard form of a language, this method would necessarily fail with Middle English even if we knew more about it than we do and, in view of the state of surviving documents, ever will. It is safe to assume that for its speakers the linguistic heterogeneity of Middle English was ordered in some way, but it was so only for continually shifting speech communities, whose number and individual geographical spread we know very little about. The scene changed dramatically in the fifteenth century: the emergence of a new standard language began to re-institute a linguistic norm for written supraregional English. This development was a natural consequence of the acceptance of English in public domains, and was speeded up by the change-over to English as the Chancery language in 1430.