Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Summary A grammar-based probabilistic parser is described, and experimental results are presented for the parser as trained and tested on a 676,000-word, highly varied treebank of unrestricted English text. Probabilistic decision trees are utilized as a means of prediction, and a grammar with about 3000 semantic-and-syntactic tags, and 1100 non-terminal node labels suppl/es detailed lingzdstic information. Further such data is supplied for prediction purposes by thousands of questions about raw words, expres~ sions, and the sentence as a whole. The rich/n.formation base used for parse prediction allows the system to parse in a domain-general, totally--open-vocabulary setting, and to output highly-detailed semantic as well as syntactic information for sentences proccessed. Finally, a statistical procedure is described for converting less-detailed into more--detailed treebank, for use in increasing parser accuracy via much larger training treeb~.nlc~.
Pronunciation by analogy (PbA) is an emerging, data-driven technique with potential application in text-to-speech (TTS) systems, as well as being an influential psychological model of reading aloud. The underlying idea is that a pronunciation for an unknown word (i.e., one not in the dictionary, or lexicon, of the human or machine “reader”) is assembled by matching substrings of the input to substrings of known, lexical words, hypothesizing a partial pronunciation for each matched substring from the lexical knowledge of the “reader,” and concatenating the partial pronunciations. This paper assesses the capability of PbA to derive pronunciations for unknown words of English. As a psychological model, PbA is “under-specified,” that is, the implementor of a simulation of the process faces detailed choices which can only be resolved by trial and error. One goal for this paper is to explore the impact of certain basic implementational choices on the performance of PbA systems. The variables studied are the specific lexical database used as the basis of the analogy process, the way of ranking/scoring candidate pronunciations, and the effect of manual versus automatic alignment of letters and phonemes. When tested with short (monosyllabic) pseudowords previously used in experimental psychology studies, the lowest error rate achieved is 14.3% (for a test set of size 70). We conclude that current PbA systems are at best poor models of pseudoword pronunciation by humans. To assess their suitability for use in a TTS application, in which multisyllabic words will be encountered, the implementations have also been tested with lexical words temporarily removed from the dictionary. The best performance obtained was 93.5% phonemes correct (corresponding to 67.9% words correct) for a 16,280-word dictionary. This is vastly superior to the 25.7% words correct obtained using a set of popular letter-to-sound rules, indicating considerable scope for analogy methods to be exploited in future TTS systems.
Whether we approve or disapprove of the style or the content of what writes in A Funerall Elegye has nothing to do with the issue of its authorship. What we believe authorship to be, however, has everything to do with it. Attribution research asks for an understanding of how an author creates, what an author leaves of himself in the work, and how that differs from the linguistic system belonging to the period. The study of authoring and authorial idiolects is now interdisciplinary and empirical. It rests on the testimony of authors and, over the past half century, on repeated experiments by neuroscientists, cognitive psychologists, linguists, and other disinterested observers on the process of uttering sentences. Once we know how the mind shapes language into speech or writing, we will be in a good position to understand how Shakespeare did so. Attribution evidence takes three forms. It can be external, found on the title page or the author's preface, and here resting in the historical events described in the poem. It can be interpretive, stemming from a reading of the meaning of the text or of the author's style. Or it can be linguistic, extracting substylistic characteristics of the writing in the hope they may distinguish the writer from any other writer: that is, a fingerprint. The external evidence for Shakespeare as is good but not incontrovertible, despite best efforts by Foster, Abrams, and others in the debate. Someone with Shakespeare's initials published an elegy in 1612 with the printer who published his sonnets in 1609. The poet in his preface says that he personally knew the subject of the poem, an Oxford graduate murdered in Exeter. Because several poets with these initials are active at this time, and none has a very strong link to the deceased young man, the identification of author remains open. Even if Shakespeare's full name were on the title page, doubt would dog this attribution. Anyone can claim authorship of a work falsely; and anyone can incorrectly attribute a work and publish it. Only the author knows for sure, and Shakespeare never left a list of his works in his own hand. Printers sold books supposedly by Shakespeare that are not. Even his friends could innocently pass off, as his, many scenes by another playwright. Henry VIII belongs to both Shakespeare and Fletcher, not just to Shakespeare, as Heminges and Condell lead us to believe. To assign authorship on the basis only of testimony by others is to judge on circumstancial evidence, which routinely leaves juries deeply worried. Abrams's reading of the poem plausibly argues that several passages self-identify the poet as an actor-playwright. Other passages contain word clusters from Shakespeare's known works, one of them then unpublished. Because the only practicing playwright in 1612 known to have the initials is Shakespeare, Abrams reasonably concludes that the authorship question can be answered. On the other hand, Foster argues that Thorpe the printer misidentified Shakespeare as W. H. in the preface to his sonnets. Critics who disagree with Foster and Abrams may say that W. S. was reversed or, less plausibly, that someone misread f for long s in the manuscript. Further, verbal parallels between Shakespeare's work and that of others are well known in scholarship. They appear even in arguments that Shakespeare did not write the works known certainly as his. Source studies use such parallel passages. The main difficulty with them relates to our fragmentary knowledge of the state of common English idiom in the period. How rare is the fixed phrase or collocation (word pair in variable order) or word cluster (combination of fixed phrase and collocation)? Foster uses a large Renaissance lexical database, Shaxicon, to establish that the phrase Court opinion appears only in A Funerall Elegye and in a late redaction of Shakespeare's lost play, Cardenio. Abrams cites a choice parallel between the poem and Richard II, but he does not say that he has searched the literature of the period to see if the cluster is elsewhere. …
A lexical modeling methodology was employed to examine how the distribution of phonemic patterns in the lexicon constrains lexical equivalence under conditions of reduced phonetic distinctiveness experienced by speechreaders. The technique involved selection of a phonemically transcribed machine-readable lexical database; definition of transcription rules based on measures of phonetic similarity; application of the transcription rules to a lexical database and formation of lexical equivalence classes; and computation of 3 metrics to examine the transcribed lexicon. The metric percent words unique demonstrated that distribution of words in the language preserves lexical uniqueness across a wide range in the number of potentially available phonemic distinctions. Expected class size demonstrated that if at least 12 phonemic equivalence classes were available, any given word would be highly similar to only a few other words. Percent information extracted provided evidence that high-frequency words tend not to reside in the same lexical equivalence classes as other high-frequency words. The steepness of the functions obtained for each metric shows that small increments in the number of visually perceptible phonemic distinctions can result in substantial changes in lexical uniqueness. (PsycINFO Database Record (c) 2016 APA, all rights reserved)
Computer self-efficacy and outcome expectancy scales were developed using 306 responses to a questionnaire distributed by a national mail survey to end users of computer systems in a variety of functional business areas. Confirmatory factor analysis using a structural equations approach was used to develop three scales. The scales were found to demonstrate satisfactory psychometric properties. The reliability coefficients for these scales were as follows: .85 for computer self-efficacy; .88 for work-related outcome expectancy; and .89 for personal outcome expectancy. The scales provide a strong foundation from which to refine the measurement of computer self-efficacy and outcome expectancy. From these refinements, empirical models that include self-efficacy and outcome expectancy as determinants of information technology acceptance at the individual level of analysis can be improved.
This paper describes the multilingual text editor MtScript developed in the framework of the MULTEXT project.MtScript enables the use of many differentwriting systems in the same document (Latin, Arabic,Cyrillic, Hebrew, Chinese, Japanese, etc.). Editingfunctions enable the insertion or deletion of textzones even if they have opposite writing directions.In addition, the languages in the text can be marked,customized keyboard input rules can be associated witheach language and different character coding systems(one or two bytes) can be combined. MtScript isbased on a portable environment (Tcl/Tk). MtScript.1.1version has been developed underUnix/X-Windows (Solaris, Linux systems) and otherversions are planned to be ported to the Windows andMacintosh environments. The current 1.1 versionpresents several limits that will be fixed in futureversions, such as the justification of bi-directionaltexts, printing support, and text import/exportsupport. Future versions will use SGML and TEI norms,which offer ways of encoding multilingual texts andare to a large extent meant for interchange.
The authors examined whether perception of emotional stimuli is normal in amnesia and whether emotional arousal has the same enhancing effect on memory in amnesic patients as it has in healthy controls. Forty standardized color pictures were presented while participants rated each picture according to emotional intensity (arousal) and pleasantness (valence). An immediate free-recall test was given for the pictures, followed by a yes-no recognition test. Arousal and valence ratings were highly similar among the amnesic patients and controls. Emotional arousal (regardless of valence) enhanced both recall and recognition of the pictures, and this enhancement was proportional for amnesic patients and controls. Results suggest that emotional perception and the enhancing effect of emotional arousal on memory are intact in amnesia.
This paper presents two groups of text encodingproblems encountered by the Brown University WomenWriters Project (WWP). The WWP is creating a full-textdatabase of transcriptions of pre-1830 printed bookswritten by women in English. For encoding our texts weuse Standard Generalized Markup Language (SGML),following the Text Encoding Initiative’s Guidelines for Electronic Text Encoding andInterchange. SGML is a powerful text encoding systemfor describing complex textual features, but a fullexpression of these may require very complex encoding,and careful thought about the intended purpose of theencoded text. We present here several possibleapproaches to these encoding problems, and analyze theissues they raise.
The reliability of the Functional Assessment Measure (FIM+FAM) is an important issue with its increased use in the measurement of neurological disability and rehabilitation outcome. Although the Motor items have good reliability ratings, the Cognitive items are more difficult to complete and their reliability is not as good. This study tests the suggestion that this might be due to the Cognitive items being more abstract. A keyword from each of four Motor items was compared with a keyword from four Cognitive items. Abstractness was measured by measuring the 'imageability' of each keyword. The Motor items were found to have a significantly higher mean imageability rating than the Cognitive items. Thus, there is support for the suggestion that abstractness contributes to the poorer reliability of the Cognitive items. These results led to the proposal that the reliability of the Cognitive items might be improved by various methods of increasing the tangibility of these measures (e.g. subdivision of broad categories of disabilities, enhancing item descriptions, training raters to increase their recognition of relevant observations, and using specific assessment tasks to elicit relevant behaviours).
The statement, ’’Results of most non-traditional authorship attribution studies are not universally accepted as definitive,'' is explicated. A variety of problems in these studies are listed and discussed: studies governed by expediency; a lack of competent research; flawed statistical techniques; corrupted primary data; lack of expertise in allied fields; a dilettantish approach; inadequate treatment of errors. Various solutions are suggested: construct a correct and complete experimental design; educate the practitioners; study style in its totality; identify and educate the gatekeepers; develop a complete theoretical framework; form an association of practitioners.
Multa renascentur quae iam cecidere cadentque quae nunc sunt in honore vocabula si volet usus quem penes arbitrium est et vis et norma loquendi. (Horace, Ars Poetica) Many a word yfalne shall eft arise And such as now bene held in hiest prise Will fall as fast, when vse and custome will Onely vmpiers of speach, for force and skill. (Puttenham's translation, 1589) 1. Introduction All through the history of dictionary-making and reflections on standard languages, usage has been one of the lexicographers' major stumbling-blocks. (1) Although it can be argued that the question of a linguistic norm is even more relevant in the fields of phonology and syntax, it is eminently important for vocabulary, too. To indicate the currency, degrees of acceptability and the geographical, social and stylistic restrictions of individual words is a problem common to both more prescriptive and more descriptive dictionaries. Johnson's aim to establish a standard for the English lexicon left but a few words unaffected in his great dictionary of 1755 and these had usage labels attached to such entries to warn readers against using them; modem dictionary editors sometimes use panels of experts to decide on the acceptability of words and their individual meanings. Whatever method for establishing usage is preferred, whether that of the editor's Sprachgefuhl or the combined judgment of experts (a method which appears to permit neatly graded degrees of acceptability), any indication of usage carries with it idiosyncratic elements as well as the restriction to a particular period, region and class-related opinion. It would be silly to complain about this, since it is the very essence of usage, that the labels used to indicate it have a certain degree of vagueness about them. The narrower the description of correct usage is, the more certain it is that the evidence will be homogeneous, and the more likely that the guidance will be taken as prescriptive. In a dictionary which has 'usage' as the first word of its title, readers will rightly expect this type of information to be central. The UDASEL project was in fact started with the explicit intention of concentrating on the fully accepted loanwords from English in 16 European languages, (2) rather than focussing on more narrowly etymological questions of word origins. As will be obvious (or become evident from my discussion), the difficulty of describing the usage of specific items in an individual language multiply when 16 sets of data are compared (which makes possible up to 240 binary comparisons). The basic decision as far as the currency of loanwords from English is concerned was not to rely on text corpora, for various reasons: a) such collections from recent texts were not available for most of the languages concerned; b) where such corpora existed, their representativeness and comparability was open to doubt; c) the frequencies elicited from them would not really be indicators of acceptability: for this the context and context would have to be investigated for each individual item, and judgments would therefore in any case be based, at least partially, on the interpreter's Sprachgefuhl. All this made it necessary, if the evidence was to be covered in a 'snapshot' manner for the early 1990s, to rely on the competence of educated users of the individual language, and leave judgments on the acceptability of English words to the specialist. This hard-won decision is supported by the fact that pilot studies with German informants (Gorlach 1994) have shown that speakers' views on degrees of acceptability conform to a greater extent than might have been expected, or feared. (3) 2. Degrees of acceptability After the pilot study was completed, the grading decided on was based on a combination of degrees of integration and of acceptability/currency: these two aspects are not always compatible, but coincide in most cases, and they were combined in order to prevent the system from becoming too complex for a general reader to take in and remember. …
Naanaam's Poems Represent One Of The Most Subversive, Eclectic, and marginal styles of contemporary Persian poetry written by exiled Iranians. His knowledge of contemporary world poetry as well as classical Persian poetry provides him with solid grounds for a peculiar style that is at once thoughtful and playful, serious and sarcastic, engaged and apathetic, violent and kind. He jolts his readers out of routine readings in a variety of ways: he uses English words in the middle of Persian writings; he parodies the well-known lines by the dominant figures of contemporary Persian poetry; he splendidly highlights the ironies and contradictions of everyday life experiences that we normally tend to take for granted; he antagonizes and subverts all that is dominant, be it poetic style, linguistic norm, Logic, or Truth.
Large tree databases as knowledge repositories become more and more important; a prominent example are the treebanks in computational linguistics: text corpora consisting of up to five million words tagged with syntactic information. Consequently, these large amounts of structured data pose the problem of fast tree retrieval: Given a database T of labeled multiway trees and a query tree q, find efficiently all trees t ∈ T that contain q as subtree. This paper presents a generalization of the classical n-gram indexing technique for supporting fast retrieval of multiway tree structures: Treegram indexing covers database trees with subtrees of fixed height; each entry of the resulting index represents such a subtree together with the database trees that contain this subtree. The evaluation of a given query q preselects those database trees that contain all of q ’s cover trees and, in turn, tests these candidates rigorously for containment of q. As an application of treegram indexing, we describe the VENONA retrieval system, which handles the BH t treebank containing 508,650 phrase structure trees found in the morphosyntactical analysis of The Old Testament with altogether 3.3 million wordforms—results of a computational-linguistics project at the Ludwig-Maximilian’s University of Munich.
In a study that used Hermans's (1987, 1988) valuation procedure, 40 participants each provided a highly valued experience of four types: science, religion, interpersonal and intrapersonal conflict between science and religion. They then rated these valuations on 30 affect terms, some of rich were later organized into categories: Positive, Negative, Self, and Other. Participants also filled out questionnaires, that were used to categorize them as low or high in scientific and religious orientation. Typical valuations of participants in these four science and religion categories are presented as qualitative idiographic information. In addition, quantitative analyses of affect ratings are presented as nomothetic information. Generally, affect ratings of scienfific experiences were more Self-directed while religious experience valuations involved equally high levels of Self and Other affect. Both scientific and religious experiences were evaluated as having Positive but not Negative affect. Interpersonal and intropersonal conflict were experienced as more Self-directed than Other-oriented. While interpersonal conflicts displayed more Negative affect than Positive, intrapersonal conflict was evaluated equally on these two measures.
We present and experimentally evaluate a new model of prounciation by analogy: the paradigmatic cascades model. Given a pronunciation lexicon, this algorithm first extracts the most productive paradigmatic mappings in the graphemic domain, and pairs them statistically with their correlate(s) in the phonemic domain. These mappings are used to search and retrieve in the lexical database the most promising analog of unseen words. We finally apply to the analogs pronunciation the correlated series of mappings in the phonemic domain to get the desired pronunciation.
Automatic text categorization is a complex and useful task for many natural language processing applications. Recent approaches to text categorization focus more on algorithms than on resources involved in this operation. In contrast to this trend, we present an approach based on the integration of widely available resources as lexical databases and training collections to overcome current limitations of the task. Our approach makes use of WordNet synonymy information to increase evidence for bad trained categories. When testing a direct categorization, a WordNet based one, a training algorithm, and our integrated approach, the latter exhibits a better perfomance than any of the others. Incidentally, WordNet based approach perfomance is comparable with the training approach one.
The Berber lexicon today is affected by two major factors. First, there is the historically strong influence of Arabic; and second, there is irregular individual neologism, in the sense that there exist many newly coined words or phrases that usually disobey the Berber linguistic norms and have not yet received general public acceptance. The desire to purify the language by eliminating the foreign loans and supplanting them by newly coined erratic terms may do more harm than good to the Berber lexicon. The main cause of this wide-ranging borrowing is Moroccan Arabic/ Berber bilingualism, which characterizes the speech of Berberophones. Moroccan Arabic is spoken or at least understood by most Berber Speakers across Central Morocco. Moroccan Arabic, however, is used by Berberophones out of necessity, especially in formal settings or when the Speakers are in contact with local authorities of public administration personnel This sociolinguistic Situation is due to the large linguistic and cultural impact of urban centres over rural areas, and to the fact that literacy is achieved through Arabic in schools. This kind of bilingualism has engendered a regression of Berber as a language of communication, but it does not imply that Berber isfaced with imminent death.
This study tested predictions derived from D.C. Glass's (1977) uncontrollability model regarding the link between control-related personality attributes and the dissociation of affective and autonomic responses to stress. Pressured drive, measured by the Jenkins Activity Survey (D. S. Krantz, D. C. Glass, & M. L. Snyder, 1974), and emotional defensiveness, measured by the Marlowe-Crowne Social Desirability Scale (D. P. Crowne & D. Marlowe, 1964), were examined in relation to cardiovascular and affective responses to mental arithmetic in 31 male and 26 female college students. Pressured drive was positively associated with cardiovascular reactivity but unrelated to affect ratings. In contrast, emotional defensiveness was unrelated to cardiovascular reactivity, but high scores were associated with smaller increases in self-reported negative affect. The findings suggest that these potentially health-damaging personality attributes may influence stress response measures through independent mechanisms for maintaining environmental control and self-control.
Automatic Text Categorization (TC) is a complex and useful task for many natural language applications, and is usually performed through the use of a set of manually classified documents, a training collection. We suggest the utilization of additional resources like lexical databases to increase the amount of information that TC systems make use of, and thus, to improve their performance. Our approach integrates WordNet information with two training approaches through the Vector Space Model. The training approaches we test are the Rocchio (relevance feedback) and the Widrow-Hoff (machine learning) algorithms. Results obtained from evaluation show that the integration of WordNet clearly outperforms training approaches, and that an integrated technique can effectively address the classification of low frequency categories.
Psychophysiological changes during long-distance driving may be associated with driving fatigue and morbidity. Measures of stress and arousal, including heart rate, blood pressure, catecholamines, cortisol, state anxiety, and self-ratings of stress and arousal were collected from 10 long-distance bus drivers during 12-hour driving shifts and at matched times on nondriving rest days. Cardiovascular and catecholamine data were elevated across the entire work day, compared with rest days. Self-reported stress and state anxiety were elevated only at the preshift measure, and these elevations were interpreted as the result of anticipatory anxiety and additional work demands at the beginning of the shift. Decelerating activation from the 9th to the 12th hours of driving were reflected in slower heart rate and lower subjective arousal ratings. Suggested explanations for these findings are that drivers experience a release of tension when they anticipate the end of the shift and therefore deactivation is a signal or precursor to the onset of fatigue in physiological adjustment mechanisms.
This paper addresses issues in automated treebank construction. We show how standard part-of-speech tagging techniques extend to the more general problem of structural annotation, especially for determining grammatical functions and syntactic categories. Annotation is viewed as an interactive process where manual and automatic processing alternate. Efficiency and accuracy results are presented. We also discuss further automation steps.
We present and experimentally evaluate a new model of prounciation by analogy: the paradigmatic cascades model. Given a pronunciation lexicon, this algorithm first extracts the most productive paradigmatic mappings in the graphemic domain, and pairs them statistically with their correlate(s) in the phonemic domain. These mappings are used to search and retrieve in the lexical database the most promising analog of unseen words. We finally apply to the analogs pronunciation the correlated series of mappings in the phonemic domain to get the desired pronunciation.
Introduction to the special issue on computational linguistics using large corpora, Kenneth W. Church and Robert L. Mercer generalized probabilistic LR parsing of natural language (corpora) with unification-based grammars, Ted Briscoe and John Carroll accurate methods for the statistics of surprise and coincidence, Ted Dunning a program for aligning sentences in bilingual corpora, William A. Gale and Kenneth W. Church structural ambiguity and lexical relations, Donald Hindle and Mats Rooth text-translation alignment, Martin Kay and Martin Roescheisen retrieving collocations from text - Xtract, Frank Smadja using register-diversified corpora for general language studies, Douglas Biber from grammar to lexicon - unsupervised learning of lexical syntax, Michael R. Brent the mathematics of statistical machine translation - parameter estimation, Peter F. Brown et al building a large annotated corpus of English - the Penn treebank, Mitchell P. Marcus et al lexical semantic techniques for corpus analysis, James Pustejovsky et al coping with ambiguity and unknown words through probabilistic models, Ralph Weischedel et al.
This technical report is an appendix to Eisner (1996): it gives superior experimental results that were reported only in the talk version of that paper. Eisner (1996) trained three probability models on a small set of about 4,000 conjunction-free, dependency-grammar parses derived from the Wall Street Journal section of the Penn Treebank, and then evaluated the models on a held-out test set, using a novel O(n^3) parsing algorithm. The present paper describes some details of the experiments and repeats them with a larger training set of 25,000 sentences. As reported at the talk, the more extensive training yields greatly improved performance. Nearly half the sentences are parsed with no misattachments; two-thirds are parsed with at most one misattachment. Of the models described in the original written paper, the best score is still obtained with the generative (top-down) "model C." However, slightly better models are also explored, in particular, two variants on the comprehension (bottom-up) "model B." The better of these has an attachment accuracy of 90%, and (unlike model C) tags words more accurately than the comparable trigram tagger. Differences are statistically significant. If tags are roughly known in advance, search error is all but eliminated and the new model attains an attachment accuracy of 93%. We find that the parser of Collins (1996), when combined with a highly-trained tagger, also achieves 93% when trained and tested on the same sentences. Similarities and differences are discussed.
Mexican Spanish is often described as a conservative variety whose distinguishing features go back only to the 19th century. This dissertation proposes that the origins of Mexican Spanish, at least with regard to verbal paradigm reduction, may be traced to the first century of the colonial period. My analysis is based on the collection of documents compiled in Documentos linguisticos de la Nueva Espana: Altiplano Central by Concepcion Company. The documents in this book are non-literary sources written during the colonial period. The following subjects were studied in this dissertation: language policy and Castilianization during the colonial period in Mexico, forms of address and the elimination of the second person plural, the -ra form and its function as an imperfect subjunctive and, finally, a historical analysis of the temporal values of the present perfect. It was found that the Mexican Spanish verbal paradigm, since the first century of the colonial period, has been characterized by the reduction of plural forms and a tendency to prefer the ending -ra form for past subjunctive, and a temporal value of an open past for the present perfect. However, there were two specific periods of time when this tendency was broken, at the beginning of the 17th and of the 19th century. At the same time, these two periods were characterized by a reassessment of the European values and imposition of its linguistic norms. This study confirms that, during the colonial period Mexican Spanish was caught between two strong opposing tendencies. One of these tendencies is the process of linguistic simplification, common in colonial territories; and the other is the process of monocentric standardization, by means of which the linguistic norm of Spain exerted pressure in the colonies.
Abstract Previous studies examining the association between social comparison processes and body image dissatisfaction have yielded inconsistent findings. This study examined whether such discrepancies are due to either the use of identical comparison targets for all subjects or variability in body mass. Specifically, 216 subjects were randomly assigned to one of three experimental conditions: self-generated upward comparison group, self-generated downward comparison group, or control group. Dependent variables were measures of body image. Results indicated that increasing body mass and trait comparison tendencies were associated with increased body dissatisfaction. However, the experimental manipulation did not affect body image ratings. Results suggest that social comparison processes may operate similarly over a range of body mass index (BMI) values.
This study was designed to compare the effects of different kinds of visual presentations and music alone on university nonmusic students' affective and cognitive responses to music. Four groups of participants were presented with excerpts from the first and fourth movements of Beethoven's Symphony No. 6 in F major (“Pastoral”). Two groups heard music excerpts only, one interpretation conducted by Stowkowski and one by Bernstein. One of the video groups viewed corresponding excerpts from the movie Fantasia while listening to the Stowkowski recording. A second group viewed and listened to a performance video of the Vienna Philharmonic filmed during the Bernstein recording. All students (N = 128) completed cognitive listening tests based on the excerpts, rated the music on Likert-type affective scales, and responded to two open-ended questions. Significant effects of presentation condition were found. Cognitive scores were higher for the performance video than the music plus animation video on both movements. Scores for the two music-only presentations were not significantly different from each other or the two video presentations. Although affective ratings were not significantly different in magnitude between the presentation groups, the animation video (Fantasia.) presentation ranked consistently higher in affect than the other presentations. Implications of these results regarding the effects of different types of visual information presented to music listeners are discussed.
We present preliminary results concerning robust techniques for resolving bridging definite descriptions. We report our analysis of a collection of 20 Wall Street Journal articles from the Penn Treebank Corpus and our experiments with WordNet to identify relations between bridging descriptions and their antecedents.
We show how a treebank can be used to cluster words on the basis of their syntactic behavior. The resulting clusters represent distinct types of behavior with much more precision than parts of speech. As an example we show how prepositions can be automatically subdivided by their syntactic behavior and discuss the appropriateness of such a subdivision. Applications of this work are also discussed. 1 Introduction The construction of classes of words, or calculation of distances between words, has frequently drawn the interest of researchers in natural language processing. Many of these studies aimed at finding classes based on co-occurrences, often combined with the aim of establishing semantic similarity between words (McMahon and Smith, 1996; Brown et al., 1992; Dagan, Markus, and Markovitch, 1993; Dagan, Pereira, and Lee, 1994; Pereira and Tishby, 1992; Grefenstette, 1992). We suggest a method for clustering words purely on the basis of syntactic behavior. We show how the necessary...
A large collection of texts may be reached through the Internet and this provides a powerful platform from which common-sense knowledge may be gathered. This paper presents a system that contains a core knowledge base structured around WordNet, a lexical database, capable of extracting contextual information from a given input text. Such context information is then used to retrieve other texts from the Internet that relate to that context. When processed by the system, these new texts bring more information that represents an enhanced domain context for the initial text. This is an incremental method for text processing that acquires domain knowledge from other texts. The paper describes the system architecture, its core knowledge base and inference engine, and the acquisition of new knowledge from corpora.
Queer theory offers insights for political economy on how humans induce categories and conflate traits in ways psychologists call "illusory correlations." A Bayesian simulation is constructed of people interacting and using probits to compare their rankings of alternatives to estimate the subjective probability i.) that others have the same tastes as them, and ii.), for each alternative, that this alternative is their best choice. These simulations are found to, at least simplistically, resemble a type of illusory correlation which gained increased prominence in the United States from 1930 to 1960, and earlier in the United Kingdom when queer panics conflated "predatory" and "traitorous" with lesbian/gay. This modeling of the social articulation of preferences leads to conjectures on the role of Michel Foucault's épistémè (1972), Barbara Ponse's principle of consistency (1978), Jeffrey Escoffier's master code (1985), Sandra Bem's schema (1981), John R. Searle's Background (1990, 1992, 1995), and Judith Butler's linguistic norms (1993). Here these are called cognitive dispositions or codes, and are seen as social structures which grow as individuals try to form homopreference networks to process information in parallel, collectively. The concept of identity or ideology entrepreneurs is used to establish the importance of institutional analysis for political economy. Such an incorporation of desire on a par with logic-"rationality"-is called post/modern and is used to overcome the silence in both neoMarxian and neoclassical political economy on queer theory and queer issues.
In two studies, pedestrians in Old and New Delhi (India) and Dhaka (Bangladesh) were asked about their reactions to three stressors common to rapidly growing urban areas in South Asia: noise, air pollution, and crowding. Results from the first study, a survey of men in Old Delhi, indicated that respondents who were more upset by noise and by crowding also reported more physical symptoms and less perceived control. In the second study, male and female pedestrians were interviewed in New Delhi and Dhaka. Results revealed consistent gender, country, and gender by country effects on measures of general affect, ratings of stressors, and coping responses. In addition, results from an experimental manipulation in Study 2 indicated that in both countries, telling pedestrians about the effects of air pollution or crowding made them feel significantly worse than they would have felt had they not been given any information.
All natural language processing systems (such as parsers, generators, taggers) need to have access to a lexicon about the words in the language. This thesis presents a lexicon architecture for natural language processing in Turkish. Given a query form consisting of a surface form and other features acting as restrictions, the lexicon produces feature structures containing morphosyntactic, syntactic, and semantic information for all possible interpretations of the surface form satisfying those restrictions. The lexicon is based on contemporary approaches like feature-based representation, inheritance, and unification. It makes use of two information sources: a morphological processor and a lexical database containing all the open and closed-class words of Turkish. The system has been implemented in SICStus Prolog as a standalone module for use in natural language processing applications.
Abstract To compare the social support behaviors of violent and nonviolent husbands, we recruited four groups of couples‐violent and distressed (VD); violent/nondistressed (VND); nonviolent/distressed (NVD);and nonviolent/nondistressed (NVND). Two systems were used to code couples’discussions of wives’personal problems. Using the Social Support Interaction Coding System (Bradbury & Pasch, 1994), no violent‐nonviolent group differences emerged; however, as listeners, NVND husbands were the most positive and tended to be the least negative. Using a coding system designed for this study (i.e., Social Support Behavior/Affect Rating System), we confirmed the hypothesis that violent husbands would offer less social support than would nonviolent husbands. Relative to nonviolent men, violent husbands were less positive, more belligerent/domineering, more contemptuous/disgusted, and more upset by the wife's problem. Relative to NVND husbands, violent husbands displayed more anger and tension, VND husbands were more critical of their wives’problem, and VD men were more critical of the possible solutions wives offered. We discuss differences in the two coding systems relevant to the detection of violent‐nonviolent group differences. Across both systems, few group differences in wife behavior emerged, suggesting that husband behavior better differentiates violent from nonviolent couples when wives are discussing personal problems.
The increasing availability of corpora annotated for linguistic structure prompts the question: if we have the same texts, annotated for phrase structure under two different schemes, to what extent do the annotations agree on structuring within the text? We suggest the term tree alignment to indicate the situation where two markup schemes choose to bracket off the same text elements. We propose a general method for determining agreement between two analyses. We then describe an efficient implementation, which is also modular in that the core of the implementation can be reused regardless of the format of markup used in the corpora. The output of the implementation on the Susanne and Penn treebank corpora is discussed.
A treebank is a corpus of tagged and bracketed sentences capturing the linguistic properties of a (sub)language in an empirical way. The CASSANDRA treebank is developed as a sideline to the GALEN-IN-USE project in which it serves to make the relationships between natural language phenomena and semantic representations of medical expressions explicit, and to assist in the quality assurance of the modelling centres. The end result is a multilingual linguistic knowledge repository from which lexicons and grammars of various types can be derived in an automatic way.
This study examined signs of mania on the Rorschach, specifically whether manic inpatients (n = 24) produce different thematic content and thought disorder than comparison groups of paranoid schizophrenic (n = 27) and schizoaffective (n = 25) inpatients. Rorschach protocols were scored by a trained rater for the Thought Disorder Index and the Schizoid-Affective Rating Scale. Results indicated that all 3 groups had moderate levels of thought disorder, but the manic inpatients produced significantly more combinatory thinking and affective content responses than the other 2 groups. The paranoid schizophrenic and schizoaffective patients did not produce significantly more schizoid content and were not different on any other types of thought disorder than the manic patients. These findings are discussed in terms of the contribution of thought disorder and affective thematic content in making the diagnosis of mania on the Rorschach.
This paper describes the novel ways in which the Orlando Project, based at the Universities of Alberta and Guelph, is using SGML to create an integrated electronic history of British women's writing in English. Unlike most other SGML-based humanities computing projects which are tagging existing texts, we are researching and writing new material, including biographies, items of historical significance, and many kinds of literary and historical interpretation, all of which incorporates sophisticated SGML encoding for content as well as structure. We have created three DTDs, for biographies, for writing-related activities and publications, and for social, political and other events. A major factor influencing the design of the DTDs was the requirement to be able to merge and restructure the entire text base in many ways in order to retrieve and index it and to reflect multiple views and interpretations. In addition a stable and well-documented system for tagging was deemed essential for a team which involves almost twenty people, including eight graduate students, in two locations.
A variety of approaches to annotating reference in corpora have been adopted. This paper reviews four approaches to the annotation of reference in corpora. Following this we present a variety of results from one annotated corpus, the UCREL anaphoric treebank, relevant to automated reference resolution.
Research based on treebanks is ongoing for many natural language applications. However, the work involved in building a large-scale treebank is laborious and time-consuming. Thus, speeding up the process of building a treebank has become an important task. This paper proposes two versions of probabilistic chunkers to aid the development of a bracketed corpus. The basic version partitions part-of-speech sequences into chunk sequences, which form a partially bracketed corpus. Applying the chunking action recursively, the recursive version generates a fully bracketed corpus. Rather than using a treebank as a training corpus, a corpus, which is tagged with part-of-speech information only, is used. The experimental results show that the probabilistic chunker has a correct rate of more than 94% in producing a partially bracketed corpus and also gives very encouraging results in generating a fully bracketed corpus. These two versions of chunkers are simple but effective and can also be applied to many natural language applications.
Two experiments examined the feasibility of psychological assessment using interactive voice response (IVR) technology and the potential sensitivity of such assessments to alcohol and fatigue effects. In Experiment 1, 10 subjects performed a 12-min battery of six IVR-administered tasks, Monday through Friday, over 2 weeks. Minimal learning effects were evident during training. Repeated administrations indicated high test-retest reliabilities. In Experiment 2 (double-blind, alcohol/placebo crossover design), 7 subjects were tested every 2 h over a 24-h period during two experimental sessions (peak blood alcohol concentrations =80 mg/dL). Several IVR-administered tasks were sensitive to alcohol impairment, but not as sensitive as laboratory-based measures specifically designed to assess alcohol impairment. Little evidence for fatigue-related impairment was obtained. The results support optimism for the potential to assess psychomotor and cognitive functioning distally via telephony; however, further refinement and validation of the methods are needed.
Video cameras provide a simple, noninvasive method for monitoring a subject’s eye movements. An important concept is that of the resolution of the system, which is the smallest eye movement that can be reliably detected. While hardware systems are available that estimate direction of gaze in real time from a video image of the pupil, such systems must limit image processing to attain real-time performance and are limited to a resolution of about 10 arc minutes. Two ways to improve resolution are discussed. The first is to improve the image processing algorithms that are used to derive an estimate. Offline analysis of the data can improve resolution by at least one order of magnitude for images of the pupil. A second avenue by which to improve resolution is to increase the optical gain of the imaging setup (i.e., the amount of image motion produced by a given eye rotation). Ophthalmoscopic imaging of retinal blood vessels provides increased optical gain and improved immunity to small head movements but requires a highly sensitive camera. The large number of images involved in a typical experiment imposes great demands on the storage, handling, and processing of data. A major bottleneck had been the real-time digitization and storage of large amounts of video imagery, but recent developments in video compression hardware have made this problem tractable at a reasonable cost. Images of both the retina and the pupil can be analyzed successfully using a basic toolbox of image-processing routines (filtering, correlation, thresholding, etc.), which are, for the most part, well suited to implementation on vectorizing supercomputers.
This paper discusses the encoding of proper names using the TEI Guidelines, describing the practice of the Women Writers Project at Brown University, and the CELT Project at University College, Cork. We argue that such encoding may be necessary to enable historical and literary research, and that the specific approach taken will depend on the needs of the project and the audience to be served. Because the TEI Guidelines provide a fairly flexible system for the encoding of proper names, we conclude that projects may need to collaborate to determine more specific constraints, to ensure consistency of approach and compatibility of data.
Contents: B.B. Kachru, Ladislav Zgusta: Illinois Years. - Bibliography of Publications by Ladislav Zgusta. - B.B. Kachru, Introduction. - I: Contextualizing Culture: D. Bartholomew, Otomi Culture from Dictionary Illustrative Sentences. - P. Corbin, Un Film, Deux Linguistes et Quelques Dictionnaires. Un Regard Particulier sur Simple Mortel de Pierre Jolivet. - J.A. Fishman, Dictionaries as Culturally Constructed and as Culture-Constructing: Reciprocity View as Seen from Yiddish Sources. - F.J. Hausmann, Allusions Litteraires et Citations Historiques Dans le Tresor de la Langue Francaise. - L.F. Lara, Towards a Theory of the Cultural Dictionary. - W.P. Lehmann, Spindle or the Distaff. - Y. Malkiel, Principal Categories of Learned Words. - E.A. Nida, Lexical Cosmetics. - II: Lexicography in Historical Context: F. Karttunen, Roots of Sixteenth-Century Mesoamerican Lexicography. - T.B.I. Creamer, Current State of Chinese Lexicography. - D.A. Kibbee, 'New Historiography', History of French and 'Le Bon Usage' in Nicot's Dictionary (1606). - N. Dinh-Hoa, On Chi-nam Ngoc-am Giai-nghia: An Early Chinese-Vietnamese Dictionary. - G. Stein, Chaucer and Lydgate in Palsgrave's Lesclarcissement. - III: Ideology, Norms and Language Use: M.A. Ezquerra,Political Considerations on Spanish Dictionaries. - D.M.T.Cr. Farina, Marrism and Soviet Lexicography. - C. Marello, Florence like Athens and Italian like Greek: An Ideologically Biased Theme in the Forewords of Some Italian Thesauri of the 19th Century. - A. Wierzbicka, Dictionaries and Ideologies: Three Examples from Eastern Europe. - R.D. Zorc, Philippine Regionalism versus Nationalism and the Lexicographer. - IV: Pluricentricity and Ethnocentricism: J. Algeo, British and American Biases in English Dictionaries. - C.W. Kim, One Language, Two Ideologies, and Two Dictionaries: Case of Korean. - B. Lewandowska-Tomaszczyk, Worldview and Verbal Senses. - C. Poirier, De la Soumission a la Prise de Parole: Le Cheminement de la Lexicographie au Quebec. - J. Whitcut, Taking it for Granted: Some Cultural Preconceptions in English Dictionaries. - V: Dictionaries across Languages and Cultures: Y. Kachru, Lexical Exponents of Cultural Contact: Speech Act Verbs in Hindi-English Dictionaries. - R.J. Steiner, Bilingual Dictionary in Cross-Cultural Contexts. - VI: Language Dynamics vs. Prescriptivism: A.P. Cowie, Learner's Dictionary in a Changing Cultural Perspective. - R.H. Gouws, Dictionaries and the Dynamics of Language Change. - F.E. Knowles, Dictionaries for the People or for People? - VII: Language Learner as the Consumer: G.M. Dalgish, Learners' Dictionaries: Keeping the Learner in Mind. - VIII: Structuring Semantics: F.F.M. Dolezal, Dictionary as Philosophy: Reconstructing the Meaning of Our Father. - M. Ritchie Key, Meaning as Derived from Word Formation in South American Indian Languages. - J.P. Louw, How Many Meanings to a Word? - IX: Ethical Issues and Lexicologists' Biases: D.L. Gold, When Religion Intrudes into Etymology (On The Word: Dictionary that Reveals the Hebrew Source of English). - T. McArthur, Culture-Bound and Trapped by Technology: Centuries of Bias in the Making of Wordbooks. - X: Terminology across Cultures: Z. Polacek, Amharic Lexicography and the Dynamics of Sociopolitical Terminology. - G.O. Richter, Grammatical Indications in Chinese Monolingual Dictionaries. - XI: Afterword: B.B. Kachru, Directions and Challenges.
This paper reports a detailed evaluation of the effectiveness of a system that has been developed for the identification and retrieval of morphological variants in searches of Latin text databases. A user of the retrieval system enters the principal parts of the search term (two parts for a noun or adjective, three parts for a deponent verb, and four parts for other verbs), this enabling the identification of the type of word that is to be processed and of the rules that are to be followed in determining the morphological variants that should be retrieved. Two different search algorithms are described. The algorithms are applied to the Latin portion of the Hartlib Papers Collection and to a range of classical, vulgar and medieval Latin texts drawn from the Patrologia Latina and from the PHI Disk 5.3 datasets. The effectiveness of these searches demonstrates the effectiveness of our procedures in providing access to the full range of classical and post-classical Latin text databases.
Mathematical and computer simulation modeling are often computationally demanding procedures, so much so that even certain parts of these procedures, such as parameter estimation, exceed the capacities and speed of the best modern computer facilities. A good deal of effort has therefore been dedicated to speeding up and making more efficient programs such as those that are meant to find a global minimum of a parameter space. Our experience, however, is that such well-explored technical procedures in fact represent some of the shortest components of the total set of procedures by which models are developed. In this article we discuss what elements take up the lion’s share of development time and speculate on what lessons can be drawn concerning the role of high-performance computing in such enterprises.