Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
Computer self-efficacy and outcome expectancy scales were developed using 306 responses to a questionnaire distributed by a national mail survey to end users of computer systems in a variety of functional business areas. Confirmatory factor analysis using a structural equations approach was used to develop three scales. The scales were found to demonstrate satisfactory psychometric properties. The reliability coefficients for these scales were as follows: .85 for computer self-efficacy; .88 for work-related outcome expectancy; and .89 for personal outcome expectancy. The scales provide a strong foundation from which to refine the measurement of computer self-efficacy and outcome expectancy. From these refinements, empirical models that include self-efficacy and outcome expectancy as determinants of information technology acceptance at the individual level of analysis can be improved.
We examined Milgram’s (1977) lost-letter technique using e-mail. In the first experiment, 79 college faculty received mock lost e-mail messages. Nineteen percent of those who received the messages responded, in all cases by returning the message to the “sender” instead of forwarding it to the “recipient.” In the second study, attitudes toward presidential candidate Ross Perot were examined by sending out two different messages to 200 randomly selected e-mail addresses in the United States. Although there was no differential response rate, examination of content revealed attitudes consistent with concurrent poll data.
This paper presents a method for developing limited-context grammar rules in order to mark up text automatically, by attaching specific text segments to a small number of well-defined and application-determined semantic categories. The Text Analysis Tool with Object Encoding (TATOE) was used in order to support the iterative process of developing a set of rules as well as for constructing and managing the lexical resources. The work reported here is part of a real-world application scenario: the automatic semantic mark up of German news messages, as provided by a German press agency, according to the SGML-based standard News Industry Text Format (NITF) to facilitate their further exchange. The implemented export mechanism of the semantic mark up into NITF is also described in the paper.
Some comments are offered on the papers given at the conference, which are divided into three groups: visual perception, models and neural networks, and data analysis. The analysis stresses the complexity faced by scientific theories in each of these three areas, and consequently why the demand for computing capacity will continue to increase, with no practical bound in sight.
No AccessAmerican Journal of Speech-Language PathologySecond Opinion1 May 1997Understanding Language DelayA Response to van Kleeck, Gillam, and Davis Rhea Paul Rhea Paul Contact author: R. Paul, PhD, Department of Speech, Portland State University, P.O. Box 751, Portland, OR 97207 Email: E-mail Address: [email protected] Portland State University Google Scholar More articles by this author https://doi.org/10.1044/1058-0360.0602.40 SectionsAboutFull TextPDF ToolsAdd to favoritesDownload CitationTrack Citations ShareFacebookTwitterLinked In References Aram, D., Ekelman, B., & Nation, J. (1984). Preschoolers with language disorders: 10 years later. Journal of Speech and Hearing Science, 27, 232–244. AbstractGoogle Scholar Aram, D., & Nation, J. (1980). Preschool language disorders and subsequent language and academic difficulties. Journal of Communication Disorders, 13, 159–170. CrossrefGoogle Scholar Bain, B., & Olswang, L. (1995). Examining readiness for learning two-word utterances by children with specific expressive language impairment: Dynamic assessment validation. American Journal of Speech-Language Pathology, 4(1), 81–91. LinkGoogle Scholar Belfiore, K. (1996). Intervention history of children with slow expressive language development. Unpublished master’s thesis, Portland State University, Portland, OR. Google Scholar Bishop, D., & Adams, C. (1990). A prospective study of the relationship between specific language impairment, phonological disorders, and reading retardation. Journal of Child Psychology and Psychiatry, 31, 1027–1050. CrossrefMedlineGoogle Scholar Bleile, K., & Miller, S. (1993). Infants and toddlers. Bernthal, J. (Ed.), Articulatory and phonological disorders in special populations. New York: Thieme. Google Scholar Bzoch, K., & League, R. (1971). The Receptive Expressive Emergent Language Scale (REEL). Gainesville, FL: Language Education Division, Computer Management Corporation. Google Scholar Capute, A., Shapiro, B., Wachtel, R., Gunther, V., & Palmer, F. (1986). The clinical linguistic and auditory milestone scale (CLAMS). American Journal of Diseases in Children, 40, 694–698. Google Scholar Casby, M. (1992). The cognitive hypothesis and its influence on speech-language services in schools. Language, Speech, and Hearing Services in Schools, 23, 198–202. LinkGoogle Scholar Coplan, J., Gleason, J. R., Ryan, R., Burke, M., & Williams, M. (1982). Validation of an early language milestone scale in a high-risk population. Pediatrics, 70, 677–683. Google Scholar Fenson, L., Dale, P., Reznick, S., Hartung, J., & Burgess, S. (1990). Norms for the MacArthur Communicative Development Inventories. Poster presented at the International Conference on Infant Studies, Montreal, Quebec. Google Scholar Fenson, L., Dale, P., Reznick, S., Thal, D., Bates, E., Hartung, J., Pethick, S., & Reilly, J. (1993). MacArthur Communicative Development Inventories. San Diego: Singular Publishing Group. Google Scholar Fey, M., Catts, H., & Larrivee, L. (1995). Preparing preschoolers for the academic and social challenges of school. Fey, M., Windsor, J., & Warren, S. (Eds.) Language intervention: Preschool through elementary years (pp. 3– 37). Baltimore: Paul H. Brookes. Google Scholar Frankenburg, W., & Dodds, J. (1967). The Denver Developmental Screening Test. Journal of Pediatrics, 71, 181. Google Scholar Gardner, M. (1981). Expressive One-Word Picture Vocabulary Test-Revised. Novato, CA: Academic Therapy Publications. Google Scholar Garvey, M., & Gordon, N. (1973). A follow-up of children with disorders of speech development. British Journal of Communication Disorders, 8, 17–28. Google Scholar Griffiths, C. (1969). A follow-up study of children with disorders of speech. British Journal of Communication Disorders, 4, 46–56. Google Scholar Hall, K., & Tomblin, J. (1978). A follow-up study of children with articulation and language disorders. Journal of Speech and Hearing Disorders, 43, 227–241. LinkGoogle Scholar Honzik, M. (1983). Measuring mental abilities in infancy: The value and limitations. Lewis, M. (Ed.), Origins of intelligence: Infancy and early childhood (2nd ed.). NY: Plenum. Google Scholar King, R., Jones, D., & Lasky, E. (1982). In retrospect: A fifteen year follow-up of speech-language disordered children. Language, Speech, and Hearing Services in Schools, 13, 24–32. LinkGoogle Scholar Lahey, M. (1990). Who shall be called language disordered? Some reflections and one perspective. Journal of Speech and Hearing Disorders, 55, 612–620. LinkGoogle Scholar Lee, L. (1974) Developmental Sentence Analysis. Evanston, IL: Northwestern University Press. Google Scholar McCall, B. (1983). A conceptual approach to early mental development. Lewis, M. (Ed.), Origins of intelligence: Infancy and early childhood (2nd ed.). NY: Plenum. Google Scholar McCarthy, D. (1972). McCarthy Scales of Children’s Abilities. NY: Psychological Corp. Google Scholar Nelson, K., Camarata, S., Welsh, J., Butkovsky, L., & Camarata, M. (1996). Effects of imitative and conversational recasting treatment on the acquisition of grammar in children with specific language impairment and younger language-normal children. Journal of Speech and Hearing Research, 39, 850–859. LinkGoogle Scholar Newcomer, P., & Hammill, D. (1988). Test of Language Development–Primary. Austin, TX: Pro-Ed. Google Scholar Nippold, M., & Schwarz, I. (1996). Children with slow expressive language development: What is the forecast for school achievement?, American Journal of Speech-Language Pathology, 5(2), 22–30. LinkGoogle Scholar Paul, R. (1993). Outcomes of early expressive language delay. Journal of Childhood Communication Disorders, 15, 7–14. Google Scholar Paul, R. (1995). Language disorders from infancy through adolescence: Assessment and Intervention. St. Louis: Mosby-Year Book. Google Scholar Paul, R. (1996). Clinical implications of the natural history of slow expressive language development. American Journal of Speech-Language Pathology, 5(2), 5–21. LinkGoogle Scholar Paul, R., & Cohen, D. (1984). Outcomes of severe disorders of language acquisition. Journal of Autism and Developmental Disorders, 14, 405–421. CrossrefMedlineGoogle Scholar Paul, R., & Elwood, T. (1991). Maternal linguistic input to toddlers with slow expressive language development. Journal of Speech and Hearing Research, 34, 982–988. LinkGoogle Scholar Paul, R., & Jennings, P. (1992). Phonological behavior in toddlers with slow expressive language development. Journal of Speech and Hearing Research, 35, 99–107. LinkGoogle Scholar Paul, R., & Shiffer, M. (1991). Communicative initiations in normal and late-talking toddlers. Applied Psycholinguistics, 12(4), 419–431. Google Scholar Paul, R., Spangle Looney, S., & Dahm, P. (1991). Communication and socialization skills at ages 2 and 3 in “late-talking” young children. Journal of Speech and Hearing Research, 34, 858–865. LinkGoogle Scholar Rauh, V., Achenbach, T., Nurcombe, B., Howell, C., & Teti, D. (1988). Minimizing adverse effects of low birthweight: Four-year results of an early intervention program. Child Development, 59, 3, 544–553. Google Scholar Rescorla, L. (1989). The Language Development Survey: A screening tool for delayed language in toddlers. Journal of Speech and Hearing Disorders, 54, 587–599. LinkGoogle Scholar Rescorla, L., & Ratner, N. (1996). Phonetic profiles of toddlers with specific expressive language impairment. Journal of Speech and Hearing Research, 39, 153–165. AbstractGoogle Scholar Rose, S., Feldman, A., Wallace, I., & McCarton, C. (1989). Infant visual attention: Relation to birth status and developmental outcome during the first five years. Developmental Psychology, 25, 560–576. Google Scholar Shriberg, L. (1993). Four new speech and prosody-voice measures for genetics research and other studies in developmental phonological disorders. Journal of Speech and Hearing Research, 36, 105–140. LinkGoogle Scholar Silva, P. (1980). The prevalence, stability and significance of developmental language delay in preschool children. Developmental Medicine and Child Neurology, 22, 768–777. CrossrefMedlineGoogle Scholar Sparrow, S., Balla, D. & Cicchetti, D. (1984). Vineland Adaptive Behavior Scales. Circle Pines, MN: AGS. Google Scholar Stevenson, J., & Richman, N. (1978). Behavior, language and development in three year old children. Journal of Autism and Childhood Schizophrenia, 8, 299–313. Google Scholar Stoel-Gammon, C. (1991). Normal and disordered phonology in two-year olds. Topics in Language Disorders, 11(4), 21–32. Google Scholar Thal, D., Oroz, M., & McCaw, V. (1995). Phonological and lexical development in normal and late-talking toddlers. Applied Psycholinguis-tics, 16, 407–424. Google Scholar van Kleeck, A., Gillam, R. B., & Davis, B. (1997). When is “watch and see” warranted? A response to Paul’s 1996 article, “Clinical implications of the natural history of slow expressive language development”. American Journal of Speech-Language Pathology, 6(2), 34–39. Google Scholar Weismer, S., Murray-Branch, J., & Miller, J. (1994). A prospective longitudinal study of language development in late talkers. Journal of Speech and Hearing Research, 37, 852–867. AbstractGoogle Scholar Wetherby, A., Yonclas, D., & Bryan, A. (1989). Communicative profiles of preschool children with handicaps: Implications for early identification. Journal of Speech and Hearing Disorders, 54, 148–158. LinkGoogle Scholar Whitehurst, G., & Fischel, J. (1994). Early developmental language delay: What, if anything, should the clinician do about it?, Journal of Child Psychology and Psychiatry, 35, 613–648. Google Scholar Whitehurst, G., Fischel, J., Lonigan, C., Valdez-Menchaca, M., Arnold, D., & Smith, M. (1991). Treatment of early expressive language delay: If, when, and how. Topics in Language Disorders, 11(4), 55–68. Google Scholar Whitehurst, G., Smith, M., Fischel, J., Arnold, D., & Lonigan, L. (1991). The continuity of babble and speech in children with early expressive language delay. Journal of Speech and Hearing Research, 34, 1121–1129. AbstractGoogle Scholar Additional Resources FiguresReferencesRelatedDetailsCited ByPerspectives on Language Learning and Education15:3 (93-100)1 Oct 2008Early Language Delay and Risk for Language ImpairmentErica M. Ellis and Donna J. ThalContemporary Issues in Communication Science and Disorders34:Fall (118-133)1 Oct 2007Issues in Research on Children With Early Language DelayIrina Tsybina and Alice Eriks-BrophyJournal of Speech, Language, and Hearing Research42:5 (1234-1248)1 Oct 1999Effects of Treatment on Linguistic and Social Skills in Toddlers With Delayed Language DevelopmentShari Brand Robertson and Susan Ellis WeismerAmerican Journal of Speech-Language Pathology7:4 (46-56)1 Nov 1998Expressive Vocabulary Development and Word Combinations of Spanish-English Bilingual ToddlersJanet L. PattersonJournal of Speech, Language, and Hearing Research41:3 (627-641)1 Jun 1998Concurrent and Predictive Validity of an Early Language Screening ProgramThomas Klee, David K. Carson, William J. Gavin, Lisa Hall, Amy Kent and Shaily ReeceAmerican Journal of Speech-Language Pathology6:2 (2-2)1 May 1997From the EditorMarc E. Fey Volume 6Issue 2May 1997Pages: 40-49 Get Permissions Add to your Mendeley library HistoryReceived: Jan 14, 1997Accepted: Mar 3, 1997 Published in issue: May 1, 1997 Metrics Downloaded 308 times Topicsasha-topicsasha-article-typesKeywordschild languagelanguage delayearly interventionlearning disabilitiesCopyright & PermissionsCopyright © 1997 American Speech-Language-Hearing AssociationLoading...
Word sense disambiguation assumes word senses. Withinthe lexicography and linguistics literature, they areknown to bevery slippery entities. The first part of the paperlooks at problemswith existing accounts of ‘word sense’ and describesthe various kinds of ways in which a word's meaning candeviate from its coremeaning. An analysis is presented in which wordsenses areabstractions from clusters of corpus citations, inaccordance withcurrent lexicographic practice. The corpus citations,not the wordsenses, are the basic objects in the ontology. Thecorpus citationswill be clustered into senses according to thepurposes of whoever or whatever does the clustering. In theabsence of suchpurposes, word senses do not exist.
The increasing availability of corpora annotated for linguistic structure prompts the question: if we have the same texts, annotated for phrase structure under two different schemes, to what extent do the annotations agree on structuring within the text? We suggest the term tree alignment to indicate the situation where two markup schemes choose to bracket off the same text elements. We propose a general method for determining agreement between two analyses. We then describe an efficient implementation, which is also modular in that the core of the implementation can be reused regardless of the format of markup used in the corpora. The output of the implementation on the Susanne and Penn treebank corpora is discussed.
Video cameras provide a simple, noninvasive method for monitoring a subject’s eye movements. An important concept is that of the resolution of the system, which is the smallest eye movement that can be reliably detected. While hardware systems are available that estimate direction of gaze in real time from a video image of the pupil, such systems must limit image processing to attain real-time performance and are limited to a resolution of about 10 arc minutes. Two ways to improve resolution are discussed. The first is to improve the image processing algorithms that are used to derive an estimate. Offline analysis of the data can improve resolution by at least one order of magnitude for images of the pupil. A second avenue by which to improve resolution is to increase the optical gain of the imaging setup (i.e., the amount of image motion produced by a given eye rotation). Ophthalmoscopic imaging of retinal blood vessels provides increased optical gain and improved immunity to small head movements but requires a highly sensitive camera. The large number of images involved in a typical experiment imposes great demands on the storage, handling, and processing of data. A major bottleneck had been the real-time digitization and storage of large amounts of video imagery, but recent developments in video compression hardware have made this problem tractable at a reasonable cost. Images of both the retina and the pupil can be analyzed successfully using a basic toolbox of image-processing routines (filtering, correlation, thresholding, etc.), which are, for the most part, well suited to implementation on vectorizing supercomputers.
Norm and variation in the lexical structure of Italian medical language: Illustrations from two Centuries
This paper addresses the question of whether it is possible tosense-tag systematically, and on a large scale, and how we shouldassess progress so far. That is to say, how to attach each occurrenceof a word in a text to one and only one sense in a dictionary – aparticular dictionary of course, and that is part of the problem. Thepaper does not propose a solution to the question, though we havereported empirical findings elsewhere (Cowie et al., 1992;Wilks et al., 1996; Wilks and Stevenson, 1997), and intend to continue andrefine that work. The point of this paper is to examine two well-knowncontributions critically: The first (Kilgarriff, 1993), which is widelytaken to show that the task, as defined, cannot be carried outsystematically by humans and, secondly (Yarowsky, 1995), which claimsstrikingly good results at doing exactly that.
Multi-Media does not, just by itself, guarantee accelerated learning and enhanced motivation unless there is a clear pedagogical progression and learning strategy. The authors describe and analyze the didactic dimensions to be considered when designing a multi-media tool, based on their own experience as software authors and language trainers.
The paper describes a morphological analyser forEstonian and how using a text corpus influenced theprocess of creating it and the resulting programitself. The influence is not limited to the lexicononly, but is also noticeable in the resulting algorithm andimplementation too. When work on the analyser began,there were no computational treatment of Estonianderivatives and compounds. After some cycles ofdevelopment and testing on the corpus, we came up withan acceptable algorithm for their treatment. Both themorphological analyser and the speller based on ithave been successfully marketed.
Contents: B.B. Kachru, Ladislav Zgusta: Illinois Years. - Bibliography of Publications by Ladislav Zgusta. - B.B. Kachru, Introduction. - I: Contextualizing Culture: D. Bartholomew, Otomi Culture from Dictionary Illustrative Sentences. - P. Corbin, Un Film, Deux Linguistes et Quelques Dictionnaires. Un Regard Particulier sur Simple Mortel de Pierre Jolivet. - J.A. Fishman, Dictionaries as Culturally Constructed and as Culture-Constructing: Reciprocity View as Seen from Yiddish Sources. - F.J. Hausmann, Allusions Litteraires et Citations Historiques Dans le Tresor de la Langue Francaise. - L.F. Lara, Towards a Theory of the Cultural Dictionary. - W.P. Lehmann, Spindle or the Distaff. - Y. Malkiel, Principal Categories of Learned Words. - E.A. Nida, Lexical Cosmetics. - II: Lexicography in Historical Context: F. Karttunen, Roots of Sixteenth-Century Mesoamerican Lexicography. - T.B.I. Creamer, Current State of Chinese Lexicography. - D.A. Kibbee, 'New Historiography', History of French and 'Le Bon Usage' in Nicot's Dictionary (1606). - N. Dinh-Hoa, On Chi-nam Ngoc-am Giai-nghia: An Early Chinese-Vietnamese Dictionary. - G. Stein, Chaucer and Lydgate in Palsgrave's Lesclarcissement. - III: Ideology, Norms and Language Use: M.A. Ezquerra,Political Considerations on Spanish Dictionaries. - D.M.T.Cr. Farina, Marrism and Soviet Lexicography. - C. Marello, Florence like Athens and Italian like Greek: An Ideologically Biased Theme in the Forewords of Some Italian Thesauri of the 19th Century. - A. Wierzbicka, Dictionaries and Ideologies: Three Examples from Eastern Europe. - R.D. Zorc, Philippine Regionalism versus Nationalism and the Lexicographer. - IV: Pluricentricity and Ethnocentricism: J. Algeo, British and American Biases in English Dictionaries. - C.W. Kim, One Language, Two Ideologies, and Two Dictionaries: Case of Korean. - B. Lewandowska-Tomaszczyk, Worldview and Verbal Senses. - C. Poirier, De la Soumission a la Prise de Parole: Le Cheminement de la Lexicographie au Quebec. - J. Whitcut, Taking it for Granted: Some Cultural Preconceptions in English Dictionaries. - V: Dictionaries across Languages and Cultures: Y. Kachru, Lexical Exponents of Cultural Contact: Speech Act Verbs in Hindi-English Dictionaries. - R.J. Steiner, Bilingual Dictionary in Cross-Cultural Contexts. - VI: Language Dynamics vs. Prescriptivism: A.P. Cowie, Learner's Dictionary in a Changing Cultural Perspective. - R.H. Gouws, Dictionaries and the Dynamics of Language Change. - F.E. Knowles, Dictionaries for the People or for People? - VII: Language Learner as the Consumer: G.M. Dalgish, Learners' Dictionaries: Keeping the Learner in Mind. - VIII: Structuring Semantics: F.F.M. Dolezal, Dictionary as Philosophy: Reconstructing the Meaning of Our Father. - M. Ritchie Key, Meaning as Derived from Word Formation in South American Indian Languages. - J.P. Louw, How Many Meanings to a Word? - IX: Ethical Issues and Lexicologists' Biases: D.L. Gold, When Religion Intrudes into Etymology (On The Word: Dictionary that Reveals the Hebrew Source of English). - T. McArthur, Culture-Bound and Trapped by Technology: Centuries of Bias in the Making of Wordbooks. - X: Terminology across Cultures: Z. Polacek, Amharic Lexicography and the Dynamics of Sociopolitical Terminology. - G.O. Richter, Grammatical Indications in Chinese Monolingual Dictionaries. - XI: Afterword: B.B. Kachru, Directions and Challenges.
Dictionary markup is one of the concerns of the Text Encoding Initiative (TEI), an international project for text encoding. In this paper, we investigate ways to use and extend the TEI encoding scheme for the markup of Korean dictionary entries. Since TEI suggestions for dictionary markup are mainly for western language dictionaries, we need to cope with problems to be encountered in encoding Korean dictionary entries. We try to extend and modify the TEI encoding scheme in the way suggested by the TEI. Also, we restrict the content model so that the encoded dictionary might be viewed as a database as well as a computerized, originally printed, dictionary.
The authors here show that machine learning techniques can be used for designing an archaeological typology, at an early stage when the classes are not yet well defined. The program (LEGAL, LEarning with GAlois Lattice) is a machine learning system which uses a set of examples and counter-examples in order to discriminate between classes. Results show a good compatibility between the classes such as the yare defined by the system and the archaeological hypotheses.
The statement, ’’Results of most non-traditional authorship attribution studies are not universally accepted as definitive,'' is explicated. A variety of problems in these studies are listed and discussed: studies governed by expediency; a lack of competent research; flawed statistical techniques; corrupted primary data; lack of expertise in allied fields; a dilettantish approach; inadequate treatment of errors. Various solutions are suggested: construct a correct and complete experimental design; educate the practitioners; study style in its totality; identify and educate the gatekeepers; develop a complete theoretical framework; form an association of practitioners.
This paper presents a new view of ExplanatiomBased Learning (EBL) of natural language parsing. Rather than employing EBL for specializing parsers by inferring new ones, this paper suggests employing EBL for learning how to reduce ambiguity only partially.We exemplify this by presenting a new EBL method that learns parsers that avoid spurious overgeneration, and we show how the same method can be used for reducing the sizes of stochastic: grammars learned from treebanks, e.g. (Bod, 1995, Charniak, 1996,
Two experiments examined the feasibility of psychological assessment using interactive voice response (IVR) technology and the potential sensitivity of such assessments to alcohol and fatigue effects. In Experiment 1, 10 subjects performed a 12-min battery of six IVR-administered tasks, Monday through Friday, over 2 weeks. Minimal learning effects were evident during training. Repeated administrations indicated high test-retest reliabilities. In Experiment 2 (double-blind, alcohol/placebo crossover design), 7 subjects were tested every 2 h over a 24-h period during two experimental sessions (peak blood alcohol concentrations =80 mg/dL). Several IVR-administered tasks were sensitive to alcohol impairment, but not as sensitive as laboratory-based measures specifically designed to assess alcohol impairment. Little evidence for fatigue-related impairment was obtained. The results support optimism for the potential to assess psychomotor and cognitive functioning distally via telephony; however, further refinement and validation of the methods are needed.
This paper describes the novel ways in which the Orlando Project, based at the Universities of Alberta and Guelph, is using SGML to create an integrated electronic history of British women's writing in English. Unlike most other SGML-based humanities computing projects which are tagging existing texts, we are researching and writing new material, including biographies, items of historical significance, and many kinds of literary and historical interpretation, all of which incorporates sophisticated SGML encoding for content as well as structure. We have created three DTDs, for biographies, for writing-related activities and publications, and for social, political and other events. A major factor influencing the design of the DTDs was the requirement to be able to merge and restructure the entire text base in many ways in order to retrieve and index it and to reflect multiple views and interpretations. In addition a stable and well-documented system for tagging was deemed essential for a team which involves almost twenty people, including eight graduate students, in two locations.
In this report two programs for statistical analysis of concordance lines are described. The programs have been developed for analyzing he lexical context of a given word. It is shown how different parameter settings influence the outcome of collocational analysis, and how the concept of collocation can be extended to allow the extraction of lines typical for a word from a set of concordance lines. Even though all the examples are for English, the software is completely language independent and only requires minimal linguistic resources.
This paper describes the LT NSL system (McKelvie et al., 1996), an architecture for writing corpus processing tools. This system is then compared with two other systems which address similar issues, the GATE system (Cunningham et al., 1995) and the IMS Corpus Workbench (Christ, 1994). In particular we address the advantages and disadvantages of an SGML approach compared with a non-sgml database approach.
This paper reports a detailed evaluation of the effectiveness of a system that has been developed for the identification and retrieval of morphological variants in searches of Latin text databases. A user of the retrieval system enters the principal parts of the search term (two parts for a noun or adjective, three parts for a deponent verb, and four parts for other verbs), this enabling the identification of the type of word that is to be processed and of the rules that are to be followed in determining the morphological variants that should be retrieved. Two different search algorithms are described. The algorithms are applied to the Latin portion of the Hartlib Papers Collection and to a range of classical, vulgar and medieval Latin texts drawn from the Patrologia Latina and from the PHI Disk 5.3 datasets. The effectiveness of these searches demonstrates the effectiveness of our procedures in providing access to the full range of classical and post-classical Latin text databases.
In Modern Japanese, three nominalizers, no, koto, and zero, are used in overlapping but slightly differing syntactic and semantic environments. This paper will discuss two types of strategy in nominalizer variation in Modern Japanese, koto versus no, on the one hand, and no versus zero, on the other. The two types of variation are shown to be very different in nature. The contrast between koto and no is primarily a semantic one characterizable in terms of greater versus lesser semantic specificity. Koto nominalization is the most semantically sensitive type of nominalization; it tends to be called for where its semantic contribution is expected, no matter how it is. No nominalization, which generally lacks semantic content, is the most unmarked nominalization strategy extensively used in Modern Japanese. The contrast between no and zero, on the other hand, is primarily a syntactic one with a diachronic dimension. In Old Japanese, which lacked the nominalizer no, a syntactic contrast existed between koto and zero. Zero nominalization, a remnant of Old Japanese, is now generally a nonproductive process limited to lexicalized expressions with varying degrees of archaism. No nominalization has long competed with zero nominalization and has successfully replaced it in many syntactic environments including some of the lexicalized expressions where zero was traditionally the norm. However, in spite of the overall dominance of no nominalization, zero nominalization has continued to be used fairly productively in some limited syntactic environments and is not expected to disappear from the syntactic inventory of Modern Japanese in the near future
This paper describes a statistics-based Chinese parser, which parses the Chinese sentences with correct segmentation and POS tagging information through the fol]owing processing stages: I) to predict constituent boundaries, 2) to match open and close brackets and produce sytactic ls, 3) to disambiguate and choose the best parse Itee. Evaluating the parser against a smaller Chinese treebank with 5573 sentences, it shows the following encouraging results: 86% precision, 86% recall, 1.1 crossing brackets per sentence and 95% labeled precision.
Automatic Text Categorization (TC) is a complex and useful task for many natural language applications, and is usually performed through the use of a set of manually classified documents, a training collection. We suggest the utilization of additional resources like lexical databases to increase the amount of information that TC systems make use of, and thus, to improve their performance. Our approach integrates WordNet information with two training approaches through the Vector Space Model. The training approaches we test are the Rocchio (relevance feedback) and the Widrow-Hoff (machine learning) algorithms. Results obtained from evaluation show that the integration of WordNet clearly outperforms training approaches, and that an integrated technique can effectively address the classification of low frequency categories.
A treebank is a corpus of tagged and bracketed sentences capturing the linguistic properties of a (sub)language in an empirical way. The CASSANDRA treebank is developed as a sideline to the GALEN-IN-USE project in which it serves to make the relationships between natural language phenomena and semantic representations of medical expressions explicit, and to assist in the quality assurance of the modelling centres. The end result is a multilingual linguistic knowledge repository from which lexicons and grammars of various types can be derived in an automatic way.
Automatic text categorization is a complex and useful task for many natural language processing applications. Recent approaches to text categorization focus more on algorithms than on resources involved in this operation. In contrast to this trend, we present an approach based on the integration of widely available resources as lexical databases and training collections to overcome current limitations of the task. Our approach makes use of WordNet synonymy information to increase evidence for bad trained categories. When testing a direct categorization, a WordNet based one, a training algorithm, and our integrated approach, the latter exhibits a better perfomance than any of the others. Incidentally, WordNet based approach perfomance is comparable with the training approach one.
Automatic text categorization is a complex and useful task for many natural language processing applications. Recent approaches to text categorization focus more on algorithms than on resources involved in this operation. In contrast to this trend, we present an approach based on the integration of widely available resources as lexical databases and training collections to overcome current limitations of the task. Our approach makes use of WordNet synonymy information to increase evidence for bad trained categories. When testing a direct categorization, a WordNet based one, a training algorithm, and our integrated approach, the latter exhibits a better perfomance than any of the others. Incidentally, WordNet based approach perfomance is comparable with the training approach one.
A lexical modeling methodology was employed to examine how the distribution of phonemic patterns in the lexicon constrains lexical equivalence under conditions of reduced phonetic distinctiveness experienced by speech-readers. The technique involved (1) selection of a phonemically transcribed machine-readable lexical database, (2) definition of transcription rules based on measures of phonetic similarity, (3) application of the transcription rules to a lexical database and formation of lexical equivalence classes, and (4) computation of three metrics to examine the transcribed lexicon. The metric percent words unique demonstrated that the distribution of words in the language substantially preserves lexical uniqueness across a wide range in the number of potentially available phonemic distinctions. Expected class size demonstrated that if at least 12 phonemic equivalence classes were available, any given word would be highly similar to only a few other words. Percent information extracted (PIE) [D. Carter, Comput. Speech Lang. 2, 1-11 (1987)] provided evidence that high-frequency words tend not to reside in the same lexical equivalence classes as other high-frequency words. The steepness of the functions obtained for each metric shows that small increments in the number of visually perceptible phonemic distinctions can result in substantial changes in lexical uniqueness.
Psychophysiological changes during long-distance driving may be associated with driving fatigue and morbidity. Measures of stress and arousal, including heart rate, blood pressure, catecholamines, cortisol, state anxiety, and self-ratings of stress and arousal were collected from 10 long-distance bus drivers during 12-hour driving shifts and at matched times on nondriving rest days. Cardiovascular and catecholamine data were elevated across the entire work day, compared with rest days. Self-reported stress and state anxiety were elevated only at the preshift measure, and these elevations were interpreted as the result of anticipatory anxiety and additional work demands at the beginning of the shift. Decelerating activation from the 9th to the 12th hours of driving were reflected in slower heart rate and lower subjective arousal ratings. Suggested explanations for these findings are that drivers experience a release of tension when they anticipate the end of the shift and therefore deactivation is a signal or precursor to the onset of fatigue in physiological adjustment mechanisms.
This paper addresses issues in automated treebank construction. We show how standard part-of-speech tagging techniques extend to the more general problem of structural annotation, especially for determining grammatical functions and syntactic categories. Annotation is viewed as an interactive process where manual and automatic processing alternate. Efficiency and accuracy results are presented. We also discuss further automation steps.
We present and experimentally evaluate a new model of prounciation by analogy: the paradigmatic cascades model. Given a pronunciation lexicon, this algorithm first extracts the most productive paradigmatic mappings in the graphemic domain, and pairs them statistically with their correlate(s) in the phonemic domain. These mappings are used to search and retrieve in the lexical database the most promising analog of unseen words. We finally apply to the analogs pronunciation the correlated series of mappings in the phonemic domain to get the desired pronunciation.
Introduction to the special issue on computational linguistics using large corpora, Kenneth W. Church and Robert L. Mercer generalized probabilistic LR parsing of natural language (corpora) with unification-based grammars, Ted Briscoe and John Carroll accurate methods for the statistics of surprise and coincidence, Ted Dunning a program for aligning sentences in bilingual corpora, William A. Gale and Kenneth W. Church structural ambiguity and lexical relations, Donald Hindle and Mats Rooth text-translation alignment, Martin Kay and Martin Roescheisen retrieving collocations from text - Xtract, Frank Smadja using register-diversified corpora for general language studies, Douglas Biber from grammar to lexicon - unsupervised learning of lexical syntax, Michael R. Brent the mathematics of statistical machine translation - parameter estimation, Peter F. Brown et al building a large annotated corpus of English - the Penn treebank, Mitchell P. Marcus et al lexical semantic techniques for corpus analysis, James Pustejovsky et al coping with ambiguity and unknown words through probabilistic models, Ralph Weischedel et al.
This technical report is an appendix to Eisner (1996): it gives superior experimental results that were reported only in the talk version of that paper. Eisner (1996) trained three probability models on a small set of about 4,000 conjunction-free, dependency-grammar parses derived from the Wall Street Journal section of the Penn Treebank, and then evaluated the models on a held-out test set, using a novel O(n^3) parsing algorithm. The present paper describes some details of the experiments and repeats them with a larger training set of 25,000 sentences. As reported at the talk, the more extensive training yields greatly improved performance. Nearly half the sentences are parsed with no misattachments; two-thirds are parsed with at most one misattachment. Of the models described in the original written paper, the best score is still obtained with the generative (top-down) "model C." However, slightly better models are also explored, in particular, two variants on the comprehension (bottom-up) "model B." The better of these has an attachment accuracy of 90%, and (unlike model C) tags words more accurately than the comparable trigram tagger. Differences are statistically significant. If tags are roughly known in advance, search error is all but eliminated and the new model attains an attachment accuracy of 93%. We find that the parser of Collins (1996), when combined with a highly-trained tagger, also achieves 93% when trained and tested on the same sentences. Similarities and differences are discussed.
Abstract Previous studies examining the association between social comparison processes and body image dissatisfaction have yielded inconsistent findings. This study examined whether such discrepancies are due to either the use of identical comparison targets for all subjects or variability in body mass. Specifically, 216 subjects were randomly assigned to one of three experimental conditions: self-generated upward comparison group, self-generated downward comparison group, or control group. Dependent variables were measures of body image. Results indicated that increasing body mass and trait comparison tendencies were associated with increased body dissatisfaction. However, the experimental manipulation did not affect body image ratings. Results suggest that social comparison processes may operate similarly over a range of body mass index (BMI) values.
This paper reports on an empirically based system that automatically resolves VP ellipsis in the 644 examples identified in the parsed Penn Treebank. The results reported here represent the first systematic corpus-based study of VP ellipsis resolution, and the performance of the system is comparable to the best existing systems for pronoun resolution. The methodology and utilities described can be applied to other discourse-processing problems, such as other forms of ellipsis and anaphora resolution. The system determines potential antecedents for ellipsis by applying syntactic constraints, and these antecedents are ranked by combining structural and discourse preference factors such as recency, clausal relations, and parallelism. The system is evaluated by comparing its output to the choices of human coders. The system achieves a success rate of 94.8%, where success is defined as sharing of a head between the system choice and the coder choice, while a baseline recency-based scheme achieves a success rate o,I:75.0 % by this measure. Other criteria for success are also examined. When success is defined as an exact, word-for-word match with the coder choice, the system performs with 76.0 % accuracy, and the baseline approach achieves only 14.6% accuracy. Analysis of the individual components of the system shows that each of the structural and discourse constraints used are strong predictors of the antecedent of VP ellipsis. 1.
Pronunciation by analogy (PbA) is an emerging, data-driven technique with potential application in text-to-speech (TTS) systems, as well as being an influential psychological model of reading aloud. The underlying idea is that a pronunciation for an unknown word (i.e., one not in the dictionary, or lexicon, of the human or machine “reader”) is assembled by matching substrings of the input to substrings of known, lexical words, hypothesizing a partial pronunciation for each matched substring from the lexical knowledge of the “reader,” and concatenating the partial pronunciations. This paper assesses the capability of PbA to derive pronunciations for unknown words of English. As a psychological model, PbA is “under-specified,” that is, the implementor of a simulation of the process faces detailed choices which can only be resolved by trial and error. One goal for this paper is to explore the impact of certain basic implementational choices on the performance of PbA systems. The variables studied are the specific lexical database used as the basis of the analogy process, the way of ranking/scoring candidate pronunciations, and the effect of manual versus automatic alignment of letters and phonemes. When tested with short (monosyllabic) pseudowords previously used in experimental psychology studies, the lowest error rate achieved is 14.3% (for a test set of size 70). We conclude that current PbA systems are at best poor models of pseudoword pronunciation by humans. To assess their suitability for use in a TTS application, in which multisyllabic words will be encountered, the implementations have also been tested with lexical words temporarily removed from the dictionary. The best performance obtained was 93.5% phonemes correct (corresponding to 67.9% words correct) for a 16,280-word dictionary. This is vastly superior to the 25.7% words correct obtained using a set of popular letter-to-sound rules, indicating considerable scope for analogy methods to be exploited in future TTS systems.
This study was designed to compare the effects of different kinds of visual presentations and music alone on university nonmusic students' affective and cognitive responses to music. Four groups of participants were presented with excerpts from the first and fourth movements of Beethoven's Symphony No. 6 in F major (“Pastoral”). Two groups heard music excerpts only, one interpretation conducted by Stowkowski and one by Bernstein. One of the video groups viewed corresponding excerpts from the movie Fantasia while listening to the Stowkowski recording. A second group viewed and listened to a performance video of the Vienna Philharmonic filmed during the Bernstein recording. All students (N = 128) completed cognitive listening tests based on the excerpts, rated the music on Likert-type affective scales, and responded to two open-ended questions. Significant effects of presentation condition were found. Cognitive scores were higher for the performance video than the music plus animation video on both movements. Scores for the two music-only presentations were not significantly different from each other or the two video presentations. Although affective ratings were not significantly different in magnitude between the presentation groups, the animation video (Fantasia.) presentation ranked consistently higher in affect than the other presentations. Implications of these results regarding the effects of different types of visual information presented to music listeners are discussed.
In a study that used Hermans's (1987, 1988) valuation procedure, 40 participants each provided a highly valued experience of four types: science, religion, interpersonal and intrapersonal conflict between science and religion. They then rated these valuations on 30 affect terms, some of rich were later organized into categories: Positive, Negative, Self, and Other. Participants also filled out questionnaires, that were used to categorize them as low or high in scientific and religious orientation. Typical valuations of participants in these four science and religion categories are presented as qualitative idiographic information. In addition, quantitative analyses of affect ratings are presented as nomothetic information. Generally, affect ratings of scienfific experiences were more Self-directed while religious experience valuations involved equally high levels of Self and Other affect. Both scientific and religious experiences were evaluated as having Positive but not Negative affect. Interpersonal and intropersonal conflict were experienced as more Self-directed than Other-oriented. While interpersonal conflicts displayed more Negative affect than Positive, intrapersonal conflict was evaluated equally on these two measures.
We show how a treebank can be used to cluster words on the basis of their syntactic behavior. The resulting clusters represent distinct types of behavior with much more precision than parts of speech. As an example we show how prepositions can be automatically subdivided by their syntactic behavior and discuss the appropriateness of such a subdivision. Applications of this work are also discussed. 1 Introduction The construction of classes of words, or calculation of distances between words, has frequently drawn the interest of researchers in natural language processing. Many of these studies aimed at finding classes based on co-occurrences, often combined with the aim of establishing semantic similarity between words (McMahon and Smith, 1996; Brown et al., 1992; Dagan, Markus, and Markovitch, 1993; Dagan, Pereira, and Lee, 1994; Pereira and Tishby, 1992; Grefenstette, 1992). We suggest a method for clustering words purely on the basis of syntactic behavior. We show how the necessary...
We introduce a novel parser based on a probabilistic version of a left-corner parser. The left-corner strategy is attractive because rule probabilities can be conditioned on both top-down goals and bottom-up derivations. We develop the underlying theory and explain how a grammar can be induced from analyzed data. We show that the left-corner approach provides an advantage over simple top-down probabilistic context-free grammars in parsing the Wall Street Journal using a grammar induced from the Penn Treebank. We also conclude that the Penn Treebank provides a fairly weak testbed due to the flatness of its bracketings and to the obvious overgeneration and undergeneration of its induced grammar.
A large collection of texts may be reached through the Internet and this provides a powerful platform from which common-sense knowledge may be gathered. This paper presents a system that contains a core knowledge base structured around WordNet, a lexical database, capable of extracting contextual information from a given input text. Such context information is then used to retrieve other texts from the Internet that relate to that context. When processed by the system, these new texts bring more information that represents an enhanced domain context for the initial text. This is an incremental method for text processing that acquires domain knowledge from other texts. The paper describes the system architecture, its core knowledge base and inference engine, and the acquisition of new knowledge from corpora.
Queer theory offers insights for political economy on how humans induce categories and conflate traits in ways psychologists call "illusory correlations." A Bayesian simulation is constructed of people interacting and using probits to compare their rankings of alternatives to estimate the subjective probability i.) that others have the same tastes as them, and ii.), for each alternative, that this alternative is their best choice. These simulations are found to, at least simplistically, resemble a type of illusory correlation which gained increased prominence in the United States from 1930 to 1960, and earlier in the United Kingdom when queer panics conflated "predatory" and "traitorous" with lesbian/gay. This modeling of the social articulation of preferences leads to conjectures on the role of Michel Foucault's épistémè (1972), Barbara Ponse's principle of consistency (1978), Jeffrey Escoffier's master code (1985), Sandra Bem's schema (1981), John R. Searle's Background (1990, 1992, 1995), and Judith Butler's linguistic norms (1993). Here these are called cognitive dispositions or codes, and are seen as social structures which grow as individuals try to form homopreference networks to process information in parallel, collectively. The concept of identity or ideology entrepreneurs is used to establish the importance of institutional analysis for political economy. Such an incorporation of desire on a par with logic-"rationality"-is called post/modern and is used to overcome the silence in both neoMarxian and neoclassical political economy on queer theory and queer issues.
In two studies, pedestrians in Old and New Delhi (India) and Dhaka (Bangladesh) were asked about their reactions to three stressors common to rapidly growing urban areas in South Asia: noise, air pollution, and crowding. Results from the first study, a survey of men in Old Delhi, indicated that respondents who were more upset by noise and by crowding also reported more physical symptoms and less perceived control. In the second study, male and female pedestrians were interviewed in New Delhi and Dhaka. Results revealed consistent gender, country, and gender by country effects on measures of general affect, ratings of stressors, and coping responses. In addition, results from an experimental manipulation in Study 2 indicated that in both countries, telling pedestrians about the effects of air pollution or crowding made them feel significantly worse than they would have felt had they not been given any information.