Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
The Kay Elemetrics nasometer measures nasalance, a parameter of speech that reflects the proportion of total acoustic energy that is emitted nasally, making it possible to infer velopharyngeal (VP) function noninvasively. Nasometric evaluation is potentially widely applicable in the clinical assessment of suspected VF impairment. For clinical use, a patient’s mean nasalance on a passage of known phonetic composition must be compared to age-appropriate population norms. Most potential clinical subjects are young; many are preliterate. Passages for which norms have been established (Zoo, Rainbow, NasalSentences) are syntactically, semantically, and lexically complex, phonetically heterogeneous, phonologically mature, and long. Individual child subjects’ nasalance scores, if obtained at all, are therefore likely to be contaminated by artifacts created through hesitation noises, filled pauses, phonetic deviance from normed target, age differences, and measurement errors induced by coaching procedures. Differences in phonetic content between normed and actual utterances are almost inevitable; they lead to uninterpretable results. This study reports a technique for obtaining clinically useful nasalance scores from young, preliterate subjects, even those evidencing phonological deficits or noncompliant behavior. Large-n norms for preschool and primary children are presented.
This paper describes a procedure for obtaining conditional accuracy functions(CAFs) from naive observers and a restricted number of trials. The method permits the experimenter to counter the subjects’ tendency to favor accuracy in tasks in which stimulus discrimination is easy. Each time a block of 12 trials contains less than three errors, observers are instructed, by means of a speed-up signal, to respond faster. The subject is continuously informed about her/his effective reaction time. The data show that the desired speed-accuracy tradeoff was obtained within each of the 7 observers. The mean percent error was around 25%.
Describes a synthesis-by-analogy system which is a model of novel-word pronunciation by humans. It uses analogy in both orthographic and phonological domains and is applied to the pronunciation of novel words in British English and German. A major part of this cross-language study concerned the impact of implementational choices on performance, where this was defined as the ability of the system to produce pronunciations in line with those given by humans. The size and content of the lexical database on which any analogy system must be based were also considered. The better performing implementations produced useful results for both British English and German. However, best results for each of the 2 languages were obtained from different implementations. The system described is also a psychological model of reading aloud. (German & French abstracts) (PsycINFO Database Record (c) 2018 APA, all rights reserved)
Although many scholars in literature currently seem mainly interested in theory, the focus on literary texts is what defines literature studies. Computer technology and the statistical methods it fosters are applicable to both the theoretical and to the interpretative issues which scholars of literature habitually address. Genette's distinction between the homodiegetic and the autodiegetic perspective in first-person narrative can be confirmed statistically. Roquentin's loneliness inLa nausée can be shown to be a formal characteristic of the type of novel he narrates, thus validating his commentary on his society. The computer can be used to deal with standard literary questions in a principled fashion, and a new orientation of literature studies on a cultural history model, which Mark Olsen recommends, is not necessary.
The study of signs is divided between those scholars who use the Saussurian binary sign (semiology) and those who prefer Charles Peirce's tripartite sign (semiotics). The common view of the opposition between the two types of signs does not take into consideration the methodological conditions of applicability of these two types of signs. This is particularly important in the field of literary studies and hence for the preparation of electronic programs for text analysis. The Peircian sign explicitly entails the discovery of a truth of meaning that claims to be universal and not reducible to a collection of opinions based on fragmented information; it also imposes the task of elucidating a transhistorical and universal significantion encoded in a text. Contrary to Peirce's view of the sign, our use of computer programs for text analysis, however, demonstrates that we implicitly treat every literary text as a set of linguistic data (letters, phonemes, syntagmatic segments, etc.) which are reducible to units that can be treated separately. A brief comparison of the results obtained from computer analyses of the French poet Stéphane Mallarmé's text, “Le Cygne,” with those obtained from two Peircian analyses (by Riffaterre and Champigny) of the same text demonstrates that our current methods of computer textual analysis are based on a Saussurian semiology, which is unidimensional and limited, and that these methods are still quite unable to produce a semiotic interpretation based on a totalizing hierarchy of the text's various discursive components.
Recent articles have noted that humanities computing techniques and methodologies remain marginal to mainstream literary scholarship. Mark Olsen's paper discusses this phenomenon and argues for large scale analyses of text databases that would incorporate a shift in theoretical orientation to include greater stress on intertextuality and sign theory. Part of Olsen's argument revolves on the need to move away from the syntactic and overt grammatical elements of textual language to more subtle semantics and meaning systems. While provocative and important, Olsen's stance remains rooted in literary theoretical constructs. Another level of language, the cognitive, offers equally interesting challenges for humanities computing, though the paradigms for this type of computer-based exploration are derived from disciplines traditionally removed from the humanities. The riddle, a nearly universal genre, offers a window onto some of the cognitive processes involved in deep level language function. By analyzing the riddling process, different methods of computational modelling can be inferred, suggesting new avenues for computing in the humanities.
Although Columbus'Diary of the first voyage to America as we know it is largely a transcription of the original diary carried out by Bartolomé de las Casas, commentators and readers often treat it as if it were Columbus' work alone. Editions published to date do not separate the explorer's narrative from that of his transcriber or editor. Since style can influence readers' perceptions of a writer's personality, it is important to determine characteristics of writing attributed to Columbus that may pertain instead to his trascriber. This study employs the computer to explore the style of Las Casas and that of Columbus. Differences in the writing of each “author” emerge with computer assistance by isolating Columbus' words from those of his transcriber and analyzing selected features of vocabulary, sentence length, and syntax.1
Lack of a critical mass of scholars involved with the computer-assisted analysis of texts (CAAT), coupled with insufficient communication among various sectors of the literary and linguistic disciplines, has led to a skewed notion of computing humanists' work among their colleagues. This paper highlights the gap through examples of misunderstood humanist needs and achievements drawn from both recent media reports and humanities conferences. It suggests that networking and less modesty in manuscript submission can be at least partial solutions. The author cites some of his own published work and work-in-progress on Stendhal and Gobineau in refuting Mark Olsen's thesis that the dominance of single- or dual-author studies must be the cause of CAAT's “failure” to make significant inroads in mainstream literary journals. The author builds a case for the use of both diachronic and synchronic lexico-statistical data in carrying out such studies successfully. He recommends a new “Synthetic Criticism” where relevant quantitative methods would not be absent.
This paper concurs with Mark Olsen's premise that computer-aided literature studies should take a different direction, one that is more suited to the computer's strength in analyzing large corpora of texts. However, the authors take issue with his conclusion that a reorientation of the notions of textual analysis is necessary in order to exploit the computer's capabilities. Contemporary medieval studies already provides us with models of textual analysis which are well suited to computer development. Though they stem from the particularities of medieval textual production, these models can perhaps be useful in the study of modern literatures.
ABSTRACT: Despite the current extensive research on many variables associated with English language testing, little attention has been focused on the inherent variability in the linguistic ‘norms’ for Standard English that are generally tested. This paper examines data from domains of Standard English in several native‐speaker (e.g. British) and non‐native (e.g. Malaysian) varieties of English, focusing on linguistic and attitudinal factors leading to changes in these norms. It then discusses implications of these changes for test construction aimed at assessing proficiency in English as a global language, particularly through such standardized instruments as the Test of English for International Communication (TOEIC).
In this paper an expert system is described which is called Lexicographer and which aims at supplying the user with diverse information about Russian words, including bibliographic information concerning individual lexical items. It is supposed that the system may be of use for a practical computational linguist and at the same time will serve as an instrument of linguistic research.
For those studying languages with rich word structures, a morphological parser is a valuable tool. PC-KIMMO is a parser for small computers that is based on Koskenniemi's two-level model of morphology. Of the many practical uses for a morphological parser such as PC-KIMMO, this article describes one: producing automatically glossed interlinear text.
This paper deals with the problem of discovering rules that govern social interactions and relations in preliteral societies. Two older computer programs are first described which can receive data, possibly incomplete and redundant, representing kinship relations among named individuals. The programs then establish a knowledge base in the form of a directed graph, which the user can query in a variety of ways. Another program, written on the “top” of these (rewritten in LISP), can form concepts of various properties, including kinship relations, of and between the individuals. The concepts are derived from the examples and non-examples of a certain social pattern, such as inheritance, succession, marriage, class (tribe, moiety, clan, etc.) membership, domination-subordination, incest and exogamy. The concepts become hypotheses about the rules, which are corroborated, modified or rejected by further examples and non-examples.
This paper describes a natural language generation system known as VINCI, which accepts as input a formal description of some subset of a natural language, and generates strings in the language. With the help of an attribute grammar formalism, the system can be used to simulate on a computer components of several current linguistic theories. The program, implemented in C, runs under a variety of operating systems, including UNIX, MS-DOS and VM/CMS. In this paper we consider not only the design of the system, but also some of its applications in linguistic modelling and second language acquisition research.
Various objections are raised against current practice in co-occurrence analysis. The use of Yule's coefficient Y is then advocated.
Our work aims at the optimization of existing tools for computer-assisted description and analysis of textual data. More specifically, we have been involved in the thematic description of clauses and clause complexes of Quebec budget speeches from 1934 to 1960. Our main objective is to enhance the work already done in this direction by elaborating the analytic framework through a study of the thematic structure of these discourses. We first set out the general context of our work by briefly explaining the research project on political discourse under the Duplessis Regime in Quebec (1936–60) and giving a brief survey of the parsing strategy applied to the corpus. Second, we present the theoretical background of thematic analysis and the operational model that we are using here. Finally, we try to illustrate the relevance of such methodological work on research data.
This study reports on computer-aided investigation of salient differences in the essay idiolects of the Mexican writers Octavio Paz and Rosario Castellanos and suggests that some of them may be linked to gender. It describes use of ready-made software and computational strategies requiring no tagging and minimal ocular scan. It suggests some parameters that can be searched and in most cases quantified to explore characteristics posited by linguistic and literary scholars, taking into consideration the particular language and culture of the authors.
Proper Names (PNs) present a problem for the automatic processing and understanding of naturally occurring text. Due to their poor coverage in existing lexical resources and the continual appearance of new names, they represent a large body of unknown lexical data. Moreover, the complexity of the constructions in which they can appear and their own internal structure make them difficult to process, even if they are initially known. Yet the successful analysis of names is often crucial to the full understanding of a text. This paper proposes a solution to the problem and describes a natural language processing (NLP) system, FUMES, which makes use of the internal structure of names and the descriptive information that regularly accompanies them to produce lexical and knowledge base entries for unknown PNs. We present some preliminary results showing the viability of this approach for the identification of proper names.
Word sense disambiguation has been recognized as a major problem in natural language processing research for over forty years. Both quantitive and qualitative methods have been tried, but much of this work has been stymied by difficulties in acquiring appropriate lexical resources. The availability of this testing and training material has enabled us to develop quantitative disambiguation methods that achieve 92% accuracy in discriminating between two very distinct senses of a noun. In the training phase, we collect a number of instances of each sense of the polysemous noun. Then in the testing phase, we are given a new instance of the noun, and are asked to assign the instance to one of the senses. We attempt to answer this question by comparing the context of the unknown instance with contexts of known instances using a Bayesian argument that has been applied successfully in related tasks such as author identification and information retrieval. The proposed method is probably most appropriate for those aspects of sense disambiguation that are closest to the information retrieval task. In particular, the proposed method was designed to disambiguate senses that are usually associated with different topics.
BOOK NOTICES 865 Syllables, tones, and verb paradigms. (Studies in Chinantec languages, 4.) Ed. by William R. Merrifield and Calvin R. Rensch. Dallas: Summer Institute of Linguistics, 1990. Pp. vii, 130. Paper $10.00. This slim paperback contains six papers written in the 1970s, originally intended to comprise the first volume in a series on the Chinantec languages. (The Chinantec languages are spoken in a northern area of the Mexican state of Oaxaca.) Due to delays in publication, this book is instead the fourth volume in SIL' s Chinantec series. In tone and quality this collection resembles a set of departmental working papers. The expositions are sketchy in places, sometimes requiring a greater familiarity with Chinantec data than one can acquire from the article at hand. The editors' introduction does not explain why they have pulled together these particular papers, which have little in common as a set other than their focus on Chinantec. The authors often cite their own previous work, as well as the work of other authors in this book. For these reasons the publication seems targeted more for 'in-house' consumption than for the attention of linguists at large. 'Comaltepec Chinantec tone' (3-20), by Judi Lynn Anderson, Isaac H. Martinez, & Wanda Pace, discusses the interaction of tone, stress, and syllable structure, with particular attention to sandhi phenomena. A stressed syllable may bear one of seven different surface tone configurations: a level tone low, mid, or high, or a contour low-mid, low-high, high-mid, or high-low. In 'Comaltepec Chinantec verb inflection ' (21-62), Wanda Pace derives the surface tones from five underlying tones (L, M, H, LM, LH), discussing in addition some of the sandhi rules that give rise to the surface tone patterns. This analysis serves as an introduction to the complex verbal system, in which person, number, aspect, and lexical class are indicated largely by variations in tone, stress, and vowel length in the verbal root. In 'The Lealeo Chinantec syllable' (63-73), James E. Rupp relates the shape of the Chinantec syllable to tone. Calvin R. Rensch, in 'Phonological realignment in Lealeo Chinantec' (75-89), traces the development of certain features from ProtoChinantec to the Lealeo dialects. 'Quiotepec Chinantec tone' (91-105), by Richard Gardner & William R. Merrifield, presents the tonology of the Quiotepec dialect. Finally, in 'Moving and arriving in the Chinantla' (107-30), David O. Westley & William R. Merrifield describe the syntax and semantics of Chinantec verbs ofmotion, focussing on the intriguing way in which deixis is grammaticalized m the verbal inflection system. The exposition in some of these papers is weakened by the use of idiosyncratic descriptive devices. It is, of course, the norm for areal studies to have a distinct lingo; but then the editors of this sort of anthology owe it to the reader to footnote some of the less common descriptors early on in the book. The theoretical underpinnings of some analyses are also unclear, and this problem is exacerbated by some sloppy rule-writing (24) and other uninsightful attempts at 'formalizing' generalizations (70). Taken together, though, this collection of papers is a fairly good source for some fascinating Chinantec data. [Brian M. Sietsema, MerriamWebster Inc. and Westfield State College.] Bridges between psychology and linguistics: A Swarthmore Festschrift for Lila Gleitman. Ed. by Donna Jo Napoli and Judy Anne Kegl. Hillsdale, New Jersey: Lawrence Erlbaum, 1991. Pp. xii, 299. Lila Gleitman is well known among linguists for her research on language and cognition in blind and deaf children, 'motherese', and reading. The 14 articles in this volume honor her four years at Swarthmore College, where she founded linguistics and psycholinguistics in 1968. Almost all of the authors are Swarthmore alumni, most have studied under Gleitman, and several have included personal acknowledgements attesting to her influence on their careers. The papers are succinctly previewed in the Introduction (vii-xii), but the inappropriateness of the title's bridge metaphor soon becomes apparent. These diverse articles may form a continuum from psychology to linguistics, but they do not explicitly address links between the two. Most do not reflect Gleitman's particular research interests, and only three of her publications are cited in the entire book (a fourth is...
This article describes an intelligent computer-assisted language instruction system that is designed to teach principles of syntactic style to students of English. Unlike conventional style checkers, the system performs a complete syntactic analysis of its input, and takes the student's stylistic intent into account when providing a diagnosis. Named STASEL for Stylistic Treatment At the Sentence Level, the system is specifically developed for the teaching of style, and makes use of artificial intelligence techniques in natural language processing to analyze free-form input sentences interactively.
The paper sets out twenty proposals for the development and evaluation of Computer Assisted Language Learning (CALL) programs. These proposals emerge from special characteristics of language instruction and of the use of computers to assist in language instruction. We combine theoretically-based assumptions with empirical findings drawn from investigation of language courseware for Hebrew speakers in Israel. We first list four unique features of language instruction: (1) the object-language-meta-language distinction; (2) computer as written medium vs. language as primary spoken medium; (3) teaching of second language skills vs. linguistics; (4) the computer as an electronic tool vs. the computer as a cognitive entity simulating the speaker. We then show how these unique characteristics of language instruction (mother-tongue and foreign language) impose special proposals on language courseware. These proposals should be observed in the development of language courseware and in the evaluation of such programs. Clearly, these proposals integrate with general courseware proposals.
Click to increase image sizeClick to decrease image size Notes1. I wish to thank Dr Lynn Williams of the University of Exeter for reading a draft of this article and making several suggestions.2. The Guernica Statute of Autonomy (1979) had effect in the three Spanish Basque provinces of Alava, Guipúzcoa and Vizcaya which became the three members of the Basque Autonomous Community. One of the aims of the 1982 Ley de normalización del uso del euskera is to protect every speaker's right to use Basque in the spheres of administration, education and in all means of communication. In Navarra, the fourth Spanish Basque province, the co-official status of Basque was recognized by the Parlamento Foral Navarro in November 1980.3. I follow Haugen's (1966) terminology here and interpret ‘vernacular’ as an underdeveloped language in the functional sense. E. Haugen, ‘Dialect, language, nation’, American Anthropologist, LXVIII (1966), 922–35.4. E. B. Ryan, ‘Why do Low-Prestige Language Varieties Persist?’, in H. Giles and R. N. St Clair (eds.), Language and Social Psychology, (Oxford: Basil Blackwell, 1979), 145–57.5. Pedro de Yrizar, Contribución a la dialectología de la lengua vasca, I (Zarauz: Caja de Ahorros Provincial de Guipúzcoa, 1981).6. Bonaparte made five trips to various parts of the Basque Country between 1856 and 1869. His extensive research allowed a detailed classification of the regional varieties; one of his major contributions was the Carte des sept provinces basques montrant la délimitation actuelle de l’Euskara et sa division en dialectes, sous-dialectes et variétés, completed in 1863 and published in London in 1866.7. Bonaparte, 98.8. K. Rotaetxe, ‘La norma vasca: codificación y desarrollo’, Revista española de lingüística, XVII, 2 (1987) 219–44.9. P. Lafitte, Grammaire basque (Navarro-labourdin littéraire) (Bayonne: 1944)10. See K. Rotaetxe, 228—30, for an account of the grammatical and orthographic reforms carried out in the process of standardization and the extent to which they eliminated the characteristics of vizcaíno.11. Rotaetxe, 240–43.12. J. I. Olabuénaga et al., La lucha del euskara en la Comunidad Autónoma Vasca (Vitoria: Servicio Central de Publicaciones del Gobierno Vasco, 1983). This work is a presentation of the results of the 1981 Census and of extensive surveys concerning language competence, use and attitudes carried out in the Autonomous Community at the beginning of the 1980s.13. K. Rotaetxe does not make a distinction here between learning Basque and learning Batua. It is probably true that in the case of most learners Batua is the norm adhered to, and this would explain why those living in Vizcaya encounter difficulties when attempting to put their newly acquired language to use. It should be remembered, however, that some establishments in Vizcaya teach a standard form of vizcaíno. It is quite possible that the low percentage of successful learners in Vizcaya includes precisely those speakers who have been educated in this standard vizcaíno.14. The higher success rate in Alava could seem rather surprising if it is remembered that the Basque spoken there is similar to that spoken in Vizcaya and is classified as a variety of vizcaíno: it could be argued that learners in Alava will be faced by the same problems of linguistic distance from Batua as those confronting their counterparts in Vizcaya. It is very likely that this is so for learners in those areas of Alava where Basque is spoken by the majority of the population. However, in Alava as a whole, the awareness of the distance between vizcaíno and Batua is not as acute as it is in Vizcaya, since the Basque-speakers form a very small minority and do not constitute a strong vizcaíno-speaking community. In Vitoria, capital of Alava and of the Basque Autonomous Community, there is a higher percentage of Basque-speakers, but they have migrated from several parts of the Basque Country and do not form a linguistically homogeneous group. The polarization vizcaíno/Batua does not, therefore, occur to the same degree.15. The grammatical calques on Castilian are not so much an intrinsic feature of Batua as a reflection of the fact that most people who write in Basque also write in Castilian, and tend to translate from Castilian when writing in Basque.16. Whereas Batua, therefore, is sometimes criticized for reproducing Castilian syntax, there are some lexical items which illustrate how Batua uses native Basque formations whilst the regional dialects use Castilian loan-words (for example eskribatu, ‘to write’ of vizcaíno from the Castilian escribir is translated as idatzi in Batua). Some native Basque speakers regard these features of the Batua lexicon as excessively purist, whereas others seem to interpret them as indications of their own linguistic inadequacy. It is worth remembering, however, that the extent of Castilian influence on native Basque-speakers’ vocabulary is possibly not as great as they themselves sometimes claim.17. Inventario de arquitectura rural alavesa (Vitoria: Diputación Foral de Alava, 1981).18. Inventario, 184.19. he informant was given the freedom to complete the forms of the Padrón either in Basque (Batua) or in Castilian. The question dealing with language competence was presented in the following way in Castilian: Conocimiento de euskara Señale con una X su nivel de comprensión, habla, lectura y escritura. 1. Nada 2. Con dificultad 3. Bien 20. The form of categorization adopted in the presentation of the Padrón results is designed to give each informant a nivel global de euskara. The informant is classified according to his own evaluation of his competence in the four skills of Comprehension, Speaking, Reading and Writing. The Classifications are: Euskaldunes alfabetizados: individuals who understand, speak, read and write Basque well. Euskaldunes parcialmente alfabetizados: individuals who understand and speak Basque well, but read and write the language with difficulty. Euskaldunes no alfabetizados: individuals who understand and speak Basque well, but are unable to read and write the language. Cuasi-euskaldunes alfabetizados: individuals who understand Basque well or with difficulty, speak Basque with difficulty, and read and write well or with difficulty. Cuasi-euskaldunes no alfabetizados: individuals who understand Basque well or with difficulty, speak Basque with difficulty, but are unable to read and write the language. Cuasi-euskaldunes pasivos: individuals who understand Basque well or with difficulty, but are unable to speak the language. Erdaldunes: individuals who are unable to understand or speak Basque. Although an attempt has been made to allow for the fact that there is not, in the case of every speaker, a constant reduction in levels of competence from Comprehension to Speaking, Speaking to Reading, and Reading to Writing (that is, there is a recognition that some individuals will be more competent in written skills than in oral skills), the seven groupings referred to do not appear to cover all the permutations provided by the informants’ evaluations of their competence in the four skills.21. The variety of words used to refer to the Basque language can lead to confusion. Euskara batua (or Batua) refers to the standard norm. In Castilian, non-standardized dialectal forms are often referred to as vasco or vascuence.22. J. I. Ruiz Olabuénaga, Atlas lingüístico vasco (Vitoria: Servicio Central de Publicaciones del Gobierno Vasco, 1984); J. I. Ruiz Olabuénaga et al., La lucha del euskara.23. The informant was given the freedom to complete the questionnaire either in Basque (Batua) or in Castilian. (All interviews, however, were conducted in Castilian.) The question dealing with language competence was presented in the following way in Castilian:24. These age-divisions were decided upon for two reasons. For purposes of comparison it was important to respect the age-cohorts used in the presentation of the results of the Padrón. At the same time the divisions attempt to take into account some of the historical factors which have affected the use and acquisition of Basque.25. When the interview was conducted in the absence of other Basque-speakers it is very unlikely that the informant felt that he was being tested: my very limited knowledge of Basque did not allow me to challenge his evaluations. However, the fact remains that it was probably more difficult for an informant to stretch the truth when confronted with the researcher than when allowed to complete the questionnaire in privacy.26. Olabuénaga et al., La lucha, 28.27. I am grateful to I. Agote (Política lingüística, Gobierno Vasco) for suggesting the possibility of adapting these evaluation methods to sociolinguistic surveys.28. J. L. M. Trim, Developing a Unit/Credit Scheme of Adult Language Learning (Oxford: Pergamon, 1980).29. M. Oskarsson, Approaches to Self-assessment in Foreign Language Learning (Oxford: Pergamon, 1980).30. When asked in my survey to state which was their first language, 11 of the 75 informants answered ‘Castilian’, 58 answered ‘Basque’ and six answered ‘Both Castilian and Basque’. Those who have Castilian as their native language are most likely to be adults who have moved to Aramayona from other parts of Spain. In some cases their offspring, although competent Basque speakers because of their exposure to the language at school, will claim Castilian to be their native tongue.31. Padrón municipal de habitantes de la Comunidad Autónoma de Euskadi 3. Educación y euskara (Vitoria: Instituto Vasco de Estadística [EUSTAT], 1988).32. It must be remembered that these tables have been compiled from questions which were originally presented to the informants in Basque and Castilian. See Notes 12 and 16 for the original terms used. Note that in the Padrón the levels of competence followed the order Nada, Con dificultad and Bien, whereas in my survey the order was the reverse (Fácilmente, Con dificultad and Nada). In order to ensure consistency, I have arranged the results of the Padrón in Table 7 in the order Bien, Con dificultad and Nada.
The word senses in a published dictionary are a valuable resource for natural language processing and textual criticism alike. In order that they can be further exploited, their nature must be better understood. Lexicographers have always had to decide where to say a word has one sense, where two. The two studies described here look into their grounds for making distinctions. The first develops a classification scheme to describe the commonly occurring distinction types. The second examines the task of matching the usages of a word from a corpus with the senses a dictionary provides. Finally, a view of the ontological status of dictionary word senses is presented.
Blanchet, Philippe - Remarks on "Peuchère" or the role of the imaginary in the evolution of languages in Provence (additions to the article 11 Aunt Portal's complex". Rousselot looks at the effect of the idealization of the norm in the shift to French by the South-Eastern bourgeoisie ("portalism"). This analysis must be tempered by saying that adapting the lexical items of Provençal to regional French can be explained by other mechanisms and this is also true of other processes (the "accent"). On the other hand, idealization of the Occitan norm ("à la française") and systematization of the regional French (Francitan), has hardly spread in Provence where regional French has taken on great importance as a vehicle of identity and prestige, while provençal is resisting the force of a coercitive norm.
The present paper summarizes the major methods and results of the multi-dimensional approach to genre variation. The approach combines the resources of computational tools, large text corpora, and multivariate statistical tools (such as factor analysis and cluster analysis). It has been used to address issues such as the relations among spoken and written genres in English, and the historical development of genres and styles. The approach has also been applied to other languages; in this regard it has been used to address broader theoretical issues, such as the extent to which genre, and style variation are comparable cross-linguistically, and the linguistic consequences of literacy.
Three models for word frequency distributions, the lognormal law, the generalized inverse Gauss-Poisson law and the extended generalized Zipf's law are compared and evaluated with respect to goodness of fit and rationale. Application of these models to frequency distributions of a text, a corpus and morphological data reveals that no model can lay claim to exclusive validity, while inspection of the extrapolated theoretical vocabulary sizes raises doubts as to whether the urn scheme with independent trials is the correct underlying model for word frequency data. The role of morphology in shaping word frequency distributions is discussed, as well as parallelisms between vocabulary richness in literary studies and morphological productivity in linguistics.
Hebrew Studies 33 (1992) 119 Reviews Though nagging. these problems do not undermine the values and virtues of the study. What a pleasure to read! Darr writes with clarity, economy. and substance. She knows scholarship and how to teach it. Synagogues. churches. and introductory college courses can benefit from her work. Not least. the graciousness of its demeanor and the generosity of its vision are a welcome gift in an age of verbal assault. Phyllis Trible Union Theological Seminary New York. NY 10027 HEBREW LINGUISTICS: A JOURNAL FOR HEBREW DESCRIPTIVE, COMPUTATION AL, AND APPLIED LINGuIsTIcs. No. 31-32. Maya Fruchtman, ed. Pp ix + 114. Ramat-Gan. Israel: Bar-Han University, 1991. Paper. This double issue contains six articles (and a response to one of them) in Hebrew. with English abstracts and one short correction-note in English regarding the inappropriateness of characterizing Jewish languages as pidgins/creoles. The emphasis is on computational linguistics and discourse analysis. The issue opens with an article by Michal Ephratt, which proposes an algorithm for recognizing linguistic jokes, such as puns, which rely on multiple interpretation of sentences, where the "punch" is attained by the gap between the normal, "least costly" reading and the least expected. "most costly" one. Hanna David, and Hillel Weiss in his comments on her article, discuss a more general computational issue involving the application of complex mathematical models to analyze and characterize literary texts versus the use of the computer as a tool for (a) storing extensive. complex bodies of literary corpora. and (b) subsequent pulling out data that are relevant to precise determination of literary hypotheses. Weiss points out that use of mathematical algorithms in computational literary analysis has become marginal and that most computational work today centers on the building up of extensive literary data bases. At the same time. he outlines his own computational algorithm for distinguishing between poetry and prose. Hebrew Studies 33 (1992) 120 Reviews Zahava Goldstein and Michael Moore test a mathematical model for predicting active vocabulary on Hebrew-speaking children. In Israel, evaluation of vocabulary has always been performed by sampling words out of a dictionary and asking what they mean-an unsophisticated, often misleading procedure. The model applied here provides for reliable prediction of individuals' active vocabulary based on frequency of distribution in written or spoken samples. Yitzhak Zadka's note is a comment on an earlier article in Hebrew Linguistics, which discusses the modal meaning of ~eyn ~el mi lifnot ("there is nobody to tum to"). Zadka points out that the modality of this structure is not restricted to existential sentences of this type; rather, it covers a variety of patterns involving an attributive infinitive. Yitzhak Roeh and Raphael Nir demonstrate how Israeli news discourse tends to employ indirect speech to assure "objectivity," while keeping direct speech transmission to a minimum. It is therefore of particular interest to study partial deviations from the indirect speech standard in the news, as manifest in the use of direct speech elements in indirect speech or in mimetic direct speech. Such departures are permitted only to the extent that their effects (whether empathy, respect, etc., or suspicion, irony, etc.) conform with the "national consensus" and mainstream ideology, reflecting notions of "appropriateness," norms and values-and their hierarchical ranking. Lea Sarig shows how discourse analysis can account for linguistic dissimilarities emerging from comparison of a translation with its source text. "Discrepancies" in employing means of cohesion in translation from Arabic to Hebrew can be attributed to the translator's attempt to improve cohesion of discourse in the target language. Thus, grammatical anaphora may be replaced by lexical repetition/variation when the distance between the antecedent and the anaphor is substantial; the opposite conversion may occur when repetition is felt to be redundant. Connectives may be deleted, replaced by other grammatical items, or added, depending on the translator 's sense of optimal cohesion in the target language. The strength of Hebrew Linguistics continues to be in its interdisciplinary nature and in its offering Hebraists and scholars from other disciplines who wish to contribute to Hebrew language study a forum for discussion and exchange. With the general increase in interdisciplinary research and the "coming of age" of the...
Confirmatory factor analysis was used to test the structure of 5-item affect rating scales designed to measure positive affect and negative affect. A proposed circumplex affect structure was the source of scales constructed to represent a cluster of positive terms, including pleasantness and activation; the negative terms represented anxiety, depression, and hostility. The hypothesized simple-structured positive and negative trait affect factors, with a moderate correlation between them, were found in all cases. Equivalent structure was confirmed for younger adults, middle-aged, and older adults of good health and above-average education. Although the hypothesized simple-structured positive and negative factors emerged for all other groups, three other tests of factor equivalence failed to be confirmed: trait and state factors in the older adult group were not identical. Factors derived from healthy and frail elders were structurally different. Variability among frail elders and variability over 30 days within the same person, when factored, also showed nonequivalence. Although the scales are extremely useful in assessing affect, comparisons across some subject groups should be made with caution.
OBJECTIVE: Since previous work indicated smaller than normal temporal lobe structures in schizophrenic patients, the authors tested the hypothesis that this abnormality might be reflected in abnormally large sylvian fissures. METHOD: The subjects were 48 schizophrenic patients and 51 normal comparison subjects matched groupwise with regard to age and sex. CSF spaces (sylvian fissures, temporal lobe sulci, temporal horns, third ventricle, lateral ventricles, and superficial cerebral sulci) were visually assessed with the magnetic resonance imaging rating protocol of the Consortium to Establish a Registry for Alzheimer's Disease (CERAD). RESULTS: The sylvian fissures of the schizophrenic patients were found to be bilaterally wider than those of the comparison subjects. There were no other significant differences. CONCLUSIONS: Schizophrenic patients appear to have larger than normal sylvian fissures, which may reflect smaller superior temporal gyri.
This paper describes the lexical database tool LOLA (Linguistic-Oriented Lexical database Approach) which has been developed for the construction and maintenance of lexicons for the machine translation system LMT. First, the requirements such a tool should meet are discussed, then LMT and the lexical information it requires, and some issues concerning vocabulary acquisition are presented. Afterwards the architecture and the components of the LOLA system are described and it is shown how we tried to meet the requirements worked out earlier. Although LOLA originally has been designed and implemented for the German-English LMT prototype, it aimed from the beginning at a representation of lexical data that can be reused for other LMT or MT prototypes or even other NLP applications. A special point of discussion will therefore be the adaptability of the tool and its components as well as the reusability of the lexical data stored in the database for the lexicon development for LMT or for other applications.
To date, no fully suitable data model for lexical databases has been proposed. As lexical databases have proliferated in multiple formats, there has been growing concern over the reusability of lexical resources. In this paper, we propose a model based on feature structures which overcomes most of the problems inherent in classical database models, and in particular enables accessing, manipulating or merging information structured in multiple ways. Because of their widespread use in the representation of linguistic information, the applicability of feature structures to lexical databases seems natural, although to our knowledge this has not yet been implemented. The use of feature structures in lexical databases also opens up the possibility of compatibility with computational lexicons.
B.A. Uspenskij's theory that there was Church Slavonic-Russian diglossia in medieval Rus' has provoked a great deal of controversy in the last few years, stimulated in particular by the publication of his Istorija russkogo literaturnogo jazyka (XI-XVII w.) in 1987. 1 The diglossie approach, which has won its share of supporters as well as detractors, differs from previous, genetically oriented treatments of the early Russian linguistic situation in that it attempts to reconstruct the normative thought of medieval bookmen. By stressing the role of linguistic consciousness, the proponents of diglossia make explicit and important claims about the pragmatic aspects of language use in medieval Rus' in particular, about how writers perceived and utilized linguistic variation. In this article, I will be concerned chiefly with these pragmatic ramifications of Uspenskij's approach. I will begin by discussing the main points of the diglossie theory and will then examine certain aspects of it that seem to be in conflict with the observable characteristics of early Russian written sources.
This paper is concerned with the question of how to extract lexical knowledge from Machine-Readable Dictionaries (MRDs) within a lexical database which integrates a lexicon development environment. Our long term objective is the creation of a large lexical knowledge base using semiautomatic techniques to recover syntactic and semantic information from MRDs. In doing so, one finds that reliance on a single MRD source induces inadequacies which could be efficiently redressed through access to combined MRD sources. In the general case, the integration of information from distinct MRDs remains a problem hard, perhaps impossible, to solve without the aid of a complete, linguistically motivated database which provides a reference point for comparison. Nevertheless, advances can be made by attempting to correlate dictionaries which are not too dissimilar. In keeping with these observations, we describe a software package for correlating MRDs based on sense merging techniques and show how such a tool can be employed in augmenting a lexical knowledge base built from a conventional MRD with thesaurus information.