Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
Most work in statistical parsing has focused on a single corpus: the Wall Street Journal portion of the Penn Treebank. While this has allowed for quantitative comparison of parsing techniques, it has left open the question of how other types of text might a#ect parser performance, and how portable parsing models are across corpora. We examine these questions by comparing results for the Brown and WSJ corpora, and also consider which parts of the parser's probability model are particularly tuned to the corpus on which it was trained. This leads us to a technique for pruning parameters to reduce the size of the parsing model. 1
We present a stochastic parsing system consisting of a Lexical-Functional Grammar (LFG), a constraint-based parser and a stochastic disambiguation model. We report on the results of applying this system to parsing the UPenn Wall Street Journal (WSJ) treebank. The model combines full and partial parsing techniques to reach full grammar coverage on unseen data. The treebank annotations are used to provide partially labeled data for discriminative statistical estimation using exponential models. Disambiguation performance is evaluated by measuring matches of predicate-argument relations on two distinct test sets. On a gold standard of manually annotated f-structures for a subset of the WSJ treebank, this evaluation reaches 79% F-score. An evaluation on a gold standard of dependency relations for Brown corpus data achieves 76% F-score.
Bien loin que l'objet precede le point de vue, on dirait que c'est le point de vue qui cree l'objet (Ferdinand de Saussure 1916:23).Traditionally, historiographers of Afrikaans have argued that a relatively uniform (spoken) vernacular existed at the Cape from the late eighteenth century, where it constituted the L(ow) variety in a diglossic situation (see Raidt 1991). Standardization of Afrikaans has been described accordingly in a “naturalistic” fashion as the codification and elaboration of this preexistent vernacular. This paper summarizes the results of a variationist study of late nineteenth- and early twentieth-century documents that shows the traditional view to be seriously flawed. While “Cape Dutch Vernacular” (or “Afrikaans,” as it came to be called) was a well-defined entity in the popular consciousness from the mid-nineteenth century, the actual patterns of language use in the historical texts indicate the existence of a complex social dialect continuum until the early twentieth century. Variation patterns described for the late eighteenth century are shown still to be productive around 1900. Linguistic standardization, understood as the reduction of variation and the emergence of a linguistic norm, was rapid and strongly marked by cultural and political nationalism.* I should like to thank Roger Lass and Paul Roberge for many stimulating discussions about Afrikaans historical linguistics and for their comments on earlier versions of the work presented in this article. The responsibility for any mistakes remains of course my own.
International audience
This paper presents a bottom-up generator that makes use of Information Retrieval techniques to rank potential generation candidates by comparing them to a data base of stored instances. We introduce two general techniques to address the search problem, expectation-driven search and dynamic grammar rule selection, and present the architecture of an implemented generation system called IGEN. Our approach uses a domain-specific generation grammar that is automatically derived from a semantically tagged treebank. We then evaluate the efficiency of our system.
This study provides the information of reliability and implications of narrative essay tests for measuring achievement motive as a part of employee selection processes. To develop the key achievement motive explicit and concrete rating criterion, the TAT scoring method was applied to data (n=100) of essay tests gathered from seven raters. Three raters were chosen from entry level workers and the other four were professional writers of verbal testing items, forming two contrast groups. The reliability of the achievement motive ratings was calculated for each group by the interrater reliability approach, resulting that there was no significant difference between the groups. Coefficients of each group's achievement motive, general mental ability and personality traits were calculated, suggesting the possibility that the achievement motive ratings thus derived from essay tests implies individuality.
Although research has examined differences in people’s responses to natural versus technologically caused disasters, research has not examined the differences in people’s attitudes toward disasters in natural versus built environments. This study examined the effects of the type of environment and awareness of the problem on attitudes toward the cleanup of oil spills. The results showed that the type of environment did not affect ratings of the importance of the environmental problem or how it should be cleaned up. However, people were more concerned about the environmental and community impacts of the cleanup process in the built environment. Awareness of the problem was a more important factor than type of environment for understanding attitudes toward the oil spill cleanup. People who were more aware of the oil spills viewed the problems as more important and were more concerned that the environments be returned to their previous states.
The authors describe a densely amnesic man who has acquired explicit semantic knowledge of famous names and vocabulary words that entered popular culture after the onset of his amnesia. This new semantic knowledge was temporally graded and existed over and above the implicit memory he demonstrated in reading speed and accuracy, familiarity ratings, and his ability to make correct guesses on unfamiliar items. However, his postmorbid knowledge was limited to verbal labels denoting famous people and words; he possessed virtually no explicit knowledge of the meaning of these words or the identities of these individuals, although there was some evidence that some of this information had been acquired at an implicit level. Findings are discussed in the context of a neural network model (J. L. McClelland, B. L. McNaughton, & R. C. O'Reilly, 1995) of semantic acquisition.
espanolEn este articulo se examina el uso de ir a + infinitivo y del futuro en -re en las doce ciudades que se incluyen en el Macrocorpus de la norma linguistica culta de las principales ciudades del mundo hispanico. El analisis cuantitativo proporcionara la frecuencia de empleo de la perifrasis y de la forma en -re cuando se utilizan para expresar futuridad, asi como las diferencias que puedan existir entre las distintas ciudades. Tambien se estudiara la incidencia del tipo de oracion, el sexo y la edad en el uso de ambas formas. Finalmente, el analisis de regresion multiple nos permitira establecer la probabilidad de que los condicionantes analizados condicionen de manera significativa la eleccion de las dos formas de futuro. EnglishThis paper examines the use of the verb group 'ir a + infinitive' and the future-tense form ending in '-re' in the twelve cities included in the Macrocorpus de la norma linguistica culta de las principales ciudades del mundo hispanico (The Macrocorpus of the cultural linguistic norm of the majar cities of the Hispanic World). The quantitative analysis will provide results conceming how often the verb group 'ir a + infinitive' and the '-re' form are employed to express future time, as well as the differences between the cities reviewed. Sentence-type, sex and age will be other variables whose frequency will also be studied with respect to the use of both forms. Finally, the multiple regression analysis will allow us establish how likely the aforementioned variables will significantly condition the choice of these future-tense forms.
AIM: In this paper the balance of affective and instrumental communication employed by nurses during the admission interview with recently diagnosed cancer patients was investigated. RATIONALE: The balance of affective and instrumental communication employed by nurses appears to be important, especially during the admission interview with cancer patients. METHODS: For this purpose, admission interviews between 53 ward nurses and simulated cancer patients were videotaped and analysed using the Roter Interaction Analysis system, in which a distinction is made between instrumental and affective communication. RESULTS: The results reveal that more than 60% of nurses' utterances were of an instrumental nature. Affective communication occurred, but was more related to global affect ratings like giving agreements and paraphrases than to discussing and exploring actively patients feelings by showing empathy, showing concern and optimism. CONCLUSION: In future, nurses should be systematically provided with (continuing) training programmes, in which they learn how to communicate effectively in relation to patients' emotions and feelings, and how to integrate emotional care with practical and medical tasks.
Grammars are valuable resources for natural language processing. We divide the process of grammar development into three tasks: selecting a formalism, defining the prototypes, and building a grammar for a particular human language. After a brief discussion about the first two tasks, we focus on the third task. Traditionally, grammars are built by hand and there are many problems with this approach. To address these problems, we built two systems that automatically generate grammars. The first system (LexOrg) solves two major problems in grammar development: namely, the redundancy caused by the reuse of structures in a grammar and the lack of explicit generalizations over the structures in a grammar. LexOrg takes several types of specification as input and combines them to automatically generate a grammar. The second system (LexTract) extracts Lexicalized Tree Adjoining Grammars (LTAGs) and Context-free Grammars (CFGs) from Treebanks, and builds derivation trees that can be used to train statistical LTAG parsers directly. In addition to creating Treebank grammars and producing training materials for parsers, LeXTract is also used to evaluate the coverage of existing hand-crafted grammars, to compare grammars for different languages, to detect annotation errors in Treebanks, and to test certain linguistic hypotheses. LexOrg and LeXTract provide two different perspectives on grammars. In LexOrg, elementary trees in an LTAG grammar are the result of combining language specifications such as tree descriptions. In LeXTract, elementary trees are building blocks of syntactic structures in a Treebank. LexOrg makes explicit the language specifications that form elementary trees, whereas LeXTract makes explicit the elementary trees that form syntactic structures. The systems provide a rich set of tools for language description and comparison that greatly enhances our ability to build and maintain grammars and Treebanks effectively.
This paper presents a statistical parser for a wide-coverage Combinatory Categorial Grammar (CCG) derived from the Penn Treebank. The Treebank is translated to a corpus of canonical CCG derivations. We de ne a generative statistical model over CCG derivations and train it on the transformed Treebank.
This paper describes a wide-coverage statistical parser that uses Combinatory Categorial Grammar (CCG) to derive dependency structures. The parser differs from most existing wide-coverage treebank parsers in capturing the long-range dependencies inherent in constructions such as coordination, extraction, raising and control, as well as the standard local predicate-argument dependencies. A set of dependency structures used for training and testing the parser is obtained from a treebank of CCG normal-form derivations, which have been derived (semi-) automatically from the Penn Treebank. The parser correctly recovers over 80% of labelled dependencies, and around 90% of unlabelled dependencies.
Manual, large scale (computational) grammar development is time consuming, expensive and requires lots of linguistic expertise. More recently, a number of alternatives based on treebank resources (such as Penn-II, Susanne, AP treebank) have been explored. The idea is to automatically ``induce'' or rather read off (P)CFG grammars from the parse annotated treebank resources and to use the treebank grammars thus obtained in (probabilistic) parsing or as a starting point for further grammar development. The approach is cheap, fast, automatic, large scale, ``data driven'' and based on real language resources.\n\nTreebank grammars typically involve large sets of lexical tags and non-lexical categories as syntactic information tends to be encoded in monadic category symbols. They feature flat rules (trees) that can ``underspecify'' attachment possibilities. Treebank grammars do not in general follow Xbar architectural design principles (this is not to say that treebank grammars do not have design principles). As a consequence, treebank grammars tend to have very large CFG rule bases (e.g. Penn-II > 17,000 CFG rules for about 1 million words of text) with often only minimally differing rules. Even though treebank grammars are large, they are still incomplete, exhibiting unabated rule accession rates. From a grammar engineering point of view, the size of the rule base poses problems for maintainability, extendability and, if a treebank grammar is to be used as a CF-base in a LFG grammar, for functional (feature-structure) annotations. From the point of view of theoretical linguistics, flat treebank trees and treebank grammars extracted from such trees do not express linguistic generalisations. From the perspective of empirical and corpus linguistics, flat trees are well-motivated as they allow underspecification of subtle and often time consuming attachment decisions. Indeed, it is sometimes doubted whether highly general Xbar schemata usefully scale to ``real'' language.\n\nIn previous work we developed methodologies for automatic feature-structure annotation of grammars extracted from treebanks. Automatic annotation of ``raw'' treebank grammars is difficult as annotation rules often need to identify subsequences in the RHSs of flat treebank rules as they explicitly encode head, complement and modifier relations. Xbar based CFG rules should substantially facilitate automatic feature-structure annotation of grammar rules.\n\nIn the present paper we conduct a number of experiments to explore a space of possible grammars based on a small fragment of the AP treebank resource. Starting with the original treebank fragment we automatically extract a CFG G. We then apply an automatic structure preserving grammar compaction step which generalises categories in the original treebank fragment and reduces the number of rules extracted, resulting in a generalised treebank fragment and in a compacted grammar Gc. The generalised fragment is then manually corrected to catch missed constituents (and the like) resulting in an automatically extracted, compacted and (effectively manually) corrected grammar Gc,m. Manual correction proceeds in the ``spirit'' of treebank grammars (we do not introduce Xbar analyses). We then explore how many of the manual correction steps on treebank trees can be achieved automatically. We develop, implement and test an automatic treebank ``grooming'' methodology which is applied to the generalised treebank fragment to yield a compacted and automatically corrected grammar Gc,a. Grammars Gc,m and Gc,a are very similar to compiled out ``flat'' LFG-82 style grammars. We explore regular expression based compaction (both manual and automatic) to relate Gc,m to a LFG-82 style grammar design. Finally, we manually recode a subsection of the generalised and manually corrected treebank fragment into ``vanilla-flavour'' XBar based trees. From these we extract a compacted, manually corrected, XBar based grammar Gc,m,x. We evaluate our grammars and methods using standard labelled bracketing measures and according to how well they perform under automatic feature-structure annotation tasks.
The Japanese language, as spoken by native Japanese, has undergone a tremendous change in recent years-more noticeably in its spoken than in its written aspect. Some people brought up in the good old days deplore this, attributing it to the ignorance or negligence of decorum in speech among the younger generation. But the fact is that more and more adults are finding themselves unwittingly committing errors in usage, which they were formerly trained at school to avoid by all means. Among these deviations from linguistic norms, there are some which look likely to be established as perfectly acceptable usage, no matter whether one favors or disfavors them. Among them the most easily observable are: (1) recurrence of a rising intonation in mid-sentences, (2) omission of a morpheme, a word, or even a phrase, which was once considered indispensable in correct usage, (3) recurrence of what looks like an "empty" (semantically meaningless) word, (4) recurrence of what amounts almost to a cliche, and (5) prevalence of "feminine" (often infantile) language over strong "masculine" language, particularly in dialogs. What these phenomena reflect is, in the view of this paper-writer a kind of enervation in the verbal culture of the Japanese in general.
Finding simple, non-recursive, base noun phrase is an important step for many natural language processing applications. This paper presents a new corpus-based approach using decision tree for that purpose. In contrast to previous methods for Base NP identification, we adopt a decision tree trained from Penn Treebank to identify Base NP. And a self-learning mechanism is further integrated into our model. Experimental results show good performances using our method. The method can also be applied to processing of any other language.
The situation with regard to Russian language resources is fragmented and disorganized. For this reason, it is important to promote for Russian the development of its basic resources in one package that could be used for development of speech products. The paper presents a design of the Russian lexical databases, corpora and supporting tools (system for construction and support of lexical databases, system for transcription, morphological analyzer and normalyzer) developed for wide usage in speech engineering.
This document describes the syntactic bracketing guidelines for the Penn Korean Treebank, which is an online corpus of Korean texts annotated with morphological and syntactic information. The corpus consists of around 54,000 words and 5,000 sentences. The Treebank uses a phrase structure style of annotation, making head/phrasal node distinctions, argument/adjunct distinctions, and identifying empty arguments and traces for moved constituents. This document is organized as follows. In section 2, the basic syntactic ingredients of a clause structure are presented. Some notational conventions are introduced in section 3, including different types of syntactic tags, such as head level tags, phrase level tags and function tags used in the Treebank. In section 4, the bracketing guidelines for various types of clauses are discussed, including simple clauses, subordinate clauses, and clauses with coordination. Several types of subcategorizaion frames found in the Treebank are then presented in section 5, followed by bracketing guidelines for various linguistic phenomena in sections 6 to 21, including guidelines for annotating punctuation. The document ends with guidelines for handling some bracketing ambiguities and for handling some confusing examples.
In statistical parsing, the probabilistic models are used to evaluate the possibility of each candidate parse tree, where the parse tree with the largest probability is deemed to be the final result of the parsing. Therefore, the core of statistical parsing is a probabilistic evaluation model. The main difference among the various probabilistic evaluation models lies in which types of features in the context are used to assign the probabilities to the parse trees. Various probabilistic evaluation models have been proposed in the field of statistical parsing, where different models use different feature types. How to evaluate a feature type's predictive power for the parsing tree? The paper proposes an information theory based feature type analysis model. Using the method, we can quantitatively analyze the power of different feature types for syntactic structure prediction from the viewpoint of information theory. The basic idea is that we use entropy and conditional entropy to measure whether a feature type grasps some of the information for syntactic structure prediction. If the average uncertainty of the syntactic structures declines apparently, the feature type is deemed to have grasped some intrinsic linguistic information in the context that has close relation to the syntactic structure. Using Penn Treebank as training and testing set, our experiment quantitatively analyze the different feature types' predictive power for syntactic structure predictive power for syntactic structure prediction in a systematic way and draws a series of conclusions which reflect the predictive power of different feature types and feature type combination for syntactic parsing.
Natural language processing technologies offer ease-of-use of computers for average users, and ease-of-access to on-line information. Natural language, however, is complex, and the traditional methods of parsing with a single grammar and parser may result in an inefficient and large system that is difficult to maintain and is fragile in dealing with language irregularities. This paper begins by reviewing an alternative effort in grammar decomposition (also known as grammar partitioning) for natural language parsing, which aims to alleviate these problems. We then propose a novel automatic approach for grammar partitioning, in comparison with a random method of partitioning. Our experiments show that syntactic GLR parsing is formidable for the Wall Street Journal corpus in the Penn Treebank when a single grammar is used. This is due to too many grammar rules for parsing table generation. However, grammar partitioning solves the problem and offers a viable alternative. Our results also show that our automatic grammar partitioning method based on the mutual information criterion fares better than a random partitioning method and exhibits efficiency in parsing as well as high parse coverage. 1
Low interrater reliability coefficients are a common problem for behavior rating scales. One hypothesis to account for this is that raters have different frames of reference from which to judge behaviors. In the present study, the interrater reliability of the Devereux Behavior Rating Scale-School Form was examined, and the hypothesis that teacher frame of reference influences ratings was explored. Special and general education teachers rated the behavior of 51 children with emotional disturbance (ED), and general education teachers inde pendently rated the behavior of 51 matched control children. Interrater reliability coeffi cients were higher for the general education sample than for the sample of children with ED. Limited support was found for the hypoth esis that frame of reference may affect ratings. Findings suggest that many factors influence ratings and that teachers may benefit from rater training.
The stability over time of serum IgG antibody levels to human papillomavirus type 16 (HPV-16) was determined by comparing the HPV-16 capsid antibody levels in serial serum samples of an age-stratified random subsample of 1656 primiparous mothers resident in Helsinki who were followed until their second pregnancy, on average 29.5 months later. The correlation between the first and second pregnancy HPV-16 serum antibody levels of the same woman was high, even when >4 years had elapsed between pregnancies (r =.822). Between negativity, indeterminate results, or quartiles of positivity, the predictive values for being classified in the same category on both occasions ranged between 42% and 91%. Correlation coefficients, predictive values, and kappa coefficients between serial samples all were comparable with those of repeat analyses of the same sample, indicating that HPV capsid antibody levels are generally stable during several years of follow-up.
We study the impact of richer syntactic dependencies on the performance of the structured language model (SLM) along three dimensions: parsing accuracy (LP/LR), perplexity (PPL) and word-error-rate (WER, N-best re-scoring). We show that our models achieve an improvement in LP/LR, PPL and/or WER over the reported baseline results using the SLM on the UPenn Treebank and Wall Street Journal (WSJ) corpora, respectively. Analysis of parsing performance shows correlation between the quality of the parser (as measured by precision/recall) and the language model performance (PPL and WER). A remarkable fact is that the enriched SLM outperforms the baseline 3-gram model in terms of WER by 10% when used in isolation as a second pass (N-best re-scoring) language model.
Se presentan los valores normativos de la adapta The norms of the Spanish adaptation for shows 9 ci6n espanola de los conjuntos 9-14 (segunda parte) del 14 (second part) of the International Affective Picture Intemational Affective Picture System (lAPS). Los resul System (lAPS) are presented. The results are highly tados muestran una alta consistencia con los obtenidos consistent with those obtained in the first part of the en la primera parte de la adaptaci6n espanola y con los Spanish adaptation and in the original USA version. The valores originales norteamericanos. La distrlbuclon de picture distribution in the bi-dimensional space, defined las diapositivas en el espacio bidimensional valencia by the ratings of valence and arousal, displays the typical arousal adopta la tipica forma de boomerang, obssrvan boomerang form. In addition, a more pronounced slope dose una menor inclinaci6n, junto con una mayor disper -and a smaller dispersion- is observed in the si6n, en el brazo que se extiende hacia el polo agrada unpleasantextreme of the boomerang than in the pleasant ble que en el brazo que se extiende hacia el polo des one. The correlations between the Noth-American and agradable. Las correlaciones entre las evaluaciones the Spanish values are all highly significant. Nevertheless, norteamericanas y espanolas son todas altamente sig the differences found in the first part of the study are nificativas. No obstante, se confirman las diferencias confirmed: Spanish people perceive the affective pictures, encontradas en la primera parte del trabajo, en el sen as a whole, as more arousing and less dominant than tido de que los aspafioles perciben las imagenes North-Americans. Similarly, our results confirm the gender afectivas, en su conjunto, con mayor nivel de activaci6n differences previously found: women rate the pictures y menor nivel de control que los norteamericanos. Asi as more arousing and less dominant than men. Although mismo, sa confirman las diferencias de genero encon no significant differences are observed in valence ratings, tradas anteriormente: las mujeres otorgan a las image as a whole, the pictures evaluated as more pleasant are nes un mayor nivel de acnvaclon y un nivel menor de clearly different for men and women. The implications of control que los varones. Aunque no exlsten diferenclas the results regarding the theoretical model underlying the significativas en las estimaciones globales de la valencia lAPS are highlighted. afectiva, las imagenes evaluadas como mas agradables por varones y mujeres son claramente diferentes. Las Key words: emotion, affective valence, arousal, implicaciones de los resultados con respecto al modelo dominance, cross-cultural differences, gender te6rico que subyace al lAPS son resaltadas. differences.
Chunk parsing has focused on the recognition of partial constituent structures at the level of individual chunks. Little attention has been paid to the question of how such partial analyses can be combined into larger structures for complete utterances. Such larger structures are not only desirable for a deeper syntactic analysis. They also constitute a necessary prerequisite for assigning function-argument structure.The present paper offers a similarity-based algorithm for assigning functional labels such as subject, object, head, complement, etc. to complete syntactic structures on the basis of prechunked input.The evaluation of the algorithm has concentrated on measuring the quality of functional labels. It was performed on a German and an English treebank using two different annotation schemes at the level of function-argument structure. The results of 89.73 % correct functional labels for German and 90.40 % for English validate the general approach.
We examine an everyday Caribbean oral gesture, kiss-teeth or (KST), exploring previously-unresolved problems of meaning. Such forms are as examples of African cultural continuity across the Diaspora, often overlooked despite continuing interest in historical links between Caribbean Creoles and African communication systems. Forms such as (KST) are typically treated as lexical items: dictionary entries provide overlapping lists of emotions or affective states (eg, “scorn, impatience”) for each of several entries (suck-teeth, chups, etc.). Such approaches are inadequate, as the meaning of (KST) is not a single semantic unit, while lists are incomplete, contingent and inadequate. We distinguish ideophones from metalinguistic labels; consider geographical distribution and diffusion with respect to both functions and particular forms; and analyze related signs as a set, with reference to shared pragmatic function. (KST) is an inherently evaluative and inexplicit oral gesture with a sound-symbolic component, and a remarkably stable set of functions across the Diaspora: an interactional resource with multiple possibilities for sequential organization, often used to negotiate moral positioning among speakers and referents, and closely linked to community norms and expectations of conduct and attitude. It participates in a system of indirect discourse, requiring co-construction of intention by speaker and hearers. Moreover, it functions in personal narratives to mark both internal and external evaluation, sometimes ambiguously. Each of the proposed functions is illustrated with data ranging from historical to contemporary, oral to literary, monologic to interactional. Esther Figueroa & Peter L Patrick
This study presents normative data for the Speed and Capacity of Language Processing (SCOLP) testfrom an older American sample. The SCOLP comprises 2 subtests: Spot-the-Word, a lexical decision task, providing an estimate of premorbid intelligence, and Speed of Comprehension, providing a measure of information processing speed. Slowed performance may resultfrom normal aging, brain damage (e.g., head injury), or dementing disorders or may represent the intact performance of someone who always performed at the low end of normal. The SCOLP enables the clinician to differentiate between these possibilities. Adequate age-appropriate norms to differentiate dementia from normal aging do not exist. We present data from 424 older community-dwelling Americans (75-94 years old). The results confirm that information processing speed slows with increasing age. By contrast, increasing age has little effect on lexical decision. Thus, our data suggest that the SCOLP shows promise as a tool to help distinguish between normal aging and the early stages of dementia.
Reviewed by: The Russian language today by Larissa Ryazanova-Clarke, Terence Wade Edward J. Vajda The Russian language today. By Larissa Ryazanova-Clarke and Terence Wade. London & New York: Routledge, 1999. Pp. xii, 369. This is the first comprehensive account of Russian language evolution devoted to the last fifteen years of the twentieth century. Although most recent changes involve vocabulary, some grammatical patterns have also entered a period of flux so that the momentous events attendant on the collapse of communism in Russia seem to have affected all layers of the language. One of the book’s strong points is its use of copious examples from contemporary literature and media sources to illustrate all of the changes it describes. The book also includes a good survey of scholarly and normative works devoted to evaluating changes in Russian language usage; these sources appear in Cyrillic without translation in the form of a final bibliography (340–58). Although the authors intend principally to give a descriptive (rather than prescriptive) account of Russian at the close of the twentieth century, they have much to say about how other specialists regard the changes taking place. They also begin their survey in 1917 rather than 1985, although linguistic developments of the communist era are already well documented in such works as The Russian language in the 20th century (Bernard Comrie, Gerald Stone, and Maria Polinsky, Oxford: Clarendon Press, 1996). Nevertheless, inclusion of this material provides a useful point of comparison for recent trends that might otherwise appear unique in the history of the language. In fact, Russian during the twentieth century has undergone several periods of rapid change, particularly in the early years of Bolshevik rule (1917–28). These years witnessed a significant renegotiation of the boundary between standard and substandard speech as well as seemingly irrevocable alterations in the status of religious and political terminology. Analogous processes are once again afoot, albeit sometimes in the opposite direction. The book is divided into two parts of roughly equal length. The two chapters of Part 1 (3–165) describe innovations in vocabulary, recounting decade by decade the adoption or rejection of vast numbers of lexical items. The past fifteen years get an entire chapter to themselves, and they deserve one as more new [End Page 397] vocabulary has entered Russian during this time than at any other since the early communist period. While English has adopted a mere handful of Russian words in this short time, it has unwittingly become the source of entire new vocabularies for post-communist Russia, donating such items as killer ‘assassin’, sejl ‘sale’, imidzh ‘public image’, electorat ‘voters’, and hundreds of others. New loans often trigger a restructuring in the function of native synonyms. These patterns, along with the unpredictable stylistic nuances the new loans themselves acquire, are explained on the basis of examples in context. Russian killer, for instance, turns out to be ‘somehow respectable, modern and even interesting’ (163) when compared to the old native ubijtsa ‘murderer’. The remaining four chapters, packaged together as Part 2, are devoted to recent structural changes ranging from derivational morphology to syntax. Ch. 3 (169–239) discusses new word-formation models and recent extensions of old ones. Ch. 4 (240–82) covers new trends in case use and syntax; chief among these are the creation of plural forms for many singularia tantum nouns, an expanding usage of the accusative for marking negated direct objects, and the use of certain transitive verbs without an object. Most of these changes appear to be receiving momentum from the easing of the strict linguistic norms once enforced for all publishing and broadcasting. Ch. 5 (283–306) discusses the merry-go-round of place name changes that has again swept the country. The final chapter, entitled ‘The state of the language’, traces the origins of innovation to factors as diverse as youth slang and the poor speaking skills of Russia’s contemporary parliamentarians. The opinions of a variety of specialists, from Alexander Solzhenitsyn to leading university grammarians, are also surveyed. The authors close with their own, rather positive assessment on the future evolution and international role of Russian. This well researched and often entertaining book is essential reading...
Arizona Journal of Hispanic Cultural Studies 271 foice us to confronr aspects of out cultural history and identity we as Americans, perhaps Norm Americans, must confront: our rapacity, racism, machoism (sexism) in our dealing with this land and its peoples. They also present us with admirable acrs of choice, as we continue to define ourselves for worse or for bertei against a backdrop rhat threatens void but promises the sublime, climbable peaks of possibility. (212) In light of this statement, the book appeals to be written for the Kit Carsons of today, the multicultural polyglots who might make a fatal mistake (you know the kind). On this didactic point, and on Canfield's evident pleasure in studying Southwestern novels and films, Mavericks on the Border is a well-intended contribution to the revision of U.S. cultural and political history. Roberto Cantú California State University, Los Angeles Variation and Change in Spanish Cambridge University Press, 2000 By Ralph Penny Evet since William Labov's seminal woik on sound changes in progress in Martha's Vineyard (1963), one of the most important contributions of variationist sociolinguistics has been the possibility of detecting linguistic change in progress. The study of variance and its correlation with stylistic and social factors reveals the very source of linguistic change, and allows for an understanding of howpaiticulai innovations spread, shedding light on the mechanisms of both changes in progress and changes that have already been completed. These ideas underlie Penny's Variation and Change in Spanish, whose main merit is the attempt to integrate synchronic and diachronic perspectives into the study of the history of Spanish. In this book, instead of following the tradition of historical manuals that oiganize theii content around abrupt phonological, morphosyntactic and lexical changes across time, Penny emphasizes the vaiiation, both geogiaphical and social, that gave rise to change in Spanish. Undei this approach, the social history of the speakers is highlighted as Penny reconstructs some of the main mechanisms undetlying variation and change that are observable in former philological studies of Spanish. He emphasizes the changes caused by leveling of irregularities and simplification of structures, and argues that these two processes are rhe main forces driving Spanish evolution as a result of dialect contact and mixing due to constant population movement since the Middle Ages. In chapter 1, "Introduction," Penny briefly sets forth the theoretical framework and tetminology derived from historical sociolinguistics. In chapter 2, "Dialect, language, variety: definitions and relationships," the differences between dialect and language are discussed, clarifying common myths about this relationship among nonlinguists. The concepts of diglossia and diasystems ate applied to chaiacterize some of the relationships between the linguistic varieties in the Iberian Peninsula. Here Penny emphasizes the "seamlessness " of social and geographic dialectal continua, and thus regards the tree model, commonly used in historical linguistics, as inadequate due to, among other reasons, its individual branches that mask the continuity of the Peninsular Romance continuum. Chapter 3, "Mechanisms of Change," aims to present the ways in which linguistic innovations travel thtough both geogiaphical and social space. Grounded in the theory that linguistic innovations "ate passed from one individual to anothei through the accommodation processes which occur in face to face conracr" (63), Penny discusses leveling and simplification in late medieval and eatly modem Spanish. According to the authot, leveling explains (1) the reduction of the six medieval Spanish sibilants to three (in central and northern Spain) or two (elsewhere), (2) the variance between the initial IhI realization and dropping, and the final /h/-less solution, and (3) the merger of the voiced labial fricative and stop that initiated in the 15di century and the final IhI 272 Arizona Journal of Hispanic Cultural Studies victoiy. Simplification, a slightly different process, is responsible for (1) the merger of the perfect auxiliaries, (2) the history of strong preterites, and (3) the neai-meigei of the -er and -«-verb classes. After exemplifying cases of hyperdialectalism, reallocation of variants, and waves, Penny turns to the social factors thar govern rhe propagation of linguistic innovations, drawing from Leslie and James Milroys (1985) work on types of social networks. According to die Milroys, diffuse networks (weak ties among sevetal people) fosrer linguistic change...
The current state of affairs is characterised as one in which general SLA models have syntax as their core and pay less and variable attention to other linguistic levels, notably lexis. In order to improve the current situation we need involvement from both the vocabulary research community and SLA model builders. It is demonstrated how the former group readily borrows key concepts from psycholinguistics and SLA theory and rethinks them from a lexical point of view. However, such borrowing and recasting is often done in a piecemeal fashion to fit specific research issues. As for SLA model builders, some examples are discussed that are regarded as serious attempts at integrating lexis into a particular acquisition model. One is L2 reading research and vocabulary acquisition through reading, which illustrates a high degree of integration with common research goals and mutual theoretical inspiration. A second example underlines the fact that there is an obvious potential for including lexis in the ‘focus on form’ movement. It is our contention that more attention to lexis should supplement the predominantly grammatical ‘focus on form’ that is the current norm.
Reviewed by: Urban voices: Accent studies in the British Isles ed. by Paul Foulkes, Gerard J. Docherty Jeffrey L. Kallen Urban voices: Accent studies in the British Isles. Ed. by Paul Foulkes and Gerard J. Docherty. London: Arnold/New York: Oxford University Press, 1999. Pp. xiii, 313; 1 audiocassette/CD. The term ‘accent studies’ used in the subtitle of this book is intended by the editors to mark out a new territory ‘which intersects (at least) dialectology, sociolinguistics, phonetics and phonology’ and in which ‘accent variation can be seen as a pursuit in its own right, rather than being an issue towards the periphery of numerous separate academic traditions’ (6). This proposed new discipline builds on concrete phonetic data but gives ample room for abstract phonological analysis. The field incorporates studies in language variation, especially as understood in urban settings where issues of conflicting prestige norms, social class, gender, age stratification, and ethnicity assume greater prominence than they might elsewhere. Investigation into language variation naturally leads to questions on language change. To include these themes in accent studies, however, is not to include everything. The field which Foulkes and Docherty delimit in their introduction would not include dialectological approaches to the lexicon, morphology, or syntax, and it de-emphasizes or excludes broader problems in sociolinguistics such as the modeling of social class and social network, conversational analysis, and the social and political context of language usage. Most of the fifteen papers are divided into two parts: an introductory phonetic overview and a more detailed treatment of a problem in accent studies. The phonetic introductions follow a roughly standard pattern. Lexical sets adapted in varying degrees from those of Wells (1982) provide a common point of reference for vowels while consonantal variation is described phonemically under orthographic labels such as T (for stops in words such as time, butter, and hat) or TH (variably realizable as θ, ð, f, v, t, d, etc. in words such as thin and breathe). Most of the overviews are necessarily brief. Their common form facilitates regional comparisons, but the format sometimes makes it difficult or impossible to present community-wide variation in sufficient detail. Given these limitations, it is the variety of methods used to investigate specialized problems that provides the most striking feature of the book. The strand in accent studies which relies on instrumental phonetics is represented by Gerard J. Docherty and Paul Foulkes in comparing data from Derby and Newcastle, Jane StuartSmith in examining voice quality in Glasgow, and James M. Scobbie, Nigel Hewlett, and Alice Turk in a critical re-appraisal of the Scottish vowel length rule that shows it to be more limited in scope than is generally assumed. The paper by D&F has far-reaching methodological implications; as it demonstrates the use of instrumental methods to detect fine phonetic differences [End Page 833] that are not auditorily very salient but which nevertheless show socially-conditioned patterns of distribution. Quantitative approaches which presuppose a correlation between accent and socially significant nonlinguistic factors such as age, gender, socioeconomic class, and ethnicity are well-represented. Many of these papers also address questions of accent leveling and divergence. Thus Dominic Watt and Lesley Milroy give a convincing quantitative analysis of the tension between local and supraregional norms in Newcastle, using age, social class, and gender as the major determinants; Anne Grethe Mathisen similarly looks at features in Sandwell in the West Midlands, considering too the role of style in conditioning the realization of linguistic variables; Ann Williams and Paul Kerswill examine dialect convergence in relation to social change and mobility in the noncontiguous areas of Milton Keynes, Reading, and Hull; Kevin McCafferty gives valuable empirical evidence on the controversial topic of ethnicity and social class in (London) Derry English (incidentally drawing attention to the need for more scholars to grapple with ethnicity and English in Britain); and Inger M. Mees and Beverley Collins present a real-time study of change in Cardiff English which focuses on female speakers and examines variation...
The author elaborates on the meanings of the adjectives „srpski“ and \n„srbijanski“ in general and in the phrases where both these adjectives refer to the \nrepublic of Serbia (both country and state) in particular. The latter can be seen in \nthe examples such as srpska (or srbijanska) vlada the government of Serbia. The \nauthor takes into consideration the existing body of literature about this subject \n(authors such as: Е. Fekete, М. Nikolič, І. Кlajn, and М. Sipka) and analyzes the \ninstances of ambiguity involving „srpski“ and „srbijanski“ (as in the example abоve). Using the lexical and semantic norms of the contemporary standard Serbian \nlanguage, the author points to the possibility of differentiating the usage of these \ntwo adjectives. Нe also emphasizes the need of further elaboration and verification \nof the linguistic standardization criteria in general and in the field of the lexicon \n(lexical meaning and usage) in particular.
Editors’ introduction Like Darnton in this volume, Coulthard is interested in the practical uses that can be made of the phenomenon of repetition in text. His concern is with textual plagiarism, which he describes as involving texts in a Matching relation that is intended to remain undetected. This is more than simply a witty choice of phrasing: as noted for example in our Introduction, Matching relations rely on repetition, and, in many cases, plagiarists repeat not only the ideas but the wordings of the texts that they plagiarise. Identifying lexical repetition between texts is thus a practical step towards identifying possible cases of plagiarism. Coulthard explores plagiarism in three different areas: literary texts, student essays, and police records of interviews with and statements by suspects. He deals with two major issues: detection and directionality. Detection of plagiarism or unauthorised collaboration between writers can be difficult when, for example, a teacher is faced with large numbers of essays to mark — and even more so when the marking may be shared out amongst different teachers. This is where the occurrence of repetition of lexical items can be exploited. Whereas for Darnton’s purposes what is important is repetition in context (essentially, the repetition — with some changes — of whole sentences rather than of individual words), Coulthard shows that for his purposes measuring the percentage of vocabulary items shared by any two texts is sufficiently revealing. This has the advantage that it can be calculated automatically by computer. When a particularly high level of sharing is noted, the texts can be pulled out and subjected to individual scrutiny to confirm whether plagiarism is indeed involved. Once plagiarism is identified, the issue of directionality may arise: that is, determining which is the original text and which is the plagiarised one. With published texts this is normally a simple matter, since the chronology can be decided by date of publication; but with student essays and police records the analyst will typically need to rely on evidence in the texts themselves. Coulthard discusses various methods by which directionality can be established. At this point, his focus switches from repetition between the texts to cohesion — repetition and conjunction — within each text: that is, to the ways in which the texts are organised and the organisation is signalled. He demonstrates that his approach can be used to illuminate the process by which a supposedly independent text has in fact been derived from another — a process which may have extremely serious implications in legal cases. Underlying any discussion of plagiarism is the question of ‘voices’: how far is it possible to identify a writer’s personal voice, or style, and to detect places where that voice is overlaid or replaced by the voice of another? Coulthard argues that one way of approaching this question is through repetition — that texts (and the body of texts produced by each writer) have their own norms in terms of the language choices that the writers make. This raises an interesting comparison with Scott’s paper in this volume: the two papers can be seen as complementary in certain respects. If Scott deals with the ‘aboutness’ of texts and highlights what texts have in common despite their diversity, Coulthard’s study might be characterised as dealing with the ‘who-ness’ of texts and highlighting essentially what makes texts distinctive despite their similarities.
The transition from socialist to market economics is typically informed by outcomes-based social welfare theory (SWT). Institutionless, intentionally valuefree SWT is ill-suited to this enterprise. The only evaluative standard to which it gives rise—efficiency—is indeterminate, and the theory is not accommodative of other dimensions of moral evaluation. By contrast, the contractarian enterprise focuses on the role and importance of formal and informal institutions, including ethical norms. Given that individuals should be treated as moral equivalents, the project assigns lexical priority to rights and regards justice as impartiality. This explicitly normative, institutional approach permits analysis of potential conflicts between informal norms and prospective, formal rules of the games. Moreover, it underscores the instrumental and intrinsic value of rights in the transition process. Finally, the emphasis on impartiality—embodied in the generality principle—facilitates analysis of constitutional constraints on behavior that is inimical to the transition process. Timothy P. Roth, Contractarian Analysis, Ethics, and Emerging Economies, Journal of Markets & Morality 4, no. 1 (Spring 2001): 55-72
The following three papers have been originally read at a panel «Buddhist (Hybrid) Sanskrit» organized in the framework of the XIIth Conference of the International Association of Buddhist Studies held in Lausanne (Switzerland) on August 24, 1999. The purpose of the panel was, as formulated in the call for papers, first, to reassess the seminal work of Franklin Edgerton which is mainly known as his monumental Buddhist Hybrid Sanskrit Grammar and Dictionary to which a series of his articles dealing with this language (labeled hereafter «Buddhist Sanskrit») are to be usefully added. The reassessment has been and still is deemed possible indeed on the basis of new analysis of the texts in Buddhist Sanskrit known to Edgerton as well of the texts discovered and published after Edgerton’s work has been completed. The second purpose was to reconsider the problem of the internal structural cohesion of the Buddhist linguistic tradition involving a thorough analysis of grammatical and lexical evidence in Buddhist Sanskrit texts. The three scholars who responded to the call and whose papers have been prepared for the present publication base their research on different data and use understandingly different approaches, but I like to stress that all are aware of the complex nature of the linguistic and literary phenomena they examine. Interestingly, two of them, S. Karashima and K. Lang, share unpremeditatingly, needless to say, several presuppositions which seem to me as fertile as promising for further research. The main point common to these two authors is that beyond the general bewildering picture of Buddhist data, commonly considered as escaping any attempt to uncover an underlying linguistic structure and norm, the authors still see at least regular phenomena following a technique which cannot be due to a haphazard use of the language material and forms. The position of R. Salomon is different as different as his evidence, as will be seen from his article.
SUMMARYAntonio de Nebrija (1444?-1522) published his Gramatica Castellana in 1492, at a time when humanist appreciation of Castilian as a cultural language had not yet advanced to a discussion of its possibilities to become an established norm. However, an analysis of Nebrija's linguistic and grammatical theories does shed some light on this question. For instance, it becomes clear that the new method which he proposes for the teaching of Latin (nova ratio Nebrissensis) presupposed a recognition of the presence of universal grammatical concepts in the pupil's mother tongue. Such a conception is possible because Nebrija accepts an essential starting point of the medieval speculative tradition: language composition may be reduced to two basic concepts: materia (lexical element submitted to 'corruption') and forma (other elements — 'accidents' — which are stable). This composition is common to all languages. Therefore, Nebrija holds that by making use of the constrastive method it is possible to study two languages such as Latin and Castilian (which also happen to be closely related). Therefore, we must not consider the Gramdtica Castellana as separate from the rest of Nebrija's scholarly production. He himself had coined the notion of 'unity in diversity' concerning his grammatical work. In order to teach the Castilian language and, starting from Castilian, Latin, Nebrija writes grammatical and lexicographical works which have an underlying unity. His general approach was exclusive to Nebrija; however, although nobody before him had worked out such an ambitious project, there is no doubt that he was continuing on the way in which grammatical tradition had been heading for some time. An example of this tradition is the so-called Grarnmatica pro-verbiandi. In this paper, the main features of this kind of medieyal grammar are analyzed. It is argued that they constitute the immediate precursor of the Nebrija's undertaking, since we find in them didactic postulates which he developed further. These postulates led Nebrija to a contrastive grammar of Latin and Castilian and the creation of a grammatical terminology for the vernacular.RESUMEAntonio de Nebrija (1444?-1522) publia sa Gramdtica Castellana en 1492. A cette epoque, les pensees des humanistes n'etaient pas encore arri-vees au point de concevoir le castilian comme langue de culture, et par consequent, de la soumettre a une norme. Neanmoins, l'etude de l'oeuvre linguistique et grammaticale d'Antonio de Nebrija apporte des lumieres sur la genese de cette idee. Nebrija part de l'idee que la nouvelle methode qu'il propose pour l'enseignement de la langue latine (la nova ratio nebris-sensis), presuppose une connaissance des concepts grammaticaux acquise par l'etude de la langue maternelle de l'eleve. L'origine de cette idee se trouve dans la tradition speculative medievale selon laquelle la composition du langage se reduit a deux concepts: materia (elements lexiques 'corrupti-bles') et forma (elements 'non-corruptibles'). Cette composition est commune a toutes les langues. Nebrija en deduit qu'il est possible de recourir a une etude contrastive de deux langues, le latin et le castilian, qui, par ail-leurs, offrent l'avantage d'etre intimement apparentees. La Gramatica Castellana ne peut donc pas etre consideree comme une oeuvre isolee dans la production de Nebrija. En effet, il insiste lui-meme sur l'idee de l'unite dans la diversite de sa production grammaticale. Personne avant lui n'avait mene a bien un projet aussi ambicieux, mais il n'y a pas de doute qu'il sui-vait par la une voie emprunte par la tradition grammaticale depuis quelque temps deja, par la grammatica proverbiandi. Dans notre etude, nous analy-sons les principaux traits caracteristiques de cette tradition grammaticale medievale qui constitue l'antecedent inmediat a l'oeuvre de Nebrija. On y retrouve les postulats didactiques que developpera notre grammairien, la mise en valeur du contraste latin-castillan et la creation d'une terminologie grammaticale en langue vernaculaire.ZUSAMMENFASSUNGAntonio de Nebrija (1444?-1552) veroffentlichte seine Gramdtica Cas-tellana im Jahre 1492. Damals waren die Uberlegungen der Humanisten noch nicht so weit gediehen, das Spanische als Kultursprache anzusehen und es folglich einer Norm zu unterwerfen. Allerdings zeigen sich bei der Untersuchung der linguistischen und grammatischen Ideen von Antonio de Nebrija manche Ansatze, welche in diese Richtung weisen. Nebrija geht von dem Gedanken aus, daB die neue Methode, welche er fur den Lateinunter-richt vorschlagt (die nova ratio Nebrissensis), eine Kenntnis von (allgemei-nen) grammatischen Grundbegriffen voraussetzt, welche beim Studium der Muttersprache des Schulers zu entwickeln sind. Ansatze zu dieser Auffas-sung finden sich bereits im Mittelalter. Danach sind bei alien Sprachen zwei Komponenten zu unterscheiden, die materia (das dem 'Verfall' unterworfe-ne lexikalische Element) und die (unveranderliche) forma. Nebrija scheint es daher moglich, zwei Sprachen wie Latein und Spanisch kontrastiv zu stu-dieren, zumal diese noch miteinander verwandt sind. Die Gramdtica Castel-lana kann daher nicht isoliert vom ubrigen Werk Nebrijas betrachtet wer-den. Er selbst weist so auch immer wieder auf die Einheitlichkeit seines Werkes hin. Niemand vor ihm hatte je ein so ehrgeiziges Projekt verwir-klicht, aber es besteht auch kein Zweifel daran, das die grammatische Tradition vor ihm schon seit geraumer Zeit diese Richtung eingeschlagen hatte, die grammatica proverbiandi. In der vorliegenden Arbeit analysieren wir die Hauptmerkmale dieses Typs mittelalterlicher Grammatikbucher, welche als die direkten Vorlaufer des Werks von Nebrija anzusehen sind, findet man in ihnen doch Forderungen an den Fremdsprachenunterricht, welche von unserem Grammatiker spater weiterentwickelt werden, den kontrastierenden Vergleich zwischen Latein und Spanisch sowie die Schaf-fung einer grammatischen Terminologie in der Muttersprache.
In this contribution, we discuss how a fuzzy querying interface can support the generation of linguistic database summaries — a special technique of data mining. Links between our approach to linguistic summaries and the well-known technique of association rules is shown. The generation of linguistic summaries is implemented by using the authors’ FQUERY for Access package.
This article compares the word frequencies of the few most commonwords in Spanish as revealed by a modern corpus of over fivethousand words with a corpus of Golden-Age Spanish texts of overa million words, and finds that although de is by far themost common word in contemporary Spanish, in the 16thand 17th Centuries it was considerably less frequent, and in many texts was less frequent than y, or quefor which shared very similar frequency figures. It is arguedthat this significant change in the Spanish language comes aboutin the 20th Century.
Cross-language information retrieval (CLIR), where queriesand documents are in different languages, has of late become one ofthe major topics within the information retrieval community. Thispaper proposes a Japanese/English CLIR system, where we combine aquery translation and retrieval modules. We currently target theretrieval of technical documents, and therefore the performance of oursystem is highly dependent on the quality of the translation oftechnical terms. However, the technical term translation is stillproblematic in that technical terms are often compound words, and thusnew terms are progressively created by combining existing basewords. In addition, Japanese often represents loanwords based on itsspecial phonogram. Consequently, existing dictionaries find itdifficult to achieve sufficient coverage. To counter the firstproblem, we produce a Japanese/English dictionary for base words, andtranslate compound words on a word-by-word basis. We also use aprobabilistic method to resolve translation ambiguity. For the secondproblem, we use a transliteration method, which corresponds wordsunlisted in the base word dictionary to their phonetic equivalents inthe target language. We evaluate our system using a test collectionfor CLIR, and show that both the compound word translation andtransliteration methods improve the system performance.
Abstract: Target‐language discourse norms ( both lexically based formulaic speech and culturally based pragmatic abilities ) receive relatively scant attention in our curricula, yet they are of paramount importance in allowing speakers to maintain smooth communication. This article starts with two premises: ( 1 ) that oral skills classes provide the ideal forum in which to address discourse issues, and ( 2 ) that target‐language discourse norms, particularly as they relate to another cultural belief system, must be taught explicitly. After defining what is meant by “discourse norms,” this article offers a three‐part pedagogical discussion. The first part expands on Schmidt's “noticing hypothesis” ( Schmidt, 1990; 1993, Schmidt & Frota, 19861, arguing for the necessity of overt instruction in facilitating discourse learning. The second segment explores issues of course content by asking, Which features are most important for students to notice, and how might they be organized into a learning sequence? The final section focuses on instructional options, suggesting activity types that are intended to raise students' awareness of targeted norms. Issues of performance objectives ( productive vs. conceptual control ) and assessment are also addressed.