Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
To assess the effects of discrepancy between two independent variables, investigators sometimes compute difference scores and correlate such scores with a criterion variable. However, the correlation of the difference with the criterion is accounted for by the correlations of the difference constituents with the criterion and the constituents’ variances. It follows that when investigators are testing a prediction that is not captured by the difference constituents’ main effects, using the difference correlation analysis may be misleading. Under these circumstances, the effects of a discrepancy between two independent variables can be assessed by a test of their interaction. The problems inherent in using difference scores and the advantage of testing the interaction are illustrated in relation to research programs on two separate topics in social psychology.
"Things seem to decline..." - Language, ethnicity and identity illustrated by material from a former Swedish colony in Misiones, Argentina and dialect material from Bjurholm, Sweden Nowadays there is a universal tendency towards convergence and simplification in several official European languages, the existence of non-codified languages is threatened and dialectal varieties are subject to levelling. If minority languages and intralinguistic varieties are to survive, this depends to a great extent on the identity of the individual speaker and the values and attitudes attached to his/her variety and speech behaviour. Besides, it is also due to the size of the language group, sharing the same values, and its cultural activities, forming part of its tradition. For that reason I have selected material from two threatened speech communities: one from a former Swedish colony in Misiones, Argentina, the other from Bjurholm, a small dialect-speaking community in the interior of Västerbotten, Sweden, in order to study the mechanisms causing the preservation or loss of the linguistic varieties as well as the cultural boundaries. This study consists of three parts and is divided into eleven chapters. The first part consists of chapter 1-4. Initially concepts of ethnicity, identity and culture are discussed in the light of different sciences, i.e. social anthropology (Barth 1969; Hylland Eriksen 1998), ethnology, the sociology of language (Fishman 1989) and sociolinguistics (Edwards 1985). In the following chapter emigrant material (narratives, letters, local history) forms part of the historical dimension: from the cultural contacts of individuals arriving in Brazil to the later Swedish settlement in Misiones, where it is appropriate to talk about an ethnic group, its collective history and Swedishness. In chapter 3 the continuity of the cultural heritage is illustrated by onomastic material (personal names) from three generations of Swedish descendants. Chapter 4 is a report of investigations carried out in the 1990's. In 1999, the Swedish language had been maintained among 20 of the 32 informants of Swedish descent, each one representing one family network. Their identity is hyphenated: they are all Argentine but of Swedish descent, and constitute an ethnic group. 24 of them had grown up with Swedish as their first language. Language attitudes had become more positive since 1988, when a similar investigation took place. Lately a new denomination has appeared: Los Nordicos, and it is discussed whether it is ethnic or not. Part two consists of chapter 5-10 and is the result of the project Dialects in change. The principal questions in this pilot study concern dialect boundaries: are they still maintained or subject to levelling? Which dialectal items are used as boundary markers and which are substituted by standard forms? A dialect boundary implies that dialects are still used in fairly genuine forms and that a local norm is prevailing, which seems to be the case in Bjurholm. Via three types of data, including inquiries, two tests: on dialectal vocabulary (50 lexical items) and translation of twelve standard sentences into dialect versions, besides type recordings of authentic speech, the author has tried to describe the local norm, based on individual micro data. In chapter 9 criteria for dialect variables used in quantitative studies are discussed. Based on this material there is evidence for dialect boundaries towards Lappland as well as towards the neighbour parish of Vindeln (former Degerfors). This material has to be extended to serve as a base for general conclusions and methods must be refined for further investigations. This will be possible, as the old dialects are still in use among the two oldest generations of adults. If they are to be used in the future depends on the younger generation (25-^4-0 years). In the final chapter data from the bilingual Swedish speakers in Misiones and the dialect speakers of Bjurholm are summarized, discussed and compared. For linguistic survival three key concepts are important: contact, prestige and identification, which can be related to ethnicity on a group level and identity on an individual level. There are striking similarities between them. In both cases there are boundaries between "us" and "them", but these are more subtle in an intralinguistic perspective. Both categories are using spoken varieties in a transitional stage, as Misiones Swedish soon will be extinct and the genuine Bjurholm dialect subject to levelling. Both varieties are also informal and diglossie in function, although codes are not always strictly kept apart. While the use of Misiones Swedish is reduced to the family sphere, the Bjurholm dialect can be extended to a wider range of domains. The great difference seems to concern the history of the varieties and the size of the group of speakers: the Bjurholm dialect can be traced back to at least 1750, maybe even to dialect splitting in the medieval time, while Misiones Swedish has been used for about 100 years by three generations of speakers, nowadays reduced to a number of approximately 150 persons.
DOAJ is a unique and extensive index of diverse open access journals from around the world, driven by a growing community, committed to ensuring quality content is freely available online for everyone.
The database Profil has been set up tooffer readers studying modern literarymanuscripts a reference tool to identifywatermarked papers. In the study of writers'drafts as in artists' sketches, the differentkinds of papers used provide valuableinformation on the genesis of a work of art andwatermarks, when they exist, are the bestvisible hint allowing us to identify paper. Amultimedia database, with digitized images moreprecise than usual traced design, seems to beappropriate to register, visualize, and comparemodern watermarked papers. Besides itsusefulness for specialists, such a databasebearing on modern manuscripts should also beconceived in a didactic perspective, as it isoriented towards literary scholars who are notparticularly familiar with the history of modern paper. In this paper we present the database Profilwhich includes a set of digitized images from acollection of betagraphies made by thereproduction service of the National FrenchLibrary. Then we explain problems of databasenormalization when human sciences areinvolved.
The paper describes a tagging scheme designed for the Russian Treebank and presents tools used for corpus creation. 1. Introductory Remarks The present paper describes a project aimed at developing the first annotated corpus of Russian texts. Large text corpora have been used in the computational linguistics community for quite a long time now; at present, over 20 large corpora for the main European languages are available, the largest of them containing hundreds of millions of words (Language Resources (1997); Marcus, Santorini and Marcinkiewicz (1993); Kurohashi, Nagao (1998)). For Russian, annotated corpora had been nonexistent until 2000 when the first part of the corpus reported here was compiled (Boguslavsky et al., 2000). Since then, Russian corpus linguistics has been evolving
Linguistically annotated corpora have nowadays a crucial theoretical as well as applicative role in Natural Language Processing. Italian still lacks such a resource. The paper describes a large scale effort to fill this gap by developing a multi-level annotated corpus, the Italian Syntactic-Semantic Treebank, which represents one of the main actions of an ongoing Italian national project (SI-TAL). The paper provides an overview the annotated corpus, its architecture and composition. For each annotation level, annotation guidelines are briefly discussed, also in the light of underlying motivations and prospective uses of the resource.
This paper presents the syntactic annotation level of a project aimed at providing a small dialog corpus with multiple levels of annotation. The syntactic annotation is based on dependency syntax. We outline the reasons for choosing dependency, and show the syntactic annotation for some constructions. We finish by describing the current state of the project. 1.
This paper reports on the development of a new eye-tracking system for noninvasive recording of eye movements. The eye tracker uses a flying-spot laser to selectivelyimage landmarks on the eye and, subsequently, measure horizontal, vertical, and torsional eye movements. Considerable work was required to overcome the adverse effects of specular reflection of the flying-spot from the surface of the eye onto the sensing elements of the eye tracker. These effects have been largely overcome, and the eye-tracker has been used to document eye movement abnormalities, such as abnormal torsional pulsion of saccades, in the clinical setting.
In this paper we analyze the problems set up in border lands, especially when a confluence of linguistic norms has taken place; an example is what happened in the Kingdom of Murcia along the Low Middle Ages, where settlers of different origins and also of different religion or race, Christians (Castilians and Catalans), Mussulmans or Jews lived together during some periods and followed one another in other time, leaving their traces on the onomastics and the toponymy. Some times the settlers' mark remained in the way of naming, but other times they reduced themselves to translate the names given by preceding settlers. With regard to onomastics, the traditions of each people remained evident and so have transmitted along the centuries; in the XIIIth. century it is very important the Catalan influence, a reflex of the repopulations; in the XIVth. century instead, because of the predominance of the Castilian model, a graphic adaptation of the family names received from preceding stages took place. Analyzing the documentation of that age permits us verify that the life together of peoples and languages enriched the toponomastic stock.
The measure of the lexical richness of literary texts as a tool in thecomparative analysis of literary style has been hampered by the problem ofthe inequality of text lengths within and between literary corpora. Thispaper proposes an empirical method of description of lexical richness byaveraging measures on multiple chunks of text of a standard lengthwithin a literary work or corpus. A workss average vocabulary richness,average portion of hapax legomenaof the corpus from which it derives,and average repetition of frequently appearing vocabulary may thencharacterize that work relative to other works partitioned along withit. This method reveals the possibility of significant variance of thesemeasures of vocabulary among works of a single authorss corpus and warnsagainst the notion of some absolute authorial stylistic character. Weapply this method of vocabulary averaging to the corpora of threeplaywrights from classical antiquity whose works are chronologicallyrankable: Euripides, Aristophanes, and Terence. We look for trends in vocabulary richness over time, which we posit functions as anindicator of progressively changing authorial ability or inclination. This method then holds the potential of predicting datesfor undateable or tenuously dated works within a corpus of otherwisesecurely dated texts. From the results derived, a relatively late date forthe composition of the redrafted version ofAristophaness Clouds appearslikely; we predict an early composition date for the redraft of TerencessHecyra (and thus are inclined to think that the playwright did verylittle redrafting); and finally we findEuripidess Electra andSupplices exhibiting vocabulary characteristics of extremely latecomposition and we predict dates much later than those assigned based onmetrical considerations.
We tested a computer-based procedure for assessing reader strategies that was based on verbal protocols that utilized latent semantic analysis (LSA). Students were given self-explanation—reading training (SERT), which teaches strategies that facilitate self-explanation during reading, such as elaboration based on world knowledge and bridging between text sentences. During a computerized version of SERT practice, students read texts and typed self-explanations into a computer after each sentence. The use of SERT strategies during this practice was assessed by determining the extent to which students used the information in the current sentence versus the prior text or world knowledge in their self-explanations. This assessment was made on the basis of human judgments and LSA. Both human judgments and LSA were remarkably similar and indicated that students who were not complying with SERT tended to paraphrase the text sentences, whereas students who were compliant with SERT tended to explain the sentences in terms of what they knew about the world and of information provided in the prior text context. The similarity between human judgments and LSA indicates that LSA will be useful in accounting for reading strategies in a Web-based version of SERT.
The specification phase is one of the most important and least supported parts of the software development process. The SAREL system has been conceived as a knowledge-based tool to improve the specification phase. The purpose of SAREL (Assistance System for Writing Software Specifications in Natural Language) is to assist engineers in the creation of software specifications written in Natural Language (NL). These documents are divided into several parts. We can distinguish the Introduction and the Overall Description as parts that should be used in the Knowledge Base construction. The information contained in the Specific Requirements Section corresponds to the information represented in the Requirements Base. In order to obtain high-quality software requirements specification the writing norms that define the linguistic restrictions required and the software engineering constraints related to the quality factors have been taken into account. One of the controls performed is the lexical analysis that verifies the words belong to the application domain lexicon which consists of the Required and the Extended lexicon. In this sense a synonym management process is needed in order to get a quality software specification. The aim of this paper is to present the synonym management process performed during the Knowledge Base construction. Such process makes use of the Spanish Wordnet developed inside the Eurowordnet project. This process generates both the Required lexicon and the Extended lexicon that will be used during the Requirements Base construction.
Editors’ introduction Like Darnton in this volume, Coulthard is interested in the practical uses that can be made of the phenomenon of repetition in text. His concern is with textual plagiarism, which he describes as involving texts in a Matching relation that is intended to remain undetected. This is more than simply a witty choice of phrasing: as noted for example in our Introduction, Matching relations rely on repetition, and, in many cases, plagiarists repeat not only the ideas but the wordings of the texts that they plagiarise. Identifying lexical repetition between texts is thus a practical step towards identifying possible cases of plagiarism. Coulthard explores plagiarism in three different areas: literary texts, student essays, and police records of interviews with and statements by suspects. He deals with two major issues: detection and directionality. Detection of plagiarism or unauthorised collaboration between writers can be difficult when, for example, a teacher is faced with large numbers of essays to mark — and even more so when the marking may be shared out amongst different teachers. This is where the occurrence of repetition of lexical items can be exploited. Whereas for Darnton’s purposes what is important is repetition in context (essentially, the repetition — with some changes — of whole sentences rather than of individual words), Coulthard shows that for his purposes measuring the percentage of vocabulary items shared by any two texts is sufficiently revealing. This has the advantage that it can be calculated automatically by computer. When a particularly high level of sharing is noted, the texts can be pulled out and subjected to individual scrutiny to confirm whether plagiarism is indeed involved. Once plagiarism is identified, the issue of directionality may arise: that is, determining which is the original text and which is the plagiarised one. With published texts this is normally a simple matter, since the chronology can be decided by date of publication; but with student essays and police records the analyst will typically need to rely on evidence in the texts themselves. Coulthard discusses various methods by which directionality can be established. At this point, his focus switches from repetition between the texts to cohesion — repetition and conjunction — within each text: that is, to the ways in which the texts are organised and the organisation is signalled. He demonstrates that his approach can be used to illuminate the process by which a supposedly independent text has in fact been derived from another — a process which may have extremely serious implications in legal cases. Underlying any discussion of plagiarism is the question of ‘voices’: how far is it possible to identify a writer’s personal voice, or style, and to detect places where that voice is overlaid or replaced by the voice of another? Coulthard argues that one way of approaching this question is through repetition — that texts (and the body of texts produced by each writer) have their own norms in terms of the language choices that the writers make. This raises an interesting comparison with Scott’s paper in this volume: the two papers can be seen as complementary in certain respects. If Scott deals with the ‘aboutness’ of texts and highlights what texts have in common despite their diversity, Coulthard’s study might be characterised as dealing with the ‘who-ness’ of texts and highlighting essentially what makes texts distinctive despite their similarities.
This paper is concerned with the investigation of the relevance and suitability of the data mining approach to serial documents. Conceptually the paper is divided into three parts. The first part presents the salient features of data mining and its symbiotic relationship to data warehousing. In the second part of the paper, historical serial documents are introduced, and the Ottoman Tax Registers (Defters) are taken as a case study. Their conformance to the data mining approach is established in terms of structure, analysis and results. A high-level conceptual model for the Defters is also presented. The final part concludes with a brief consideration of the implication of data mining for historical research.
Abstract: Target‐language discourse norms ( both lexically based formulaic speech and culturally based pragmatic abilities ) receive relatively scant attention in our curricula, yet they are of paramount importance in allowing speakers to maintain smooth communication. This article starts with two premises: ( 1 ) that oral skills classes provide the ideal forum in which to address discourse issues, and ( 2 ) that target‐language discourse norms, particularly as they relate to another cultural belief system, must be taught explicitly. After defining what is meant by “discourse norms,” this article offers a three‐part pedagogical discussion. The first part expands on Schmidt's “noticing hypothesis” ( Schmidt, 1990; 1993, Schmidt & Frota, 19861, arguing for the necessity of overt instruction in facilitating discourse learning. The second segment explores issues of course content by asking, Which features are most important for students to notice, and how might they be organized into a learning sequence? The final section focuses on instructional options, suggesting activity types that are intended to raise students' awareness of targeted norms. Issues of performance objectives ( productive vs. conceptual control ) and assessment are also addressed.
This study presents normative data for the Speed and Capacity of Language Processing (SCOLP) testfrom an older American sample. The SCOLP comprises 2 subtests: Spot-the-Word, a lexical decision task, providing an estimate of premorbid intelligence, and Speed of Comprehension, providing a measure of information processing speed. Slowed performance may resultfrom normal aging, brain damage (e.g., head injury), or dementing disorders or may represent the intact performance of someone who always performed at the low end of normal. The SCOLP enables the clinician to differentiate between these possibilities. Adequate age-appropriate norms to differentiate dementia from normal aging do not exist. We present data from 424 older community-dwelling Americans (75-94 years old). The results confirm that information processing speed slows with increasing age. By contrast, increasing age has little effect on lexical decision. Thus, our data suggest that the SCOLP shows promise as a tool to help distinguish between normal aging and the early stages of dementia.
This study aims to understand the meaning and function of degree adverbs. There has been no reasonable basis to set up the list of degree adverbs. It means that there is a lack of understanding the meaning and function of these categories. Degree adverbs must be used as a term that indicates adverbs with the semantic feature ‘degree’: It is intra-lexical distinctive, multivalent, and a kind of relative concept. It necessarily accompanies ‘norms of judgment’ and also grade ‘degree scales’. The degree adverbs that this paper has treated such as maeu, mucheok, gajang, etc. do not have ‘degree’. Therefore, these are not appropriate to have the name like degree adverbs. We need to use more appropriate degree adverbs than the existing terms and it depends on their meanings and functions.
This document describes the Part-of-Speech (POS) tagging guidelines for the Penn Korean Treebank Project. The corpus used for this project consists of around 54,000 words and 5,000 sentences. This document starts with a summary of the tagset used in the Penn Korean Treebank, followed by a more detailed discussion of each tag with examples. Then pairs of tags that are easily confused with each other are discussed and guidelines on how to distinguish one from the other for a given base forms and inflections are presented. The document concludes with a list of specific problematic examples with guidelines on how to handle such cases.
Reviewed by: The Russian language today by Larissa Ryazanova-Clarke, Terence Wade Edward J. Vajda The Russian language today. By Larissa Ryazanova-Clarke and Terence Wade. London & New York: Routledge, 1999. Pp. xii, 369. This is the first comprehensive account of Russian language evolution devoted to the last fifteen years of the twentieth century. Although most recent changes involve vocabulary, some grammatical patterns have also entered a period of flux so that the momentous events attendant on the collapse of communism in Russia seem to have affected all layers of the language. One of the book’s strong points is its use of copious examples from contemporary literature and media sources to illustrate all of the changes it describes. The book also includes a good survey of scholarly and normative works devoted to evaluating changes in Russian language usage; these sources appear in Cyrillic without translation in the form of a final bibliography (340–58). Although the authors intend principally to give a descriptive (rather than prescriptive) account of Russian at the close of the twentieth century, they have much to say about how other specialists regard the changes taking place. They also begin their survey in 1917 rather than 1985, although linguistic developments of the communist era are already well documented in such works as The Russian language in the 20th century (Bernard Comrie, Gerald Stone, and Maria Polinsky, Oxford: Clarendon Press, 1996). Nevertheless, inclusion of this material provides a useful point of comparison for recent trends that might otherwise appear unique in the history of the language. In fact, Russian during the twentieth century has undergone several periods of rapid change, particularly in the early years of Bolshevik rule (1917–28). These years witnessed a significant renegotiation of the boundary between standard and substandard speech as well as seemingly irrevocable alterations in the status of religious and political terminology. Analogous processes are once again afoot, albeit sometimes in the opposite direction. The book is divided into two parts of roughly equal length. The two chapters of Part 1 (3–165) describe innovations in vocabulary, recounting decade by decade the adoption or rejection of vast numbers of lexical items. The past fifteen years get an entire chapter to themselves, and they deserve one as more new [End Page 397] vocabulary has entered Russian during this time than at any other since the early communist period. While English has adopted a mere handful of Russian words in this short time, it has unwittingly become the source of entire new vocabularies for post-communist Russia, donating such items as killer ‘assassin’, sejl ‘sale’, imidzh ‘public image’, electorat ‘voters’, and hundreds of others. New loans often trigger a restructuring in the function of native synonyms. These patterns, along with the unpredictable stylistic nuances the new loans themselves acquire, are explained on the basis of examples in context. Russian killer, for instance, turns out to be ‘somehow respectable, modern and even interesting’ (163) when compared to the old native ubijtsa ‘murderer’. The remaining four chapters, packaged together as Part 2, are devoted to recent structural changes ranging from derivational morphology to syntax. Ch. 3 (169–239) discusses new word-formation models and recent extensions of old ones. Ch. 4 (240–82) covers new trends in case use and syntax; chief among these are the creation of plural forms for many singularia tantum nouns, an expanding usage of the accusative for marking negated direct objects, and the use of certain transitive verbs without an object. Most of these changes appear to be receiving momentum from the easing of the strict linguistic norms once enforced for all publishing and broadcasting. Ch. 5 (283–306) discusses the merry-go-round of place name changes that has again swept the country. The final chapter, entitled ‘The state of the language’, traces the origins of innovation to factors as diverse as youth slang and the poor speaking skills of Russia’s contemporary parliamentarians. The opinions of a variety of specialists, from Alexander Solzhenitsyn to leading university grammarians, are also surveyed. The authors close with their own, rather positive assessment on the future evolution and international role of Russian. This well researched and often entertaining book is essential reading...
Reviewed by: New horizons in the study of language and mind by Noam Chomsky D. Terence Langendoen New horizons in the study of language and mind. By Noam Chomsky. Cambridge: Cambridge University Press, 2000. Pp. xvii, 230. This is a collection of seven essays based on lectures and articles by Noam Chomsky from 1992 to the present, together with a foreword by Neil Smith. C has published a number of books like this one over the years, which attack the empiricist philosophy of language of Quine, Putnam, Davidson, and others and which defend his own ‘naturalist’ and ‘internalist’ views. This book also traces developments in the philosophy of language from the time of Sir Isaac Newton, and thus picks up where Cartesian linguistics (1966) leaves off. C points out that the problem of reconciling the ‘mental’ with the ‘physical’ was fundamentally altered by Newton’s demonstration that Cartesian mechanism is untenable. The ultimate solution to the ‘mind-body problem’, if it is found at all, is not likely to involve a reduction of the mental to the physical. Rather, the mental should be studied just like the physical, using whatever tools, methods, and insights are available, without arbitrary stipulations such as those of the philosophers mentioned above who limit the study of language in particular to correlations with observable behavior. Many, if not most linguists, C observes, ignore the strictures of these eminent philosophers, so that their efforts amount to nothing more than the harassment of the practitioners of an emerging science. [End Page 583] Since ‘natural language’ is what develops naturally in the course of language acquisition without instruction, the internalist and naturalist study of language does not consider those aspects of language which result from the imposition of community norms nor does it consider specialized uses which must be explicitly taught. For example, the common mass noun water does not mean ‘H2O’ in any natural language (thus rendering irrelevant to the study of natural languages such thought experiments as Putnam’s 1975 ‘twin earth’ thought experiments), and the consideration of what water does mean in a natural language leads to the conclusion that its reference cannot be determined extensionally. The same is true for every referring expression in a natural language, including proper nouns. Further, C maintains that the meanings of most lexical items in a natural language are far more elaborate than what is normally recorded in dictionaries and suggests that lexical structure is best explored within a decompositional framework such as that of Moravcsik 1990 or Pustejovsky 1995 as a kind of abstract syntax. In the first essay, C traces the evolution of his own conception of grammar, beginning with transformational-generative grammar in its various forms; continuing with the ‘principles and parameters’ framework, which he considers a more significant ‘revolution’ than transformationalgenerative grammar, the latter being a continuation of both traditional and structuralist ideas; and culminating in the ‘minimalist program’. In the principles and parameters framework, an internalized grammar (an I-language) is considered, in the words of the fifth essay, to be ‘an instantiation of the initial state [with the parameters fixed], idealizing from the actual states of the language faculty’, which are ‘the result of the interaction of a great many factors, only some of which are relevant to the inquiry into the nature of language’ (123). The development of the minimalist program was motivated by two closely related questions. First, ‘to what extent [can] the principles themselves... be reduced to deeper and natural properties of computation’ (123); and second, ‘to what extent [is] language... a “good solution” to the legibility conditions imposed by the external systems with which it interacts’ (9)? As the descriptor ‘minimalist’ suggests, C seeks a theory which is stripped to bare essentials. A language must contain phonetic and semantic features, a way of bundling these together into lexical items, and a way of combining lexical items together into larger expressions. It must also interact with other systems of the mind/brain which are responsible for producing and recognizing its expressions both phonetically and conceptually. An ideal or ‘perfect’ I-language is one whose computational apparatus consists only of entities and operations that are necessary to insure...
Intercorrelations among stylistic and emotional variables and constructvalidity deduced from relationships to other ratings of U.S. presidentssuggest that power language (language that is linguistically simple,emotionally evocative, highly imaged, and rich in references to Americanvalues) is an important descriptor of inaugural addresses. Attempts topredict the use of power language in inaugural addresses from variablesrepresenting the times (year, media, economic factors) and the man(presidential personality) lead to the conclusion that time-basedfactors are the best predictors of the use of such language (81%prediction of variance in the criterion) while presidential personalityadds at most a small amount of prediction to the model. Changes in powerlanguage are discussed as the outcome of a tendency to opt for breadthof communication over depth.
Arizona Journal of Hispanic Cultural Studies 271 foice us to confronr aspects of out cultural history and identity we as Americans, perhaps Norm Americans, must confront: our rapacity, racism, machoism (sexism) in our dealing with this land and its peoples. They also present us with admirable acrs of choice, as we continue to define ourselves for worse or for bertei against a backdrop rhat threatens void but promises the sublime, climbable peaks of possibility. (212) In light of this statement, the book appeals to be written for the Kit Carsons of today, the multicultural polyglots who might make a fatal mistake (you know the kind). On this didactic point, and on Canfield's evident pleasure in studying Southwestern novels and films, Mavericks on the Border is a well-intended contribution to the revision of U.S. cultural and political history. Roberto Cantú California State University, Los Angeles Variation and Change in Spanish Cambridge University Press, 2000 By Ralph Penny Evet since William Labov's seminal woik on sound changes in progress in Martha's Vineyard (1963), one of the most important contributions of variationist sociolinguistics has been the possibility of detecting linguistic change in progress. The study of variance and its correlation with stylistic and social factors reveals the very source of linguistic change, and allows for an understanding of howpaiticulai innovations spread, shedding light on the mechanisms of both changes in progress and changes that have already been completed. These ideas underlie Penny's Variation and Change in Spanish, whose main merit is the attempt to integrate synchronic and diachronic perspectives into the study of the history of Spanish. In this book, instead of following the tradition of historical manuals that oiganize theii content around abrupt phonological, morphosyntactic and lexical changes across time, Penny emphasizes the vaiiation, both geogiaphical and social, that gave rise to change in Spanish. Undei this approach, the social history of the speakers is highlighted as Penny reconstructs some of the main mechanisms undetlying variation and change that are observable in former philological studies of Spanish. He emphasizes the changes caused by leveling of irregularities and simplification of structures, and argues that these two processes are rhe main forces driving Spanish evolution as a result of dialect contact and mixing due to constant population movement since the Middle Ages. In chapter 1, "Introduction," Penny briefly sets forth the theoretical framework and tetminology derived from historical sociolinguistics. In chapter 2, "Dialect, language, variety: definitions and relationships," the differences between dialect and language are discussed, clarifying common myths about this relationship among nonlinguists. The concepts of diglossia and diasystems ate applied to chaiacterize some of the relationships between the linguistic varieties in the Iberian Peninsula. Here Penny emphasizes the "seamlessness " of social and geographic dialectal continua, and thus regards the tree model, commonly used in historical linguistics, as inadequate due to, among other reasons, its individual branches that mask the continuity of the Peninsular Romance continuum. Chapter 3, "Mechanisms of Change," aims to present the ways in which linguistic innovations travel thtough both geogiaphical and social space. Grounded in the theory that linguistic innovations "ate passed from one individual to anothei through the accommodation processes which occur in face to face conracr" (63), Penny discusses leveling and simplification in late medieval and eatly modem Spanish. According to the authot, leveling explains (1) the reduction of the six medieval Spanish sibilants to three (in central and northern Spain) or two (elsewhere), (2) the variance between the initial IhI realization and dropping, and the final /h/-less solution, and (3) the merger of the voiced labial fricative and stop that initiated in the 15di century and the final IhI 272 Arizona Journal of Hispanic Cultural Studies victoiy. Simplification, a slightly different process, is responsible for (1) the merger of the perfect auxiliaries, (2) the history of strong preterites, and (3) the neai-meigei of the -er and -«-verb classes. After exemplifying cases of hyperdialectalism, reallocation of variants, and waves, Penny turns to the social factors thar govern rhe propagation of linguistic innovations, drawing from Leslie and James Milroys (1985) work on types of social networks. According to die Milroys, diffuse networks (weak ties among sevetal people) fosrer linguistic change...
We consider several perceptual issues in the context of machine recognition ofmusic patterns. It is argued that a successful implementation of a musicrecognition system must incorporate perceptual information and error criteria.We discuss several measures of rhythm complexity which are used fordetermining relative weights of pitch and rhythm errors. Then, a new methodfor determining a localized tonal context is proposed. This method is based onempirically derived key distances. The generated key assignments are then usedto construct the perceptual pitch error criterion which is based on noterelatedness ratings obtained from experiments with human listeners.
The following three papers have been originally read at a panel «Buddhist (Hybrid) Sanskrit» organized in the framework of the XIIth Conference of the International Association of Buddhist Studies held in Lausanne (Switzerland) on August 24, 1999. The purpose of the panel was, as formulated in the call for papers, first, to reassess the seminal work of Franklin Edgerton which is mainly known as his monumental Buddhist Hybrid Sanskrit Grammar and Dictionary to which a series of his articles dealing with this language (labeled hereafter «Buddhist Sanskrit») are to be usefully added. The reassessment has been and still is deemed possible indeed on the basis of new analysis of the texts in Buddhist Sanskrit known to Edgerton as well of the texts discovered and published after Edgerton’s work has been completed. The second purpose was to reconsider the problem of the internal structural cohesion of the Buddhist linguistic tradition involving a thorough analysis of grammatical and lexical evidence in Buddhist Sanskrit texts. The three scholars who responded to the call and whose papers have been prepared for the present publication base their research on different data and use understandingly different approaches, but I like to stress that all are aware of the complex nature of the linguistic and literary phenomena they examine. Interestingly, two of them, S. Karashima and K. Lang, share unpremeditatingly, needless to say, several presuppositions which seem to me as fertile as promising for further research. The main point common to these two authors is that beyond the general bewildering picture of Buddhist data, commonly considered as escaping any attempt to uncover an underlying linguistic structure and norm, the authors still see at least regular phenomena following a technique which cannot be due to a haphazard use of the language material and forms. The position of R. Salomon is different as different as his evidence, as will be seen from his article.
In this paper, some electronically gathered data arepresented and analyzed about the presence of the pastin newspaper texts. In ten large text corpora of sixdifferent languages, all dates in the form of yearsbetween 1930 and 1990 were counted. For six of thesecorpora this was done for all the years between 1200and 1993. Depicting these frequencies on the timeline,we find an underlying regularly declining curve,deviations at regular places and culturally determinedpeaks at irregular points. These three phenomena areanalyzed.
In this paper, we propose a statistical method to automaticallyextract collocations from Korean POS-tagged corpus. Since a large portion of language is represented by collocation patterns, the collocational knowledge provides a valuable resource for NLP applications. One difficulty of collocation extraction is that Korean has a partially free word order, which also appears in collocations. In this work, we exploit four statistics, ‘frequency’,‘randomness’, ‘convergence’, and ‘correlation' in order to take into account the flexible word order of Korean collocations. We separate meaningful bigrams using an evaluation function based on the four statistics and extend the bigrams to n-gram collocations using a fuzzy relation. Experiments show that this method works well for Korean collocations.
This article describes the challenges posed by optical musicrecognition – a topic in computer science that aims to convert scannedpages of music into an on-line format. First, the problem is described;then a generalised framework for software is presented that emphasises keystages that must be solved: staff line identification, musical objectlocation, musical feature classification, and musical semantics. Next,significant research projects in the area are reviewed, showing how eachfits the generalised framework. The article concludes by discussingperhaps the most open question in the field: how to compare the accuracy and success of rival systems, highlighting certain steps thathelp ease the task.
Internet search engines allow access to online information from all over the world. However, there is currently a general assumption that users are fluent in the languages of all documentsthat they might search for. This has for historical reasons usually been a choice between English and the locally supported language. Given the rapidly growing size of the Internet, it is likely that future users will need to access information in languages in which they are not fluent or have no knowledge of at all. This papershows how information retrieval and machine translation can becombined in a cross-language information access frameworkto help overcome the language barrier. We presentencouraging preliminary experimental results using English queries toretrieve documents from the standard Japanese language BMIR-J2retrieval test collection. We outline the scope and purpose ofcross-language information access and provide an example applicationto suggest that technology already exists to provide effective andpotentially useful applications.
Research on recognition and generation of signed languages and the gestural component of spoken languages has been held back by the unavailability of large-scale linguistically annotated corpora of the kind that led to significant advances in the area of spoken language. A major obstacle has been the lack of computational tools to assist in efficient analysis and transcription of visual language data. Here we describe SignStream, a computer program that we have designed to facilitate transcription and linguistic analysis of visual language. Machine vision methods to assist linguists in detailed annotation of gestures of the head, face, hands, and body are being developed. We have been using SignStream to analyze data from native signers of American Sign Language (ASL) collected in our new video collection facility, equipped with multiple synchronized digital video cameras. The video data and associated linguistic annotations are being made publicly available in multiple formats.
An algorithm for analyzing difference scaling results is described. Frequency data on ordered categories that represent perceived differences for a unidimensional psychological attribute are modeled according to Thurstone’s judgment scaling model. The algorithm applies the gradient method for the maximum likelihood estimation of the model parameters. Two ways to calculate the start configuration for the model parameters are elaborated. The algorithm also provides asymptotic values for the standard errors of the estimates and three measures for the goodness of the model fit. An additional feature of DifScal is that it is suited to analyze incomplete data.
Information access methods must be improved to overcome the information overload that most professionals face nowadays. Text classification tasks, like Text Categorization, help the users to access to the great amount of text they find in the Internet and their organizations. TC is the classification of documents into a predefined set of categories. Most approaches to automatic TC are based on the utilization of a training collection, which is a set of manually classified documents. Other linguistic resources that are emerging, like lexical databases, can also be used for classification tasks. This article describes an approach to TC based on the integration of a training collection (Reuters-21578) and a lexical database (WordNet 1.6) as knowledge sources. Lexical databases accumulate information on the lexical items of one or several languages. This information must be filtered in order to make an effective use of it in our model of TC. This filtering process is a Word Sense Disambig)
In an attempt to establish a possible ‘norm’ for the distribution of translation modalities in English → Portuguese translational relationship, a varied sample of three different text typologies (legal, technical, and corporate) with six representative texts of each typology was compared, producing a total of 9,000 lexical items. By applying Vinay & Darbelnet’s and Aubert’s models, it was possible to obtain a basic pattern of the distribution of the most used translation modalities as well as to verify certain variables, such as the correlation between a higher or lower fluctuation in the frequency of the modalities and different types of texts. From three different levels of data analysis, we observed a translation hierarchy in relation to the three most frequent categories: literal translation, transposition and modulation and also that legal texts on the one hand, and technical and corporate texts on the other, seemed to organise themselves into two major groups. We also obtained some elements which would enable us to sketch a correlation between the modalities of literal translation and technical and corporate texts, as well as a correlation between the modalities of modulation and transposition with modulation and legal texts.
We examine an everyday Caribbean oral gesture, kiss-teeth or (KST), exploring previously-unresolved problems of meaning. Such forms are as examples of African cultural continuity across the Diaspora, often overlooked despite continuing interest in historical links between Caribbean Creoles and African communication systems. Forms such as (KST) are typically treated as lexical items: dictionary entries provide overlapping lists of emotions or affective states (eg, “scorn, impatience”) for each of several entries (suck-teeth, chups, etc.). Such approaches are inadequate, as the meaning of (KST) is not a single semantic unit, while lists are incomplete, contingent and inadequate. We distinguish ideophones from metalinguistic labels; consider geographical distribution and diffusion with respect to both functions and particular forms; and analyze related signs as a set, with reference to shared pragmatic function. (KST) is an inherently evaluative and inexplicit oral gesture with a sound-symbolic component, and a remarkably stable set of functions across the Diaspora: an interactional resource with multiple possibilities for sequential organization, often used to negotiate moral positioning among speakers and referents, and closely linked to community norms and expectations of conduct and attitude. It participates in a system of indirect discourse, requiring co-construction of intention by speaker and hearers. Moreover, it functions in personal narratives to mark both internal and external evaluation, sometimes ambiguously. Each of the proposed functions is illustrated with data ranging from historical to contemporary, oral to literary, monologic to interactional. Esther Figueroa & Peter L Patrick
This article reports on some data of a psycholinguistic study of first language attrition in german first generation immigrants. On the basis of the individual variation in performance evidenced by the data, I claim that L1 attrition in late bilinguals is not only the consequence of lack of L1 use. A comparison of the performance of three selected German-English bilinguals rather suggests that, among other factors, contact with other immigrants – as is the case in immigrant communities – might generate changes in linguistic competence. In this case it would be necessary to distinguish to types of intra-generational L1 attrition: (a) attrition in isolated immigrants who never use L1 in the host country, which mainly yields processing difficulties and problems in lexical retrieval, and (b) attrition in members of immigrant communities where changes of the linguistic norm within the community can take place, resulting in modifications of linguistic competence.
The Gsearch system allows the selection of sentences by syntacticcriteria from text corpora, even when these corpora contain no priorsyntactic markup. This is achieved by means of a fast chart parser,which takes as input a grammar and a search expression specified by theuser. Gsearch features a modular architecture that can be extendedstraightforwardly to give access to new corpora. The Gsearcharchitecture also allows interfacing with external linguistic resources(such as taggers and lexical databases). Gsearch can be used withgraphical tools for visualizing the results of a query.
Several Web animation methods were independently assessed on fast and slow systems running two popular Web browsers under MacOS and Windows. The methods assessed included those requiring programming (Authorware, Java, Javascript/Jscript), browser extensions (Flash and Authorware), or neither (animated GIF). The number of raster scans that an image in an animation was presented for was counted. This was used as an estimate of the minimum presentation time for the image when the software was set to update the animation as quickly as possible. In a second condition, the image was set to be displayed for 100 msec, and differences between observed and expected presentations were used to assess accuracy. In general, all the methods except Java deteriorated as a function of the speed of the computer system, with the poorest temporal resolutions and greatest variability occurring on slower systems. For some animation methods, poor performance was dependent on browser, operating system, system speed, or combinations of these.
The aim of this paper is to describe a technique for identifying the sourcesof several types of syntactic ambiguity in Arabic Sentences with a singleparse only. Normally, any sentence with two or more structuralrepresentations is said to be syntactically ambiguous. However, Arabicsentences with only one structural representation may be ambiguous. Ourtechnique for identifying Syntactic Ambiguity in Single-Parse ArabicSentences (SASPAS) analyzes each sentence and verifies the conditionsthat govern the existence of certain types of syntactic ambiguities in Arabicsentences. SASPAS is integrated with the syntactic parser, which is basedon Definite Clause Grammar (DCG) formalism. The system accepts Arabicsentences in their original script.
The current state of affairs is characterised as one in which general SLA models have syntax as their core and pay less and variable attention to other linguistic levels, notably lexis. In order to improve the current situation we need involvement from both the vocabulary research community and SLA model builders. It is demonstrated how the former group readily borrows key concepts from psycholinguistics and SLA theory and rethinks them from a lexical point of view. However, such borrowing and recasting is often done in a piecemeal fashion to fit specific research issues. As for SLA model builders, some examples are discussed that are regarded as serious attempts at integrating lexis into a particular acquisition model. One is L2 reading research and vocabulary acquisition through reading, which illustrates a high degree of integration with common research goals and mutual theoretical inspiration. A second example underlines the fact that there is an obvious potential for including lexis in the ‘focus on form’ movement. It is our contention that more attention to lexis should supplement the predominantly grammatical ‘focus on form’ that is the current norm.
In this paper a number of issues relating to theapplication of string processing techniques on musicalsequences are discussed. A brief survey of somemusical string processing algorithms is given and someissues of melodic representation, abstraction,segmentation and categorisation are presented. Thispaper is not intended to provide solutions tostring processing problems but rather tohighlight possible stumbling-block areas andraise awareness of primarily music‐elatedparticularities that can cause problems in matchingapplications.
The Gsearch system allows the selection of sentences by syntactic criteria from text corpora, even when these corpora contain no prior syntactic markup. This is achieved by means of a fast chart parser, which takes as input a grammar and a search expression specified by the user. Gsearch features a modular architecture that can be extended straightforwardly to give access to new corpora. The Gsearch architecture also allows interfacing with external linguistic resources (such as taggers and lexical databases). Gsearch can be used with graphical tools for visualizing the results of a query.