Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
Finding simple, non recursive, base noun phrase is an important step for many natural language processing applications. This paper presents a new corpus based approach using decision tree for that purpose. In contrast to previous methods for Base NP identification, we adopt a decision tree trained from Penn Treebank to identify Base NP. And a self learning mechanism is further integrated into our model. Experimental results show good performances using our method. The method can also be applied to processing of any other language.
This paper presents a statistical parser for a wide-coverage Combinatory Categorial Grammar (CCG) derived from the Penn Treebank. The Treebank is translated to a corpus of canonical CCG derivations. We de ne a generative statistical model over CCG derivations and train it on the transformed Treebank.
We aim at finding the minimal set of fragments which achieves maximal parse accuracy in Data Oriented Parsing. Experiments with the Penn Wall Street Journal treebank show that counts of almost arbitrary fragments within parse trees are important, leading to improved parse accuracy over previous models tested on this treebank (a precision of 90.8% and a recall of 90.6%). We isolate some dependency relations which previous models neglect but which contribute to higher parse accuracy.
Manual, large scale (computational) grammar development is time consuming, expensive and requires lots of linguistic expertise. More recently, a number of alternatives based on treebank resources (such as Penn-II, Susanne, AP treebank) have been explored. The idea is to automatically ``induce'' or rather read off (P)CFG grammars from the parse annotated treebank resources and to use the treebank grammars thus obtained in (probabilistic) parsing or as a starting point for further grammar development. The approach is cheap, fast, automatic, large scale, ``data driven'' and based on real language resources.\n\nTreebank grammars typically involve large sets of lexical tags and non-lexical categories as syntactic information tends to be encoded in monadic category symbols. They feature flat rules (trees) that can ``underspecify'' attachment possibilities. Treebank grammars do not in general follow Xbar architectural design principles (this is not to say that treebank grammars do not have design principles). As a consequence, treebank grammars tend to have very large CFG rule bases (e.g. Penn-II > 17,000 CFG rules for about 1 million words of text) with often only minimally differing rules. Even though treebank grammars are large, they are still incomplete, exhibiting unabated rule accession rates. From a grammar engineering point of view, the size of the rule base poses problems for maintainability, extendability and, if a treebank grammar is to be used as a CF-base in a LFG grammar, for functional (feature-structure) annotations. From the point of view of theoretical linguistics, flat treebank trees and treebank grammars extracted from such trees do not express linguistic generalisations. From the perspective of empirical and corpus linguistics, flat trees are well-motivated as they allow underspecification of subtle and often time consuming attachment decisions. Indeed, it is sometimes doubted whether highly general Xbar schemata usefully scale to ``real'' language.\n\nIn previous work we developed methodologies for automatic feature-structure annotation of grammars extracted from treebanks. Automatic annotation of ``raw'' treebank grammars is difficult as annotation rules often need to identify subsequences in the RHSs of flat treebank rules as they explicitly encode head, complement and modifier relations. Xbar based CFG rules should substantially facilitate automatic feature-structure annotation of grammar rules.\n\nIn the present paper we conduct a number of experiments to explore a space of possible grammars based on a small fragment of the AP treebank resource. Starting with the original treebank fragment we automatically extract a CFG G. We then apply an automatic structure preserving grammar compaction step which generalises categories in the original treebank fragment and reduces the number of rules extracted, resulting in a generalised treebank fragment and in a compacted grammar Gc. The generalised fragment is then manually corrected to catch missed constituents (and the like) resulting in an automatically extracted, compacted and (effectively manually) corrected grammar Gc,m. Manual correction proceeds in the ``spirit'' of treebank grammars (we do not introduce Xbar analyses). We then explore how many of the manual correction steps on treebank trees can be achieved automatically. We develop, implement and test an automatic treebank ``grooming'' methodology which is applied to the generalised treebank fragment to yield a compacted and automatically corrected grammar Gc,a. Grammars Gc,m and Gc,a are very similar to compiled out ``flat'' LFG-82 style grammars. We explore regular expression based compaction (both manual and automatic) to relate Gc,m to a LFG-82 style grammar design. Finally, we manually recode a subsection of the generalised and manually corrected treebank fragment into ``vanilla-flavour'' XBar based trees. From these we extract a compacted, manually corrected, XBar based grammar Gc,m,x. We evaluate our grammars and methods using standard labelled bracketing measures and according to how well they perform under automatic feature-structure annotation tasks.
Treebanks are of two types according to their annotation schemata: phrase-structure Treebanks such as the English Penn Treebank [8] and dependency Treebanks such as the Czech dependency Treebank [6]. Long before Treebanks were developed and widely used for natural language processing, there had been much discussion of comparison between dependency grammars and context-free phrase-structure grammars [5]. In this paper, we address the relationship between dependency structures and phrase structures from a practical perspective; namely, the exploration of different algorithms that convert dependency structures to phrase structures and the evaluation of their performance against an existing Treebank. This work not only provides ways to convert Treebanks from one type of representation to the other, but also clarifies the differences in representational coverage of the two approaches.
This paper describes a framework for examining the effects of the cognitive complexity of tasks on language production and learner perceptions of task difficulty, and for motivating sequencing decisions in task-based syllabuses. Results of a study of the relationship between task complexity, difficulty, and production show that increasing the cognitive complexity of a direction-giving map task significantly affects speaker-information-giver production (more lexical variety on a complex version and greater fluency on a simple version) and hearer-information-receiver interaction (more confirmation checks on a complex version). Cognitive complexity also significantly affects learner perceptions of difficulty (e.g. a complex version is rated significantly more stressful than a simple version). Task role significantly affects ratings of difficulty, though task sequencing (simple to complex versus the reverse sequence) does not. However, sequencing does affect the accuracy and fluency of speaker production. Implications of the findings for task-based syllabus design and further research into task complexity, difficulty, and production interactions are discussed.
In this paper we discuss the need for corpora with a variety of annotations to provide suitable resources to evaluate different Natural Language Processing systems and to compare them. A supervised machine learning technique is presented for translating corpora between syntactic formalisms and is applied to the task of translating the Penn Treebank annotation into a Categorial Grammar annotation. It is compared with a current alternative approach and results indicate annotation of broader coverage using a more compact grammar.
Reviewed by: Urban voices: Accent studies in the British Isles ed. by Paul Foulkes, Gerard J. Docherty Jeffrey L. Kallen Urban voices: Accent studies in the British Isles. Ed. by Paul Foulkes and Gerard J. Docherty. London: Arnold/New York: Oxford University Press, 1999. Pp. xiii, 313; 1 audiocassette/CD. The term ‘accent studies’ used in the subtitle of this book is intended by the editors to mark out a new territory ‘which intersects (at least) dialectology, sociolinguistics, phonetics and phonology’ and in which ‘accent variation can be seen as a pursuit in its own right, rather than being an issue towards the periphery of numerous separate academic traditions’ (6). This proposed new discipline builds on concrete phonetic data but gives ample room for abstract phonological analysis. The field incorporates studies in language variation, especially as understood in urban settings where issues of conflicting prestige norms, social class, gender, age stratification, and ethnicity assume greater prominence than they might elsewhere. Investigation into language variation naturally leads to questions on language change. To include these themes in accent studies, however, is not to include everything. The field which Foulkes and Docherty delimit in their introduction would not include dialectological approaches to the lexicon, morphology, or syntax, and it de-emphasizes or excludes broader problems in sociolinguistics such as the modeling of social class and social network, conversational analysis, and the social and political context of language usage. Most of the fifteen papers are divided into two parts: an introductory phonetic overview and a more detailed treatment of a problem in accent studies. The phonetic introductions follow a roughly standard pattern. Lexical sets adapted in varying degrees from those of Wells (1982) provide a common point of reference for vowels while consonantal variation is described phonemically under orthographic labels such as T (for stops in words such as time, butter, and hat) or TH (variably realizable as θ, ð, f, v, t, d, etc. in words such as thin and breathe). Most of the overviews are necessarily brief. Their common form facilitates regional comparisons, but the format sometimes makes it difficult or impossible to present community-wide variation in sufficient detail. Given these limitations, it is the variety of methods used to investigate specialized problems that provides the most striking feature of the book. The strand in accent studies which relies on instrumental phonetics is represented by Gerard J. Docherty and Paul Foulkes in comparing data from Derby and Newcastle, Jane StuartSmith in examining voice quality in Glasgow, and James M. Scobbie, Nigel Hewlett, and Alice Turk in a critical re-appraisal of the Scottish vowel length rule that shows it to be more limited in scope than is generally assumed. The paper by D&F has far-reaching methodological implications; as it demonstrates the use of instrumental methods to detect fine phonetic differences [End Page 833] that are not auditorily very salient but which nevertheless show socially-conditioned patterns of distribution. Quantitative approaches which presuppose a correlation between accent and socially significant nonlinguistic factors such as age, gender, socioeconomic class, and ethnicity are well-represented. Many of these papers also address questions of accent leveling and divergence. Thus Dominic Watt and Lesley Milroy give a convincing quantitative analysis of the tension between local and supraregional norms in Newcastle, using age, social class, and gender as the major determinants; Anne Grethe Mathisen similarly looks at features in Sandwell in the West Midlands, considering too the role of style in conditioning the realization of linguistic variables; Ann Williams and Paul Kerswill examine dialect convergence in relation to social change and mobility in the noncontiguous areas of Milton Keynes, Reading, and Hull; Kevin McCafferty gives valuable empirical evidence on the controversial topic of ethnicity and social class in (London) Derry English (incidentally drawing attention to the need for more scholars to grapple with ethnicity and English in Britain); and Inger M. Mees and Beverley Collins present a real-time study of change in Cardiff English which focuses on female speakers and examines variation...
Information access methods must be improved to overcome the information overload that most professionals face nowadays. Text classification tasks, like Text Categorization, help the users to access to the great amount of text they find in the Internet and their organizations. TC is the classification of documents into a predefined set of categories. Most approaches to automatic TC are based on the utilization of a training collection, which is a set of manually classified documents. Other linguistic resources that are emerging, like lexical databases, can also be used for classification tasks. This article describes an approach to TC based on the integration of a training collection (Reuters-21578) and a lexical database (WordNet 1.6) as knowledge sources. Lexical databases accumulate information on the lexical items of one or several languages. This information must be filtered in order to make an effective use of it in our model of TC. This filtering process is a Word Sense Disambig)
The Gsearch system allows the selection of sentences by syntactic criteria from text corpora, even when these corpora contain no prior syntactic markup. This is achieved by means of a fast chart parser, which takes as input a grammar and a search expression specified by the user. Gsearch features a modular architecture that can be extended straightforwardly to give access to new corpora. The Gsearch architecture also allows interfacing with external linguistic resources (such as taggers and lexical databases). Gsearch can be used with graphical tools for visualizing the results of a query.
Editors’ introduction Like Darnton in this volume, Coulthard is interested in the practical uses that can be made of the phenomenon of repetition in text. His concern is with textual plagiarism, which he describes as involving texts in a Matching relation that is intended to remain undetected. This is more than simply a witty choice of phrasing: as noted for example in our Introduction, Matching relations rely on repetition, and, in many cases, plagiarists repeat not only the ideas but the wordings of the texts that they plagiarise. Identifying lexical repetition between texts is thus a practical step towards identifying possible cases of plagiarism. Coulthard explores plagiarism in three different areas: literary texts, student essays, and police records of interviews with and statements by suspects. He deals with two major issues: detection and directionality. Detection of plagiarism or unauthorised collaboration between writers can be difficult when, for example, a teacher is faced with large numbers of essays to mark — and even more so when the marking may be shared out amongst different teachers. This is where the occurrence of repetition of lexical items can be exploited. Whereas for Darnton’s purposes what is important is repetition in context (essentially, the repetition — with some changes — of whole sentences rather than of individual words), Coulthard shows that for his purposes measuring the percentage of vocabulary items shared by any two texts is sufficiently revealing. This has the advantage that it can be calculated automatically by computer. When a particularly high level of sharing is noted, the texts can be pulled out and subjected to individual scrutiny to confirm whether plagiarism is indeed involved. Once plagiarism is identified, the issue of directionality may arise: that is, determining which is the original text and which is the plagiarised one. With published texts this is normally a simple matter, since the chronology can be decided by date of publication; but with student essays and police records the analyst will typically need to rely on evidence in the texts themselves. Coulthard discusses various methods by which directionality can be established. At this point, his focus switches from repetition between the texts to cohesion — repetition and conjunction — within each text: that is, to the ways in which the texts are organised and the organisation is signalled. He demonstrates that his approach can be used to illuminate the process by which a supposedly independent text has in fact been derived from another — a process which may have extremely serious implications in legal cases. Underlying any discussion of plagiarism is the question of ‘voices’: how far is it possible to identify a writer’s personal voice, or style, and to detect places where that voice is overlaid or replaced by the voice of another? Coulthard argues that one way of approaching this question is through repetition — that texts (and the body of texts produced by each writer) have their own norms in terms of the language choices that the writers make. This raises an interesting comparison with Scott’s paper in this volume: the two papers can be seen as complementary in certain respects. If Scott deals with the ‘aboutness’ of texts and highlights what texts have in common despite their diversity, Coulthard’s study might be characterised as dealing with the ‘who-ness’ of texts and highlighting essentially what makes texts distinctive despite their similarities.
This article reports on some data of a psycholinguistic study of first language attrition in german first generation immigrants. On the basis of the individual variation in performance evidenced by the data, I claim that L1 attrition in late bilinguals is not only the consequence of lack of L1 use. A comparison of the performance of three selected German-English bilinguals rather suggests that, among other factors, contact with other immigrants – as is the case in immigrant communities – might generate changes in linguistic competence. In this case it would be necessary to distinguish to types of intra-generational L1 attrition: (a) attrition in isolated immigrants who never use L1 in the host country, which mainly yields processing difficulties and problems in lexical retrieval, and (b) attrition in members of immigrant communities where changes of the linguistic norm within the community can take place, resulting in modifications of linguistic competence.
Imagine discourse between the arts in which the conventions of what we might call ordinary cognition do not apply, on site of intense lobbying neither tethered by history or cultural integrity, nor, frequently, concerned with social cohesion or communicative norms. It will be discourse in which the categories of an imperial culture are abrogated (however temporarily) by an indigenous one, yet it will undoubtedly also be site of intense colonization. On it, likewise, there will be an appropriation of language on an unprecedented scale. Past experience will play little part. Memory short and episodic, rather than semantic. It primal discourse. Primal in that it the site of first contact. Primal also in that it most often considered be the meeting of primitive culture and an advanced. Primal, likewise, in behavioral sense: in it, the satisfaction of physiological needs tantamount. Indeed, body and mind here are in state of kinetic unrest. This scene of prolonged immat urity, yet ontological and epistemological questions held in private language are encouraged be made public. Here the verbal arts have no canon. Literature has no prevailing cultural standard of merit. Questions of the popular and the high cultural are not naturalized and the fictional and nonfictional carry the same degree of verisimilitude as works of propaganda, rhetoric, and didacticism. In modem times, the West has become the site of this tenacious yet frequently unacknowledged imperial discourse, the discourse between multifarious forms of artistic representation win the attention of children. It in such discourse that the picture book located. 1 Arguing the need for critical language for the discussion of children's picture books, Peter Hunt suggests that to pictures into the same mould as words seems be potentially unproductive, except in terms of establishing conventions, when, of course, it is, by definition, necessary (181). It impossible, however, conventionalize pictorial representation the same degree as linguistic representation. Linguistic systems are mastered painstakingly, piece by piece, referent by referent, word by word. Pictorial systems, by contrast, are mastered all at once; they involve what Flint Schier has called natural generativity and are therefore much less conventional than linguistic systems. Each system, nevertheless, relies on general agreement and on willingness engage in communicative activity: the pictorial system on deep recognitional capacities that link object and its picture, the linguistic system on lexical and syntactical regularities and rules. In the media-saturated culture of the contemporary West, the commonalities and differences of our separate but shared experiences are frequently offered up in televisual or hypertextual format in which what Hunt describes as force set of discursive practices that address and interpellate both adults and children as potential viewers or listeners. The linguistic and the pictorial are frequently experienced as synergistic or polylogic systems bound up in this mass media, media whose intention, according Jean Baudrillard, transcribe the complexity of contemporary life into an ongoing procession of meaningless simulacra, hyperreal, a real without origin or reality (2). Baudrillard's disenchanted vision of postmodernity, articulated most profoundly in the late 1970s and 1980s, produced an interesting ontological metaphor. Disneyland, he claimed, is there conceal the fact that it the 'real' country, all of 'real' America, which Disneyland (just as prisons are there conceal the fact that it the social, in its entirety, in its banal omnipresence, which carceral). Disneyland presented as imaginary in order make us believe that the rest real, when in fact all of Los Angeles and the America surrounding it are no longer real, but of the order of the hyperreal and of simulation (25). …
Results of the noun–verb pair comprehension and production tests from the Test Battery for Auslan Morphology and Syntax (A. Schembri et al., 2000) are presented, reanalyzed, and compared to data from 2 other cases dealing with noun–verb pairs: the Auslan lexical database and a comparison of Auslan and American Sign Language (ASL) signs. The data confirm the existence of formationally related noun–verb pairs in Auslan in which the verb displays a single movement and the noun displays a repeated movement. The data also suggest that the best exemplars of noun–verb pairs of this type in Auslan form a distinct set of iconic (mimetic) signs archetypically based on inherently reversible actions (such as opening and shutting). This strong iconic link perhaps explains why the derivational process appears to be of limited productivity, though it does appear to have 'spread' to a number of signs that appear to have no such iconicity. There appears to be considerable variability in the use of the derivational markings, particularly in connected discourse, even for signs of the 'open and shut' variety. Overall, the derivational process is apparently still closely linked to an iconic base, is incipient in the grammar of Auslan, and is best described as only partially grammaticalized. (PsycINFO Database Record (c) 2016 APA, all rights reserved)
Internet search engines allow access to online information from all over the world. However, there is currently a general assumption that users are fluent in the languages of all documentsthat they might search for. This has for historical reasons usually been a choice between English and the locally supported language. Given the rapidly growing size of the Internet, it is likely that future users will need to access information in languages in which they are not fluent or have no knowledge of at all. This papershows how information retrieval and machine translation can becombined in a cross-language information access frameworkto help overcome the language barrier. We presentencouraging preliminary experimental results using English queries toretrieve documents from the standard Japanese language BMIR-J2retrieval test collection. We outline the scope and purpose ofcross-language information access and provide an example applicationto suggest that technology already exists to provide effective andpotentially useful applications.
Research on recognition and generation of signed languages and the gestural component of spoken languages has been held back by the unavailability of large-scale linguistically annotated corpora of the kind that led to significant advances in the area of spoken language. A major obstacle has been the lack of computational tools to assist in efficient analysis and transcription of visual language data. Here we describe SignStream, a computer program that we have designed to facilitate transcription and linguistic analysis of visual language. Machine vision methods to assist linguists in detailed annotation of gestures of the head, face, hands, and body are being developed. We have been using SignStream to analyze data from native signers of American Sign Language (ASL) collected in our new video collection facility, equipped with multiple synchronized digital video cameras. The video data and associated linguistic annotations are being made publicly available in multiple formats.
We present a rule--based shallow--parser compiler, which allows to generate a robust shallow-parser for any language, even in the absence of training data, by resorting to a very limited number of rules which aim at identifying constituent boundaries. We contrast our approach to other approaches used for shallow--parsing (i.e. finite-state and probabilistic methods). We present an evaluation of our tool for English (Penn Treebank) and for French (newspaper corpus "LeMonde") for several tasks (NP-chunking & "deeper" parsing).
Reviewed by: New horizons in the study of language and mind by Noam Chomsky D. Terence Langendoen New horizons in the study of language and mind. By Noam Chomsky. Cambridge: Cambridge University Press, 2000. Pp. xvii, 230. This is a collection of seven essays based on lectures and articles by Noam Chomsky from 1992 to the present, together with a foreword by Neil Smith. C has published a number of books like this one over the years, which attack the empiricist philosophy of language of Quine, Putnam, Davidson, and others and which defend his own ‘naturalist’ and ‘internalist’ views. This book also traces developments in the philosophy of language from the time of Sir Isaac Newton, and thus picks up where Cartesian linguistics (1966) leaves off. C points out that the problem of reconciling the ‘mental’ with the ‘physical’ was fundamentally altered by Newton’s demonstration that Cartesian mechanism is untenable. The ultimate solution to the ‘mind-body problem’, if it is found at all, is not likely to involve a reduction of the mental to the physical. Rather, the mental should be studied just like the physical, using whatever tools, methods, and insights are available, without arbitrary stipulations such as those of the philosophers mentioned above who limit the study of language in particular to correlations with observable behavior. Many, if not most linguists, C observes, ignore the strictures of these eminent philosophers, so that their efforts amount to nothing more than the harassment of the practitioners of an emerging science. [End Page 583] Since ‘natural language’ is what develops naturally in the course of language acquisition without instruction, the internalist and naturalist study of language does not consider those aspects of language which result from the imposition of community norms nor does it consider specialized uses which must be explicitly taught. For example, the common mass noun water does not mean ‘H2O’ in any natural language (thus rendering irrelevant to the study of natural languages such thought experiments as Putnam’s 1975 ‘twin earth’ thought experiments), and the consideration of what water does mean in a natural language leads to the conclusion that its reference cannot be determined extensionally. The same is true for every referring expression in a natural language, including proper nouns. Further, C maintains that the meanings of most lexical items in a natural language are far more elaborate than what is normally recorded in dictionaries and suggests that lexical structure is best explored within a decompositional framework such as that of Moravcsik 1990 or Pustejovsky 1995 as a kind of abstract syntax. In the first essay, C traces the evolution of his own conception of grammar, beginning with transformational-generative grammar in its various forms; continuing with the ‘principles and parameters’ framework, which he considers a more significant ‘revolution’ than transformationalgenerative grammar, the latter being a continuation of both traditional and structuralist ideas; and culminating in the ‘minimalist program’. In the principles and parameters framework, an internalized grammar (an I-language) is considered, in the words of the fifth essay, to be ‘an instantiation of the initial state [with the parameters fixed], idealizing from the actual states of the language faculty’, which are ‘the result of the interaction of a great many factors, only some of which are relevant to the inquiry into the nature of language’ (123). The development of the minimalist program was motivated by two closely related questions. First, ‘to what extent [can] the principles themselves... be reduced to deeper and natural properties of computation’ (123); and second, ‘to what extent [is] language... a “good solution” to the legibility conditions imposed by the external systems with which it interacts’ (9)? As the descriptor ‘minimalist’ suggests, C seeks a theory which is stripped to bare essentials. A language must contain phonetic and semantic features, a way of bundling these together into lexical items, and a way of combining lexical items together into larger expressions. It must also interact with other systems of the mind/brain which are responsible for producing and recognizing its expressions both phonetically and conceptually. An ideal or ‘perfect’ I-language is one whose computational apparatus consists only of entities and operations that are necessary to insure...
This study presents normative data for the Speed and Capacity of Language Processing (SCOLP) testfrom an older American sample. The SCOLP comprises 2 subtests: Spot-the-Word, a lexical decision task, providing an estimate of premorbid intelligence, and Speed of Comprehension, providing a measure of information processing speed. Slowed performance may resultfrom normal aging, brain damage (e.g., head injury), or dementing disorders or may represent the intact performance of someone who always performed at the low end of normal. The SCOLP enables the clinician to differentiate between these possibilities. Adequate age-appropriate norms to differentiate dementia from normal aging do not exist. We present data from 424 older community-dwelling Americans (75-94 years old). The results confirm that information processing speed slows with increasing age. By contrast, increasing age has little effect on lexical decision. Thus, our data suggest that the SCOLP shows promise as a tool to help distinguish between normal aging and the early stages of dementia.
The distinction of two types of intransitive verbs—unergatives (with underlying subjects) and unaccusatives (with underlying objects)—may not exist at early stages of L2 acquisition, both being syntactically represented as unergatives. This idea, referred to here as the Unaccusative Trap Hypothesis, provides an elegant developmental account for a variety of seemingly unrelated syntactic phenomena in L2 English, Japanese, and Chinese. Target language input, structural constraints on natural language linking rules, and linguistic properties of a learner's L1s shape stages in the reorganization of the lexical and syntactic components of interlanguage grammars. Although nonnative grammars may initially override the structural constraints postulated as the Unaccusative Hypothesis (Burzio, 1986; Perlmutter, 1978) and the Uniformity of Theta Assignment Hypothesis (Baker, 1988), at later developmental stages some may still achieve conformity with the norms of natural languages.
The transition from socialist to market economics is typically informed by outcomes-based social welfare theory (SWT). Institutionless, intentionally valuefree SWT is ill-suited to this enterprise. The only evaluative standard to which it gives rise—efficiency—is indeterminate, and the theory is not accommodative of other dimensions of moral evaluation. By contrast, the contractarian enterprise focuses on the role and importance of formal and informal institutions, including ethical norms. Given that individuals should be treated as moral equivalents, the project assigns lexical priority to rights and regards justice as impartiality. This explicitly normative, institutional approach permits analysis of potential conflicts between informal norms and prospective, formal rules of the games. Moreover, it underscores the instrumental and intrinsic value of rights in the transition process. Finally, the emphasis on impartiality—embodied in the generality principle—facilitates analysis of constitutional constraints on behavior that is inimical to the transition process. Timothy P. Roth, Contractarian Analysis, Ethics, and Emerging Economies, Journal of Markets & Morality 4, no. 1 (Spring 2001): 55-72
Arizona Journal of Hispanic Cultural Studies 271 foice us to confronr aspects of out cultural history and identity we as Americans, perhaps Norm Americans, must confront: our rapacity, racism, machoism (sexism) in our dealing with this land and its peoples. They also present us with admirable acrs of choice, as we continue to define ourselves for worse or for bertei against a backdrop rhat threatens void but promises the sublime, climbable peaks of possibility. (212) In light of this statement, the book appeals to be written for the Kit Carsons of today, the multicultural polyglots who might make a fatal mistake (you know the kind). On this didactic point, and on Canfield's evident pleasure in studying Southwestern novels and films, Mavericks on the Border is a well-intended contribution to the revision of U.S. cultural and political history. Roberto Cantú California State University, Los Angeles Variation and Change in Spanish Cambridge University Press, 2000 By Ralph Penny Evet since William Labov's seminal woik on sound changes in progress in Martha's Vineyard (1963), one of the most important contributions of variationist sociolinguistics has been the possibility of detecting linguistic change in progress. The study of variance and its correlation with stylistic and social factors reveals the very source of linguistic change, and allows for an understanding of howpaiticulai innovations spread, shedding light on the mechanisms of both changes in progress and changes that have already been completed. These ideas underlie Penny's Variation and Change in Spanish, whose main merit is the attempt to integrate synchronic and diachronic perspectives into the study of the history of Spanish. In this book, instead of following the tradition of historical manuals that oiganize theii content around abrupt phonological, morphosyntactic and lexical changes across time, Penny emphasizes the vaiiation, both geogiaphical and social, that gave rise to change in Spanish. Undei this approach, the social history of the speakers is highlighted as Penny reconstructs some of the main mechanisms undetlying variation and change that are observable in former philological studies of Spanish. He emphasizes the changes caused by leveling of irregularities and simplification of structures, and argues that these two processes are rhe main forces driving Spanish evolution as a result of dialect contact and mixing due to constant population movement since the Middle Ages. In chapter 1, "Introduction," Penny briefly sets forth the theoretical framework and tetminology derived from historical sociolinguistics. In chapter 2, "Dialect, language, variety: definitions and relationships," the differences between dialect and language are discussed, clarifying common myths about this relationship among nonlinguists. The concepts of diglossia and diasystems ate applied to chaiacterize some of the relationships between the linguistic varieties in the Iberian Peninsula. Here Penny emphasizes the "seamlessness " of social and geographic dialectal continua, and thus regards the tree model, commonly used in historical linguistics, as inadequate due to, among other reasons, its individual branches that mask the continuity of the Peninsular Romance continuum. Chapter 3, "Mechanisms of Change," aims to present the ways in which linguistic innovations travel thtough both geogiaphical and social space. Grounded in the theory that linguistic innovations "ate passed from one individual to anothei through the accommodation processes which occur in face to face conracr" (63), Penny discusses leveling and simplification in late medieval and eatly modem Spanish. According to the authot, leveling explains (1) the reduction of the six medieval Spanish sibilants to three (in central and northern Spain) or two (elsewhere), (2) the variance between the initial IhI realization and dropping, and the final /h/-less solution, and (3) the merger of the voiced labial fricative and stop that initiated in the 15di century and the final IhI 272 Arizona Journal of Hispanic Cultural Studies victoiy. Simplification, a slightly different process, is responsible for (1) the merger of the perfect auxiliaries, (2) the history of strong preterites, and (3) the neai-meigei of the -er and -«-verb classes. After exemplifying cases of hyperdialectalism, reallocation of variants, and waves, Penny turns to the social factors thar govern rhe propagation of linguistic innovations, drawing from Leslie and James Milroys (1985) work on types of social networks. According to die Milroys, diffuse networks (weak ties among sevetal people) fosrer linguistic change...
This article compares the word frequencies of the few most commonwords in Spanish as revealed by a modern corpus of over fivethousand words with a corpus of Golden-Age Spanish texts of overa million words, and finds that although de is by far themost common word in contemporary Spanish, in the 16thand 17th Centuries it was considerably less frequent, and in many texts was less frequent than y, or quefor which shared very similar frequency figures. It is arguedthat this significant change in the Spanish language comes aboutin the 20th Century.
In the last decade, given the availability of corpora in several distinct languages, research on multilingual part-of-speech tagging started to grow. Amongst the novelties there is mWANN-Tagger (multilingual weightless artificial neural network tagger), a weightless neural part-of-speech tagger capable of being used for mostly-suffix-oriented languages. The tagger was subjected to corpora in eight languages of quite distinct natures and had a remarkable accuracy with very low sample deviation in every one of them, indicating the robustness of weightless neural systems for part-of-speech tagging tasks. However, mWANN-Tagger needed to be tuned for every new corpus, since each one required a different parameter configuration. For mWANN-Tagger to be truly multilingual, it should be usable for any new language with no need of parameter tuning. This article proposes a study that aims to find a relation between the lexical diversity of a language and the parameter configuration that would produce the best performing mWANN-Tagger instance. Preliminary analyses suggested that a single parameter configuration may be applied to the eight aforementioned languages. The mWANN-Tagger instance produced by this configuration was as accurate as the language-dependent ones obtained through tuning. Afterwards, the weightless neural tagger was further subjected to new corpora in languages that range from very isolating to polysynthetic ones. The best performing instances of mWANN-Tagger are again the ones produced by the universal parameter configuration. Hence, mWANN-Tagger can be applied to new corpora with no need of parameter tuning, making it a universal multilingual part-of-speech tagger. Further experiments with Universal Dependencies treebanks reveal that mWANN-Tagger may be extended and that it has potential to outperform most state-of-the-art part-of-speech taggers if better word representations are provided.
The Japanese language, as spoken by native Japanese, has undergone a tremendous change in recent years-more noticeably in its spoken than in its written aspect. Some people brought up in the good old days deplore this, attributing it to the ignorance or negligence of decorum in speech among the younger generation. But the fact is that more and more adults are finding themselves unwittingly committing errors in usage, which they were formerly trained at school to avoid by all means. Among these deviations from linguistic norms, there are some which look likely to be established as perfectly acceptable usage, no matter whether one favors or disfavors them. Among them the most easily observable are: (1) recurrence of a rising intonation in mid-sentences, (2) omission of a morpheme, a word, or even a phrase, which was once considered indispensable in correct usage, (3) recurrence of what looks like an "empty" (semantically meaningless) word, (4) recurrence of what amounts almost to a cliche, and (5) prevalence of "feminine" (often infantile) language over strong "masculine" language, particularly in dialogs. What these phenomena reflect is, in the view of this paper-writer a kind of enervation in the verbal culture of the Japanese in general.
In statistical parsing, the probabilistic models are used to evaluate the possibility of each candidate parse tree, where the parse tree with the largest probability is deemed to be the final result of the parsing. Therefore, the core of statistical parsing is a probabilistic evaluation model. The main difference among the various probabilistic evaluation models lies in which types of features in the context are used to assign the probabilities to the parse trees. Various probabilistic evaluation models have been proposed in the field of statistical parsing, where different models use different feature types. How to evaluate a feature type's predictive power for the parsing tree? The paper proposes an information theory based feature type analysis model. Using the method, we can quantitatively analyze the power of different feature types for syntactic structure prediction from the viewpoint of information theory. The basic idea is that we use entropy and conditional entropy to measure whether a feature type grasps some of the information for syntactic structure prediction. If the average uncertainty of the syntactic structures declines apparently, the feature type is deemed to have grasped some intrinsic linguistic information in the context that has close relation to the syntactic structure. Using Penn Treebank as training and testing set, our experiment quantitatively analyze the different feature types' predictive power for syntactic structure predictive power for syntactic structure prediction in a systematic way and draws a series of conclusions which reflect the predictive power of different feature types and feature type combination for syntactic parsing.
The complex sentence structure of English is a bottleneck to our practical machi ne translation system. The simplification of English subordinate clauses will gr eatly relieves the burden of parsing and other grammatical or semantic analysis of a complex sentence, thus improves the output quality of the MT system. But th ere have not any satisfactory research achievements reported in this field up t o now as we know. In this paper, author's work on a corpus-based approach to English subordinate clause identification is reported. The approach integrate s rule-base d and statistical methods to get the left and right boundaries of the subordinat e clauses. The Penn Treebank corpus is used as the training standard. The precis ion and recall ratios of subordinate clause identification are tested on both cl osed and open corpora. A result of 92.9% precision and 91.26% recall is obtained for the closed test and the open test result is 80.34% precision and 83.93% rec all. This algorithm has been integrated into our machine translation system. The method can also be applied to processing of any other language.
The present study examined the relationships between quantitative volume estimates of mesial temporal lobe structures based on structural magnetic resonance imaging (MRI) and the memory and emotional functioning of individuals with temporal lobe epilepsy (TLE). Twenty individuals identified as having TLE and 24 control participants were administered a test battery that included an experimental recognition memory test incorporating both verbal and nonverbal stimuli, an experimental test of emotional functioning that measured both subjective report and skin conductance response (SCR) to emotionally salient stimuli, and a battery of standardized tests and questionnaires assessing attention, personality, and emotion perception. Patients also completed standardized measures assessing intellectual function, memory, and quality of life. The patient group demonstrated deficits on tests of memory, attention, and emotion perception. Patients also demonstrated reduced SCR, however this result was found in response to both emotional and nonemotional stimuli and so is not necessarily indicative of deficits in emotional arousal. Inconsistent with expectations, patients reported normal experiential states of arousal in response to emotionally salient stimuli. MRI data were used to measure left- and right-hemisphere volumes of the hippocampus and amygdala in the patient group, and these volumes were used as predictors of performance on behavioral measures in multiple regression analyses. Consistent with predictions, reduced amygdala volume predicted lower arousal ratings to positive emotional stimuli. However, a similar relationship was not found for arousal ratings of negative stimuli. Other predicted relationships were not demonstrated. Amygdala volume did not show a relationship with SCR, and hippocampal volume did not show a relationship with memory performance. Additional hypotheses regarding the lateralization of hippocampal and amygdala function were not supported. Results of standardized tests suggested some potential relationships between hippocampal volume and attention, and between amygdala volume and psychological characteristics, although further research would be needed to establish the degree to which these results could be generalized to a larger population. Study findings support the continued development of NM morphometric techniques to predict patterns of strengths and weaknesses demonstrated by individuals with TLE.
Interest in large-scale, grammar-based parsing has recently seen a large increase, in response to the complexities of language-based application tasks such as speech-to-speech translation, and enabled by efforts in large-scale, collaborative grammar engineering and in the induction of statistical grammars/parsers from treebanks. Parser throughput is an important consideration in real-world applications, since throughput tends to degrade as grammar size increases. Investigation of efficient approaches to parsing is therefore an important topic of research.
Finding simple, non-recursive, base noun phrase is an important step for many natural language processing applications. This paper presents a new corpus-based approach using decision tree for that purpose. In contrast to previous methods for Base NP identification, we adopt a decision tree trained from Penn Treebank to identify Base NP. And a self-learning mechanism is further integrated into our model. Experimental results show good performances using our method. The method can also be applied to processing of any other language.
CL Research's word-sense disambiguation (WSD) system is part of the DIMAP dictionary software, designed to use any full dictionary as the basis for unsupervised disambiguation. Official SENSEV AL-2 results were generated using WordNet, and separately using the New Oxford Dictionary of English (NODE). The disambiguation functionality exploits whatever information is made available by the lexical database. Special routines examined multiword units and contextual clues (both collocations, definition and example content words, and subject matter analyses); syntactic constraints have not yet been employed. The official coarsegrained precision was 0.367 for the lexical sample task and 0.460 for the all-words task (these are actually recall, with actual precision of 0.390 and 0.506 for the two tasks). NODE definitions were automatically mapped into WordNet, with precision of0.405 and 0.418 on 75 % and 70 % mapping for the lexical sample and all-words tasks, respectively, comparable to WordNet. Bug fixes and implementation of incomplete routines have increased the precision for the lexical sample to 0.429 (with many improvements still likely).
All spiritual cultures and material cultures, together with values and the way of life resulted from the two categories of cultures, are sure to play a vital role in conditioning people's statements and actions.So certain linguistic expressions and fixed forms which reflect values and life-style come to be the linguistic norm and rhetoric norm that ordinary people observe.
This study provides the information of reliability and implications of narrative essay tests for measuring achievement motive as a part of employee selection processes. To develop the key achievement motive explicit and concrete rating criterion, the TAT scoring method was applied to data (n=100) of essay tests gathered from seven raters. Three raters were chosen from entry level workers and the other four were professional writers of verbal testing items, forming two contrast groups. The reliability of the achievement motive ratings was calculated for each group by the interrater reliability approach, resulting that there was no significant difference between the groups. Coefficients of each group's achievement motive, general mental ability and personality traits were calculated, suggesting the possibility that the achievement motive ratings thus derived from essay tests implies individuality.
Abstract An analysis was made of 22 supervision groups in two psychotherapy training programmes at different levels. Its main focus concerned role patterns based on self-image ratings and changes over time. The results showed no significant differences between the two categories of supervisees, whereas the differences between the supervisors and the supervisees, independent of level of training, were highly significant. The results indicate that it is just as difficult to find one's voice and role in a supervision group at an advanced as at a basic level. For the supervisors the result was interpreted in terms of their roles in relation to the supervisees and the aim of the supervision.
Parsing a natural language with its substantial structural complexity and ambiguity has turned out to be a puzzler. While the most of attempts in this area so far has relied on hand-generated parsers, difficulties inherent in the manual construction of natural language grammar lead up to efforts to induce the grammar automatically. Our approach to the automatic grammar induction presented in this paper has resulted in design and implementation of the system GRIND (Grammar Induction), which is capable to learn a sequence of context- -dependent parse actions from a given corpus of labelled derivation trees. To this end, GRIND combines two established methods of machine learning: transformation-based learning (TBL) and inductive logic programming (ILP). Being trained and tested on corpus SUSANNE, GRIND reached the accuracy of 96 % and the recall of 68 %. Keywords: grammar induction, inductive logic programming, transformation- -based learning 1
Although research has examined differences in people’s responses to natural versus technologically caused disasters, research has not examined the differences in people’s attitudes toward disasters in natural versus built environments. This study examined the effects of the type of environment and awareness of the problem on attitudes toward the cleanup of oil spills. The results showed that the type of environment did not affect ratings of the importance of the environmental problem or how it should be cleaned up. However, people were more concerned about the environmental and community impacts of the cleanup process in the built environment. Awareness of the problem was a more important factor than type of environment for understanding attitudes toward the oil spill cleanup. People who were more aware of the oil spills viewed the problems as more important and were more concerned that the environments be returned to their previous states.
Natural language processing technologies offer ease-of-use of computers for average users, and ease-of-access to on-line information. Natural language, however, is complex, and the traditional methods of parsing with a single grammar and parser may result in an inefficient and large system that is difficult to maintain and is fragile in dealing with language irregularities. This paper begins by reviewing an alternative effort in grammar decomposition (also known as grammar partitioning) for natural language parsing, which aims to alleviate these problems. We then propose a novel automatic approach for grammar partitioning, in comparison with a random method of partitioning. Our experiments show that syntactic GLR parsing is formidable for the Wall Street Journal corpus in the Penn Treebank when a single grammar is used. This is due to too many grammar rules for parsing table generation. However, grammar partitioning solves the problem and offers a viable alternative. Our results also show that our automatic grammar partitioning method based on the mutual information criterion fares better than a random partitioning method and exhibits efficiency in parsing as well as high parse coverage. 1
In this paper a new similarity-based learning algorithm, inspired by string edit-distance (Wagner and Fischer, 1974), is applied to the problem of bootstrapping structure from scratch. The algorithm takes a corpus of unannotated sentences as input and returns a corpus of bracketed sentences. The method works on pairs of unstructured sentences or sentences partially bracketed by the algorithm that have one or more words in common. It finds parts of sentences that are interchangeable (i.e. the parts of the sentences that are different in both sentences). These parts are taken as possible constituents of the same type. While this corresponds to the basic bootstrapping step of the algorithm, further structure may be learned from comparison with other (similar) sentences. We used this method for bootstrapping structure from the flat sentences of the Penn Treebank ATIS corpus, and compared the resulting structured sentences to the structured sentences in the ATIS corpus. Similarly, the alg...
Reviews Dunn, J. A. (ed.). Language andSociety in Post-Communist Europe. Selected Papers from theFifth WorldCongress of Central andEast European Studies,Warsaw, 1995. General editor, Ronald J. Hill. Macmillan, Basingstoke and London and St Martin's Press, New York, I999. xi + I77 pp. Notes. Indexes.?40?00? THEsix chaptersof thisbook byJ. A. Dunn, LudmilaFerm,V. M. Mokienko, Wolf Moskovich, Boris Norman and Lara Ryazanova-Clarke offer original and telling insights into the changing Russian language. The articles of Alexander Krouglov and Seppo Lallukkarecall that languages other than Russianare involvedin change, while those of Stang&-Zhirovova's (on French loan-words)and Wulfhild Ziel's (on the work of Potebnja),while less central to the topic, are interestingand, in Ziel'scase, erudite. Dunn shows how the Russian language has been transformedin the postSoviet period by the disintegration of the Soviet political language and by Westernization, especially in the fields of economics, politics and social activity. The language has been infiltrated,not only by loans and slang, but by archaisms,wordplay and puns, while the principlesof politicalcorrectness are applied in a limited fashion only, as the language has struggledto borrow or createthe terminologyof a Western-typesociety. Ferm considers the development of political metaphor in the post-Soviet period, distinguishingmetaphorsthat have been used in a political sense only post I985 (especially those involving schvatka'skirmish' and probuksovka 'stalling'), metaphors with new referents, and completely new metaphors, manybasedon thehumanorganismanditsdiseases,and(perhapssymbolizing transitionto a new society)on transport. Mokienko detects a democratic spirit of spontaneity in contemporary Russian, in contrast to Soviet purism. Russian has not been alone in experiencing an influxof Americanisms,in fact the internationalizationof the language has been going on for centuries. Present change involves the resurrectionor re-evaluationof obsolete words and phrases,while neologisms have invadedthemassmedia. Linguisticnormsareimperilled,but 'pessimistic forecasts which place the present-day Russian language "on the brink of disaster"have no linguisticfoundation'(p. 82). Moskovich describes the language of nationalistic groupings in Russia, much of it anti-semitic,with calquing from Nazi German (bezzidizm 'absence of Jews') and some survivals from Soviet propaganda (siono-nacisty 'ZionoNazis '). One device isto useHebrew words,anotherinvolvesword-play,while rusaki,rusiciand nasi'our people' are used as stylistically-loadedvariants to russkie and non-Russians are termed cernye 'blacks', nerus''non-Russians' or zapadency (fromWesternUkraine). Norman'spaper on language games in contemporaryRussiandistinguishes a number of categories:unusualword formations(tennisizacija 'tennisization'), metaphors (jascik'box, TV set'), ellipses, and condensed sentences (vockax i v nedoumenii 'in glasses and in bewilderment'). Language play is portrayed as REVIEWS 309 typical of periods of upheaval, when language can be transformed -even if the world cannot. Ryazanova-Clarke defines a new genre: Western-style persuasive advertising, examining the subtleties of pronominal usage, rhetorical questions, imperatives (more regularly used than in the West: Reguljarno mojtegolovu 'Wash your hair regularly'), ellipses (Komplektes"e$ prakticnee'The set is even more practical'), informal and vernacular forms and aphorisms and repetitions. Ryazanova-Clarke concludes that the genre 'has created a new configuration of semantic resources' (p. I32). Krouglov depicts Soviet Russian as a tool designed to maintain power in the Republics, with linguistic norms imposed by the CPSU. By the I970S to the I98os Ukrainian had degenerated into 'the language of the "lower" strata of the population' (p. 38). Now attempts are being made to return to preSoviet times, and neologisms based on native roots are commoner in Ukrainian than in Russian. Lallukka's analysis of the status of Komi-Permiak epitomizes the fate of minority languages in the FSU, with Komi-Permiak relegated to colloquial registers. By the I980s Komi-Permiak had ceased to be a language of school instruction, and many parents favoured Russian as the language of social advance, while a society for the promotion of Komi-Permiak has been less than well supported in the I99os. Nadia Stang-Zhirovova examines the use of French loan-words in the language of Russian immigrants in Francophone Belgium. Most affected is the spoken language of the educated middle and upper class (Ja zvonila im, no u nichrepondeur automatique [for Russian avtootvetcik] 'I rang them but they have an answerphone'). The collection concludes with Ziel's piece on the fundamental question of the relationship between language...
GlossLexer is a multi-user sign language lexical database integrating digital video that has been designed to support the compilation process for specialist dictionaries from data collection to production. Sign entries are identified by HamNoSys notations as well as glosses, but the user always has immediate access to video clips showing the signs as uttered by the informants.
REVIEWS 309 typical of periods of upheaval, when language can be transformed -even if the world cannot. Ryazanova-Clarke defines a new genre: Western-style persuasive advertising, examining the subtleties of pronominal usage, rhetorical questions, imperatives (more regularly used than in the West: Reguljarno mojtegolovu 'Wash your hair regularly'), ellipses (Komplektes"e$ prakticnee'The set is even more practical'), informal and vernacular forms and aphorisms and repetitions. Ryazanova-Clarke concludes that the genre 'has created a new configuration of semantic resources' (p. I32). Krouglov depicts Soviet Russian as a tool designed to maintain power in the Republics, with linguistic norms imposed by the CPSU. By the I970S to the I98os Ukrainian had degenerated into 'the language of the "lower" strata of the population' (p. 38). Now attempts are being made to return to preSoviet times, and neologisms based on native roots are commoner in Ukrainian than in Russian. Lallukka's analysis of the status of Komi-Permiak epitomizes the fate of minority languages in the FSU, with Komi-Permiak relegated to colloquial registers. By the I980s Komi-Permiak had ceased to be a language of school instruction, and many parents favoured Russian as the language of social advance, while a society for the promotion of Komi-Permiak has been less than well supported in the I99os. Nadia Stang-Zhirovova examines the use of French loan-words in the language of Russian immigrants in Francophone Belgium. Most affected is the spoken language of the educated middle and upper class (Ja zvonila im, no u nichrepondeur automatique [for Russian avtootvetcik] 'I rang them but they have an answerphone'). The collection concludes with Ziel's piece on the fundamental question of the relationship between language and society in the writings of Aleskandr Potebnja,to markthe resumptionof the studyof neglected nineteenth-century figures. The proof-readinghas been thorough, with the odd misprint:'Institute'for 'Institut', 'Griefswald' for 'Greifswald' (p. xi), 'adverting' for 'advertising' (P. 131) 'Potenbja' for 'Potebnja' (twice, p. 159). There are occasional problemswith the umlaut('Stzrmer'for 'Sturmer',p. I64). Russianwordsare translatedin some papers. The book, particularly chapters two to nine, represents an illuminating contribution to the currentdebate on languages, more especially Russian, in post-CommunistEurope. Department ofModernLanguages TERENCE WADE University ofStrathclyde Schruba, Manfred. Studien zu denburlesken Dichtungen V.I. Majkovs. Slavistische Veroffentlichungen, 83. Harrassowitz Verlag, Wiesbaden, I997. Viii + I78 pp. Bibliography.Index. DM 78.oo. IN the circle of young writerswho frequented M. M. Kheraskov'shouse in Moscowin the I76os andfromI 770 in St Petersburg, V. I. Maikovstooda 310 SEER, 79, 2, 2001 littleapart.Forone thing he was ten to fifteenyearsolderthan the otherguests and fiveyearsseniorto his host (hewas alreadythirty-threewhen, in 176I, he retiredfromthe armyand settledin Moscow). Secondly, althoughlikemost of them he was a dvorianin, unlike them he had been sketchily educated. His teenage years from 1742 to I747 had been spent not in some establishment for the young nobility in the capital, but on his father'sestate in the Iaroslavl' guberniia, where there was no one qualified to teach him. His ignorance of foreignlanguages,particularlyofFrenchand German,constantlyembarrassed him. Thirdly, his lack of influentialfriendsand relationsmeant that, when in 1747 he joined his regiment, the Semenovskii Guards, he had to serve long years in the ranksand as ajunior officer.However, once out of the army, he quicklyfell into the patternof life customaryfor Russianwritersof the second half of the eighteenth century: a succession of posts in the civil service, involvement with the theatre, collaboration in literaryjournals, and masonic activities. His subsequentlife reflectedits comparativelyunprivilegedbeginnings.For one thing it revealedin him a practicalentrepreneurialstreak,so thatwhen in 1770 a shortageof high-qualitysail-clothcame to light in the Russianfleet, he founded a factory in Moscow to produce it. There were also significant consequences for his literarycareer. A youth spent in the provinces followed by fourteenyearsin the guardsbroughthim into contact with the peasant and lower merchant classes, rural, urban and military. He observed them in the fields, the taverns and the barracks.He listened to their conversation, tales and songs,and absorbedtheirwaysof speech. It wasthisrangeof experiences, unusualfor a young nobleman, which equipped him to become Russia'sfirst poet of the burlesque. Dr Schruba examines Maikov's three burlesque narrative poems: Igrok lombera (The Playerof Ombre),firstpublishedin I 763...
In this paper a new similarity-based learning algorithm, inspired by string edit-distance (Wagner and Fischer, 1974), is applied to the problem of bootstrapping structure from scratch. The algorithm takes a corpus of unannotated sentences as input and returns a corpus of bracketed sentences. The method works on pairs of unstructured sentences or sentences partially bracketed by the algorithm that have one or more words in common. It finds parts of sentences that are interchangeable (i.e. the parts of the sentences that are different in both sentences). These parts are taken as possible constituents of the same type. While this corresponds to the basic bootstrapping step of the algorithm, further structure may be learned from comparison with other (similar) sentences. We used this method for bootstrapping structure from the flat sentences of the Penn Treebank ATIS corpus, and compared the resulting structured sentences to the structured sentences in the ATIS corpus. Similarly, the algorithm was tested on the OVIS corpus. We obtained 86.04 % non-crossing brackets precision on the ATIS corpus and 89.39 % non-crossing brackets precision on the OVIS corpus.