Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
A Swedish-English dictionary for the Internet is transformed into an English-Swedish counterpart by computationally reversing the Swedish and the English lexical database. Dictionary reversal of an existing bilingual dictionary is a possible solution of obataining working material quickly. However, depending on the complexity of the original database, the new reversed dictionary may have to be edited extensively. In the reversed dictionary, the original English target language becomes source language data. One outcome of reversing source and target language information is that the new English items are not necessarily translation equivalents. Internet dictionaries present new possibilities and challenges: they can easily be updated by allowing users to contribute new headwords, but lexicographers may also have to reconsider traditional methods in dictionary design. The article concludes by discussing the possible use of parallel corpus examples to illustrate language use in bilingual dictionaries.
In this paper we analyze the problems set up in border lands, especially when a confluence of linguistic norms has taken place; an example is what happened in the Kingdom of Murcia along the Low Middle Ages, where settlers of different origins and also of different religion or race, Christians (Castilians and Catalans), Mussulmans or Jews lived together during some periods and followed one another in other time, leaving their traces on the onomastics and the toponymy. Some times the settlers' mark remained in the way of naming, but other times they reduced themselves to translate the names given by preceding settlers. With regard to onomastics, the traditions of each people remained evident and so have transmitted along the centuries; in the XIIIth. century it is very important the Catalan influence, a reflex of the repopulations; in the XIVth. century instead, because of the predominance of the Castilian model, a graphic adaptation of the family names received from preceding stages took place. Analyzing the documentation of that age permits us verify that the life together of peoples and languages enriched the toponomastic stock.
A novel eye-movement-contingent method is presented. It builds on and extends established eye-movement-contingent visual display change methods in that it uses movements of the eyes to control the presentation of acoustic information during sentence reading. In one implementation, an irrelevant spoken word is presented when the eyes cross a predetermined spatial boundary before they move on to a selected visual target word. The relationship between the spoken word and the visual target is manipulated, and the pattern of interference, caused by the presentation of the spoken word, is used to determine the nature and time course of activated representations. Results from three recently completed experiments in which the technique was used show that a word’s phonological code remains active after it has been read and that the activated code has speech-like properties.
This paper describes the CTB Coreference Annotation Guidelines for annotating pronominal anaphoric expressions in the Penn Chinese Treebank. The goals of the annotation are: to provide training data for learning-based pronoun resolution tools, and to provide a "gold" standard to be used in the evaluation of pronoun resolution algorithms. The choices that were made concerning the coindexing of pronominal anaphors and their antecedents are discussed, as are some questions that arose in trying to categorize those pronominal expressions that did not refer to specific nominal entities in the text.
^j;=5!? ITHIN the rich corpus of metrical psalms comtm itt^ta * posed during Spain's Golden Age, Fray Luis de Aivi bl T 0_Le6n's versions are universally accorded the high;@ ^ VV |@ est praise. As heir to a literary tradition that ex,.A s iGne * tended back to the late Middle Ages and early.f4Li J Renaissance, the Salamancan scholar and poet revolutionized Spain's engagement with the Psalter, establishing the lira or estrofa alirada as the dominant verse form for vernacular psalm translations (Rivers 112; Nufiez 357), and making close lexical parallelism and philological accuracy, rather than interpretive digression, the norm for most of his followers. As may be expected from the great Augustinian's role as el primer poeta humanista espaniol en lengua vulgar (A. Blecua 97), Fray Luis was widely imitated, especially among disciples of his own order, and questions of authorship and dating of the many psalm versions attributed to him continue to trouble literary historians (Nufiez 357-8; J. M. Blecua Poesia completa, 41-2). Jose Manuel Blecua, in his 1990 edition of the Poesia completa, based on all extant manuscripts, includes as genuine the following poems: Psalm 1 Beatus vir, 4 Cum invocarem, 6 ne in furore, 9 Salvum me fac, 12 Usquequo, Domine (2 versions), 17 Diligam te, 18 Coeli enarrant, 24 Ad te, Domine, levavi,
We present an algorithm which translates the Penn Treebank into a corpus of Combinatory Categorial Grammar (CCG) derivations. To do this we have needed to make several systematic changes to the Treebank which have to effect of cleaning up a number of errors and inconsistencies. This process has yielded a cleaner treebank that can potentially be used in any framework. We also show how unary type-changing rules for certain types of modifiers can be introduced in a CCG grammar to ensure a compact lexicon without augmenting the generative power of the system. We demonstrate how the combination of preprocessing and type-changing rules minimizes the lexical coverage problem. 1.
530 SEER, 8o, 3, 2002 I999). Zubova is well aware of the metatextual qualities of Russian postmodernism, and points to the intertextualgames of varioustexts. In addition to the more familiarnames, Zubova introducesher audience to lesser known authors, such as Vladimir Strochkov, Ian Satunovskii, and Vladimir Erl'.Unfortunately, Zubova's studydoes not contain any biographical detailsof the authorsshe quotes. Such an appendixwould be a usefultool for assessingthe spreadof linguisticdeviations, from the point of view of age groups, regional variations and the aesthetic preferences of the poets. It is difficultto assess whether some of the deviations from established linguistic norms were intentional, or derive from the contemporary sloppy usage of Russianlanguage that isparticularlynoticeable in post-Soviet Russianmedia. Krivulinand Shvarts,for example, are philologistsby training,and therefore are more inclined to have playful appropriationof some idioms, or absurd examplesof Soviet newspeak. As Zubova's study demonstrates, numerous poetic experiments reflect on the fluid state of the Russian language itself. In this respect, Zubova's discussion of the satirical elements relating to the concept of gender in contemporary Russian poetry is particularlyrewarding. Zubova's examples from Russian poetry reveal, for example, the uncertaintiesrelatingto gender of animals. Thus, some poets use the feminine form of the noun koshka (cat) with the additional note that it is used in their poem as a noun of masculine gender. Such examples are both amusing and obscure. Zubova suggeststhat contemporary Russian poets are struggling to revive the use of the neuter gender that otherwise has been steadily disappearing from the standard language (p. 301). Many poems quoted by Zubova use Church Slavonic constructions such as esi, bekh,izhe, byst', and sut', to name just a few (pp. 208-39). Another archaicelement that occurs in contemporarypoetry is the Double Nominative Case discussedon pages 368-70. Zubova also refers to the influence of the English language on the contemporary Russian language in relationto the use of nouns as describingwords (pp. 362-68): for example, Rus'-zemlia, shved-koroleva. Zubova'sbook mightbe seen as an attemptto mergelinguisticanalysiswith cultural anthropology (especially Durkheim's theory), since it claims that various linguistic experiments express an archetypal collective conscience (p. 7). The extensive bibliographywill be highly appreciatedby readersof the book, as well as Zubova's reassuring message that the poetic experiments reveal the dynamic process of language evolution and expose the mnemonic abilities of semiotic signs that enable the user to restore forgotten forms and contexts. Department ofFrench andRussian ALEXANDRA SMITH University ofCanterbugy, NewZealand Menzel,Birgit.Biirgerkrieg um Worte. Die russische Literaturkritik derPerestroika. Bohlau Verlag, Cologne, Weimar, Vienna, 200I. xi + 420 pp. Illustrations. Notes. Bibliography.Tables. Index. DM 89.80:?4s.9 I. How do you investigate a type of text like literary criticism? Not literary criticism in the general sense, but in a sense which native English speakers REVIEWS 53I normally do not use, and which the Russians, among others, do? Literary criticism in this particular sense means topical, applied writing that greatly influencesthe generalpublic and explicitlyevaluatesthe worksdiscussed. Two possibilities present themselves. We can analyse the contents for consistencyand substanceof the profferedargumentsand evaluations.Or we can focus on pragmatic aspects of literary criticism (for example, by underscoring, collating, and evaluating statistics illustrating the effects of criticismon the purchasingbehaviourof readers). In her Rostock habilitation thesis, Birgit Menzel chooses neither of the above possibilities.And with good reason, for she does not focus merely on the work of a single literary critic, or on the treatment of a single work by various schools or camps of literarycriticism. Rather, she assignsherselfthe more comprehensive task of presenting an 'overview of Russian literary criticism between I986 and I993' (p. I). The reader hoping for a detailed criticism of criticism in Menzel's book will thereforebe disappointed. What she does offer is a survey of groupings and argument-patternsalong with traditions and developments. In other words, Menzel seeks to grasp the changing structuresof literary criticism as an important form of literary communication underthe conditions of the disintegratingSoviet empire. The enticing path of marketingresearchremainsunfeasiblefor the simple reason that salesrecordsfor the period of the Soviet planned economy and the years which immediatelyfollowed do not at all provide an accuratepictureof what readersreallywanted. In her statisticalassessments,Menzel thereforeconfines herselfto extremely illuminatinginformationon developments regardingthe circulationnumbersfor the...
1.1 The importance of developing sociocultural competence If we were to meet an adult native speaker who had grown up in a place where there were no other people, but sufficient language input, through, for example, tapes, for that person to be in linguistic terms a fluent speaker, then it seems reasonable to say that this person would in all likelihood be regarded as socially dysfunctional. Our unfortunate would not know how to deal with the most simple situations, and unless he or she were protected and educated, a sorry end may well be just around the corner. While such a case is fantastical, the non-native speaker (NNS) who arrives in an alien culture which is markedly different from their own and who lacks sociocultural knowledge is in a position with certain parallels to that of a socially inadequate individual (Furnham 1993). While knowledge could be transferred from the native culture, there is no way of guessing correctly what the possible cultural differences or similarities are. Native speakers (NSs) would be unaware of the visitor’s lack of sociocultural knowledge (Blum-Kulka 1997), and both NNSs and NSs may even be unaware that cultures can vary as much as they do (Hinkel 2001). NSs are also likely to be find behaviour that runs counter to their society’s beliefs or norms unacceptable, and to react accordingly. After perusing Celce-Murcia et al’s (1995) list of sociocultural factors (see appendix 1), it is not difficult to see how inappropriacy in any of the listed areas could lead to problems. The acceptable length of a silence varies across cultures, and one possible reason for some students’ perceived reticence in ESL contexts could be caused by the fact that in certain cultures, people are comfortable with longer response times than is the case in English. Gestures vary across cultures, and are used to express abstract ideas (McCafferty and Ahmed 2000); potential for confusion is therefore plentiful and plain. In a liberal Western country such as England, men coming from a more patriarchal society could easily find themselves being rebuked or criticised, and might feel at a loss as to why. When and to whom the words ‘Thank you’ are required to be said in England is a notoriously confusing area, and a source of much resentment among the inhabitants of towns where there is a constant influx of language learners. While the above examples show the significance of sociocultural factors in communication, the key question is how this knowledge relates to and is formulated in language, and in particular a second language. Pavlenko and Lantolf (2000) argue that traditional models of second language acquisition account for the way we acquire lexical, phonological and grammatical units of knowledge, but that in order to understand language use in context, and therefore the pervasiveness of culture in communication, a model which accounts for learning as participation is necessary. In this model, the learner develops skills which enable him or her to engage with contextual and cultural factors of communication. Although the two models are not mutually exclusive but in fact complementary, the latter is far more appropriate for understanding language as socialisation, as an ongoing process of engagement
The aim of this work is the construction of a syntactically annotated treebank for Basque. In this paper we present first, the basis of the annotation. After examining several options we chose the scheme presented in (Carrol et al., 1998). It follows the EAGLES standards and it is based on the idea of adding to each sentence in the corpus a series of grammatical relations specifying the dependencies between modifiers and their nucleus. After the formalism has been presented, we will describe the problems we have found and the decisions we have taken to solve them. Next we present an example showing the application of the scheme to an initial corpus. Finally, we present the main conclusions about the applicability to Basque and future work.
This paper describes a national cooperational project STO, which has the aim of developing a large-scale Danish lexical database for computational use. We discuss some organisational aspects of the project and present the current project structure and main activities. Further, we discuss in more detail some of the linguistic issues that have required thorough consideration before a large-scale encoding could be initiated, encompassing topics such as the morphological encoding of compounds and proper names, as well as the syntactic encoding of there-constructions, phrasal verbs and reflexive verbs.
We summarize here the results of a series of evaluations of the annotators’ assignments of tectogrammatical (i.e. underlying syntactic) tree structures and of the values of the edges as well as the values of the attribute representing the topic-focus articulation of the sentences, within the large-scale project of the Prague Dependency Treebank.
The Papillon project aims at building a multilingual lexical database for extracting dictionaries. This paper describes the Papillon monolingual lexie structure with an example and propose some changes for spotted problems.
The Centre for Language Technology (Center for Sprogteknologi, CST) is in charge of a national project developing a large-scale Danish lexicon for HLT and NLP applications. The short name of the project is STO, which stands for SprogTegnologisk Ordbase (Lexical Database for Language Technology). The project is inspired by principles and methods applied in the multilingual LE-PAROLE project (1996-98) the aim of which was to develop harmonised written language resources for 12 EU languages. The Danish PAROLE lexicon was produced by CST and the STO project highly benefits from the experience acquired from the work mentioned. This paper deals with a few central tasks of the ongoing project. It discusses the development of a smaller lexical resource produced in a multilingual environment into a large-scale, monolingual resource. Two different methods of increasing the vocabulary will be presented in detail; the extension of the linguistic coverage and the refinement of the linguistic description by including more detailed language-specific information. Finally, some exploitation perspectives and the development of an internet-based user-interface will be presented. The STO project gets funding from the Danish Ministry for Science, Technology and Development for a period of three years (2001-2004).
Many NLP systems are based on lexical data. The development costs of such data are a major drawback in such NLP systems. In order to cut these costs, we adopt a strategy inspired from "open-source" projects to allow volunteers to collaborate in the creation of a multilingual lexical database.For this, we had to specify and develop tools to manage a lexical database containing information complete and detailed enough to be usable for a wide range of applications.This paper presents our project and details the tools, frameworks and structures used to manage such a database. We will also show some research problems still to be addressed in this context.
The performance of cursive script recognition systems may be improved by applying higher level knowledge in the form of syntax or semantics. A fundamental part of such an approach is the creation of a lexical database containing the relevant information. However, to create a semantic lexicon by hand for a large vocabulary is a considerable task, which is a major reason why so many semantic theories fail to scale up from the small, artificial domains in which they were developed. An alternative approach is to use existing sources of semantic information, such as machine-readable dictionaries (which contain definitions and domain information) and text corpora (from which collocations and domain information may be derived). The development of techniques for acquiring semantic knowledge from such resources and applying it to large vocabulary cursive script recognition is described.< <ETX xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">></ETX>
Medical records have been evolving from the traditional paper-based records to digital ones, from the method of dictating reports and transcription to voice recognition systems. The transition to digital operations will not be complete until we have the ability to combine voice recognition with automated indexing of texts. This paper introduces the methods we used to evaluate existing voice recognition software programs and presents NOMINDEX, a system that turns a medical text into MeSH codes, using the French ADM lexical database. Those systems were applied to 28 patient discharge summaries in French, produced after a coronarography, and extracted from the MENELAS corpus of texts. Using the best configuration for voice recognition, the rate of accurate recognition exceeds 98 percent. Among the indexing concepts assigned by NOMINDEX, 25 percent were not pertinent and 12 percent of the relevant concepts were missing. Most errors were related to confusion between common language and medical language, and to the coverage of the ADM lexical database. Best results would be expected with a more comprehensive lexical resource In addition, only 3 percent of the errors generated by inadequate voice recognition that remained in the configuration that performed better, impacted on automatic indexing by NOMINDEX.
It is common knowledge that the creation of language resources for Language Engineering (LE) applications is a time-consuming, and hence expensive, enterprise. From this knowledge stems the demand for the re-usability of resources, which always remains essential. In this paper we will, however, concentrate on another, complementary, aspect, namely that of combining and extending existing resources by a variety of means and with a minimum of manual interaction. The resources to be discussed below consist of (i) a large lexical database, (ii) a formalized computational lexicon, and (iii) a sense-tagged corpus for Swedish. Some results concerning the semi-automatic annotation of the corpus and examples of a variety of phenomena analysed, such as compounding, will also be given. The annotation has been performed within the framework of the SemTag project, while part of this material has been successfully used in the SENSEVAL-2 exercise. In addition to these three resources, it can be added the background material of the Swedish Language Bank (some hundred million words) that forms the basis for the creation of (i) and partly (ii). Having been developed at our department, the lexical resources can easily be accessed, and, more importantly, can be systematically improved where necessary. It should be noted that this type of work requires close cooperation between specialists in lexicography and language technology. 1.
ABSTRACT: What links the formation of the state, the formation of linguistic norms, and the development of a grammatical tradition of the state language? In France governmental institutions not only help to spread the state language, but also inspire the first analyses of that language. The model for the grammatisation of French, and the context to promote such a process, is the formation of a national judicial system, through the demand that customary law be written down, and then approved by central authorities. Establishing local texts written in the national language required the analysis of that language, and based that analysis on concepts drawn from legal language. Establishing a fixed national norm depended on the creation of a judicial system founded on writing, and on the royal bureaucracy that resulted from that change.
REVIEWS 529 Zubova, L. V. Sovremennaia russkaia poeziia v kontekste istorni iazyka.Novoe literaturnoe obozrenie, Moscow, 2000. 432 pp. Notes. Bibliography. Index. Priceunknown. ZUBOVA'S study offers a detailed and provocative analysis of Russian poetry of the I96o-9os. It discusses almost three hundred authors, including leading postmodernist figures such as Joseph Brodsky, Viktor Krivulin, Genrikh Sapgir, Sergei Stratanovskii, Dmitrii Prigov, Elena Shvarts and Viktor Sosnora. Zubova highlights playful and innovative aspects of Russian postmodernistpoetry, arguingthat many linguisticexperimentsembedded in the texts under scrutiny in the present study explore the shortcomings and inadequacyof contemporaryRussianlanguage. Zubova arguesthat linguistic games of Russian postmodernistpoets offeralternativeways of development for phonetic, semantic and grammar structures of the Russian language. Zubova holds an optimisticview that deviations from the standardlanguage, as observed in the texts she studies, do not destroy the language, but help to preserve it, especially because they resurrect from oblivion some forgotten linguistic norms from the past (p. 399). Zubova's main thesis is based on the belief that language 'is a self-correctingsystem as well as a combination of options to express various meanings' (p. 399). The book will be of great interestto linguistsand to studentsof Russianpoetry, since it offersimportant insightsinto today'sstateof the Russianlanguage. The book comprises seven chapters, including an introduction and conclusion. Chapter one outlines the main theoretical frameworkwhich is applied throughout the book; chapter two discusses phonetic aspects of contemporary poetic experiments; chapter three investigates etymological innovations; chapter four is devoted to lexical changes; chapter five talks about archaic aspects of various grammaticaldeviations;chapter six offersa detailed analysis of the linguistic games based around gender; and chapter seven analysessyntacticalstructuresof the texts. In addition, the book offersa briefsummaryof the main conclusions (pp. 398-99), bibliographyand index. To some extent, all the texts Zubova discussesmight be viewed as hypertext, with no significantdifferentiationbetween the language of the i 960S and of the I990s. Furthermore,Zubova revealspostmodernisttendencies in Russian poetry of this period by demonstratinghow the linguisticexpressionis linked to the postmodernist worldview of the authors she discusses. Zubova encouragesher readersto considersome seeminglybad poems and treatthem with a sensitivityto the irony and parodic intentions they display.As Zubova states, 'the most constructivedevices of postmodernisttext include irony, selfirony and linguistic game' (p. i i). Zubova also highlights the authors' balancing acts between high and low cultures. In this respect, such poets as Prigov, Sosnora, Shvartsand Krivulin appear to be particularlyimaginative in their use of language, exploring precarious borders between modern and archaic forms of communication, and between elitist and popular forms of expression. Zubova's findings illustratewell the intrinsicbond between postSoviet poetry and Soviet undergroundliterature.Zubova'sstudyis a welcome additionto the extensiveanalysisof postmodernistfictionundertakenby Mark Lipovetsky (RussianPostmodernist Fiction.Dialoguewith Chaos,Armonk, NY, 530 SEER, 8o, 3, 2002 I999). Zubova is well aware of the metatextual qualities of Russian postmodernism, and points to the intertextualgames of varioustexts. In addition to the more familiarnames, Zubova introducesher audience to lesser known authors, such as Vladimir Strochkov, Ian Satunovskii, and Vladimir Erl'.Unfortunately, Zubova's studydoes not contain any biographical detailsof the authorsshe quotes. Such an appendixwould be a usefultool for assessingthe spreadof linguisticdeviations, from the point of view of age groups, regional variations and the aesthetic preferences of the poets. It is difficultto assess whether some of the deviations from established linguistic norms were intentional, or derive from the contemporary sloppy usage of Russianlanguage that isparticularlynoticeable in post-Soviet Russianmedia. Krivulinand Shvarts,for example, are philologistsby training,and therefore are more inclined to have playful appropriationof some idioms, or absurd examplesof Soviet newspeak. As Zubova's study demonstrates, numerous poetic experiments reflect on the fluid state of the Russian language itself. In this respect, Zubova's discussion of the satirical elements relating to the concept of gender in contemporary Russian poetry is particularlyrewarding. Zubova's examples from Russian poetry reveal, for example, the uncertaintiesrelatingto gender of animals. Thus, some poets use the feminine form of the noun koshka (cat) with the additional note that it is used in their poem as a noun of masculine gender. Such examples are both amusing and obscure. Zubova suggeststhat contemporary Russian poets are struggling to revive the use of the neuter gender that otherwise has been steadily disappearing from the standard language (p. 301...
DOAJ is a unique and extensive index of diverse open access journals from around the world, driven by a growing community, committed to ensuring quality content is freely available online for everyone.
The aim of this work is the construction of a syntactically annotated treebank for Basque. In this paper we present first, the basis of the annotation. After examining several options we chose the scheme presented in (Carrol et al., 1998). It follows the EAGLES standards and it is based on the idea of adding to each sentence in the corpus a series of grammatical relations specifying the dependencies between modifiers and their nucleus. After the formalism has been presented, we will describe the problems we have found and the decisions we have taken to solve them. Next we present an example showing the application of the scheme to an initial corpus. Finally, we present the main conclusions about the applicability to Basque and future work.
This paper presents the syntactic annotation level of a project aimed at providing a small dialog corpus with multiple levels of annotation. The syntactic annotation is based on dependency syntax. We outline the reasons for choosing dependency, and show the syntactic annotation for some constructions. We finish by describing the current state of the project. 1.
We presented the development process and the technical specifications of K-CDA IG. We explored how the results can be used as interoperability criteria in the national EHR systems certification program. Finally, we provided recommendations that could guide other entities planning their HIE programs.
The LinGO Redwoods initiative is a seed activity in the design and development of a new type of treebank. While several medium- to large-scale treebanks exist for English (and for other major languages), pre-existing publicly available resources exhibit the following limitations: (i) annotation is mono-stratal, either encoding topological (phrase structure) or tectogrammatical (dependency) information, (ii) the depth of linguistic information recorded is comparatively shallow, (iii) the design and format of linguistic representation in the treebank hard-wires a small, predefined range of ways in which information can be extracted from the treebank, and (iv) representations in existing treebanks are static and over the (often year- or decade-long) evolution of a large-scale treebank tend to fall behind the development of the field. LinGO Redwoods aims at the development of a novel treebanking methodology, rich in nature and dynamic both in the ways linguistic data can be retrieved from the treebank in varying granularity and in the constant evolution and regular updating of the treebank itself. Since October 2001, the project is working to build the foundations for this new type of treebank, to develop a basic set of tools for treebank construction and maintenance, and to construct an initial set of 10,000 annotated trees to be distributed together with the tools under an open-source license.
This paper reports on the development of a new eye-tracking system for noninvasive recording of eye movements. The eye tracker uses a flying-spot laser to selectivelyimage landmarks on the eye and, subsequently, measure horizontal, vertical, and torsional eye movements. Considerable work was required to overcome the adverse effects of specular reflection of the flying-spot from the surface of the eye onto the sensing elements of the eye tracker. These effects have been largely overcome, and the eye-tracker has been used to document eye movement abnormalities, such as abnormal torsional pulsion of saccades, in the clinical setting.
This paper discusses issues in building a 54-thousand-word Korean Treebank using a phrase structure annotation, along with developing annotation guidelines based on the morpho-syntactic phenomena represented in the corpus.Various methods that were employed for quality control and the evaluation on the Treebank are also presented.'This word count is computed on tokenized texts and includes symbols.
The aim of this volume is to showcase the range of corpus-based linguistic research currently being carried out on languages other than English. The papers included report on work carried out on Arabic, Bulgarian, Czech, Dutch, French, German, Biblical Greek, Biblical Hebrew, Medieval Irish, Korean, Romanian and Swedish, including a number of regional and social variants. They also address a range of areas as diverse as corpus design, corpus annotation, register analysis, syntax, and quantitative linguistics. The papers in this volume will leave the reader in no doubt that corpus-based research is now being conducted for a whole rainbow of languages.
This paper discusses issues in building a 54-thousand-word Korean Treebank using a phrase structure annotation, along with developing annotation guidelines based on the morphosyntactic phenomena represented in the corpus. Various methods that were employed for quality control are presented. The evaluation on the quality of the Treebank and some of the NLP applications under development using the Treebank are also presented.
This paper presents the annotation of the German TIGER Treebank. First, issues concerning the annotation, representation as well as querying of the treebank are discussed. Within this context, the annotation tool ANNOTATE, the export and XML formats of the TIGER Treebank and the TIGER search tool are briefly introduced. Secondly, the developments of the TIGER annotation scheme and their realization in the corpus are introduced focussing on the differences between the underlying NEGRA annotation scheme and the further developed TIGER annotation scheme. The main differences are concerned with verb-subcategorization, coordination, appositions and parentheses as well as proper nouns. Thirdly, the annotation scheme is assessed through an evaluation and a problem discussion of the above mentioned changes. For this purpose, inter-annotator agreement in the TIGER project has been analyzed focussing on exactly these changes. This analysis shows where the annotators&apos; decision problems are. These difficulties are discussed in greater detail on the basis of annotation examples. The paper concludes with some suggestions for the improvement of the TIGER annotation scheme.
'I~-eebanks, such as the Penn Treebank (PTB), offer a simple approach to obtaining a broad (:overage grammar: one can simply read the g rammar off the parse trees in the treebank. While such a g rammar is easy to obtain, a square-root rate of growth of the rule set with corpus size suggests that the derived grammar is far fi'om complete and that much more treebanked text would be required to obtain a complete grammar, if one exists at some limit. However, we offer an alternative explanation in terms of the underspecification of structures within the treebank. This hypothesis is explored by applying an algorithm to compact the derived grammar by eliminating redundant rules rules whose right hand sides can be parsed by other rules. The size of the resulting compacted grammar, which is significantly less than that of the full t reebank grammar, is shown to approach a limit. However, such a compacted grammar does not yield very good performance figures. A version of the compaction algorithm taking rule probabilities into account is proposed, which is argued to be more linguistically motivated. Combined with simple thresholding, this method can be used to give a 58% reduction in g rammar size without significant change in parsing performance, and can produce a 69% reduction with some gain in recall, but a loss in precision. 1 I n t r o d u c t i o n The Penn Treebank (PTB) (Marcus et al., 1994) has been used for a ra ther simple approach to deriving large grammars automatically: one where the g rammar rules are simply 'read off' the parse trees in the corpus, with each local subtree providing the left and right hand sides of a rule. Charniak (Charniak, 1996) reports precision and recall figures of around 80% for a parser employing such a grammar. In this paper we show that the huge size of such a treebank grammar (see below) can be reduced in size without appreciable loss in performance, and, in fact, an improvement in recall can be achieved. Our approach can be generalised in terms of Data-Oriented Parsing (DOP) methods (see (Bonnema et al., 1997)) with the tree depth of 1. However, the number of trees produced with a general DOP method is so large that Bonnema (Bonnema et al., 1997) has to resort to restricting the tree depth, using a very domain-specific corpus such as ATIS or OVIS, and parsing very short sentences of average length 4.74 words. Our compaction algorithm can be easily extended for the use within the DOP framework but, because of the huge size of the derived grammar (see below), we chose to use the simplest PCFG framework for our experiments. We are concerned with the nature of the rule set extracted, and how it can be improved, with regard both to linguistic criteria and processing efficiency. In what tbllows, we report the worrying observation that the growth of the rule set continues at a square root rate throughout processing of the entire t reebank (suggesting, perhaps tha t the rule set is far from complete). Our results are similar to those reported in (Krotov et al., 1994). 1 We discuss an alternative possible source of thi,~ rule growth phenomenon, partial bracketting, and suggest that it can be alleviated by compaction, where rules that are redundant (in a sense to be defined) are eliminated from the grammar. Our experiments on compacting a PTB tree1For the complete investigation of the grammar extracted from the Penn Treebank II see (Gaizauskas, 1995)
Oxymonads are a morphologically well-characterized and highly diverse lineage of protists. They are, however, under sampled at a molecular level. It has recently been demonstrated that a genus of oxymonads, Pyrsonympha, is phylogenetically related to the excavate taxon Trimastix. Here, we addressed issues of internal oxymonad evolution. Pyrsonympha and Dinenympha are shown, by fluorescent in situ hybridization and phylogenetic evidence, to be separate genera and not morphotypes of the same organism. We demonstrated that three genera of oxymonads, Dinenympha, Pyrsonympha, and Oxymonas are each monophyletic and together form a clade which excludes other known eukaryotes. We have presented a taxonomic scheme of oxymonads taking into account their sisterhood with Trimastix and speculated on morphological evolution of oxymonads, particularly of their attachment apparatuses. Our biogeographical analysis with Japanese and Canadian Pyrsonympha and Dinenympha suggests that these genera diverged before the separation of termites that inhabit Eastern Asia and Western North America.
Note: This is a personal view of the experience gathered, and does not necessarily reflect the opinion of the Floresta team The Floresta Sinta(c)tica http://acdc.linguateca.pt/treebank/ • A collaboration project between VISL (Southern Denmark University) and Linguateca (SINTEF); project leaders: Eckhard Bick & Diana Santos • Bosque: 1,427 syntactically analysed and revised trees (1,405 distinct sentences, 36,408 tokens, ca. 34,256 words), automatically created • Started October 2000, stopped December 2001, some research still being done as of today See Afonso et al. (2002) at LREC'2002
Prediction of Chinese phrase boundary location is the base of shallow parsing or chunk parsing.It is also very important for processing real texts.With the support of our Chinese treebank including 64426 words, this paper designs and implements a method for automatic prediction of Chinese phrase boundary location based on neural network. The preliminary results show that the precision is 93.24%(close testing) and 92.56%(open testing) respectively.
The development of large coverage, rich unification- (constraint-) based grammar resources is very time consuming, expensive and requires lots of linguistic expertise. In this paper we report initial results on a new methodology that attempts to partially automate the development of substantial parts of large coverage, rich unification-(constraint-) based grammar resources. The method is based on a treebank resource (in our case Penn-II) and an automatic f-structure annotation algorithm that annotates treebank trees with proto-fstructure information. Based on these, we present two parsing architectures: in our pipeline architecture we firstextract a PCFG from the treebank following the method of [Charniak, 1993; Charniak, 1996], use the PCFG to parse new text, automatically annotate the resulting trees with our f-structure annotation algorithm and generate proto-f-structures. By contrast, in the integrated architecture we firstautomatically annotate the treebank trees with fstructure information and then extract an annotated PCFG (A-PCFG) from the treebank. We then use the A-PCFG to parse new text to generate proto-fstructures. Currently
The Prague Dependency Treebank (PDT, as described, e.g., in (Hajic, 1998) or more recently in (Hajic, Pajas and Vidova Hladka, 2001)) is a project of linguistic annotation of approx. 1.5 million word corpus of naturally occurring written Czech on three levels (“layers”) of complexity and depth: morphological, analytical, and tectogrammatical. The aim of the project is to have a reference corpus annotated by using the accumulated findings of the Prague School as much as possible, while simultaneously showing (by experiments, mainly of statistical nature) that such a framework is not only theoretically interesting but possibly also of practical use. In this contribution we want to show that the deepest (tectogrammatical) layer of representation of sentence structure we use, which represents “linguistic meaning” as described in (Sgall, Hajicova and Panevova, 1986) and which also records certain aspects of discourse structure, has certain properties that can be effectively used in machine translation1 for languages of quite different nature at the transfer stage. We believe that such representation not only minimizes the “distance” between languages at this layer, but also delegates individual language phenomena where they belong to whether it is the analysis, transfer or generation processes, regardless of methods used for performing these steps.
The basic argument of this paper is that multiple language use in learner output is not always and exclusively indicative of the deficient nature (V. Cook, 1999; Firth and Wagner, 1997; Kramsch, 1997, 1998) of the language learner with respect toan idealized monolingual second language (L2) linguistic norm. Multilingual written learner texts and learners’ explications of these texts are examined in detail. These data suggest that learners conceptualize themselves as multicompetent (V. Cook, 1991, 1992, 1999) speakers who regularly, playfully and creatively decouple conventionalized L2 form-meaning pairings in order to produce and use their own locally relevant L2 signs. Within mainstream Second Language Acquisition (SLA) research practices and correctness-oriented Foreign Language Teaching (FLT) methodologies, this departure from norm approximation may be interpreted to indicate the reduced (Harder, 1980) or deficient nature of the learner. Within a Vygotskian approach to the psychology of mind and language learning (e.g., Lantolf, 2000a; Lantolf and Appel, 1994; Vygotsky, 1978, 1986), however, the learner’s playful use of multiple linguistic codes may index resourceful, creative and pleasurable displays of multicompetence (V. Cook, 1991,1992).
SLIM is a prototype interactive multimedia self-learning linguistic software for foreign language students at beginner-false beginner level. It allows students to work both in an autonomous self-directed mode or in a way of programmed learning in which the process of self-instruction is pre-programmed and monitored. In this latter mode it incorporates assessment and evaluation tools in order to behave as an automatic tutor. It is organized into three basic components: audiovisual materials; a linguistic database recording all language material in text format; the supervisor. Audiovisual materials are partially taken from commercially available courses; the linguistic database is a highly sophisticated classification of all words and utterances of the course, both in written and spoken form, from all possible linguistic aspects. The supervisor is both an attractive, enjoyable and strongly pedagogically based software that allows the user to work on language materials. The most outstanding feature of SLIM is the use of speech analysis and recognition which is a fundamental aspect of all second language learning programmes. We also assume that a learning model can be represented by a finite state automaton made up by a fixed number of possible states – corresponding to the macro and microlevels at which the student's competence may be modelled – each one being internally constituted by the actual linguistic objects of knowledge of the language that make it up.
This paper would like to introduce the reader into those aspects of the Arabic language which require some special treatment compared to languages Europeans are more familiar with. In spite of having fresh experience in building the Prague Arabic Dependency Treebank, the authors try to take a broader view of the problems encountered under way. The topics discussed include linguistic data retrieval, morphology and morphotactics modelling, and description of the language on the analytical level.
XML(eXtensible Markup Language)is a standard which was issued by W3C(wolld wide web consortium) in February,1998.It defines the data structure by means of an open self-description.Also a group of rules of XML can be used to set up Markup Language in accordance with specific applied fields.Furthermore,based on XML,cnXML is a linguistic norm of electronic business affairs,which is keeping with the commercial habits,tradition and circuit in the continent of China.cnXML supplies a set of unified,flexible,open and extensible exchange pattern of data,which makes all trade partners including such commercial organizations as buyers,sellers,runners and intermediary etc.carry out commercial activities conveniently through Internet.So cnXML is not only a standardized norm but also a premise of electronic business development in China.