Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
This paper describes an approach to treebank development which relies on the manual development of annotation tools. The overall process of tree annotation is described, and a special emphasis is put on the description of the last tool which has been built, i.e. a dependency-based robust chunk parser. The modularization of the parser and the central role of verbal subcategorization is presented. Some experimental results, carried
This paper presents work which extends previous corpus-based work on training Machine Learning Algorithms to perform Prepositional Phrase attachment. Besides recreating others' experiments to see how algorithms' performance changes with the number of training examples and using n-fold cross-validation to produce more accurate error rates, we implemented our own vanilla Machine Learning Algorithms as a comparison. We also had people perform exactly the same task as the Machine Learning Algorithms to indicate whether the way forward lies in improving Machine Learning Algorithms or in improving the data sets used to train Machine Learning Algorithms. The results from all these experiments feed into our other work transforming the Penn TreeBank into a more useful resource for training Machine Learning Algorithms to do Prepositional Phrase attachment.
BACKGROUND: The age-related decline of dehydroepiandrosterone (DHEA) has prompted research on its experimental replacement in women. Although no relationship to sexual functioning in healthy women has been shown to date, DHEA replacement has potential for affecting sexual response. METHODS: To investigate DHEA effects, 16 sexually functional postmenopausal women participated in a randomized, double-blind, crossover protocol in which oral administration of DHEA (300 mg) or placebo occurred 60 minutes before the presentation of an erotic video segment. Blood DHEA sulfate (DHEAS) changes, subjective and physiological sexual responses, as well as affective responses were measured in response to videotaped neutral and erotic video segments. RESULTS: The concentration of DHEAS increased 2-5-fold following DHEA administration in all 16 women. Subjective ratings across DHEA and placebo conditions showed significantly greater mental (p < 0.016) and physical (p < 0.036) sexual arousal to the erotic video with DHEA vs. placebo. Positive affect also increased during the erotic video across drug conditions. Vaginal pulse amplitude (VPA) and vaginal blood volume (VBV) demonstrated a significant increase (p < 0.001) between neutral and erotic film segments within both conditions (DHEA and placebo) but did not differentiate drug conditions. CONCLUSION: In sum, increases in mental and physical sexual arousal ratings significantly increased in response to an acute dose of DHEA in postmenopausal women.
This paper starts by discussing the reasons why linguists should be interested in parallel corpora. I outline the questions that parallel corpora enable us to ask, and relate them to traditional questions in linguistics and translation theory. The paper then suggests a method for arriving at answers to some of these questions. The proposed method builds on the notion of “ modulation” from Vinay and Darbelnet (1958) and attempts to put this notion on a sounder theoretical and empirical basis. It also includes a method of sharing data from different language pairs in a “ Contrastive Linguistic Database”.
The paper deals with current lexical databases that are seen as a basis for broad-coverage general-purpose ontologies. Various extensions and refinements of existing multi-lingual lexical knowledge bases are proposed with the aim of improving the capabilities of these resources. The main goal lies in the effort to gain a better lexical knowledge representation, which is crucial to coping with the requirements of the Semantic Web. The final section discusses the question of how lexical knowledge bases can be shared and combined. It presents the designed and implemented system WOMANISER that is able to merge independently developed parts of ontologies, check inconsistencies and report errors. The paper ends with the future directions of this research.
While the corpus-based research relies on human annotated corpora, it is often said that a non-negligible amount of errors remain even in frequently used corpora such as Penn Treebank. Detection of errors in annotated corpora is important for corpus-based natural language processing. In this paper, we propose a method to detect errors in corpora using support vector machines (SVMs). This method is based on the idea of extracting exceptional elements that violate consistency. We propose a method of using SVMs to assign a weight to each element and to find errors in a POS tagged corpus. We apply the method to English and Japanese POS-tagged corpora and achieve high precision in detecting errors.
BACKGROUND: Prostatodynia is a common and often disabling condition that affects males and has the characteristics of a somatoform pain disorder. It presents with urogenital pain and urinary symptoms. Failure of conventional treatment and a successful uncontrolled pilot study with fluvoxamine in this condition prompted this study. METHOD: In a randomized double-blind trial, 42 patients with prostatodynia were assigned to receive either fluvoxamine (N = 21) or placebo (N = 21) for up to 8 weeks. Doses were adjusted according to therapeutic need. The median dose of fluvoxamine was 150 mg (range, 50-300 mg). Self-rated pain scores, urinary flow rates, and depression and anxiety scores were measured at baseline and several times throughout the study period. RESULTS: The groups were similar at baseline, and the results were examined by intent-to-treat analysis either using the last observation carried forward or, in the case of dichotomous measures, counting treatment dropouts as treatment failures. Fluvoxamine was significantly more likely to reduce pain intensity (p =.01) and normalize urinary flow rates (p =.03) with a clinically significant number needed to treat value of 1.5 (confidence interval = 1.12 to 5.50). This therapeutic effect could not be attributed to change in mood, as the 2 groups did not differ with respect to affective ratings at the end of the study. The fluvoxamine-treated group had significantly lower (p =.02) final scores on the General Health Questionnaire, indicating an overall benefit from pain relief. CONCLUSION: Fluvoxamine is a viable treatment for prostatodynia. Dose-ranging studies and longer trials are needed to evaluate this agent further.
<h3>Introduction</h3><br> Rhetorical Structure Theory (RST) Discourse Treebank was developed by researchers at the Information Sciences Institute (University of Southern California), the US Department of Defense and the Linguistic Data Consortium (LDC). It consists of 385 Wall Street Journal articles from the <a href="http://catalog.ldc.upenn.edu/LDC99T42" rel="nofollow">Penn Treebank</a> annotated with discourse structure in the RST framework along with human-generated extracts and abstracts associated with the source documents. <br> In the RST framework (Mann and Thompson, 1988), a text's discourse structure can be represented as a tree in four aspects: (1) the leaves correspond to text fragments called <em>elementary discourse units</em> (the mininal discourse units); (2) the internal nodes of the tree correspond to contiguous text <em>spans</em>; (3) each node is characterized by its <em>nuclearity</em>, or essential unit of information; and (4) each node is also characterized by a <em>rhetorical relation</em> between two or more non-overlapping, adjacent text spans. <br> <h3>Data</h3><br> The data in this release is divided into a training set (347 documents) and a test set (38 documents). All annotations were produced using a discourse annotation tool that can be downloaded from <a href="http://www.isi.edu/~marcu/discourse" rel="nofollow">http://www.isi.edu/~marcu/discourse</a>. <br> Human-generated material in the corpus includes (1) long and short abstracts for 30 documents that were intended to convey the essential information and the main topic of the article, respectively; and (2) long, short and informative extracts for 180 documents, some of which were created from scratch and some of which were derived from the humanly-producted abstracts indicated above. <br> <h3>Samples</h3><br> Please view this <a href="desc/addenda/LDC2002T07.txt">sample</a>. <br> <h3>Updates</h3><br> There are no updates at this time. </br> Portions © 1987-1989 Dow Jones & Company, Inc., © 1995, 1999, 2002 Trustees of the University of Pennsylvania
In the field of empirical natural language processing, researchers constantly deal with large amounts of marked-up data; whether the markup is done by the researcher or someone else, human nature dictates that it will have errors in it. This paper will more fully characterise the problem and discuss whether and when (and how) to correct the errors. The discussion is illustrated with specific examples involving function tagging in the Penn treebank.
The development of large coverage, rich unification- (constraint-) based grammar resources is very time consuming, expensive and requires lots of linguistic expertise. In this paper we report initial results on a new methodology that attempts to partially automate the development of substantial parts of large coverage, rich unification-(constraint-) based grammar resources. The method is based on a treebank resource (in our case Penn-II) and an automatic f-structure annotation algorithm that annotates treebank trees with proto-fstructure information. Based on these, we present two parsing architectures: in our pipeline architecture we firstextract a PCFG from the treebank following the method of [Charniak, 1993; Charniak, 1996], use the PCFG to parse new text, automatically annotate the resulting trees with our f-structure annotation algorithm and generate proto-f-structures. By contrast, in the integrated architecture we firstautomatically annotate the treebank trees with fstructure information and then extract an annotated PCFG (A-PCFG) from the treebank. We then use the A-PCFG to parse new text to generate proto-fstructures. Currently
This mainly technological paper first provides a description of the web site called PapiLex. This first part shows how a file containing XML-structured lexical entries can be managed as a lexical database by using the Document Object Model (DOM) API. PapiLex offers the three essential management functions: creation, modification and deletion of a lexical entry. In a second part, two tools for entering Unicode-formatted text are presented: one for browsers having HTML 4 and JavaScript 1.2 capability and one for Microsoft Word. Such tools can be necessary for the minority languages which have no virtual keyboard embedded in the operating systems. 1 Starting point for building a lexical base Inside the Papillon project, the construction of a lexical base for a new language may take several different ways depending on where the author has to start. The following situations may occur regarding the availability of lexical resources1,2: • no dictionary exists, • a paper dictionary exists, • an electronic form of a dictionary exists, • a lexical database exists. In the last two cases, the question is to re-work the existing data so they meet the Papillon format and to fill the remaining fields. Tools are available for recycling electronic dictionary, (e.g. Nguyen 1998). Here, we will suppose that there is no preexisting dictionary or that its existence is limited to a paper dictionary. In such cases, the lexical entries have to be typed entirely. Among the 1: In addition to the existence of lexical resource, the script used for the language has also to be in Unicode and a font has to exist for it. Actually, the scripts of a number of minority languages are not in Unicode at the moment (e.g. Shan, Tai Dam, Mon). For some of them, fonts that really work are still missing as it is the case for Khmer. 2: In case there are existing data, property rights have to be looked at to say the resource is available. different ways in which this question can be handled, we chose a particular approach that consists in creating directly the Papillon formatted base by using generic and multiplatform Internet browsers. Section 2 will present how a standard browser can be used for this task3 (PapiLex mockup) and section 3 will show that a simple JavaScript program can provide a virtual keyboard that produces Unicode text. In section 4, another issue, less directly related to Papillon, will also be presented as it provides a very practical alternative for creating Unicodeencoded entries. It addresses a Windowsspecific tool for typing Unicode text in Microsoft Word when no standard keyboard is existing yet. The software was developed for the Lao language but can be applied to others. 2 The PapiLex mockup
Mobility is the most effective leveller of dialect and accent, and mobility constitutes a powerful linguistic force today. The sociolinguistics of mobility unites several disparate threads in my own research. First, immigration represents extreme mobility, and societies with profuse immigration differ in partly predictable ways linguistically and culturally from those with little or no immigration. Second, dialect acquisition by the children of newcomers provides new perspectives on critical period effects and influences, including the Ethan Experience, in which the nativization of children is abetted by their imperception of foreign‐accent features in their parents’ speech. Third, identification of relatively recently‐arrived people from other dialect regions allows comparisons of their linguistic norms with the communal norms, and a measure of their linguistic influence. From the cumulative results, we are in a position to frame hypotheses about linguistic variables in terms of their susceptibility to change and their resistance to it, and the identities of inhibitors and accelerators. All these threads should ultimately form integral aspects of the dynamics of dialect convergence.
The Specific Affect Coding System (SPAFF; J. M. Gottman & L. J. Krokoff, 1989) has led to conclusions about which types of dyadic affect predict positive and negative outcomes in marriage, yet the lack of information about collinearity among the codes limits interpretation of SPAFF results. Psychometric properties of SPAFF were examined by assessing the interactions of 172 newlywed couples with SPAFF and with an affect rating system developed for this study. For husbands and wives, factor analysis indicated 4 distinct factors of affect, representing anger/contempt, sadness, anxiety, and humor/affection. Anger/contempt and humor/affection were associated with marital satisfaction, relationship beliefs, and appraisals of the interactions. Correlations were in the expected directions. The strengths, limitations, and implications of the data are discussed.
Ambiguity resolution in the parsing of natural language requires a vast repository of knowledge to guide disambiguation. An effective approach to this problem is to use machine learning algorithms to acquire the needed knowledge and to extract generalizations about disambiguation decisions. Such parsing methods require a corpus-based approach with a collection of correct parses compiled by human experts. Current statistical parsing models suffer from sparse data problems, and experiments have indicated that more labeled data will improve performance. In this dissertation, we explore methods that attempt to combine human supervision with machine learning algorithms to try and extend accuracy beyond what is possible with the use of limited amounts of labeled data. In each case we do this by exposing a machine learning algorithm to unlabeled data in addition to the existing labeled data. Most recent research in parsing has shown the advantage of having a lexicalized model, where the word relationships mediate knowledge about disambiguation decisions. We use Lexicalized Tree Adjoining Grammars (TAGs) as the basis of our machine learning algorithm since they arise naturally from the lexicalization of Context Free Grammars (CFGs). We show in this dissertation that probability measures applied to TAGs retain the simplicity of probabilistic CFGs along with its elegant formal properties and that while PCFGs need additional independence assumptions to be useful in statistical parsing, no such changes need to be made to probabilistic TAGs. The main results presented in this dissertation are: (1) We extend the Co-Training algorithm (Yarowsky 1995; Blum and Mitchell 1998), a machine learning technique for combining labeled and unlabeled data previously used with classifiers with 2/3 labels to the more complex problem of statistical parsing. Using empirical results based on parsing the Wall Street Journal corpus we show that training a statistical parser on the combined labeled and unlabeled data strongly outperforms training only on the labeled data. (2) We present a machine learning algorithm that can be used to discover previously unknown subcategorization frames. The algorithm can then be used to label dependents of a verb in a treebank as either arguments or adjuncts. We use this algorithm to augment the Czech Dependency Treebank with argument/adjunct information. (3) We extend a supervised classifier for automatically identifying verb alternation classes for a set of verbs so that it can be used on minimally annotated data. Previous work (Merlo and Stevenson 2001) provided a classifier for this task that used automatically parsed text. With the use of learning of subcategorization frames we construct the same type of classifier which now requires text annotated with part-of-speech tags and phrasal chunks. In each of these results we use some existing linguistic resource that has been annotated by humans and add some further significant linguistic annotation by applying statistical machine learning algorithms.
Dictionaries can be used as a basis for lexicon development for NLP applications. However, it often takes a lot of pre-processing before they are usable. In the last 5 years a product-independent database of formal word features has been developed on the basis of the Van Dale dictionaries for Dutch. The database has proven to be useful in various NLP applications. This paper describes the history, some advantages and the constraints in the development of this database.
This article focuses on technical translation and the demands imposed on subject-field expert translators who must decide how they can reconcile the linguistic constraints imposed by a particular language with the communicative expectations found in a particular domain. Our main hypothesis is that experts typically resort to syntactic calquing to render phraseological units such as lexical collocations. By so doing they reinforce their role as language planners not only by introducing terminology into the TL but also by imposing SL norms within the expert community. Hence, the role of translation during the process of term documentation should be enhanced.
SLIM is a prototype interactive multimedia self-learning linguistic software for foreign language students at beginner-false beginner level. It allows students to work both in an autonomous self-directed mode or in a way of programmed learning in which the process of self-instruction is pre-programmed and monitored. In this latter mode it incorporates assessment and evaluation tools in order to behave as an automatic tutor. It is organized into three basic components: audiovisual materials; a linguistic database recording all language material in text format; the supervisor. Audiovisual materials are partially taken from commercially available courses; the linguistic database is a highly sophisticated classification of all words and utterances of the course, both in written and spoken form, from all possible linguistic aspects. The supervisor is both an attractive, enjoyable and strongly pedagogically based software that allows the user to work on language materials. The most outstanding feature of SLIM is the use of speech analysis and recognition which is a fundamental aspect of all second language learning programmes. We also assume that a learning model can be represented by a finite state automaton made up by a fixed number of possible states – corresponding to the macro and microlevels at which the student's competence may be modelled – each one being internally constituted by the actual linguistic objects of knowledge of the language that make it up.
Ever since the widespread availability of the Penn Treebank [9], there have been numerous, statistical parsers developed for English, e.g. [8, 5, 3]. To varying degrees, these parsers and others---while very successful at the tasks for which they were designed---had the following limitations:
The PAPILLON project aims at creating a cooperative, free, permanent, web-oriented and personalizable environment for the development and the consultation of a multilingual lexical database. The initial motivation is the lack of dictionaries, both for humans and machines, between French and many Asian languages. In particular, although there are large F-J paper usage dictionaries, they are usable only by Japanese literates, as they never contain both original (kanji/kana) and romaji writing. This applies as well to Thai, Vietnamese, Lao, etc.
Introduction to WordNet: an on-line lexical database. International Journal of Lexicography, 3(4):235--44. Miller, G. and Charles, W. (1991). Contextual correlates of semantic similarity. Language and Cognitive Processes, 6(1):1--28. Moon, R., editor (2000). Collins Cobuild Dictionary of Phrasal Verbs. Harper Collins. M.P. Marcus, B. S. and Marcinkiewicz, M. (1993). Building a large annotated corpus of english: the penn treebank. Computational Linguistics, 19(2):313--30. Pearce, D. (2002). A comparative evaluation of collocation extraction techniques. In Proceedings of Third International Conference on Language Resources and Evaluation. Pedersen, T. (2002). Distance 0.1. Pulman, S. G. (1993). The recognition and interpretation of idioms. In Cacciari, C. and Tabossi, P., editors, Idioms: Processing, Structure and Interpretation, chapter 11. Lawrence Erlbaum Associates, Hillsdale, NJ. Sag, I., Baldwin, T., Bond, F., Copestake, A., and Flickinger, D. (2002). Multiword expressions: A
We report in this paper on an experiment on automatic extraction of a Tree Adjoining Grammar from the WSJ corpus of the Penn Treebank. We use an automatic tool developed by (Xia, 2001) properly adapted to our particular need. Rather than addressing general aspects of the automatic extraction we focus on the problems we have found to extract a linguistically (and computationally) sound grammar and approaches to handle them.
The annotation of the Prague Dependency Treebank is realized in two sub-collections which differ in the subtlety of annotation (the large collection and the model collection). In the present paper, we focus on deletions of complementations of verbs, postverbal nouns and adjectives, from the point of view of the annotators of the model collection. We inquire into the issues of deletions of participants of verbs with respect to the coreferential relations between the restored node and its antecedent, and we introduce the new type of deletion where the deleted node has no concrete antecedent (we call these restored nodes vague, unspecified anaphoric elements). We also specify the differences between the newly introduced lemma Unsp(ecified) and the lemma Gen(eral), used for deletions of General Participants. After analysing the process of nominalization, we extend the description of the particular types of deletions also to the postverbal nouns denoting action. 1. Three-layer system of tags The annotation of the Prague Dependency Treebank (PDT in the sequel) is basically conceived of in accordance with the theoretical assumptions of the Functional Generative Description (FGD in the sequel, see Sgall, Hajicova, and Panevova, 1986; Hajicova, 1993). The present paper deals with the current phase of annotation of PDT at the so-called tectogrammatical level (which captures underlying syntactic structures of sentences). This level of annotation, in contrast to the two preceding phases of annotation, namely the morphemic and the so-called analytical levels (for a description of the annotation scheme of PDT see e.g. Hajic et al., 2001, Hajicova et al., 2001), is realized in two steps (in two subcollections) which differ in the subtlety of annotation. The first, basic step of annotation is represented by the so-called basic or large collection (LC; today it contains about 25,000 sentences). The second step of annotation, the so-called model collection (MC), should provide full, detailed information about the underlying structure of the sentence. Due to the fact that this way of annotation is more detailed, therefore more elaborate, the model collection contains only about 400 sentences today. 2. General principles of annotation at the tectogrammatical level Tectogrammatical tree structures (TGTSs) are based on dependency syntax; they have the shape of a dependency tree with the verb as the root of the tree and its daughter nodes representing nodes depending on the governor (on each layer of the tree). The two dimensions of the tree represent the syntactic structure of the sentence (the vertical dimension) and the topic-focus articulation of the sentence, based on the underlying word order (the horizontal dimension). The tagging at the tectogrammatical level can be described by the following principles: (a) a single node of a TGTS may be a representation of more than one word; only autosemantic words have a node of their own, while the correlates of functional words (auxiliaries, prepositions etc.) are attached to the autosemantic words to which they belong (auxiliary verbs and subordinating conjunctions to the verbs, prepositions to nouns, etc.); (b) in the cases of deletion in the surface shape of the sentence, nodes are introduced into the tectogrammatical tree to 'recover' a deleted word; The Prague Bulletin of Mathematical Linguistics 78, 2002 38 (c) no non-projective structures are admitted at the tectogrammatical level (projectivity is a counterpart to continuity of constituents; non-projectivity appearing in the surface shape of some sentences is supposed to be handled by movement rules between the tectogrammatical tree and the morphemic string); (d) not only the direction of the dependence on the governing node (dependence to the left, dependence to the right) is taken into account, but also sister nodes are ordered (from left to right); their order reflects the scale of communicative dynamism which is relevant for the description of topic-focus articulation.
1 Jonas Kuhn was at the University of Stuttgart when the main work reported in this paper was performed. Creation of high-quality treebanks requires expert knowledge and is extremely time consuming. Hence applying an already existing grammar in treebanking is an interesting alternative. This approach has been pursued in the syntactic annotation of German newspaper text in the TIGER project. We utilized the large-scale German LFG grammar of the PARGRAM project for semi-automatic creation of TIGER treebank annotations. The symbolic LFG grammar is used for full parsing, followed by semi-automatic disambiguation and automatic transfer into the treebank format. The treebank annotation format is a ‘hybrid ’ representation structure which combines constituent analysis and functional dependencies. Both types of information are provided by the LFG analyses. Although the grammar and the treebank representations coincide in core aspects, e.g. the encoding of grammatical functions, there are mismatches in analysis details that are comparable to translation mismatches in natural language translation. This motivates the use of transfer technology from machine translation. The German LFG grammar analyzes on average 50 % of the sentences, roughly 70 % thereof are assigned a correct parse; after OT-filtering, a sentence gets 16.5 analyses on average (median: 2). We argue that despite the limits in corpus coverage the applications of the grammar in treebanking is useful especially for reasons of consistency. Finally, we sketch future extensions and applications of this approach, which include partial analyses, coverage extension, annotation of morphology, and consistency checks.
This study investigated the human eyeblink startle reflex as a measure of alcohol cue reactivity. Alcohol-dependent participants early (n = 36) and late (n = 34) in abstinence received presentations of alcohol and water cues. Consistent with previous research, greater salivation and higher ratings of urge to drink occurred in response to the alcohol cues. Differential salivary and urge responding to alcohol versus water cues did not vary as a function of abstinence duration. Of special interest was the finding that startle response magnitudes were relatively elevated to alcohol cues, but only in individuals early in abstinence. Affective ratings of alcohol cues suggested that alcohol cues were perceived as aversive. Methodological and theoretical implications of the findings are discussed.
This paper discusses issues in building a 54-thousand-word Korean Treebank using a phrase structure annotation, along with developing annotation guidelines based on the morphosyntactic phenomena represented in the corpus. Various methods that were employed for quality control are presented. The evaluation on the quality of the Treebank and some of the NLP applications under development using the Treebank are also presented.
To assess the effects of discrepancy between two independent variables, investigators sometimes compute difference scores and correlate such scores with a criterion variable. However, the correlation of the difference with the criterion is accounted for by the correlations of the difference constituents with the criterion and the constituents’ variances. It follows that when investigators are testing a prediction that is not captured by the difference constituents’ main effects, using the difference correlation analysis may be misleading. Under these circumstances, the effects of a discrepancy between two independent variables can be assessed by a test of their interaction. The problems inherent in using difference scores and the advantage of testing the interaction are illustrated in relation to research programs on two separate topics in social psychology.
Lexical-Functional Grammar f-structures are abstract syntactic representations approximating basic predicate-argument structure. Treebanks annotated with f-structure information are required as training resources for stochastic versions of unification and constraint-based grammars and for the automatic extraction of such resources. In a number of papers (Frank, 2000; Sadler, van Genabith and Way, 2000) have developed methods for automatically annotating treebank resources with f-structure information. However, to date, these methods have only been applied to treebank fragments of the order of a few hundred trees. In the present paper we present a new method that scales and has been applied to a complete treebank, in our case the WSJ section of Penn-II (Marcus et al, 1994), with more than 1,000,000 words in about 50,000 sentences.
While initial treebanks and treebank parsers primarily involved surface analysis, recent work focuses on predicate argument (PA) structure. PA structure provides means to regularize variants (e.g., actives/passives) of sentences so that individual patterns may have better
^j;=5!? ITHIN the rich corpus of metrical psalms comtm itt^ta * posed during Spain's Golden Age, Fray Luis de Aivi bl T 0_Le6n's versions are universally accorded the high;@ ^ VV |@ est praise. As heir to a literary tradition that ex,.A s iGne * tended back to the late Middle Ages and early.f4Li J Renaissance, the Salamancan scholar and poet revolutionized Spain's engagement with the Psalter, establishing the lira or estrofa alirada as the dominant verse form for vernacular psalm translations (Rivers 112; Nufiez 357), and making close lexical parallelism and philological accuracy, rather than interpretive digression, the norm for most of his followers. As may be expected from the great Augustinian's role as el primer poeta humanista espaniol en lengua vulgar (A. Blecua 97), Fray Luis was widely imitated, especially among disciples of his own order, and questions of authorship and dating of the many psalm versions attributed to him continue to trouble literary historians (Nufiez 357-8; J. M. Blecua Poesia completa, 41-2). Jose Manuel Blecua, in his 1990 edition of the Poesia completa, based on all extant manuscripts, includes as genuine the following poems: Psalm 1 Beatus vir, 4 Cum invocarem, 6 ne in furore, 9 Salvum me fac, 12 Usquequo, Domine (2 versions), 17 Diligam te, 18 Coeli enarrant, 24 Ad te, Domine, levavi,
Quantitative evaluation of parsers has traditionally centered around the PARSEVAL measures of crossing brackets, (labeled) precision, and (labeled) recall. However, it is well known that these measures do not give an accurate picture of the quality of the parsers output. Furthermore, we will show that they are especially unsuited for partial parsers. In recent years, research has concentrated on dependencybased evaluation measures. We will show in this paper that such a dependency-based evaluation scheme is particularly suitable for partial parsers. TüBa-D, the treebank used here for evaluation, contains all the necessary dependency information so that the conversion of trees into a dependency structure does not have to rely on heuristics. Therefore, the dependency representations are not only reliable, they are also linguistically motivated and can be used for linguistic purposes.
Korean Combinatory Categorial Grammar (KCCG) is an extendedcombinatory categorial grammar formalism to capture thesyntax and interpretation of a relative freess word order, longdistance scrambling, and other specific characteristics of Korean.KCCG formalism can uniformly handle word order variations amongarguments and adjuncts within a clause, as well as in complexclauses and across clause boundaries, i.e. long distancescrambling. The approach we develop takes advantage of the ability of CCGfor type raising and composition along with the ability of variablecategories and unordered argument modeling for relatively freeword order treatment (Lee et al., 1994; Lee et al., 1997).We apply a probability model and heuristics using Koreancharacteristics to our KCCG parser.Results of the experiments on varioustext genre show that the KCCG parser performsat 87.67/87.03% constituent precision/recall.
This paper describes the CTB Coreference Annotation Guidelines for annotating pronominal anaphoric expressions in the Penn Chinese Treebank. The goals of the annotation are: to provide training data for learning-based pronoun resolution tools, and to provide a "gold" standard to be used in the evaluation of pronoun resolution algorithms. The choices that were made concerning the coindexing of pronominal anaphors and their antecedents are discussed, as are some questions that arose in trying to categorize those pronominal expressions that did not refer to specific nominal entities in the text.
In pace with the success of corpus-based approaches to theoretical and computational linguistics, the collocation of corpora has evolved into a research activity in its own. As the currently available corpora either lack annotation depth or closure, more data will be annotated in the future, preferably with minimal human intervention. This paper tries to approach the problem of treebank development from a logic-based learning perspective, applying several alternative forms of inference in order to assess their potential for automatically generalizing from a seed corpus annotated by hand to a corpus of POS annotated sentences, in order to automatically produce syntactic annotations that are good enough to use as training material for a parser. We shall show that syntactic annotations can be created automatically in large quantity via deductive and abductive explanation-based learning (EBL). Although these automatically created structures are not statistically representative with respect to many quantitative aspects of the treebank, the annotations may provide useful qualitative and quantitative data which might be extracted and reinvested into a parser. We shall compare the benefits and investments of automatically created structures to that of human-annotated structures and suggest some possible strategies how EBL approaches can be combined with manual annotation.
We present a new approach to topological parsing of German which is corpus-based and built on a simple model of probabilistic CFG parsing. The topological field model of German provides a linguistically motivated, flat macro structure for complex sentences. Besides the practical aspect of developing a robust and accurate topological parser for hybrid shallow and deep NLP, we investigate to what extent topological structures can be handled by context-free probabilistic models. We discuss experiments with systematic variants of a topological treebank grammar, which yield competitive results.
This paper describes a lexicalized tree adjoining grammar (LTAG) based parsing system for Korean which combines corpus-based morphological analysis and tagging with a statistical parser. Part of the challenge of statistical parsing for Korean comes from the fact that Korean has free word order and a complex morphological system. The parser uses an LTAG grammar which is automatically extracted using LexTract (Xia et al., 2000) from the Penn Korean TreeBank (Han et al., 2002). The morphological tagger/analyzer is also trained on the TreeBank. The tagger/analyzer obtained the correctly disambiguated morphological analysis of words with 95.78/95.39% precision/recall when tested on a test set of 3,717 previously unseen words. The parser obtained an accuracy of 75.7% when tested on the same test set (of 425 sentences). These performance results are better than an existing off-the-shelf Korean morphological analyzer and parser run on the same data
The LinGO Redwoods initiative is a seed activity in the design and development of a new type of treebank. While several medium- to large-scale treebanks exist for English (and for other major languages), pre-existing publicly available resources exhibit the following limitations: (i) annotation is mono-stratal, either encoding topological (phrase structure) or tectogrammatical (dependency) information, (ii) the depth of linguistic information recorded is comparatively shallow, (iii) the design and format of linguistic representation in the treebank hard-wires a small, predefined range of ways in which information can be extracted from the treebank, and (iv) representations in existing treebanks are static and over the (often year- or decade-long) evolution of a large-scale treebank tend to fall behind the development of the field. LinGO Redwoods aims at the development of a novel treebanking methodology, rich in nature and dynamic both in the ways linguistic data can be retrieved from the treebank in varying granularity and in the constant evolution and regular updating of the treebank itself. Since October 2001, the project is working to build the foundations for this new type of treebank, to develop a basic set of tools for treebank construction and maintenance, and to construct an initial set of 10,000 annotated trees to be distributed together with the tools under an open-source license.
Is there a general model that can predict the perceived phrase structure in language and music? While it is usually assumed that humans have separate faculties for language and music, this work focuses on the commonalities rather than on the differences between these modalities, aiming at finding a deeper 'faculty'. Our key idea is that the perceptual system strives for the simplest structure (the 'simplicity principle'), but in doing so it is biased by the likelihood of previous structures (the 'likelihood principle'). We present a series of data-oriented parsing (DOP) models that combine these two principles and that are tested on the Penn Treebank and the Essen Folksong Collection. Our experiments show that (1) a combination of the two principles outperforms the use of either of them, and (2) exactly the same model with the same parameter setting achieves maximum accuracy for both language and music. We argue that our results suggest an interesting parallel between linguistic and musical structuring.
You have accessThe ASHA LeaderFeature1 Nov 2002AAC, Literacy and Bilingualism Ovetta L. Harrison-Harris Ovetta L. Harrison-Harris Google Scholar https://doi.org/10.1044/leader.FTR2.07202002.4 SectionsAbout ToolsAdd to favorites ShareFacebookTwitterLinked In Children who use augmentative and alternative communication (AAC) have historically been challenged in their attainment of literacy skills. These challenges are even greater for AAC users who are bilingual. AAC users in the United States comprise large numbers of individuals from culturally and linguistically diverse backgrounds. Current demographic trends indicate that linguistic diversity will continue to intensify. During the 12 years between 1986 and 1998, the number of U. S. children who were identified as limited English proficient increased from 1.6 million to 9.9 million (see Tucker 1999). It is estimated that, by the year 2050, 40% of school-aged children in the United States will come from homes where English is not the first language. Individuals who use AAC systems surely will be represented in this group. The fact that many children in the United States, including those who use AAC systems, live amidst a sea of languages has captured national attention and has influenced our educational system. The new thrust to achieve educational equality represents a historic change. Many bilingual or monolingual schools that taught in languages other than English existed before World War II. For example, many German-only schools could be found in the northern Midwest. Afterward, a pattern of English-only instruction dominated our education system. As recognition of the cultural and linguistic diversity of the United States grew, a need to provide effective and appropriate education for bilingual children arose. Educators, parents, and researchers have challenged the notion of an English-only education for children from linguistically diverse backgrounds. Research supports the notion of education for limited-English-proficient children, including those relying on AAC systems, to be introduced in their first language, providing a transition to stronger second-language usage. This is logical given the fact that literacy attainment depends on language. Language learning, including reading and writing, is always culturally based. Reading and writing involve particular ways of using and thinking about written language that go beyond finding meaning in text and include the construction of sociocultural viewpoints or ways of understanding the world around us. It is important to realize the sociocultural and communicative nature of literacy, because of the possible therapeutic impact when working with bilingual AAC users. Writing, similarly, is a contextualized social event. It is a transactional, circular process created from a person's linguistic resources and interaction with past experiences. Viewing literacy learning as it is socially constructed through language provides a nice perspective of the need to educate linguistically and culturally diverse children from their first-language knowledge base. Research Challenges The challenges of literacy attainment for both monolingual and bilingual children who use AAC have become an area of focus for special educators, speech-language pathologists, parents, and researchers. Although research in reading and writing development of AAC users has increased steadily over the past 10 years, only recently have researchers turned their attention to reading and writing development of bilingual AAC users. Many of these children are unsuccessful in developing literacy, yet there is increasing recognition that this group is capable of developing sophisticated reading and writing skills. Bilingual AAC users who are highly successful in developing these skills make tremendous gains in overall language development and in use of their AAC systems. Acquisition of more vocabulary and the ability to compose text are just two advantages that literacy attainment brings to their receptive and expressive language development. Major focus has been brought to the topic of literacy attainment for bilingual AAC users because of its particular importance for this population. Attainment of literacy allows bilingual and monolingual AAC users, like all students, to be able to prepare messages to be used at a later time, produce exact messages, and learn vocabulary with which they can spontaneously spell out messages. But Light and McNaughton (l993) give three reasons why literacy development holds additional importance for AAC communicators. First, their face-to-face communication skills are often severely limited. Communication can be quite slow. Often the able-bodied message receiver doesn't have time to participate in communication interaction with an AAC communicator. Research shows that, in interactions between a person who is using an AAC system and a speaking person, the speaking person often dominates the interaction, and the person using the AAC system may not have opportunities to initiate topics or converse fully. Literacy gives an AAC communicator the opportunity to overcome many of the restrictions of face-to-face interaction, especially those imposed by slow AAC systems. Through writing it is possible for individuals to communicate more fully, to express themselves in more detail, and to circumvent some of the time limitations that they would normally experience in face-to-face interactions. The second aspect of school literacy importance for individuals who use AAC systems is that those who are preliterate are often limited to an ideographic literacy system. Some of these graphic systems force AAC communicators to use a closed vocabulary set and do not allow them to generate words to communicate new ideas. For example, an AAC communicator may operate a system composed of just 50 pictures or 100–200 ideographic symbols. They do not have access to the many thousands of concepts and ideas that they need in order to communicate fully and effectively. The use of orthographic literacy skills can be one way to open up access to a full range of concepts and vocabulary to students who use AAC. The literate AAC communicator, using traditional orthography, may spell words that are not printed on their communication boards or indicate first letters of words to which they don't have access on their communication system. In this way, they can use literacy skills to communicate in face-to-face interactions. The literacy development of augmentative communicators also may provide them with a means to participate in society by using written communication (as others also use written communication) to express opinions and give information. Using literacy as others do may help the bilingual AAC communicator advocate for bilingual education and acquire a sense of belonging to society as well as a stronger sense of value. The third way that literacy development carries added importance for bilingual AAC users involves vocational opportunities. In North America, there are very few individuals who use AAC systems who are competitively employed. The number holding white-collar jobs is few. The range of job opportunities available to individuals who have physical disabilities in general is restricted. AAC communicators are not usually employed in jobs requiring manual labor. Thus, they may need highly developed literacy skills for jobs involving, for example, data entry or word processing. Given limited vocational opportunities, the role of literacy in job preparation for bilingual AAC communicators is critical. Yvonne's Story AAC users must rely on innovative and sometimes creative strategies to learn to read, write, and monitor their understanding of what they are reading. Literacy-learning strategies for bilingual AAC users have not received as much attention as those of monolingual users. Some of the unique struggles and successes of literacy attainment can be seen in the story of Yvonne, a young Puerto Rican AAC user. Yvonne provides a wonderful example of the importance of first-language support and the use of specific literacy-learning strategies for bilingual AAC users. Yvonne is a 10-year-old girl with cerebral palsy of the spastic quadriplegic variety. She is nonambulatory and limited-speaking secondary to cerebral palsy. Her hearing and vision are within normal limits. During my initial contact with Yvonne, her intellectual functioning had not been formally determined. Yvonne's family immigrated to the United States one year before my initial contact with them. She is an only child. The primary language of the home is Spanish. Her father had limited English proficiency and her mother spoke no English at the time of my initial contact, although over the course of the school year they gained more proficiency. Another important characteristic of this family was the fact that the parents decided not to have any other children in order to devote total attention to Yvonne's education and health needs. Although no extended family lived in the area, they resided in a supportive neighborhood with other Puerto Ricans. Yvonne communicated primarily through use of an eye-gaze communication board. She used Mayer-Johnson Symbols and usually had a maximum of six symbols on her board. Other methods of communication included a smile/frown, yes/no response. A smile meant yes and a frown meant no. Yvonne also communicated by directing her eyes toward people or items that she wanted. Yvonne was not reading or writing very much in English when we first met. She may have recognized some English words that she encountered daily such as the names of her school, teacher, and classmates, and she had limited environmental vocabulary. I was not sure of her exact reading proficiency in Spanish; however, she did not demonstrate the ability to independently read upper-elementary-graded text w ritten in Spanish and answer basic content questions. Her listening comprehension for stories read to her in Spanish was good. We were not able to assess written language use because the classroom lacked the technology for text composition. Yvonne had a strong desire to learn to read more proficiently. Yvonne was a student in a general elementary school located in western Massachusetts. Her classroom was nongraded, but the students, all classified as special needs, were of comparable ages to those of fourth grade. The room was self-contained and designated by the school system as a special education classroom. The special need categories included physically and cognitively impaired. Half of the class comprised other Puerto Rican children. My role was that of AAC literacy consultant, but I also brought my expertise in the area of multiculturalism in speech-language pathology. My initial meeting with Yvonne occurred early in the school year, in her classroom with the classroom teacher and instructional aide. Yvonne immediately greeted me with a welcoming smile because she appeared to know that I was there especially to help her learn. During my initial meeting I was able to informally assess that Yvonne had good cognitive skills. She used her voice to initiate communication to bring attention to matters of need or interest. She laughed appropriately at jokes, her eyes followed speakers in a conversation, and she spontaneously used her eyes to appropriately answer yes/no questions. All of the conversations around her and directed to her by her teacher were in English. Yvonne obviously acquired some English proficiency, although she may not have understood everything. I had formal training in Spanish and worked some years earlier in a predominately Mexican-American school district in Southern California where I used the language daily. Although I lacked confidence in my use of Spanish, I greeted Yvonne and introduced myself in Spanish. Approaching her using Spanish set a tone for Yvonne that I was supportive of her background and language usage. She recognized that I needed help using the dialect of Spanish that she was familiar with as a primary way of communicating with her. We learned quickly to work together around the use of a language system. Honoring her first language was important to our working together. Another important factor was Yvonne's desire and willingness to learn English, which contributed significantly to her rapid acquisition of stronger English proficiency. On my second day of visiting the classroom, I was extremely pleased to meet the school SLP assigned to Yvonne. This wonderfully competent, energetic clinician just happened to be bilingual in English and Spanish. With a bilingual SLP and my knowledge of literacy-learning techniques for AAC users, Yvonne blossomed over the course of that academic year in her English proficiency and particularly in her ability to read and spell. A Successful Technique I first introduced a spelling/word-level reading technique to Yvonne that proved to be highly successful and allowed her to gain 10–12 new words in reading recognition and spelling each week. Upper-elementary-aged bilingual AAC users with profiles similar to Yvonne should start with whole-word-level reading aimed at teaching recognition of entire words such as swim, pool, the, or cap. Instruction of whole words leads to success in reading phrases and simple sentences quickly. Phonetic instruction should occur as well. The Words on the Wall technique, which can be used with monolingual as well as bilingual AAC users, begins by the teacher selecting approximately 3–5 new words that the student needs to learn. These should be words relevant to familiar situations and not spelling words from a spelling book. For example, Yvonne went swimming each week in school and thus, during her first week, she learned the words swimming, towel, pool, water, and splash. These words were initially introduced in Spanish only. The next step in this technique is to make the word accessible by writing it in large print on a sentence strip and attaching it to the wall. The word may initially be paired with a symbol, with the symbol being phased out over time leaving just the written word. The student and the teacher define the word and talk about events involving the target word. After all of the target words are discussed and displayed on the wall, the teacher asks the student to identify each word one at a time as in a spelling test. Yvonne used eye gaze to identify her target words. During the next day or week, depending on how well the student masters each set of words, introduce more words (1–3 a day). Leave all words on the wall for the school year, increasing the number of words each week. Review old and new words. After enough words are mastered, have the student begin to read simple sentences. Introduce words such as a and the to allow formation of sentences. The school SLP delivered all of the training to Yvonne in Spanish first and followed it with English only after she knew that Yvonne understood the word in Spanish. Because this literacy-learning technique is based primarily at the word level, it is easier to transition from the Spanish to the English word. The school SLP also kept in close contact with Yvonne's parents, phoning them and sending home each week the word that Yvonne was working on. Yvonne's communication reflected her increased vocabulary. A board in Spanish was sent home and used with her parents and an English board was used at school initially. As Yvonne's parents gained more English proficiency, they requested to have the English communication board as well. During the school year we piloted different types of high-tech AAC devices and switches with Yvonne. We also explored technology for writing purposes during this year. Assessment Words on the Wall lends itself to a Maze Reading Assessment technique once a student has acquired reading of simple sentences. This technique involves the deletion of target words in a sentence leaving a blank space. The student should be provided with three alternative words in random order at each blank (correct choice, incorrect choice of the same part of speech, incorrect choice of a different part of speech). For example: The boy ate a ______ (truck, this, banana). This technique can be used easily with many AAC users. Yvonne's eye gazed to her chosen word using this technique. The scale of reading proficiency most often used for informal reading assessments such as this is 90% accuracy indicating that the student is reading at an independent level, 60%–80% accuracy relating to a level where more instruction is needed, and below 60% is equivalent to a frustration level. For Yvonne, the Maze technique was delivered in English because she already had mastered the words on the wall and read simple sentences in English. The Words on the Wall technique and a Maze Reading Assessment Technique are two techniques that can be culturally and linguistically sensitive and used well with AAC users. Voice output is not required for these techniques, and the words are derived from the students' existing linguistic bases or contextual experiences. Other techniques also can be used to facilitate literacy development with bilingual AAC users. Techniques that contextualize instruction in the experiences of the home and first language are desirable. For young bilingual literacy-language learners, it is important to use interactive learning techniques that involve the teacher, peers, and the AAC user. Techniques that allow students to demonstrate competence in using language and literacy throughout the school day in all instructional activities are greatly beneficial. Techniques that use narratives such as storytelling, listening to stories, or writing are good for content development. These narratives should be delivered in the language that will allow the child to gain academic skill while learning English. My first year with Yvonne was a successful one. She gained approximately 10 new words a week over the course of the school year. For AAC users similar to Yvonne in age and cognitive ability, this is an expected rate of growth. There is no typical rate of growth for all AAC users because this population is so diverse in skill and ability. The Next Year I returned to visit Yvonne the next year when she had been promoted to a new class and school. The successful learning environment that she had previously experienced had come to an abrupt end. There was a lack of continuity with her education from the previous year. Yvonne was in a new school with all new staff. There was no Spanish language support. The literacy-learning methods had been abandoned. Communication with Yvonne was a problem. There was limited communication between the school and home. I spent the first few days in Yvonne's classroom as a participant observer and quickly assessed the social and literacy-learning needs of everyone involved in Yvonne's schooling. The goals of my intervention with Yvonne during this second school year included elimination of the communication problem between the school and the family and establishment of better trust and communication, reestablishment of appropriate instructional methods, eliminating AAC barriers, and supporting cultural identity through literacy lessons/interactions and development of a more efficient communication system. The lack of Spanish support and having to demonstrate and convince the new teachers of Yvonne's literacy-learning capabilities resulted in lost time in her development. Strong first-language support and knowledge of specific literacy-learning techniques for bilingual AAC users led to a successful outcome for Yvonne. She enjoys reading and had a strong desire to continue reading and learning English. This was compatible to the wishes of her parents. Like Yvonne, not all bilingual AAC users have significant difficulties learning to read and write; however, many of them do. Therefore, it becomes important to communication disorders specialists to identify variables of language that are predictive of later reading difficulties. Researchers and other professionals from different fields of study are combining their interests to close the knowledge/information gap that exists between what is already known about bilingual AAC users' acquisition and development and the information needed to help develop intervention strategies for successful written language. Strategies for Monolingual Clinicians: A Postscript Although I did have formal training in Spanish in high school and college and had worked in a predominately Spanish-speaking community in Southern California, I still lacked confidence to converse with Yvonne in Spanish when I first met her. I knew that there were many dialects of Spanish, and I initially did not know enough about the Spanish that she and her family used. Clinicians who are monolingual or who lack information about a second-language-speaking student must do the research to find linguistic information particular to that student. Such knowledge is also helpful in understanding the contexualized uses of literacy in the home that will complement those used in the classroom. General professional development in the area of bilingual literacy learning is highly recommended, as is professional development in AAC. Understanding policies in educating bilingual students that are implemented in your school district is important. Clinicians should understand how policy affects access to instruction for bilingual students. Social, cultural, and economic issues that affect student learning and instruction also should be well understood. It is helpful to gain information from parents, other teachers, and community members about ways that they find helpful in instructing bilingual AAC users. Ovetta Harrison-Harris is chair of the department of communication sciences and disorders at Howard University. She is project director for a U.S. Department of Education Office of Special Education and Rehabilitative Services-funded graduate training program in AAC with an emphasis in multiculturalism and literacy development. For More Information Light J., Binger C., & Smith A.K. (1994). Story reading interactions between pre-schoolers who use AAC and their mothers. Augmentative and Alternative Communication, 10, 225–268. CrossrefGoogle Scholar Light J., & McNaughton D. (1993). Literacy and Augmentative and Alternative Communication (AAC): Expectations and Priorities of Parent and Teachers. Topics and Language Disorders, 13(2), 33–46. CrossrefGoogle Scholar Light J., & Smith A.K. (1993). Home literacy experience of pre-schoolers who use augmentative communication systems and their non disabled peers. Augmentative and Alternative Communication, 9, 10–25. CrossrefGoogle Scholar Pearson B.Z., Fernandez S., & Oller D.K. (1993a). Lexical developmental in simultaneous bilingual infants: Comparison to monolinguals. Language Learning, 43, 93–120. CrossrefGoogle Scholar Pearson B.Z., Fernandez S.C., & Oller D.K. (1993b). Lexical development in bilingual infants and toddlers: Comparison to monolingual norms. Language Learning, 43(1), 93–120. CrossrefGoogle Scholar Pearson B.Z., Fernandez S., & Oller D.K. (1995). Cross-language synonyms in the lexicons of bilingual infants: One language or two?, Journal of Child Language, 22, 345–68. CrossrefGoogle Scholar Pearson B.Z., Oller D.K., Umbel V.M., & Fernandez M.C. (1996, October). The Relationship of Lexical Knowledge to Measures of Literacy and Narrative Discourse in Monolingual and Bilingual Children. Paper presented at the Second Language Research Forum, Tucson, Google Scholar Tucker A perspective on and bilingual education Google Scholar Ovetta L. is chair of the department of communication sciences and disorders at Howard University. She is project director for a U.S. Department of Education Office of Special Education and Rehabilitative Services-funded graduate training program in AAC with an emphasis in multiculturalism and literacy development. of the ASHA Special Augmentative and Alternative Communication, and a for With Communication to your in Nov &
A Swedish-English dictionary for the Internet is transformed into an English-Swedish counterpart by computationally reversing the Swedish and the English lexical database. Dictionary reversal of an existing bilingual dictionary is a possible solution of obataining working material quickly. However, depending on the complexity of the original database, the new reversed dictionary may have to be edited extensively. In the reversed dictionary, the original English target language becomes source language data. One outcome of reversing source and target language information is that the new English items are not necessarily translation equivalents. Internet dictionaries present new possibilities and challenges: they can easily be updated by allowing users to contribute new headwords, but lexicographers may also have to reconsider traditional methods in dictionary design. The article concludes by discussing the possible use of parallel corpus examples to illustrate language use in bilingual dictionaries.
This paper discusses issues in building a 54-thousand-word Korean Treebank using a phrase structure annotation, along with developing annotation guidelines based on the morpho-syntactic phenomena represented in the corpus.Various methods that were employed for quality control and the evaluation on the Treebank are also presented.'This word count is computed on tokenized texts and includes symbols.
The aim of this volume is to showcase the range of corpus-based linguistic research currently being carried out on languages other than English. The papers included report on work carried out on Arabic, Bulgarian, Czech, Dutch, French, German, Biblical Greek, Biblical Hebrew, Medieval Irish, Korean, Romanian and Swedish, including a number of regional and social variants. They also address a range of areas as diverse as corpus design, corpus annotation, register analysis, syntax, and quantitative linguistics. The papers in this volume will leave the reader in no doubt that corpus-based research is now being conducted for a whole rainbow of languages.