Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
I present data from Tambora, a now extinct language of central Sumbawa, and argue from the lexical data and the inferred phonology, compared with areal norms, that it was a Papuan language spoken by a trading population of southern Indonesia. The existence into historical times of a large and nonreclusive Papuan political entity this far west forces a major revision of our ideas about the linguistic macrohistory of Eastern Indonesia.
Some of the speech databases and large spoken language corpora that have been collected during the last fifteen years have been (at least partly) annotated with a broad phonetic transcription. Such phonetic transcriptions are often validated in terms of their resemblance to a handcrafted reference transcription. However, there are at least two methodological issues questioning this validation method. First, no reference transcription can fully represent the phonetic truth. This calls into question the status of such a transcription as a single reference for the quality of other phonetic transcriptions. Second, phonetic transcriptions are often generated to serve various purposes, none of which are considered when the transcriptions are compared to a reference transcription that was not made with the same purpose in mind. Since phonetic transcriptions are often used for the development of automatic speech recognition (ASR) systems, and since the relationship between ASR performance and a transcription’s resemblance to a reference transcription does not seem to be straightforward, we verified whether phonetic transcriptions that are to be used for ASR development can be justifiably validated in terms of their similarity to a purpose-independent reference transcription. To this end, we validated canonical representations and manually verified broad phonetic transcriptions of read speech and spontaneous telephone dialogues in terms of their resemblance to a handcrafted reference transcription on the one hand, and in terms of their suitability for ASR development on the other hand. Whereas the manually verified phonetic transcriptions resembled the reference transcription much closer than the canonical representations, the use of both transcription types yielded similar recognition results. The difference between the outcomes of the two validation methods has two implications. First, ASR developers can save themselves the effort of collecting expensive reference transcriptions in order to validate phonetic transcriptions of speech databases or spoken language corpora. Second, phonetic transcriptions should preferably be validated in terms of the application they will serve because a higher resemblance to a purpose-independent reference transcription is no guarantee for a transcription to be better suited for ASR development.
Language deviation is a linguistic device of purposeful violation of language norms,which serves not only as the necessity but also as the inevitability of language development.There are various forms of language deviation,including phonological deviation,lexical deviation,grammatical deviation,semantic deviation,graphological deviation,deviation of register and figurative deviation.It is the reflection of variety of language and its social nature.Language deviation does bring richer cultural connotation to every language.
This study employs corpus-based and computer-aided methodology to investigate Polish advanced learners' use of the past progressive in L2 English. The analysis is performed within the framework of the Aspect Hypothesis and it is based on narrative and descriptive essays (the total of 35,319 tokens) drawn from the PELCRA learner corpus. The learner data is contrasted with two sections of the FLOB corpus comprising native speakers' texts of a comparable genre. The results have demonstrated that, although Polish advanced learners on the whole can dissociate the use of tense-aspect morphology from its prototypical combinations with the lexical aspect and their use of the past progressive across the situation types reflects the native patterns, they tend to overuse past progressive forms in comparison to the native norm. This tendency can be explained by L1 influence, but it is also reinforced by writing instruction, which fails to sensitise students to the stylistic effects of employing preterite versus past progressive forms to describe the background situation in a narrative.
This article examines the corpus of multinationals’ codes of conduct on CSR issues which has been collated by the ILO. Through lexical software analysis we identify three main points of reference in CSR codes of conduct: respect for ILO norms, discussion of the company’s relationship to society, and reinforcement of its internal discipline and organisation. Surprisingly, the issue of corporate responsibility itself constitutes a small part of the text of the codes. Their main targets are employees, who are charged with a dual task: to ensure the implementation of the principles stated in the codes, and to protect the assets of the company. In a reflexive dimension, codes of conduct help us to understand the key characteristics of the companies which made them.
07–484 Aceto, Michael (East Carolina U, USA; acetom@ecu.edu ), Statian Creole English: An English-derived language emerges in the Dutch Antilles. World Englishes (Blackwell) 25.3 & 4 (2006), 411–435. 07–485 Anchimbe, Eric A. (U Munich, Germany), World Englishes and the American tongue. English Today (Cambridge University Press) 22.4 (2006), 3–9. 07–486 Bartha, Csilla & Anna Borbély (Hungarian Academy of Sciences, Budapest, Hungary; bartha@nytud.hu ), Dimensions of linguistic otherness: Prospects of minority language maintenance in Hungary. Language Policy (Springer) 5.3 (2006), 337–365. 07–487 Coetzee-Van Rooy, Susan (North-West U, Potchefstroom, South Africa; basascvr@puk.ac.za ), Integrativeness: Untenable for world Englishes learners? World Englishes (Blackwell) 25.3 & 4 (2006), 437–450. 07–488 Gooskens, Charlotte (U Groningen, The Netherlands; c.s.gooskens@rug.nl ) & Renée van Bezooijen, Mutual comprehensibility of written Afrikaans and Dutch: Symmetrical or asymmetrical? Literary and Linguistic Computing (Oxford University Press) 21.4 (2006), 543–557. 07–489 Gooskens, Charlotte & Wilbert Heeringa (U Groningen, The Netherlands; c.s.gooskens@rug.nl ), The relative contribution of pronunciational, lexical, and prosodic differences to the perceived distances between Norwegian dialects. Literary and Linguistic Computing (Oxford University Press) 21.4 (2006), 477–492. 07–490 Guilherme, Manuela (U De Coimbra, Portgual), English as a Global language and education for cosmopolitan citizenship. Language and International Communication (Multilingual Matters) 7.1 (2007), 72–90. 07–491 Koscielecki, Marek (The Open U, Hongk Kong, China). Japanized English, its context and socio-historical background. English Today (Cambridge University Press) 22.4 (2006), 25–31. 07–492 Meilin, Chen (Three Gorges University, China) & Hu Xiaoqiong, Towards the acceptability of China English at home and abroad. English Today (Cambridge University Press) 22.4 (2006), 44–52. 07–493 Mesthrie, Rajend (U Cape Town, South Africa; raj@humanities.uct.ac.za ), World Englishes and the multilingual history of English. World Englishes (Blackwell) 25.3 & 4 (2006), 381–390. 07–494 Poole, Brian (Ministry of Manpower, Muscat, the Sultanate of Oman), Some effects of Indian English on the language as it is used in Oman. English Today (Cambridge University Press) 22.4 (2006), 21–24. 07–495 Robinson, Ian (U Calabria, Italy), Genre and loans: English words in an Italian newspaper. English Today (Cambridge University Press) 22.4 (2006), 9–20. 07–496 Ross, Kathryn (U Oxford, UK; kathryn.ross@trinity.ox.ac.uk ), Status of women in highly literate societies: The case of Kerala and Finland. Literacy (Blackwell) 40.3 (2006), 171–178. 07–497 Sala, Bonaventure M. (Cameroon), Does Cameroonian English have grammatical norms? English Today (Cambridge University Press) 22.4 (2006), 59–64. 07–498 Wei-Yu Chen, Cheryl (National Taiwan Normal U, Taiwan; wychen66@hotmail.com ), The mixing of English in magazine advertisements in Taiwan. World Englishes (Blackwell) 25.3 & 4 (2006), 467–478. 07–499 Wong, Jock (National U Singapore, Singapore; jockonn@hotmail.com ), Contextualizing aunty in Singaporean English. World Englishes (Blackwell) 25.3 & 4 (2006), 451–466. 07–500 Xiaoxia, Cui (Yunnan U, China), An understanding of ‘China English’ and the learning and use of the English language in China. English Today (Cambridge University Press) 22.4 (2006), 40–43. 07–501 Young, Ming Yee Carissa (Macao U Science & Technology, Macau; myyoung@must.edu.mo ), Macao students' attitudes toward English: A post-1999 survey. World Englishes (Blackwell) 25.3 & 4 (2006), 479–490.
This paper deals with a multimodal annotation scheme dedicated to the study of gestures in interpersonal communication, with particular regard to the role played by multimodal expressions for feedback, turn management and sequencing. The scheme has been developed under the framework of the MUMIN network and tested on the analysis of multimodal behaviour in short video clips in Swedish, Finnish and Danish. The preliminary results obtained in these studies show that the reliability of the categories defined in the scheme is acceptable, and that the scheme as a whole constitutes a versatile analysis tool for the study of multimodal communication behaviour.
This is a study of German compound nouns, used metaphorically to refer to three life style manifestations characteristic of postmodern society: consumption, health and fitness orientation, and the pursuit of pleasure. The lexical items are classified according to the source domains of the respective metaphors, thus demonstrating the variety of perspectives that come into play and providing an insight into the ways people think as well as their attitudes and values with respect to the phenomena focussed on in the study. Common to all lexemes is an element of excess, which suggests that the phenomena in question are regarded as violations of societal norms.
Both classroom instruction and lexical database development stand to benefit from applied research on sign language, which takes into consideration American Sign Language rules, pedagogical issues, and teacher characteristics. In this study of technical science signs, teachers' experience with signing and, especially, knowledge of content, were found to be essential for the identification of signs appropriate for instruction. The results of this study also indicate a need for a systematic approach to examine both sign selection and its impact on learning by deaf students. Recommendations are made for the development of lexical databases and areas of research for optimizing the use of sign language in instruction. (PsycINFO Database Record (c) 2016 APA, all rights reserved)
This paper presents a scheme for ranking of spelling error corrections for Urdu. Conventionally spell-checking techniques do not provide any explicit ranking mechanism. Ranking is either implicit in the correction algorithm or corrections are not ranked at all. The research presented in this paper shows that for Urdu, phonetic similarity between the corrections and the erroneous word can serve as a useful parameter for ranking the corrections. This combined with a new technique Shapex that uses visual similarity of characters for ranking gives an improvement of 23% in the accuracy of the one-best match compared to the result obtained when the ranking is done on the basis of word frequencies only.
In this paper, we describe the application of a bidirectional dependency parser trained on the Turin University Treebank.
We introduce MaltParser, a data-driven parser generator for dependency parsing. Given a treebank in dependency format, MaltParser can be used to induce a parser for the language of the treebank. MaltParser supports several parsing algorithms and learning algorithms, and allows user-defined feature models, consisting of arbitrary combinations of lexical features, part-of-speech features and dependency features. MaltParser is freely available for research and educational purposes and has been evaluated empirically on Swedish, English, Czech, Danish and Bulgarian. 1.
We describe a test-time score normalization technique (T-Norm) for text-dependent speaker verification that is robust to lexical mismatch. The main challenge to the deployment of T-Norm in a text-dependent task is the mismatch between the lexicon of the target speaker model in the application and that of the cohort speaker models. We show the negative effect of that mismatch in controlled experiments and propose a hybrid scoring scheme (T-Norm and background model) to remedy it. In a lexically mismatched scenario, which is inherent to the deployment of T-Norm in a text-dependent system, we show a 31% relative error rate reduction using the hybrid scoring over T-Norm alone. A 22% relative error rate reduction is measured over the baseline (no T-Norm) system.
This Research Discusses about the interference of Betawi Melayu language in Indonesia cmguage by the witters of Journal Hai. Interference is a kind of deviation in using of the norms which existing as the effect of language contact or mastery> more than one language. Beside that:t also discribes about the cause of appearing the interference the gendre of interference which - onsisting of morphology, lexical, and grammatical level.
The standard language norm fulfils two basic requirements: stability of language and its development.The former covers replacing of foreign terms with Croatian equivalents or at least their adaptation according to the rules of the Croatian language.The latter implies fulfilling new lexical needs.The economic power of the United States of America is reflected in the influence of the English language on term-formation in Croatian.Acceptance of lexical innovations is primarily gained due to thelfnguage of the media.
Reduced speech fluency is frequent in clinical paediatric populations, an unexplained finding. To investigate age related effects on speech fluency variables, we analysed samples of narrative speech (picture description) of 308 healthy children, aged 5 to 17 years, and studied its relation with verbal fluency tasks. All studied measures showed significant developmental effects. Speech rate and verbal fluency scores increased, while pauses, repetitions and locution time declined with age. Speech rate correlated with semantic fluency tasks suggesting that it also depends upon the efficacy of lexical retrieval. These results indicate that the interpretation of disorders of speech fluency in childhood must incorporate age appropriate norms.
The Dutch spelling system, like other European spelling systems, represents a certain balance between preserving the spelling of morphemes (the morphological principle) and obeying letter-to-sound regularities (the phonological principle). We present experimental results with artificial learners that show a competition effect between the two principles: adhering more to one principle leads to more violations of the other. The artificial learners, memory-based learning algorithms, are trained (1) to convert written words to their phonemic counterparts and (2) to analyze written words on their morphological composition, based on data extracted from the CELEX lexical database. As an exception to the competition effect we show that introducing the schwa as a letter in the spelling system causes both morphology and phonology to be learnt better by the artificial learners. In general we argue that artificial learning studies are a tool in obtaining objective measurements on a spelling system that may be of help in spelling reform processes. (PsycINFO Database Record (c) 2016 APA, all rights reserved)
The present dissertation addresses a set of questions about processes involved in lexical access and literacy and psycholinguistic factors (such as word learning age, word frequency etc.) that affect them in monolingual and bilingual speakers. Four experiments examined these issues for nouns in native English speakers and bilingual Hindī- English speakers with a developmental perspective. Experiments 1 and 2 were conducted with English monolinguals in San Diego. In experiment 1, age of acquisition norms were collected from college-age adults. In experiment 2, online picture naming data was collected from four age groups of English monolinguals (5-7, 8-10, 11-13 and college-age adults). Experiments 3 and 4 were conducted on Hindī-English bilinguals in India. In experiment 3, age of acquisition and word frequency norms were collected from college-age adults. In experiment 4, online picture naming and word reading data were collected from three age groups of Hindī-English bilinguals (8-10, 11-13 and college-age adults). Comparisons of performance on two lexical access tasks (on-line picture naming and word reading) in monolingual English speakers and bilingual Hindī-English speakers, were conducted. Results and discussion are aimed at addressing issues of language processing, lexical access and development. Overall, results indicate that there is developmental improvement on the lexical access tasks. In addition, the predictor- outcome relationships are generally similar for both monolinguals and bilinguals. Age of acquisition is the most consistent predictor of both picture naming and word reading behavior, in both monolingual and bilingual speaker. There are differential effects of frequency in the languages of the bilingual in the word reading task, with orthographic differences interacting with frequency effects. However, there are interesting differences that arise between the monolinguals and within the bilinguals, because of language dominance and proficiency. Results and discussion focus on quantitative analyses, examining lexical access processes in monolinguals and bilinguals, and examining the relationship between the psycholinguistic variables (such as age of acquisition, frequency, and syllable length) and performance on the language production tasks. Future directions focus on highlighting some limitations of this research, in addition to discussing the need for more in depth qualitative analyses, and extending these paradigms to clinical populations
This paper describes experiments carried out utilizing a variety of machine-learning methods (the k-nearest neighborhood, decision list, maximum entropy, and support vector machine), and using six machine-translation (MT) systems available on the market for translating tense, aspect, and modality. We found that all these, including the simple string-matching-based k-nearest neighborhood used in a previous study, obtained higher accuracy rates than the MT systems currently available on the market. We also found that the support vector machine obtained the best accuracy rates (98.8%) of these methods. Finally, we analyzed errors against the machine-learning methods and commercially available MT systems and obtained error patterns that should be useful for making future improvements.
Wordnets, which are repositories of lexical semantic knowledge containing semantically linked synsets and lexically linked words, are indispensable for work on computational linguistics and natural language processing. While building wordnets for Hindi and Marathi, two major Indo-European languages, we observed that the verb hierarchy in the Princeton Wordnet was rather shallow. We set to constructing a verb knowledge base for Hindi, which arranges the Hindi verbs in a hierarchy of is-a (hypernymy) relation. We realized that there are unique Indian language phenomena that bear upon the lexicalization vs. syntactically derived choice. One such example is the occurrence of conjunct and compound verbs (called Complex Predicates) which are found in all Indian languages. This paper presents our experience in the construction of lexical knowledge bases for Indian languages with special attention to Hindi. The question of storing versus deriving complex predicates has been dealt with linguistically and computationally. We have constructed empirical tests to decide if a combination of two words, the second of which is a verb, is a complex predicate or not. Such tests provide a principled way of deciding the status of complex predicates in Indian language wordnets.
Spoken dialogue systems (SDSs) can be used to operate devices, e.g. in the automotive environment. People using these systems usually have different levels of experience. However, most systems do not take this into account. In this paper, we present a method to build a dialogue system in an automotive environment that automatically adapts to the user’s experience with the system. We implemented the adaptation in a prototype and carried out exhaustive tests. Our usability tests show that adaptation increases both user performance and user satisfaction. We describe the tests that were performed, and the methods used to assess the test results. One of these methods is a modification of PARADISE, a framework for evaluating the performance of SDSs [Walker MA, Litman DJ, Kamm CA, Abella A (Comput Speech Lang 12(3):317–347, 1998)]. We discuss its drawbacks for the evaluation of SDSs like ours, the modifications we have carried out, and the test results.
Syntactic parsing requires a fine balance between expressivity and complexity, so that naturally occurring structures can be accurately parsed without compromising efficiency. In dependency-based parsing, several constraints have been proposed that restrict the class of permissible structures, such as projectivity, planarity, multi-planarity, well-nestedness, gap degree, and edge degree. While projectivity is generally taken to be too restrictive for natural language syntax, it is not clear which of the other proposals strikes the best balance between expressivity and complexity. In this paper, we review and compare the different constraints theoretically, and provide an experimental evaluation using data from two treebanks, investigating how large a proportion of the structures found in the treebanks are permitted under different constraints. The results indicate that a combination of the well-nestedness constraint and a parametric constraint on discontinuity gives a very good fit with the linguistic data.
The current paper has a twofold objective. On the one hand, it describes the creation and the features of the Szeged Treebank, which is currently the largest manually processed Hungarian textual database serving as a reference material for research in natural language processing. On the other hand, detailed information is given about different experiments that aimed at the automatic recognition of syntactic structures with the use of machine learning algorithms. In order to provide comparable results, we applied methods of different categories, namely a rule-based, a logic and a numeric learner to pre-defined parsing problems. The aforementioned Szeged Treebank was used for the training and the testing of the algorithms.
Transforming syntactic representations in order to improve parsing accuracy has been exploited successfully in statistical parsing systems using constituency-based representations. In this paper, we show that similar transformations can give substantial improvements also in data-driven dependency parsing. Experiments on the Prague Dependency Treebank show that systematic transformations of coordinate structures and verb groups result in a 10% error reduction for a deterministic data-driven dependency parser. Combining these transformations with previously proposed techniques for recovering non-projective dependencies leads to state-of-the-art accuracy for the given data set.
Linguists use treebanks as resource for collecting evidence of phenomena which cannot be easily recovered from data that is annotated at word level only, this includes collecting quantitative data, getting non-categorical information such as heaviness or finding natural sounding counter examples 1 (e.g. Uszkoreit et al. (1998); Arnold et al. (2000); Bresnan et al. (to appear)) 2. Tools such as TIGERSearch allow us easy access to the encoded information. 3 This poster presents work on the Tübinger Baumbank deutscher Zeitungssprache (Tüba-D/Z). It describes the encoding of coordination phenomena in the treebank and gives a qualitative and quantitative survey. 2 The TüBa-D/Z Treebank It is a corpus of newspaper texts which currently comprises about 22 000 sentences (more than 381 000 tokens) taken from the Wissenschafts-CD of ’die tageszeitung ’ (taz). The annotation combines information on inflectional morphology, part of speech, phrase structure (or rather recursive chunking), grammatical dependencies and topological fields. In addition, it includes marking of named entities and annotation of anaphoric and coreference relations (cf. Hinrichs et al. (2004)).
We present the Heart of Gold middleware by demonstrating three XML-based integration scenarios where multi-dimensional markup produced online by multilingual natural language processing (NLP) components is combined to deliver rich, robust linguistic markup for use in NLP-based applications like information extraction, question answering and semantic web. The scenarios include (1) robust deep-shallow integration, (2) shallow processing cascades, and (3) treebank storage of multi-dimensionally annotated texts.
Data-driven grammatical function tag assignment has been studied for English using the Penn-II Treebank data. In this paper we address the question of whether such methods can be applied successfully to other languages and treebank resources. In addition to tag assignment accuracy and f-scores we also present results of a task-based evaluation. We use three machine-learning methods to assign Cast3LB function tags to sentences parsed with Bikel's parser trained on the Cast3LB treebank. The best performing method, SVM, achieves an f-score of 86.87% on gold-standard trees and 66.67% on parser output - a statistically significant improvement of 6.74% over the baseline. In a task-based evaluation we generate LFG functional-structures from the function-tag-enriched trees. On this task we achive an f-score of 75.67%, a statistically significant 3.4% improvement over the baseline.
Treebanks are linguistically annotated corpora with some previously established scheme of grammatical analysis. In any domain and in the (bio-) medical field in particular, such resources constitute a fundamental piece of knowledge for empirically-based, data-driven language processing, human language technologies and linguistic research and have attracted an increased interest during recent years. The interest for treebanks in biomedicine is guided by the fact that information extraction and (bio-) text mining research is shifting focus from the extraction and annotation of named entities to the extraction and annotation of relations and interactions between entities. This is usually associated by the extraction of verbal – alias predicate-argument – structures (cf. Kulick et al. [1], Tateisi et al. [2]). Semantic relations (e.g. between entities) and role extraction and labelling (e.g. agent, object) constitute a considerable challenge for automatic tools, although recent evaluation competitions such as the PASBio (Wattarujeekrit et al., [3]) and the BioCreAtIvE (Hirschman et al., [4]) revealed that some systems could present significant progress in performance in this area. In this paper, we present our current activities towards the compilation and the multi-layered annotation of a domain-dependent corpus for Swedish in the area of medicine. The focus of the paper is based on the description of the constituent structure and functionally oriented annotation of the corpus. Moreover, the annotation scheme adopted, which incorporates three main layers of linguistic processing, lexical analysis, shallow semantic analysis and syntactic processing, will be exemplified. For the syntactic analysis we use a cascaded finite-state parser, aware of the shallow semantic annotations produced. The result of this analysis, including syntactic parsing and shallow semantic analysis, is transformed into the TIGER-XML interchange format ([5]). Our goal is to produce a large, rich in annotations, medical treebank suitable for both corpus-based grammar learning systems, for semantic relation extraction and for linguistic exploration of theoretical nature. Motivation for this work is given in Section 2. Background work in the area of biomedical syntactic analysis and treebanking is presented in Section 3. Section 4 gives a brief description of the corpus used in this work, while Section 5 deals with the pre-processing steps applied into a sample of the corpus. Section 6 presents evaluation results based on this sample, while Section 7 summarizes the paper and proposes directions for future work.
Approximately 60 kinds of lexical relations have been recognized in languages of the world (Grimes & Grimes, 1993). In this paper, I present evidence for a wide range of lexical relations in the Ilokano language. Grimes explains the meaning of lexical relations in terms of the way two words are related but differ in meaning, giving examples such as write and writer, row and rower. I will exemplify some of the types of lexical relations attested in Ilokano: 1) Verbs with an incorporated nominal, e.g. ag-diram'os 'to wash one's face' where the implied noun is 'face'; aginnaw 'wash dishes', implied noun, 'dishes'. There is no word for 'face' in agdiram'os, nor word for 'dishes' in aginnaw. 2) Derived nouns expressing an agentive relation, e.g. from the verb agsugal 'to gamble' the derived noun is mannugal 'gambler', agsurat 'to write', mannurat 'writer'. 3) Reduplication of a noun describing a condition of that noun, e.g. saka 'foot', saka-saka 'barefoot'; ima 'hand', ima-ima 'emptyhanded'. 4) Derived verbs denoting animal vocalization, e.g. aso 'dog', agtaol (phonation) 'to bark', ul'ul'ol (onomatopoeia); 5) Derived verbs denoting a quantum, e.g. sangalilig a sua 'one section of a pomelo'; 6) Complements, e.g. biag ken patay 'life and death'; 7) Derived verbs denoting function, e.g. karayan 'river', agayos 'to flow', sabong flower', agukrad 'to bloom'. I will then show how these derivations are handled in the Ilokano Lexical Database where they are listed making use of the band format.
Deterministic parsing guided by treebank-induced classifiers has emerged as a simple and efficient alternative to more complex models for data-driven parsing. We present a systematic comparison of memory-based learning (MBL) and support vector machines (SVM) for inducing classifiers for deterministic dependency parsing, using data from Chinese, English and Swedish, together with a variety of different feature models. The comparison shows that SVM gives higher accuracy for richly articulated feature models across all languages, albeit with considerably longer training times. The results also confirm that classifier-based deterministic parsing can achieve parsing accuracy very close to the best results reported for more complex parsing models.
This paper describes the construction of a dependency bank gold standard for Arabic, DCU 250 Arabic Dependency Bank (DCU 250), based on the Arabic Penn Treebank Corpus (ATB) (Bies and Maamouri, 2003; Maamouri and Bies, 2004) within the theoretical framework of Lexical Functional Grammar (LFG). For parsing and automatically extracting grammatical and lexical resources from treebanks, it is necessary to evaluate against established gold standard resources. Gold standards for various languages have been developed, but to our knowledge, such a resource has not yet been constructed for Arabic. The construction of the DCU 250 marks the first step \ntowards the creation of an automatic LFG f-structure annotation algorithm for the ATB, \nand for the extraction of Arabic grammatical and lexical resources.
We present the implementation of a system which extracts not only lexicalized grammars but also feature-based lexicalized grammars from Korean Sejong Treebank. We report on some practical experiments where we extract TAG grammars and tree schemata. Above all, full-scale syntactic tags and well-formed morphological analysis in Sejong Treebank allow us to extract syntactic features. In addition, we modify Treebank for extracting lexicalized grammars and convert lexicalized grammars into tree schemata to resolve limited lexical coverage problem of extracted lexicalized grammars.
We describe several improvements to the method of treebank-based LFG induction for Spanish from the Cast3LB treebank (O’Donovan et al., 2005). We discuss the different categories of problems encountered and present the solutions adopted. Some of the problems involve a simple adoption of existing linguistic analyses, as in our treatment of clitic doubling and null subjects. In other cases there is no standard LFG account for the phenomenon \nwe wish to model and we adopt a compromise, conservative solution. This is exemplified by our treatment of Spanish periphrastic constructions. In yet another case, the less configurational nature of Spanish means that the LFG annotation algorithm has to rely mostly on Cast3LB function tags, and consequently a reliable method of adding those tags to parse trees had to be developed. This method achieves over 6% improvement over the baseline for the \nCast3LB-function-tag assignment task, and over 3% improvement over the baseline for LFG f-structure construction from function-tag-enriched trees.
The Hamburg implementation of the Weighted Constraint Dependency Grammar formalism (WCDG) includes an example grammar with comprehensive coverage for written German. This manual is the annotation guideline that was used to define the goals of the grammar and to create the Hamburg Dependency Treebank also published in the course of this project.
Each year the Conference on Computational Natural Language Learning (CoNLL) features a shared task, in which participants train and test their systems on exactly the same data sets, in order to better compare systems. The tenth CoNLL (CoNLL-X) saw a shared task on Multilingual Dependency Parsing. In this paper, we describe how treebanks for 13 languages were converted into the same dependency format and how parsing performance was measured. We also give an overview of the parsing approaches that participants took and the results that they achieved. Finally, we try to draw general conclusions about multi-lingual parsing: What makes a particular language, treebank or annotation scheme easier or harder to parse and which phenomena are challenging for any dependency parser?
In this paper, current dependencybased treebanks are introduced and analyzed.The methods used for building the resources, the annotation schemes applied, and the tools used (such as POS taggers, parsers and annotation software) are discussed.
This report explores the question of compatibility between annotation projects including translating annotation formalisms to each other or to common forms. Compatibility issues are crucial for systems that use the results of multiple annotation projects. We hope that this report will begin a concerted effort in the field to track the compatibility of annotation schemes for part of speech tagging, time annotation, treebanking, role labeling and other phenomena.
Sentence similarity measures play an increasingly important role in text-related research and applications in areas such as text mining, Web page retrieval, and dialogue systems. Existing methods for computing sentence similarity have been adopted from approaches used for long text documents. These methods process sentences in a very high-dimensional space and are consequently inefficient, require human input, and are not adaptable to some application domains. This paper focuses directly on computing the similarity between very short texts of sentence length. It presents an algorithm that takes account of semantic information and word order information implied in the sentences. The semantic similarity of two sentences is calculated using information from a structured lexical database and from corpus statistics. The use of a lexical database enables our method to model human common sense knowledge and the incorporation of corpus statistics allows our method to be adaptable to different domains. The proposed method can be used in a variety of applications that involve text knowledge representation and discovery. Experiments on two sets of selected sentence pairs demonstrate that the proposed method provides a similarity measure that shows a significant correlation to human intuition