Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
Reviews1 43 Merriam-Webster'sAdvanced Learner's English Dictionary. 2008. Springfield: Merriam-Webster. Pp. 2016. JL b 1IiC world of dictionaries for advanced learners of English has long been dominated by two British publishers: Oxford University Press, which published the first modern learner's dictionary, die learner's Dictionary of Current F.nglkh (now the Oxford Advanced learner's Dictionary, or OAlJ)), in 1948; and Longman, which entered the market with die Longman Dictionary of Contemporary English (IJ)OCF) in 1 978. Both publishers have followed dieir flagship products with dictionaries for intermediate and basic learners, with picture dictionaries, and with electronic products. In 1981, Longman published an American version of its intermediate-level dictionary as the Ixmgman Dictionary ofAmerican English, and for more than a decade this dictionary was about all diat was available to learners who preferred American English—die huge market of students living and working in die United States, as well as learners outside die US (particularly in Japan and many countries in die Americas). American dictionary publishers did not specialize in dictionaries for learners and were relative latecomers in spotting die market potential in ESL products. When diey did decide to enter die fray in die 1990s, diey badly miscalculated. Some, such as Merriam-Webster and Webster's New World, created "basic" dictionaries by adapting an existing data set; others, such as Heinle, Random House, and NTC, opted for single-audior works. These books showed little or no evidence diat dieir compilers were aware ofwhat by dien was more dian forty years' wordi of research and innovation in die production of learner's dictionaries—of die importance of using a controlled defining vocabulary, of showing pronunciations in the International Phonetic Alphabet, ofshowing grammatical patterns, of examining corpora oflearners' essays to aid in die writing ofusage notes, ofusing corpora to determine die relative frequency ofwords, patterns, and expressions so as to aid in making inclusion decisions. In die meantime, bodi Cambridge University Press and Longman produced new corpus-based editions of intermediate-level American dictionaries, and Longman published the lemgman Advanced American Dictionary, essentially an Americanized version of the Ixmgman Dictionary ofContemporary Englkh. With the exception of die Newbury House dictionaries published by Heinle, which have had some success, the home-grown American products have not presented a real challenge to the dominance of the American English learner's dictionary market by British publishers. A decade ago, die publishers at Merriam-Webster saw, quite righüy, that there was room for a strong American contender, and set out to produce one. The resulting dictionary, Merriam-Webster's Advanced learner's F.nglkh Dictionary (MWAlJJ)), largely avoids the clear weaknesses of its predecessors, and represents a solid, well-edited, and authoritative resource. In reviewing the text, I will make reference to its chiefcompetitors, the Dictionaries: journal ofthe Dictionary Society ofNorth America 30 (2009), 143-150 144Reviews Macmillan F'.nglkh Dictionaryfor Advanced Learners ofAmerican F.nglkh (MFD) and the Longman Advanced American Dictionary (IAAl)). Headwords and the Organization of Entries The number ofdiscrete lexical items—"100,000 words and phrases"—that are defined in MWAIFJ) is comparable to that of its competitors. Most learner's dictionaries list each part ofspeech as a separate homograph; MWAUDAso gives separate homographs for etymologically distinct words, e.g., calf the baby cow and the calf of your leg. It includes some British English terms and expressions, much as British dictionaries include some American ones. But sometimes the choice of inclusion is odd, and I suspect this may be because while the MerriamWebster citation resources are voluminous, a citation bank does not have the characteristics of a corpus that make it possible for a lexicographer tojudge centrality and frequency when considering whether to include a given term. Thus MWAIFJ) has an entry for cloud-cuckoo-land while the British edition of MFJ) does not include the American equivalent, la-la land; yet MWAIJD fails to provide an entry for earlier, even though this comparative form is frequent and its use should be illustrated for learners. MWIAFJ) gives the pronunciation of headwords in the International Phonetic Alphabet, which is now the undisputed norm for a learner's dictionary. Unfortunately, stress patterns are not shown...
This paper studies the nature of the BEI-construction in Cantonese, with Mandarin as the standard language of comparison. Although the BEI-construction has been much studied in Mandarin, the same in not true for Cantonese. Although this construction has traditionally been termed a "passive", I will show that it can have a different range of semantic interpretations in Cantonese. I argue that BEI is not confined to passive, but is used under certain circumstances to form a causative construction as well. The differences in behaviour between passive-BEI and causative-BEI can be seen in tests with anaphoric binding. I conclude that while the passive structure is mono-clausal, the causative structure must be bi-clausal. The Cantonese BEI-constructions have an obligatory agent-phrase which cannot be dropped. This differs from Mandarin and the challenge is to find an account for this phenomenon, especially if we are to claim that this construction is a passive. The optionality of the agent phrase is characteristic of passives and yet Cantonese deviates from this norm. I argue that passive in Cantonese is a syntactic process and predict that only transitive verbs may participate in this construction. I utilize the universal v-VP structure on transitive verbs, proposed by Chomsky (1995), to guarantee that the external theta role must be retained. I also examine the much debated status of BEI which is used in the BEI-construction. Although this construction can be used to derive both a passives and a causatives, it does not necessarily mean that two separate BEIs must be posited. I conclude that BEI can be treated as a category-neutral element which can interact in both causative and passive structures. To support this proposal I appeal to the functional versus lexical distinction of categories and projections.
Written text is one of the fundamental manifestations of human language, and the study of its universal regularities can give clues about how our brains process information and how we, as a society, organize and share it. Among these regularities, only Zipf's law has been explored in depth. Other basic properties, such as the existence of bursts of rare words in specific documents, have only been studied independently of each other and mainly by descriptive models. As a consequence, there is a lack of understanding of linguistic processes as complex emergent phenomena. Beyond Zipf's law for word frequencies, here we focus on burstiness, Heaps' law describing the sublinear growth of vocabulary size with the length of a document, and the topicality of document collections, which encode correlations within and across documents absent in random null models. We introduce and validate a generative model that explains the simultaneous emergence of all these patterns from simple rules. As a r)
We present a framework for interfacing a PCFG parser with lexical information from an external resource following a different tagging scheme than the treebank. This is achieved by defining a stochastic mapping layer between the two resources. Lexical probabilities for rare events are estimated in a semi-supervised manner from a lexicon and large unannotated corpora. We show that this solution greatly enhances the performance of an unlexicalized Hebrew PCFG parser, resulting in state-of-the-art Hebrew parsing results both when a segmentation oracle is assumed, and in a real-word parsing scenario of parsing unsegmented tokens.
We present several algorithms for assigning heads in phrase structure trees, based on different linguistic intuitions on the role of heads in natural language syntax. Starting point of our approach is the observation that a head-annotated treebank defines a unique lexicalized tree substitution grammar. This allows us to go back and forth between the two representations, and define objective functions for the unsupervised learning of head assignments in terms of features of the implicit lexicalized tree grammars. We evaluate algorithms based on the match with gold standard head-annotations, and the comparative parsing accuracy of the lexicalized grammars they give rise to. On the first task, we approach the accuracy of hand-designed heuristics for English and inter-annotation-standard agreement for German. On the second task, the implied lexicalized grammars score 4% points higher on parsing accuracy than lexicalized grammars derived by commonly used heuristics.
This letter examines how high rates of churn - the continuous process of node arrival and departure - affect rating mechanisms for peer-to-peer (P2P) networks. In particular, short peer lifetimes mean reputations are often generated from a small number of transactions, and thus are few reliable. To understand this relationship, this letter introduces an analytical model which determines the optimal transaction rate and the expected time to produce a reliable reputation, under both exponential and Pareto lifetime distributions.
We present in this paper an approach to assessing student paraphrases in the intelligent tutoring system iSTART. The approach is based on measuring the semantic similarity between a student paraphrase and a reference text, called the textbase. The semantic similarity is estimated using knowledge-based word relatedness measures. The relatedness measures rely on knowledge encoded in Word-Net, a lexical database of English. We also experiment with weighting words based on their importance. The word importance information was derived from an analysis of word distributions in 2,225,726 documents from Wikipedia. Performance is reported for 12 different models which resulted from combining 3 different relatedness measures, 2 word sense disambiguation methods, and 2 word-weighting schemes. Furthermore, comparisons are made to other approaches such as Latent Semantic Analysis and the Entailer.
The paper presents a strategy for deriving English to Urdu translation using English to Hindi MT system. The English-Hindi lexical database is used to collect all possible Hindi words and phrases. These are further aug-mented by including their morphological vari-ations and attaching all possible post-positions. This list is used to provide mapping from Hindi to Urdu. There may be change in gender and a word or a word group may be of multiple parts of speech. These are resolved using information available from English-Hindi MT. As Urdu is structurally very close to Hindi using similar post-positions, the out-put obtained is as acceptable as the Hindi translation. 1
Syntactic parsing is a central problem and a challenge in the field of natural language processing. It attracts many studies and consequently there exists the effective parsers for several popular languages such as English and Chinese. For Vietnamese parsing, there have been a few studies focusing on this problem, these studies lack of applying modern techniques, and no popular parser has been released. This paper presents the first study on developing a Vietnamese wide coverage parser based on lexicalized probabilistic context free grammar (LPCFG) and using a standard parsed corpus (similar to Penn Treebank). In this paper the Bikel's parser is modified to analyze Vietnamese. We also provide a comparison based on investigating different parsing models and different linguistic features. The best configuration achieves around 78% of F-score.
This paper proposes an extension of Dependency Tree Semantics (DTS), an underspecified logic originally proposed in [20], that uniformily implements constraints on Nested Quantification, Island Constraints and logical Redundancy. Unfortunately, this extension makes the complexity exponential in the number of NPs, in the worst cases. Nevertheless, we conducted an experiment on the Turin University Treebank [6], a Treebank of italian sentences annotated in a syntactic dependency format, whose results seem to indicate that these cases are very rare in real sentences.
This paper presents an on-going effort which aims to annotate the Wall Street Journal sections of the Penn Treebank with the help of a hand-written large-scale and wide-coverage grammar of English. In doing so, we are not only focusing on the various stages of the semi-automated annotation process we have adopted, but we are also showing that rich linguistic annotations, which can apart from syntax also incorporate semantics, ensure that the treebank is guaranteed to be a truly sharable, re-usable and multi-functional linguistic resource.
This paper presents a set of experiments performed on parsing the Basque Dependency Treebank. We have applied feature propagation to dependency parsing, experimenting the propagation of several morphosyntactic feature values. In the experiments we have used the output of a parser to enrich the input of a second parser. Both parsers have been generated by Maltparser, a freely data-driven dependency parser generator. The transformations, combined with the pseudoprojective graph transformation, obtain a LAS of 77.12% improving the best reported results for Basque.
In many languages general syntactic cues are insufficient to disambiguate crucial relations in the task of Parsing. In such cases semantics is necessary. In this paper we show the effect of minimal semantics on parsing. We did experiments on Hindi, a morphologically rich free word order language to show this effect. We conducted experiments with the two data-driven parsers MSTPaser and MaltParser. We did all the experiments on a part of Hyderabad Dependency Treebank. With the introduction of minimal semantics we achieved an increase of 1.65% and 2.01% in labeled attachment score and labeled accuracy respectively over state-of-the-art data driven dependency parser.
This paper calls for a broadening of the discussion of English language teaching (ELT) practices in Japan. We review issues associated with the global spread of English and link this discussion to the present “standard” English model of ELT in Japan. We propose three major benefits that would follow from an inclusion of non-“standard” (i.e., non-American/British) Englishes in Japanese EFL classrooms. First, familiarity with different varieties could increase learners’ confidence when interacting with other nonnative speakers (NNSs). Second, we review literature that shows that NNS-NNS interactions actually help learners improve their language skills. Finally, recognition of non-“standard” varieties of English would help Japanese learners challenge monolithic western-centric worldviews that marginalize regional, cultural, and linguistic norms and values. We connect this theory to practice by suggesting some possible changes to ELT in Japan.
Grammar rules for Clausal Coordinate Ellipsis (CCE) are based nearly exclusively on linguistic judgments (intuitions).For German, the extent to which grammar rules based on this type of empirical evidence generate all and only CCE structures that populate text corpora, has only been explored with the TIGER treebank of written newspaper text.How well these rules fit spoken German is unknown.In this paper, we study the applicability of judgment-based CCE rules to spontaneously spoken German by means of the TBa-D/S treebank, which is based on dialogues for appointment scheduling and travel planning from the VERBMOBIL project.The judgment-based CCE rules are shown to hold nearly equally well for spoken as for written text: The proportion of deviations from the rules are virtually identical-less than 3% of the utterances/sentences that include a clausal coordination (compared to about 1% in the TIGER treebank).Moreover, the relative frequencies in VERBMOBIL of four main CCE types distinguished in the literature reveal a pattern that resembles the pattern observed in CGN2.0, the Corpus of Spoken Dutch.
This paper presents a novel application of incorporating Alternating Structure Optimization (ASO) to conduct the task of text chunking of Semantic Role Labeling (SRL) in Chinese texts. ASO is a competent linear algorithm based on the theory of multi-task learning. In this paper, by constructing several SRL tasks to constitute a multi-task, we are able to encode the inference obtained by ASO algorithm as additional feature to further boost the performance of the target task employing Conditional Random Fields (CRFs). To our knowledge, our method is the first that incorporates multi-task learning into a statistical model in SRL for Chinese texts. We evaluate our approach on Penn Treebank data sets and obtain encouraging result.
This paper presents a tool for extracting multi-word expressions from corpora in Modern Greek, which is used together with a parallel concordancer to augment the lexicon of a rule-based machine-translation system. The tool is part of a larger extraction system that relies, in turn, on a multilingual parser developed over the past decade in our laboratory. The paper reviews the various NLP modules and resources which enable the retrieval of Greek multi-word expressions and their translations: the Greek parser, its lexical database, the extraction and concordancing system.
Enhanced sensitivity to information of negative (compared to positive) valence has an adaptive value, for example, by expediting the correct choice of avoidance behavior. However, previous evidence for such enhanced sensitivity has been inconclusive. Here we report a clear advantage for negative over positive words in categorizing them as emotional. In 3 experiments, participants classified briefly presented (33 ms or 22 ms) masked words as emotional or neutral. Categorization accuracy and valence-detection sensitivity were both higher for negative than for positive words. The results were not due to differences between emotion categories in either lexical frequency, extremeness of valence ratings, or arousal. These results conclusively establish enhanced sensitivity for negative over positive words, supporting the hypothesis that negative stimuli enjoy preferential access to perceptual processing.
Broad-coverage annotated treebanks necessary to train parsers do not exist for many resource-poor languages. The wide availability of parallel text and accurate parsers in English has opened up the possibility of grammar induction through partial transfer across bitext. We consider generative and discriminative models for dependency grammar induction that use word-level alignments and a source language parser (English) to constrain the space of possible target trees. Unlike previous approaches, our framework does not require full projected parses, allowing partial, approximate transfer through linear expectation constraints on the space of distributions over trees. We consider several types of constraints that range from generic dependency conservation to language-specific annotation rules for auxiliary verb analysis. We evaluate our approach on Bulgarian and Spanish CoNLL shared task data and show that we consistently outperform unsupervised methods and can outperform supervised learning for limited training data.
The Arabic language has a very rich/complex morphology. Each Arabic word is composed of zero or more prefixes, one stem and zero or more suffixes. Consequently, the Arabic data is sparse compared to other languages such as English, and it is necessary to conduct word segmentation before any natural language processing task. Therefore, the word-segmentation step is worth a deeper study since it is a preprocessing step which shall have a significant impact on all the steps coming afterward. In this article, we present an Arabic mention detection system that has very competitive results in the recent Automatic Content Extraction (ACE) evaluation campaign. We investigate the impact of different segmentation schemes on Arabic mention detection systems and we show how these systems may benefit from more than one segmentation scheme. We report the performance of several mention detection models using different kinds of possible and known segmentation schemes for Arabic text: punctuation separation, Arabic Treebank, and morphological and character-level segmentations. We show that the combination of competitive segmentation styles leads to a better performance. Results indicate a statistically significant improvement when Arabic Treebank and morphological segmentations are combined.
In this paper, we propose a modular cascaded approach to data driven dependency parsing. Each module or layer leading to the complete parse produces a linguistically valid partial parse. We do this by introducing an artificial root node in the dependency structure of a sentence and by catering to distinct dependency label sets that reflect the function of the set internal labels vis-a¿-vis a distinct and identifiable linguistic unit, at different layers. The linguistic unit in our approach is a clause. Output (partial parse) from each layer can be accessed independently. We applied this approach to Hindi, a morphologically rich free word order language using MST parser. We did all our experiments on a part of Hyderabad Dependency Treebank. The final results show an increase of 1.35% in unlabeled attachment and 1.36% in labeled attachment accuracies over state-of-the-art data driven Hindi parser.
Articles in the Penn TreeBank were identified as being reviews, summaries, letters to the editor, news reportage, corrections, wit and short verse, or quarterly profit reports. All but the latter three were then characterised in terms of features manually annotated in the Penn Discourse TreeBank --- discourse connectives and their senses. Summaries turned out to display very different discourse features than the other three genres. Letters also appeared to have some different features. The two main findings involve (1) differences between genres in the senses associated with intra-sentential discourse connectives, inter-sentential discourse connectives and inter-sentential discourse relations that are not lexically marked; and (2) differences within all four genres between the senses of discourse relations not lexically marked and those that are marked. The first finding means that genre should be made a factor in automated sense labelling of non-lexically marked discourse relations. The second means that lexically marked relations provide a poor model for automated sense labelling of relations that are not lexically marked.
We adapt a semantic role parser to the domain of goal-directed speech by creating an artificial treebank from an existing text tree-bank. We use a three-component model that includes distributional models from both target and source domains. We show that we improve the parser's performance on utterances collected from human-machine dialogues by training on the artificially created data without loss of performance on the text treebank.
DeSR is a statistical transition-based dependency parser which learns from annotated corpora which actions to perform for building parse trees while scanning a sentence. We describe recent improvements to the parser, in particular stacked parsing, exploiting a beam search strategy and using a Multilayer Perceptron classifier. For the Evalita 2009 Dependency Parsing task DesR was configured to use a combination of stacked parsers. The stacked combination achieved the best accuracy scores in both the main and pilot subtasks. The contribution to the result of various choices is analyzed, in particular for taking advantage of the peculiar features of the TUT Treebank.\n\nKeywords: parser, dependency parsing, perceptron, classifier, natural language.
Purpose The purpose of this paper is to investigate the characteristics of social network comments to give a broad overview to serve as a baseline for future research. Design/methodology/approach English comments from a representative sample of public MySpace profiles were examined with a collection of exploratory analyses, using automatic data processing, quantitative techniques and content analyses. Findings Comments were normally for general friendship maintenance and were typically short, with 95 per cent having 57 or fewer words. They contained a combination of standard spelling, apparently accidental mistakes, slang, sentence fragments, “typographic slang” and interjections. Several new creative spelling variants derived from previous forms of computer‐mediated communication have become extremely common, including u, ur,:), haha and lol. The vast majority of comments (97 per cent) contained at least one non‐standard language feature, suggesting that members almost universally recognise the informal nature of this kind of messaging. Research limitations/implications The investigation only covered MySpace and only analysed English comments. Practical implications MySpace comments should not be written in, or judged by, standard linguistic norms and may cause special problems for information retrieval. Originality/value This is the first large‐scale study of language in social network comments.
Cognitive factors such as catastrophic thoughts regarding pain, and conversely, one's acceptance of that pain, may affect emotional functioning among persons with chronic pain conditions. The aims of the present study were to examine the effects of both catastrophizing and acceptance on affective ratings of experimentally induced ischemic pain and also self-reports of depressive symptoms. Sixty-seven individuals with chronic back pain completed self-report measures of catastrophizing, acceptance, and depressive symptoms. In addition, participants underwent an ischemic pain induction procedure and were asked to rate the induced pain. Catastrophizing showed significant effects on sensory and intensity but not affective ratings of the induced pain. Acceptance did not show any significant associations, when catastrophizing was also in the model, with any form of ratings of the induced pain. Catastrophizing, but not acceptance, was also significantly associated with self-reported depressive symptoms when these two variables were both included in a regression model. Overall, results indicate negative thought patterns such as catastrophizing appear to be more closely related to outcomes of perceived pain severity and affect in persons with chronic pain exposed to an experimental laboratory pain stimulus than does more positive patterns as reflected in measures of acceptance.
Abstract. This paper presents a method to incorporate statistical information into a rule-based parser to resolve syntactic ambiguities. We extract the statistical information from the Penn Treebank, and apply the information to the rule-based parser. For the extraction of the statistical information the tag conversion is needed because of the disagreement of the tags and the bracketing style. We will show the effect of the tag conversion with experiments. The final result shows about 7 % error rate reduction in the dependency evaluation. We will also show how much each type of statistical information affects the parsing performance.
In the decade of the sixties, the main concern of the experts in Basque language was how to unite the speech community. However, after forty years of Standard Basque, the objective is to develop a language that is well adapted to each register. The lexis is the key in the functional development that the adaptation to each register requires, and that is exactly what we shall study in this work. We shall analyse, from the point of view of functional diversity, the methodology of the Standard Basque Dictionary, the purpose of which is to organize the standard Basque lexicon. We reach the conclusion that Standard Basque has taken certain steps to get closer to the functional variability desideratum. However, we have also detected methodological practices that may be an obstacle in this functional development. As a result, we draw up the outline of the methodological changes that would be required to accommodate linguistic norms to the variation at the register level.
This paper describes an empirical study of high-performance dependency parsers based on a semi-supervised learning approach. We describe an extension of semi-supervised structured conditional models (SS-SCMs) to the dependency parsing problem, whose framework is originally proposed in (Suzuki and Isozaki, 2008). Moreover, we introduce two extensions related to dependency parsing: The first extension is to combine SS-SCMs with another semi-supervised approach, described in (Koo et al., 2008). The second extension is to apply the approach to second-order parsing models, such as those described in (Carreras, 2007), using a two-stage semi-supervised learning approach. We demonstrate the effectiveness of our proposed methods on dependency parsing experiments using two widely used test collections: the Penn Treebank for English, and the Prague Dependency Tree-bank for Czech. Our best results on test data in the above datasets achieve 93.79% parent-prediction accuracy for English, and 88.05% for Czech.
Semantic processing represents the new challenge for all applications that require text understanding, as for instance Q/A. In this paper we will highlight the need to couple statistical approaches with deep linguistic processing and will focus on ldquoimplicitrdquo or lexically unexpressed linguistic elements that are nonetheless necessary for a complete semantic interpretation of a text. We will address the following types of ldquoimplicitrdquo entities and events: - grammatical ones, as suggested by a linguistic theories like LFG or similar generative theories; - semantic ones suggested in the FrameNet project, i.e. CNI, DNI, INI; - pragmatic ones: here we will present a theory and an implementation for the recovery of implicit entities and events of (non-) standard implicatures. In particular we will show how the use of commonsense knowledge may fruitfully contribute in finding relevant implied meanings. We will also briefly explore the subject of point of view which is computed by semantic informational structure and contributes the intended entity from whose point of view is expressed a given subjective statement. We also present an evaluation based on section 24 of Penn Treebank as encoded by LFG people in the PARC-700 treebank where lexically unexpressed are adequately classified and diversified.
Parallel treebanks provide a systematic way of expressing the structural relationships between source and target texts. In this paper, we present the general design principles behind the Copenhagen Dependency Treebanks, a set of parallel treebanks for Danish, English, German, Italian and Spanish with a unified annotation of morphology, syntax, discourse, and tranlational equivalence. Finally, we suggest some hypotheses about morphology and discourse, and describe how we plan to explore them empirically on the basis of the treebanks.
Jointly parsing two languages has been shown to improve accuracies on either or both sides. However, its search space is much bigger than the monolingual case, forcing existing approaches to employ complicated modeling and crude approximations. Here we propose a much simpler alternative, bilingually-constrained monolingual parsing, where a source-language parser learns to exploit reorderings as additional observation, but not bothering to build the target-side tree as well. We show specifically how to enhance a shift-reduce dependency parser with alignment features to resolve shift-reduce conflicts. Experiments on the bilingual portion of Chinese Treebank show that, with just 3 bilingual features, we can improve parsing accuracies by 0.6% (absolute) for both English and Chinese over a state-of-the-art baseline, with negligible (~6%) efficiency overhead, thus much faster than biparsing.
One of the biggest challenges in compiling a dictionary of a minority language is managing the large quantity of lexical data. Decisions about the format and content of the dictionary or the orthography typically evolve over the years that such projects usually take. This results in inconsistencies between older and newer entries. Revising the data for publication as a dictionary introduces further inconsistencies as does having multiple contributors and/or editors. Proofreading a lexical database takes a great deal of time and the richer its structure the more this is the case. The tools described in this presentation significantly reduce this effort. Tools developed for checking the consistency of the lexical database in the Iu Mien—Chinese—English dictionary project have proven extremely helpful. Two basic approaches are used: 1) use of a program written to check for likely errors that scans the lexical database and produces an error report that is used by a lexicographer to make appropriate corrections. 2) outputting the lexical data in alternate forms that make it easier for the lexicographer to spot problem areas. These alternative forms include the reverse indexes and views structured according to semantic domains. The Iu Mien—Chinese—English dictionary project, like many minority language dictionary projects, uses SIL's Toolbox software. It is very flexible software but its capabilities to enforce consistency are quite limited. Some parts of the approach described here are specific to MDF (Multi-Dictionary Formatter) lexical databases in Toolbox but will be equally useful for other MDF databases. Other parts are specific to each of the three languages involved but will be useful for non-Toolbox lexical databases. Every dictionary is unique and this applies not only to content of the entries but also the decisions about how entries should be arranged to suit the languages involved. Other decisions about the structure are likely to be made differently even in other dictionaries of the same languages. It is the way that each dictionary combines themes that are found in many dictionaries that makes them unique, e.g. to be root based or not, to have include subentries. Therefore our approach is to use a toolkit based approach to curating lexical databases. This allows checking techniques to be mixed and matched to suit the unique aspects of a lexical project. The checking software is written in Python and relies on the toolbox module in NLTK (The Natural Language Toolkit http://nltk.sourceforge.net).
This research concerns linguistic variation and Portuguese teaching at school. It is assumed that linguistic pattern is conceived by teachers as an homogeneous norm, so that it is incompatible with linguistic norms students face with in usual text reading and writing activities. Taking into consideration official evaluation of didactic books, scholar reading activities, and sociolinguistic results that prove school interference at students writing performance, this article proposes that Portuguese classes should present variation as complex continua in which is displayed a plurality of norms.
This paper proposes a novel method to refine the grammars in parsing by utilizing semantic knowledge from HowNet. Based on the hierarchical state-split approach, which can refine grammars automatically in a data-driven manner, this study introduces semantic knowledge into the splitting process at two steps. Firstly, each part-of-speech node will be annotated with a semantic tag of its terminal word. These new tags generated in this step are semantic-related, which can provide a good start for splitting. Secondly, a knowledge-based criterion is used to supervise the hierarchical splitting of these semantic-related tags, which can alleviate overfitting. The experiments are carried out on both Chinese and English Penn Treebank show that the refined grammars with semantic knowledge can improve parsing performance significantly. Especially with respect to Chinese, our parser achieves an F 1 score of 87.5%, which is the best published result we are aware of.
Generative lexicalized parsing models, which are the mainstay for probabilistic parsing of English, do not perform as well when applied to languages with different language-specific properties such as free(r) word order or rich morphology. For German and other non-English languages, linguistically motivated complex treebank transformations have been shown to improve performance within the framework of PCFG parsing, while generative lexicalized models do not seem to be as easily adaptable to these languages.
The aim of the paper is to show that a subset of Text Encoding Initiative Guidelines is a reasonable choice as a standard for stand-off XML encoding of syntactically annotated corpora. The proposed TEI schema — actually employed in the National Corpus of Polish — is compared to other such candidate standards, including TIGER-XML, SynAF and PAULA. 1
The aim of Evalita Parsing Task is at defining and extending Italian state of the art parsing by encouraging the application of existing models and approaches. As in the Evalita'07, the Task is organized around two tracks, i.e. Dependency Parsing and Constituency Parsing. As a main novelty with respect to the previous edition, the Dependency Parsing track has been articulated into two subtasks, differing at the level of the used treebanks, thus creating the prerequisites for assessing the impact of different annotation schemes on the parsers performance. In this paper, we describe the Dependency Parsing track by presenting the data sets for development and testing, reporting the test results and providing a first comparative analysis of these results, also with respect to state of the art parsing technologies.
Az eladasban a Szeged Treebank fuggsegi fa formatumra torten atalakitasanak folyamatat mutatjuk be. Az eredetileg frazisstrukturalt treebankbl automatikus konverzio eredmenyekeppen letrejott fuggsegi fakat kezi uton ellenriztuk es javitottuk, letrehozva ezzel az els magyar nyelv kezzel annotalt dependenciakorpuszt. Jelenleg az uzleti hireket, ujsaghireket es jogi szovegeket tartalmazo alkorpuszok annotacioja fejezdott be, de terveink kozott szerepel a teljes korpusz atalakitasa fuggsegi fa formatumra. Az elkeszult adatbazis hasznosithato tobbek kozott az informaciokinyeresben es a gepi forditasban is.
Abstract This article investigates probability distributions of the dependency relation extracted from a Chinese dependency treebank. The author shows the frequency distributions of dependency type, of word class both as a dependent and a governor, of verb as a governor, and of noun as a dependent. The fitting results reveal that most of the investigated distributions are excellently fitted with a modified right-truncated Zipf-Alekseev distribution. In the analysis of exponential regressions, most of the determination coefficients R 2 are very good, which is an alternative evidence that the investigated distributions are fitted well.
Abstract This chapter discusses the theme of this volume which is about violence in the language of Victorian novels. It analyzes the works of several notable Victorian writers including Charles Dickens, Anne Brontë, George Eliot and Thomas Hardy using narratography. It explains that narratography is the apprehension of mediated narrative increments as traced out in prose or image by the analytic act of reading. This chapter argues that novel violence violates not the literary community but the linguistic norm through their calculated deviance.
This study evaluated evidence for 2 forms of emotional abnormality in posttraumatic stress disorder (PTSD): numbing and heightened negative emotionality. Forty-nine male veterans with PTSD and 75 without the disorder rated their emotional responses to photographs that depicted scenes of Vietnam combat or were drawn from the International Affective Picture System (Lang et al., 2005). Images varied in their trauma-relatedness and affective qualities. A series of repeated measures ANOVAs revealed that Vietnam combat veterans with PTSD responded to unpleasant images with greater negative emotionality (i.e., enhanced arousal and lower valence ratings) than those without the disorder and this effect was modified by the trauma-relatedness of the image with stronger effects for trauma-related images. In contrast, the 2 groups showed equivalent patterns of responses to pleasant images. Findings raise questions about the sensitivity of the International Affective Picture System rating protocol for the assessment of PTSD-related emotional numbing.
This dissertation examines the relationship between gender and language in Japanese through the often ignored lens of sexuality. Although linguists are increasingly examining these issues for American gay, lesbian, and bisexual speakers, little similar research has been done in Japan. Lesbians, in particular, are relatively invisible in Japanese society. Examining these women, who do not fit neatly into the hegemonic gender ideology, illuminates how speakers can project a specific identity by displaying or rejecting prescriptive gender-specific linguistic norms of Japanese.I analyzed data recorded from interviews with both Japanese lesbian/bisexual and heterosexual women, looking for differences in frequency and range of use of pronouns and sentence-final particles and for phonetic differences in terms of average pitch height and width. I also considered the results of a perception experiment undertaken to investigate the effect of pitch height and width on Japanese speakers' perceptions of sexuality.Although Japanese speakers were generally unable to identify a cohesive lesbian stereotype, especially in terms of language use, the perception experiment indicated that both average pitch height and width significantly affect judgments on whether a voice sounds lesbian or heterosexual. Tokens judged to be lesbian were also judged to be more masculine and less emotional than those judged to be heterosexual. Analysis of the interview data showed that lesbian participants produced an average pitch height that was significantly lower than that of heterosexual participants. In terms of gendered morphemes, lesbians were significantly more likely to use masculine morphemes than heterosexual women, both for sentence-final particles and first-person pronouns, and were significantly less likely to use the feminine first-person pronoun atashi. Finally, correlations showed that speakers who instantiate gender through the use of gendered-morphemes also do so through manipulations of pitch.Although Japanese lesbians are still fairly closeted and interviewees maintained that there are no cultural stereotypes for this group, significant differences in pitch and gendered-morpheme usage were still apparent. These lesbian/bisexual women did not appear to be mimicking men's language, but instead seemed to be rejecting hegemonic femininity and many of the cultural and linguistic stereotypes that accompany it.