Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
1 Jonas Kuhn was at the University of Stuttgart when the main work reported in this paper was performed. Creation of high-quality treebanks requires expert knowledge and is extremely time consuming. Hence applying an already existing grammar in treebanking is an interesting alternative. This approach has been pursued in the syntactic annotation of German newspaper text in the TIGER project. We utilized the large-scale German LFG grammar of the PARGRAM project for semi-automatic creation of TIGER treebank annotations. The symbolic LFG grammar is used for full parsing, followed by semi-automatic disambiguation and automatic transfer into the treebank format. The treebank annotation format is a ‘hybrid ’ representation structure which combines constituent analysis and functional dependencies. Both types of information are provided by the LFG analyses. Although the grammar and the treebank representations coincide in core aspects, e.g. the encoding of grammatical functions, there are mismatches in analysis details that are comparable to translation mismatches in natural language translation. This motivates the use of transfer technology from machine translation. The German LFG grammar analyzes on average 50 % of the sentences, roughly 70 % thereof are assigned a correct parse; after OT-filtering, a sentence gets 16.5 analyses on average (median: 2). We argue that despite the limits in corpus coverage the applications of the grammar in treebanking is useful especially for reasons of consistency. Finally, we sketch future extensions and applications of this approach, which include partial analyses, coverage extension, annotation of morphology, and consistency checks.
The redundancy is a method which strengthens the transmitted message in language. And the redundancy has two ways: expanding syllable and explanatory repetition. Well, the repeated faulty wording is a wordiness of the sentence structure, which should be deleted. So, retaining the redundancy is a seeking for superfluity, and it attributes the success to the sufficient principle of transmitting message. Leaving the redundancy out is a seeking for simplicity, and it attributes the success to the economical principle of expression. So it is a problem of how to hold the right simplicity. Nowadays there is a tendency that the demand is too strict for the linguistic norm. Now it may begin with the distinction between the redundancy and the repeated faulty wording in two ways: one is linguistic sensation and another is how to express. But it sometimes is difficult to differentiate redundancy from repetition.
International audience
Medical records have been evolving from the traditional paper-based records to digital ones, from the method of dictating reports and transcription to voice recognition systems. The transition to digital operations will not be complete until we have the ability to combine voice recognition with automated indexing of texts. This paper introduces the methods we used to evaluate existing voice recognition software programs and presents NOMINDEX, a system that turns a medical text into MeSH codes, using the French ADM lexical database. Those systems were applied to 28 patient discharge summaries in French, produced after a coronarography, and extracted from the MENELAS corpus of texts. Using the best configuration for voice recognition, the rate of accurate recognition exceeds 98 percent. Among the indexing concepts assigned by NOMINDEX, 25 percent were not pertinent and 12 percent of the relevant concepts were missing. Most errors were related to confusion between common language and medical language, and to the coverage of the ADM lexical database. Best results would be expected with a more comprehensive lexical resource In addition, only 3 percent of the errors generated by inadequate voice recognition that remained in the configuration that performed better, impacted on automatic indexing by NOMINDEX.
This paper describes Grammar Learning by Partition Search, a general method for automatically constructing grammars for a range of parsing tasks. Given a base grammar, a training corpus, and a parsing task, Partition Search constructs an optimised probabilistic context-free grammar by searching a space of nonterminal set partitions, looking for a partition that maximises parsing performance and minimises grammar size. The method can be used to optimise grammars in terms of size and performance, or to adapt existing grammars to new parsing tasks and new domains. This paper reports an example application to optimising a base grammar extracted from the Wall Street Journal Corpus. Partition Search improves parsing performance by up to 5.29%, and reduces grammar size by up to 16.89%. Parsing results are better than in existing treebank grammar research, and compared to other grammar compression methods, Partition Search has the advantage of achieving compression without loss of grammar coverage.
In this paper, we introduce a new European project named OrienTel. The aim of OrienTel is to enable the project's participants to design and develop multilingual interactive communication services for the Mediterranean and the Middle East, ranging from Morocco in the West to the Gulf states in the East, including Turkey and Cyprus. These multilingual applications will be largely speech-based and will typically be implemented on mobile and multi-modal platforms such as cellular GSM or UMTS phones, personal digital assistants (PDAs) or combinations of the two. Applications of the kind targeted are unified messaging, information retrieval, customer care, banking, WAP and service portals. To achieve this aim, the consortium will produce various surveys of the OrienTel region, compile a set of 22 linguistic databases, conduct research into ASR-related problems the OrienTel languages hold and develop demonstrator applications bearing evidence of OrienTel's multilingual orientation.
This study investigated the human eyeblink startle reflex as a measure of alcohol cue reactivity. Alcohol-dependent participants early (n = 36) and late (n = 34) in abstinence received presentations of alcohol and water cues. Consistent with previous research, greater salivation and higher ratings of urge to drink occurred in response to the alcohol cues. Differential salivary and urge responding to alcohol versus water cues did not vary as a function of abstinence duration. Of special interest was the finding that startle response magnitudes were relatively elevated to alcohol cues, but only in individuals early in abstinence. Affective ratings of alcohol cues suggested that alcohol cues were perceived as aversive. Methodological and theoretical implications of the findings are discussed.
This paper deals with up-translation - a process of lexical data transformation from any source format to the XML document. Relevant aspects of the XML format and many related technologies are surveyed first. Then, information content enhancement of existing lexical resources is discussed. The last part brings information about up-translation ofthe Dictionary of Literary Czech Language and the way of efficient storage and retrieval of data.
Language deviation is a form that deviates from linguistic norms. Language deviation is very common in language communication it is a creative use of language and is important in rhetoric. This paper shows the formation of language deviation and its humorous effect.
Ever since the widespread availability of the Penn Treebank [9], there have been numerous, statistical parsers developed for English, e.g. [8, 5, 3]. To varying degrees, these parsers and others---while very successful at the tasks for which they were designed---had the following limitations:
We present a new approach to topological parsing of German which is corpus-based and built on a simple model of probabilistic CFG parsing. The topological field model of German provides a linguistically motivated, flat macro structure for complex sentences. Besides the practical aspect of developing a robust and accurate topological parser for hybrid shallow and deep NLP, we investigate to what extent topological structures can be handled by context-free probabilistic models. We discuss experiments with systematic variants of a topological treebank grammar, which yield competitive results.
The Specific Affect Coding System (SPAFF; J. M. Gottman & L. J. Krokoff, 1989) has led to conclusions about which types of dyadic affect predict positive and negative outcomes in marriage, yet the lack of information about collinearity among the codes limits interpretation of SPAFF results. Psychometric properties of SPAFF were examined by assessing the interactions of 172 newlywed couples with SPAFF and with an affect rating system developed for this study. For husbands and wives, factor analysis indicated 4 distinct factors of affect, representing anger/contempt, sadness, anxiety, and humor/affection. Anger/contempt and humor/affection were associated with marital satisfaction, relationship beliefs, and appraisals of the interactions. Correlations were in the expected directions. The strengths, limitations, and implications of the data are discussed.
In a cooperative project between Uppsala University, the bus and truck manufacturing company Scania CV AB, and the translation company Explicon AB, issues of scaling up the transfer-based machine translation prototype MULTRA for industrial use is beeing investigated. The project is limited to one domain, automotive service literature, and one translation direction, Swedish to English, but issues concerning the change of domain, translation direction and language pair are also considered. Three focal points of the project work have been the design and implementation of the new MATS system, including the redesign, porting and integration of MULTRA, the redesign and implementation of the dictionaries of the language modules as a lexical database, and the scaling up of the dictionaries and the grammars. The system is currently trained on a corpus of aligned bitexts from the automotive service domain. The coverage of the lexical data is almost complete, and validated by professional translators, but the grammars are still limited. Despite the incomplete state of the grammars, the system already translates more than a third of the segments in the corpus. Preliminary evaluations of system performance and coverage have been made, and further development of evaluation methods and metrics are in progress. 1.
We present an algorithm which translates the Penn Treebank into a corpus of Combinatory Categorial Grammar (CCG) derivations. To do this we have needed to make several systematic changes to the Treebank which have to effect of cleaning up a number of errors and inconsistencies. This process has yielded a cleaner treebank that can potentially be used in any framework. We also show how unary type-changing rules for certain types of modifiers can be introduced in a CCG grammar to ensure a compact lexicon without augmenting the generative power of the system. We demonstrate how the combination of preprocessing and type-changing rules minimizes the lexical coverage problem. 1.
We report in this paper on an experiment on automatic extraction of a Tree Adjoining Grammar from the WSJ corpus of the Penn Treebank. We use an automatic tool developed by (Xia, 2001) properly adapted to our particular need. Rather than addressing general aspects of the automatic extraction we focus on the problems we have found to extract a linguistically (and computationally) sound grammar and approaches to handle them.
Medical records have been evolving from the traditional paper-based records to digital ones, from the method of dictating reports and transcription to voice recognition systems. The transition to digital operations will not be complete until we have the ability to combine voice recognition with automated indexing of texts. This paper introduces the methods we used to evaluate existing voice recognition software programs and presents NOMINDEX, a system that turns a medical text into MeSH codes, using the French ADM lexical database. Those systems were applied to 28 patient discharge summaries in French, produced after a coronarography, and extracted from the MENELAS corpus of texts. Using the best configuration for voice recognition, the rate of accurate recognition exceeds 98 percent. Among the indexing concepts assigned by NOMINDEX, 25 percent were not pertinent and 12 percent of the relevant concepts were missing. Most errors were related to confusion between common language and medical language, and to the coverage of the ADM lexical database. Best results would be expected with a more comprehensive lexical resource In addition, only 3 percent of the errors generated by inadequate voice recognition that remained in the configuration that performed better, impacted on automatic indexing by NOMINDEX.
The performance of cursive script recognition systems may be improved by applying higher level knowledge in the form of syntax or semantics. A fundamental part of such an approach is the creation of a lexical database containing the relevant information. However, to create a semantic lexicon by hand for a large vocabulary is a considerable task, which is a major reason why so many semantic theories fail to scale up from the small, artificial domains in which they were developed. An alternative approach is to use existing sources of semantic information, such as machine-readable dictionaries (which contain definitions and domain information) and text corpora (from which collocations and domain information may be derived). The development of techniques for acquiring semantic knowledge from such resources and applying it to large vocabulary cursive script recognition is described.< <ETX xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">></ETX>
One of the major problems in the implementation ofNatural Language Processing (NLP) or MachineTranslation(MT) is a complete lexicon: the place wherethe systems information about words is stored. There aredifficulties in deciding what information should bestored in a lexicon and even greater difficulties inacquiring this information in proper form. OriNetsystem designed to incorporate multiple lexical databaseand tools under one consistent functional interface inorder to facilitate systems requiring syntactic, semanticand lexical information of Oriya language. We dividethe whole work into two independent task. One task isto write the source file that contains the basic lexicaldata and the content of those files are the lexicalsubstance of OriNet. Lexicographer did the major workof this task. In the second task was to create a set ofprograms those would accept the source files andprocessing it ultimately to display for the user. Thispaper describes an ongoing work on designing an ObjectOriented model for OriNet system. The technology ofObject Oriented programming in particular the richlibrary of classes and programming principles in whichJava offers. It also provides a convenient tool toconceptualise the process of OriNet system. Thistechnique also allows flexibility and extensibility of thesystem with more robustness.
This paper aims at providing a general description for DIADORIM, a lexical database for Brazilian Portuguese. DIADORIM is said to successively merge two very different previous application-oriented dictionaries, increasing their user-friendliness, the reusability of their entries and their capability of incorporating new features. Besides improving the structure of the previous databases, DIADORIM also preserves their performance and functionality, as indicated in the real use and simulation tests carried out during its evaluation.
Research into the relationship between language and gender challenges group psychotherapy to pay attention to the significance of gender in shaping the language used by members and conductors. Language is a major resource in the creation of our gendered sense of self, with styles stereotypically associated with male and female. The linguistic culture of the group has stereotypical `male' and `female' aspects. The language of the therapist is critical in establishing linguistic norms, challenging or reinforcing gender stereotypes. The movement from these stereotypical styles, with the ability to draw upon both `male' and `female' characteristics, is a therapeutic movement. The absence of critical analysis of these aspects of language and gender in group theory witnesses to the power of the `social unconscious'.
In this paper, we observe various syntactic information for Korean parsing and propose a method to learn constraints and improve the efficiency of a parsing model by using the constraints. The proposed method has the following three characteristics. First, it improves the parsing efficiency since we use constraints that can prevent the parser from generating unsuitable candidates. Second, it is robust on a given Korean sentence because the attributes for the constraints are selected based on the syntactic and lexical idiosyncrasy of Korean. Third, it is easy to acquire constraints automatically from a treebank by using a decision tree learning algorithm. The experimental results show that the parser using acquired constraints can reduce the number of overgenerated candidates up to 1/2~1/3 of candidates and it runs 2~3 times faster than the one without any constraints.
There are two types... In this paper, we will describe evaluation and comparison of three different data-driven algorithms when applied to shallow parsing of Swedish. Additionally, based on the results, we propose a method for a fast and efficient development of a treebank.
SLIM is a prototype interactive multimedia self-learning linguistic software for foreign language students at beginner-false beginner level. It allows students to work both in an autonomous self-directed mode or in a way of programmed learning in which the process of self-instruction is pre-programmed and monitored. In this latter mode it incorporates assessment and evaluation tools in order to behave as an automatic tutor. It is organized into three basic components: audiovisual materials; a linguistic database recording all language material in text format; the supervisor. Audiovisual materials are partially taken from commercially available courses; the linguistic database is a highly sophisticated classification of all words and utterances of the course, both in written and spoken form, from all possible linguistic aspects. The supervisor is both an attractive, enjoyable and strongly pedagogically based software that allows the user to work on language materials. The most outstanding feature of SLIM is the use of speech analysis and recognition which is a fundamental aspect of all second language learning programmes. We also assume that a learning model can be represented by a finite state automaton made up by a fixed number of possible states – corresponding to the macro and microlevels at which the student's competence may be modelled – each one being internally constituted by the actual linguistic objects of knowledge of the language that make it up.
This mainly technological paper first provides a description of the web site called PapiLex. This first part shows how a file containing XML-structured lexical entries can be managed as a lexical database by using the Document Object Model (DOM) API. PapiLex offers the three essential management functions: creation, modification and deletion of a lexical entry. In a second part, two tools for entering Unicode-formatted text are presented: one for browsers having HTML 4 and JavaScript 1.2 capability and one for Microsoft Word. Such tools can be necessary for the minority languages which have no virtual keyboard embedded in the operating systems. 1 Starting point for building a lexical base Inside the Papillon project, the construction of a lexical base for a new language may take several different ways depending on where the author has to start. The following situations may occur regarding the availability of lexical resources1,2: • no dictionary exists, • a paper dictionary exists, • an electronic form of a dictionary exists, • a lexical database exists. In the last two cases, the question is to re-work the existing data so they meet the Papillon format and to fill the remaining fields. Tools are available for recycling electronic dictionary, (e.g. Nguyen 1998). Here, we will suppose that there is no preexisting dictionary or that its existence is limited to a paper dictionary. In such cases, the lexical entries have to be typed entirely. Among the 1: In addition to the existence of lexical resource, the script used for the language has also to be in Unicode and a font has to exist for it. Actually, the scripts of a number of minority languages are not in Unicode at the moment (e.g. Shan, Tai Dam, Mon). For some of them, fonts that really work are still missing as it is the case for Khmer. 2: In case there are existing data, property rights have to be looked at to say the resource is available. different ways in which this question can be handled, we chose a particular approach that consists in creating directly the Papillon formatted base by using generic and multiplatform Internet browsers. Section 2 will present how a standard browser can be used for this task3 (PapiLex mockup) and section 3 will show that a simple JavaScript program can provide a virtual keyboard that produces Unicode text. In section 4, another issue, less directly related to Papillon, will also be presented as it provides a very practical alternative for creating Unicodeencoded entries. It addresses a Windowsspecific tool for typing Unicode text in Microsoft Word when no standard keyboard is existing yet. The software was developed for the Lao language but can be applied to others. 2 The PapiLex mockup
In this paper we examine the construction of long-range language models using log-linear interpolation and how this can be achieved effectively. Particular attention is paid to the efficient computation of the normalisation in the models. Using the Penn Treebank for experiments we argue that the perplexity performance demonstrated recently in the literature using grammar-based approaches can actually be achieved with an appropriately smoothed 4-gram language model. Using such a model as the baseline, we demonstrate how further improvements can be obtained using loglinear interpolation to combine distance word and class models. We also examine the performance of similar model combinations for rescoring word lattices on a medium-sized vocabulary Wall Street Journal task. 1.
This paper reviews the first year of the creation of a publicly available treebank for Portuguese, Floresta Sintá(c)tica, a collaboration project between the VISL and the Computational Processing of Portuguese projects. After briefly describing the main goals and the organization of the project, the creation of the annotated objects is presented in detail: preparing the text to be annotated, applying the Constraint Grammar based PALAVRAS parser, revising its output manually in a two-stage process, and carefully documenting the linguistic options. Some examples of the kind of interesting problems dealt with are presented, and the paper ends with a brief description of the tools developed, the project results so far, and a mention to a preliminary inter-annotator test and what was learned from it. 1. Introduction: Motivation and
This paper describes new default unification, lenient default unification. It works efficiently, and gives more informative results because it maximizes the amount of information in the result, while other default unification maximizes it in the default. We also describe robust processing within the framework of HPSG. We extract grammar rules from the results of robust parsing using lenient default unification. The results of a series of experiments show that parsing with the extracted rules works robustly, and the coverage of a manually-developed HPSG grammar for Penn Treebank was greatly increased with a little overgeneration.
In this paper, I will examine some of the difficulties faced by the linguistic fieldworker who is attempting to observe and record "natural" conversations, and I will reconsider the long-held sociolinguistic notion of the observer's paradox by recasting it within Bell's (1984) framework of audience design theory. Using data gathered during my own fieldwork, I will once again call into question the idea of a single, unmarked, unperformed vernacular, the access to which is supposedly blocked by the observer's paradox. Finally, I will demonstrate that "performed" or "self-conscious" speech produced for the fieldworker can be useful in systematic linguistic analysis, and in gaining insights into local language ideologies and linguistic norms.
The specification phase is one of the most important and least supported parts of the software development process. The SAREL system has been conceived as a knowledge-based tool to improve the specification phase. The purpose of SAREL (Assistance System for Writing Software Specifications in Natural Language) is to assist engineers in the creation of software specifications written in Natural Language (NL). These documents are divided into several parts. We can distinguish the Introduction and the Overall Description as parts that should be used in the Knowledge Base construction. The information contained in the Specific Requirements Section corresponds to the information represented in the Requirements Base. In order to obtain high-quality software requirements specification the writing norms that define the linguistic restrictions required and the software engineering constraints related to the quality factors have been taken into account. One of the controls performed is the lexical analysis that verifies the words belong to the application domain lexicon which consists of the Required and the Extended lexicon. In this sense a synonym management process is needed in order to get a quality software specification. The aim of this paper is to present the synonym management process performed during the Knowledge Base construction. Such process makes use of the Spanish Wordnet developed inside the Eurowordnet project. This process generates both the Required lexicon and the Extended lexicon that will be used during the Requirements Base construction.
BACKGROUND: The age-related decline of dehydroepiandrosterone (DHEA) has prompted research on its experimental replacement in women. Although no relationship to sexual functioning in healthy women has been shown to date, DHEA replacement has potential for affecting sexual response. METHODS: To investigate DHEA effects, 16 sexually functional postmenopausal women participated in a randomized, double-blind, crossover protocol in which oral administration of DHEA (300 mg) or placebo occurred 60 minutes before the presentation of an erotic video segment. Blood DHEA sulfate (DHEAS) changes, subjective and physiological sexual responses, as well as affective responses were measured in response to videotaped neutral and erotic video segments. RESULTS: The concentration of DHEAS increased 2-5-fold following DHEA administration in all 16 women. Subjective ratings across DHEA and placebo conditions showed significantly greater mental (p < 0.016) and physical (p < 0.036) sexual arousal to the erotic video with DHEA vs. placebo. Positive affect also increased during the erotic video across drug conditions. Vaginal pulse amplitude (VPA) and vaginal blood volume (VBV) demonstrated a significant increase (p < 0.001) between neutral and erotic film segments within both conditions (DHEA and placebo) but did not differentiate drug conditions. CONCLUSION: In sum, increases in mental and physical sexual arousal ratings significantly increased in response to an acute dose of DHEA in postmenopausal women.
ABSTRACT: What links the formation of the state, the formation of linguistic norms, and the development of a grammatical tradition of the state language? In France governmental institutions not only help to spread the state language, but also inspire the first analyses of that language. The model for the grammatisation of French, and the context to promote such a process, is the formation of a national judicial system, through the demand that customary law be written down, and then approved by central authorities. Establishing local texts written in the national language required the analysis of that language, and based that analysis on concepts drawn from legal language. Establishing a fixed national norm depended on the creation of a judicial system founded on writing, and on the royal bureaucracy that resulted from that change.
The Penn Treebank encodes valuable information such as grammatical function, semantic roles, and identification of traces. The addition of such information was intended to facilitate the process of predicate-argument extraction. However, even with the enriched annotation this task is far from trivial and, to our knowledge, no complete set of predicate argument structures derived from the Treebank exists. Our paper describes a method for retrieving predicate-argument structures that circumvents the complexity of the tree structures in the corpus, while employing few template rules. Our system operates on a flattened, morphologically enriched version of the corpus. This flattened representation allows access to all levels of the tree simultaneously and thus enables the detection of the main sentence constituents by means of simple template rules. A small number of rules apply to identify the head words of each constituent and the latter fill in the constituent templates, to build the logical forms representative of the predicate argument structure. The system is robust in the face of incomplete syntactic coverage.
We describe the background and motivation for an e-learning project—IT-based Collaborative Learning in Grammar—where NLP resource reuse has become an important issue. The resources are of several kinds: POS-tagged and syntactically annotated corpora (treebanks), parsing systems and grammar writer’s workbenches, and visulization and manipulation tools for linguistically annotated corpora. Our experience thus far has been that although there are a number of such resources available e.g. on the Web, as a rule, numerous incompatibilities and lack of standardization at all levels—markup formats, linguistic annotation schemes, grammatical framework, software APIs, etc.—make the reuse of these resources into a non-trivial endeavor. 0. Preamble: the Setting It is generally acknowledged that the goal of teaching grammar—especially at the university level—should not primarily be that students memorize definitions of concepts and grammatical constructions, but rather that they understand and learn to recognize different structural patterns. This can hardly be achieved without giving students practical training in the skill of grammatical analysis. Research
Abstract There seems to be an increase in the use of foreignisms in Finnish business translation: a new linguistic norm, which incorporates foreignisms appears to be developing alongside traditional Finnish usage. Yet how justified is the use of foreignisms in business translations? Here is an analysis of a Nokia document: the author explores the status of this text as translation and offers criteria for justified and unjustified foreignisms in Finnish business translation.
Cet article a pour objectif d’interroger les relations existant entre le système linguistique (entendu au sens large, comme englobant l’ensemble des règles ou régularités qui sous-tendent la production et l’interprétation des énoncés attestés), et la culture, et plus précisément les normes communicatives en vigueur dans une société donnée. Après avoir envisagé un certain nombre de faits langagiers (unités lexicales, formes honorifiques, actes de langage et formules rituelles) qui portent manifestement la trace de ces normes culturelles sous-jacentes, l’auteure en conclut qu’il est dans une certaine mesure possible de reconstituer à partir de ces traces l’« ethos communicatif » de la société considérée, mais que ce travail de reconstitution ne va pas sans rencontrer un certain nombre de difficultés, qui sont passées en revue.
Comprehensive computational lexicons areessential to practical natural languageprocessing (NLP). To compile such computationallexicons by automatically acquiring lexicalinformation, however, we previously requiresufficiently large corpora. This study aims atpredicting the ideal size of suchautomatic-lexical-acquisition oriented corpora,focusing on six specific factors: (1) specificversus general purpose prediction, (2)variation among corpora, (3) base forms versus inflected forms, (4) open class items,(5) homographs, and (6) unknown words.Another important and related issue withregard to predictability has something to dowith data sparseness. Research using theTOTAL Corpus reveals serious datasparseness in this corpus. This, again, pointstowards the importance and necessity ofreducing data sparseness to a satisfactorylevel for the automatic lexical acquisition andreliable corpus predictions. The functions ofpredicting the number of tokens and lemmas in acorpus are based on the piecewisecurve-fitting algorithm. Unfortunately, thepredicted size of a corpus for automaticlexical acquisition is too astronomicalto compile it by using presently existingcompiling strategies. Therefore, we suggest apractical and efficient alternative method. Weare confident that this study will shed newlight on issues such as corpus predictability,compiling strategies and linguisticcomprehensiveness.
The term ‘standard language ideology’ as described by Milroy and Milroy (1998) and Lippi-Green (1994) characterises a particular set of beliefs about language. Such beliefs are typically held by populations of economically developed nation states where processes of standardisation have operated over a considerable time to produce an abstract set of norms-lexical, grammatical and (in spoken language) phonological-popularly described as constituting a standard language. The same beliefs also emerge, somewhat transformed by local histories and conditions, in these states’ colonies and ex-colonies. For example, in all the countries discussed by contributors to Cheshire (1991) where English has been imported (I confine my comments in this chapter to English-speaking hegemonies), beliefs about language similar to those discussed below can be found. These are reported in a sizeable literature; for example Platt and Weber (1980) and Gupta (1994) both describe the operation of a British-style standard language ideology in Singapore. Although debates about standard English are a staple of the British press (in the United States the most contentious ideological debates are usually slightly differently oriented, as we shall see), experts and laypersons alike have just about as much success in locating a specific agreed spoken standard variety in either Britain or the United States as have generations of children in locating the pot of gold at the end of the rainbow.
This paper describes a general-purpose sentence generation system that can achieve both broad scale coverage and high quality while aiming to be suitable for a variety of generation tasks. We measure the coverage and correctness empirically using a section of the Penn Treebank corpus as a test set. We also describe novel features that help make the generator flexible and easier to use for a variety of tasks. To our knowledge, this is the first empirical measurement of coverage reported in the literature, and the highest reported measurements of correctness.
Treebank formats and associated software tools are proliferating rapidly, with little consideration for interoperability. We survey a wide variety of treebank structures and operations, and show how they can be mapped onto the annotation graph model, and leading to an integrated framework encompassing tree and non-tree annotations alike. This development opens up new possibilities for managing and exploiting multilayer annotations.
During the analysis of the protagonist's (namely, Sharik's and Sharikov's) way of speaking in Bulgakov's story Heart of a Dog, an abrupt contrast or even a complete oppositeness of the constituents becomes apparent. In its turn, this provides the base and evidence for this ultimate oppositeness of the protagonists in the story in general. Sharikov's speech is mainly characterized by the following features: 1) absence of skills of monological speech manifested by the violation of norms of constructing sentences and by the tendency towards using short and concise sentences, 2) violation of lexical and grammatical norms, 3) abundance in colloquialisms, 4) frequency of generalized and demagogic constructions, 5) presence of officialese and ideological clichés. It is Sharikov's speech and his way of speaking that enables the reader to make conclusions about his figure in general, and determine the most important characteristic features of his inner self which are as follows: 1) low cultural level, 2) aggressiveness and growing confidence in his own right, 3) belonging to the layer of uneducated, uncivilized and often declassed people.