Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
The article shows the wealth of colloquial language features in the city environment through the presence of texts in the city reality which are designed for collective receivers/recipients (eg. sign-board, information advertising, price labels, etc.). In the research, the components revealing descanting in the urban language (dialectal and sociolectal features) were found, as well as the associated evaluation of objects and phenomena, colloquiality or even familiarity of idea transfer, and free realisation of orthographic and stylistic norms. Urban texts bear testimony of frequent language taboo breaking in the original sphere as well as in the area violating tactfulness and politeness canons, up to violation of decency and modesty. In the thesis, the changes in the sphere of native words meaning (neologisms and neosemantisms) and examples of introducing allogenic lexemes (orientalisms) are discussed. The important feature of the examples analysed is ambiguity, present in the lexical area as well as in the global apprehension of the message, which could decide about the language game played with receivers.
espanolEn la Argentina, desde 1870 se inicio una prolifica produccion de instrumentos lexicograficos que registraban singularidades lexicas. La conciencia de tal peculiaridad condujo a confeccionar, continuando con la tradicion hispanoamericana, diccionarios complementarios y contrastivos de diferentes modalidades. Por un lado, se publicaron obras descriptivas que recogian ruralismos, indigenismos, regionalismos (tanto americanismos como provincialismos o localismos) y argentinismos. Por otro, algunas normativas que recolectaban barbarismos y censuraban su uso, tomando como parametro la norma del castellano peninsular. En este trabajo, analizamos puntualmente un dominio del discurso lexicografico: los mecanismos de citacion y ejemplificacion. Primero, expondremos las diversas clases de ejemplos y sus funcionamientos. Luego, examinaremos nuestro corpus centrandonos en dos aspectos: a) las condiciones del proceso de diccionarizacion; b) los modos de funcionamiento discursivo de los ejemplos en la lexicografia monolingue argentina. Apuntamos a mostrar que dicho dominio, tanto como el paratexto, la nomenclatura y la microestructura, permite vincular el discurso lexicografico con el imaginario nacional. EnglishSince 1870, a prolific production of lexicographical instruments registering lexical singularities began in Argentina. The awareness of this peculiarity led to the elaboration following a Hispanic American tradition of complementary and contrastive dictionaries of different modalities. On one hand, descriptive works that collected ruralisms, indigenisms, regionalisms (both Americanisms and provincialisms or localisms), and Argentinisms were published. On the other hand, some normative works that gathered barbarisms and condemned their use emerged, using as a parameter the peninsular Spanish norm. In this paper, we analyze specifically a domain of lexicographical discourse: the quotation and exemplification mechanisms. First, we expose different types of examples and their functioning. Then, we examine our corpus focusing on two aspects: a) the conditions of the dictionarization process; and b) discursive functioning of examples in Argentine monolingual lexicography. We aim to show that this domain, as well as paratext, nomenclature and microstructure, allows us to link lexicographical discourse with national imaginary.
Abstract Especially for non-experts, translating legal texts is a complicated and multifaceted process, demanding much linguistic and technical competence from translators. Legal systems differ from one another, and each one has a specific set of norms, especially reflected at the lexical level – that is, in terminology. Understanding two legal systems is not easy, even for experts; of course, it is even more difficult for translators, who are usually not legal experts. This article focuses on how to quickly and transparently provide translators with some (basic) technical knowledge using ontologies, which in recent years have found considerable application in synthesizing and visualizing knowledge.
From a global perspective, bilingual language acquisition can be considered the norm rather than the exception. In bilingual communities around the world, infants exposed from birth to two different languages, or even dialects, succeed in the task of simultaneously learning their two native languages. Infants growing up in this type of environments are exposed to a complex input that contains information relative to two different phonological systems. Early in development bilingual-to-be infants must be able to differentiate the sound patterns of their two languages and start building languagespecific phonetic categories. Research on young bilinguals’ phonetic categorization and perceptual reorganization processes by the end of the first year of life has revealed interesting differences between consonant and vowel categories. Once in the lexical stage, phonetic categories already established will turn into the contrastive categories that form the phonological systems for each of the ambient languages. This is by no means an automatic process. Data from studies with monolingual toddlers participating in word learning tasks have revealed that minimal pair word labels, differing in their initial stop consonant, such as [bih] and [dih], cannot be easily learned at 14 months of age, even though /b/ and /d/ contrastive sounds can be discriminated with no difficulty at the same age (Stager & Werker, 1997). In the case of bilingual toddlers, engaged in the process of establishing two lexicons based on two distinct phonological systems, the situation is even more challenging. There are still relatively few studies specifically focusing on bilinguals’ setting up the phonetic and phonological categories of their native languages (see Werker & Byers-Heinlein, 2008, for a review). Experimental data come mostly from three research groups settled in areas where bilingual populations are available for participation in speech perception studies: J. Werker group at the University of British Columbia in Vancouver (Canada), L. Polka group at McGill University in Montreal (Canada) and the group at the University of Barcelona (Spain) whose main findings will be described in the following sections. Researchers from the above mentioned groups, dealing with bilingual infants and toddlers from various language communities and exposed to different pairs of languages, have all contributed to shed light on the adaptability of the speech processing system to cope with different types of linguistic input. What previous research in bilingual language development had told us, from a general perspective, was that the pattern of acquisition in bilinguals was rather similar to the pattern of acquisition that had been described for monolingual infants: an early language differentiation was suggested as words in both of the ambient languages were present in their initial expressive lexicons (Genesee, Nicoladis & Paradis, 1995; Pearson, Fernandez & Oller, 1995) and they followed the same steps as monolinguals’ in reaching the key milestones in the language acquisition process (Oller, Eilers, Urbano & Cobo-Lewis, 1997). From a phonological acquisition perspective, however, input to bilinguals has specific properties and clearly differs from monolingual input, not only in complexity (two lexicons, two phonologies), but also in quantity and quality of exposure to each language. Moreover, the degree of proximity between the specific lexical, phonological and morpho-syntactical properties of the two ambient languages is also a relevant factor to be taken into consideration. The complex and variable nature of the input to bilingual infants and toddlers can determine minor time-course differences in reaching specific sound discrimination abilities or in stabilizing certain phonetic categories when comparing bilingual and monolingual infants. But, more interestingly, similarities or differences in the phonetic and phonological properties of the two languages in the input can result in differences in perception/discrimination abilities observed in groups of bilinguals from different linguistic environments. Language differentiation processes, the setting up of language-specific phonetic categories, phonological representation of sounds in the lexicon, might differ when comparing bilinguals from different pairs of languages.
Electronic dictionaries covering all natural language levels are very relevant for the human use as well as for the automatic processing use, namely those constructed with respect to international standards. Such dictionaries are characterized by a complex structure and an important access time when using a querying system. However, the need of a user is generally limited to a part of such a dictionary according to his domain and expertise level which corresponds to a specialized dictionary. Given the importance of managing a unified dictionary and considering the personalized needs of users, we propose an approach for generating personalized views starting from a normalized dictionary with respect to Lexical Markup Framework LMF-ISO 24613 norm. This approach provides the re-use of already defined views for a community of users by managing their profiles information and promoting the materialization of the generated views. It is composed of four main steps: (i) the projection of data categories controlled by a set of constraints (related to the user's profiles), (ii) the selection of values with consistency checking, (iii) the automatic generation of the query's model and finally, (iv) the refinement of the view. The proposed approach was consolidated by carrying out an experiment on an LMF normalized Arabic dictionary.
As an abstract system of symbols, language can be realized in its spoken and written form. The writing process, as the basis of written language development, represents one of the four language skills – listening, speaking, reading, and writing (European Commission, 2005). Standard Croatian language consists of 32 sounds which are noted down in gajica, a Latin script composed of 22 basic and 5 derived one-letter symbols (c, c, đ, s, ž) and three two-letter symbols (dž, lj, nj), arranged in the alphabetic sequence (Hrvatski skolski pravopis, 2005). Institutional L1 acquisition in the Republic of Croatia is composed of the following institutions: preschool – elementary school – high school. However, systematic learning and teaching begins when children start school (initial reading and writing period). The development of written language includes physical activity (motor and visual) and psycho-cognitive activity (acquiring the grapheme standard, learning the rules of orthography, mastering the grammatical and lexical structure of that language). According to the graphomotor criterion, first grade pupils are expected to master all four types of Croatian grapheme system (lower and upper case block letters, and lower and upper case cursive letters); in accordance with the orthographic criterion, pupils need to master the basic orthographic principles and rules as outlined in the Croatian language curriculum (HNOS, 2005); according to the creativity criterion, pupils are expected to form and write words, structure simple sentences and create short texts. Initial writing has been emphasised as an important educational achievement in the process of functional learning and tuition in Europe as well (Bildungsplan fur die Grundschule, 1994). Unfortunately, there is a great discrepancy between theory and practice in the Croatian educational system. There are some ten primers with different approaches to teaching literacy. This fact points to the lack of standardisation in the initial process of writing and also makes one wonder about the appropriateness of teaching in the initial writing process (Bežen, 2005). All that has been mentioned points to the necessity of research in the field of initial reading and writing. The authors of this paper present the results of a research carried out in first grades of primary schools (big town, suburb, small town, and village) in the Republic of Croatia, in 2007. The aim was to establish the level of proficiency in writing as the basis of the acquisition of orthographic competence in the early language discourse. On the sample N = 301 the usage of block and cursive letters has been investigated as well as the level of acquisition of writing, formation and writing of letters, sentences and text, and testing of the knowledge of orthographic norms. The data collected were analysed by means of the SPSS statistics software. The instruments were t-test and variance analysis which were used to establish the existence of a statistically relevant difference in the examined contents, in accordance with the results at the tests of linguistic and communicative competences (Pavlicevic-Franic, 2005).
This paper describes an attempt to build a lexical database for the Yami language, an Austronesian endangered language. As the Yami language documentation and conservation projects have produced substantial corpora, we are now ready to construct the Yami online knowledge database based on the knowledge we have accumulated in the language. In this paper, we propose a model to build the Word Net-like Yami lexical semantics and database. The model is first described in detail, followed by an illustration of an ontology of fish in the implementation phase.
The aim of this article is to analyze the formation of Old English adverbs (A-Y) as retrieved from the lexical database of Old English Nerthus within the theoretical framework of the Layered Structure of the Word. Firstly, a critical review of the literature on Old English adverb formation is offered in order to emphasize the necessity of an exhaustive and theoretically up-to-date study that distinguishes clearly synchronic from diachronic aspects on the one hand, and inflectional from derivational aspects on the other. Secondly, an exhaustive analysis of the derivation of adverbs in Old English by means of different word-formation processes (zero derivation, conversion, affixation and compounding) is given. In the theoretical part, the conclusion reached is that conversion requires a Complex Word structure and that a distinction has to be drawn between syntactic exocentricity and morphological exocentricity.
This article discusses the treatment of collocations in the context of along-term project on the development of multilingual NLP tools. Besides“classical” two-word collocations, we will focus on the case of complexcollocations (3 words or more) for which a recursive design is presented in theform of collocation of collocations. Although comparatively less numerous thantwo-word collocations, the complex collocations pose important challenges forNLP. The article discusses how these collocations are retrieved from corpora,inserted and stored in a lexical database, how the parser uses such knowledgeand what are the advantages offered by a recursive approach to complexcollocations.
uni-tuebingen.de This paper describes a CoNLL-style chunk representation for the Tübingen Treebank of Written German, which assumes a flat chunk structure so that each word belongs to at most one chunk. For German, such a chunk definition causes problems in cases of complex prenominal modification. We introduce a flat annotation that can handle these structures via a stranded noun chunk. 1
The article analyzes 97 elementary schoolbooks in Buenos Aires to determine which social representations about linguistic norm underlie in these school materials. The paper reviews -especially in the defi nitions of categories, and exercises and activities- the concepts of linguistic variety, standard language and español neutro. Based on these variables, this article sees the possible repercussions in social representations that students and teachers can develop from point of view of the publishing companies.
This paper describes the transfer component of a syntax-based Example-based Machine Translation system. The source sentence parse tree is matched in a bottom-up fashion with the source language side of a parallel example treebank, which results in a target forest which is sent to the target language generation component. The results on a 500 sentences test set are compared with a top-down approach to transfer of the same system, with the bottom-up approach yielding much better results. 1
We describe a process for converting the Penn Arabic Treebank into the CCG formalism. Previous efforts have yielded CCGbanks in English, German, and Turkish, thus opening these languages to the sophisticated computational tools developed for CCG and enabling further cross-linguistic development. Conversion from a context free grammar treebank to a CCGbank is a four stage process: head finding, argument classification, binarization, and category conversion. In the process of implementing a basic CCGbank conversion algorithm, we reveal properties of Arabic grammar that interfere with conversion, such as subject topicalization, genitive constructions, relative clauses, and optional pronominal subjects. All of these problematic phenomena can be resolved in a variety of ways- we discuss advantages and disadvantages of each in their respective sections. We detail these and describe our categorial analysis of each of these Arabic grammatical phenomena in depth, as well as technical details on their integration into the conversion algorithm. 1.
For centuries, scholars have explored the deep links among human languages. In this paper, we present a class of probabilistic models that use these links as a form of naturally occurring supervision. These models allow us to substantially improve performance for core text processing tasks, such as morphological segmentation, part-of-speech tagging, and syntactic parsing. Besides these traditional NLP tasks, we also present a multilingual model for the computational decipherment of lost languages. 1. Overview Electronic text is currently being produced at a vast and unprecedented scale across the languages of the world. Natural Language Processing (NLP) holds out the promise of automatically analyzing this growing body of text. However, over the last several decades, NLP research efforts have focused on the English language, often neglecting the thousands of other languages of the world (Bender, 2009). Most of these languages are currently beyond the reach of NLP technology due to several factors. One of these is simply the lack of the kinds of hand-annotated linguistic resources that have helped propel the performance of English language systems. For complex tasks of linguistic analysis, hand-annotated corpora can be prohibitively time-consuming and expensive to produce. For example, the most widely used annotated corpus in the English language, the Penn Treebank (Marcus et al., 1994), took years for a team of professional linguists to produce. It is unrealistic to expect such resources to ever exist for the majority of the world’s languages.
Two of the main corpora available for training discourse relation classifiers are the RST Discourse Treebank (RST-DT) and the Penn Discourse Treebank (PDTB), which are both based on the Wall Street Journal corpus. Most recent work using discourse relation classifiers have employed fully-supervised methods on these corpora. However, certain discourse relations have little labeled data, causing low classification performance for their associated classes. In this paper, we attempt to tackle this problem by employing a semi-supervised method for discourse relation classification. The proposed method is based on the analysis of feature cooccurrences in unlabeled data. This information is then used as a basis to extend the feature vectors during training. The proposed method is evaluated on both RST-DT and PDTB, where it significantly outperformed baseline classifiers. We believe that the proposed method is a first step towards improving classification performance, particularly for discourse relations lacking annotated data.
In this methodological investigation, we examined the influence of cultural background on viewers' interpretations of visual stimuli and verbs elicited by these materials. French and Mandarin native speakers' interpretations of seventeen short movies, produced by French speakers, depicting various state-changing actions were collected by a 25-item cultural protocol. A slight difference in the familiarity rating of movies is found between French and Mandarin participants. We also found that Mandarin speakers used more general verbs when describing actions depicted by movies with low familiarity rating and children used more conventional forms with movies of higher familiarity. Hierarchical cluster analyses were conducted in selecting movies that were matched in action-interpretations by both language groups.
Abstract Experiences with racism and age negatively affect how Afro-Brazilians in Salvador and São Paulo rate democracy. Older cohorts are more likely to rate democracy high compared to younger cohorts who rate it as low. Respondents in Salvador tend to rate democracy lower than respondents in São Paulo. Moreover, interviews reveal that as citizens believe they are not accorded full rights, they do not agree that Brazil's political system is fully democratic. Studies examining democracy in Brazil and racial politics throughout the diaspora would benefit from examining racialized experiences of citizens, rather than simply including the demographic variable of race. It is these experiences that affect rating of democracy rather than ascribed notions of race.
This study explored the conceptual framework of dieticians' intentions to recommend functional food and the mediating role of consumption frequency. A web-based survey was designed using a self-administered questionnaire. A sample of Korean dieticians (N=233) responded to the questionnaire that included response efficacy, risk perception, consumption frequency, and recommendation intention for functional foods. A structural equation model was constructed to analyze the data. We found that response efficacy was positively related to frequency of consumption of functional foods and to recommendation intention. Consumption frequency also positively influenced recommendation intention. Risk perception had no direct influence on recommendation intention; however, the relationship was mediated completely by consumption frequency. Dieticians' consumption frequency and response efficacy were the crucial factors in recommending functional foods. Dieticians may perceive risks arising from the use of functional foods in general, but the perceived risks do not affect ratings describing dieticians' intentions to recommend them. The results also indicated that when dieticians more frequently consume functional foods, the expression of an intention to recommend functional foods may be controlled by the salience of past behaviors rather than by attitudes.
In this paper, we argue for and demonstrate the use of Prolog as a tool to query annotated corpora. We present a case study based on the German TüBa-D/Z Treebank to show that flexible and efficient corpus querying can be started with a minimal amount of effort. We end this paper with a brief discussion of performance, that suggests that the approach is both fast enough and scalable. 1
Objective: To investigate whether interviewer personality, sex or being of the same sex as the interviewee, and training account for variance between interviewers’ ratings in a medical student selection interview. Design, setting and participants: In 2006 and 2007, data were collected from cohorts of each year's interviewers (by survey) and interviewees (by interview) participating in a multiple mini-interview (MMI) process to select students for an undergraduate medical degree in Australia. MMI scores were analysed and, to account for the nested nature of the data, multilevel modelling was used. Main outcome measures: Interviewer ratings; variance in interviewee scores. Results: In 2006, 153 interviewers (94% response rate) and 268 interviewees (78%) participated in the study. In 2007, 139 interviewers (86%) and 238 interviewees (74%) participated. Interviewers with high levels of agreeableness gave higher interview ratings (correlation coefficient [r] = 0.26 in 2006; r = 0.24 in 2007) and, in 2007, those with high levels of neuroticism gave lower ratings (r = − 0.25). In 2006 but not 2007, female interviewers gave higher overall ratings to male and female interviewees (t = 2.99, P = 0.003 in 2006; t = 2.16, P = 0.03 in 2007) but interviewer and interviewee being of the same sex did not affect ratings in either year. The amount of variance in interviewee scores attributable to differences between interviewers ranged from 3.1% to 24.8%, with the mean variance reducing after skills-based training (20.2% to 7.0%; t = 4.42, P = 0.004). Conclusion: This study indicates that rating leniency is associated with personality and sex of interviewers, but the effect is small. Random allocation of interviewers, similar proportions of male and female interviewers across applicant interview groups, use of the MMI format, and skills-based interviewer training are all likely to reduce the effect of variance between interviewers.
Nivre’s method was improved by enhancing deterministic dependency parsing through application of a tree-based model. The model considers all words necessary for selection of parsing actions by including words in the form of trees. It chooses the most probable head candidate from among the trees and uses this candidate to select a parsing action. In an evaluation experiment using the Penn Treebank (WSJ section), the proposed model achieved higher accuracy than did previous deterministic models. Although the proposed model’s worst-case time complexity is O(n 2), the experimental results demonstrated an average parsing time not much slower than O(n). 1
This dissertation considers the language socialization of law students. One message that the law students encounter is that legal Swedish is an entirely new language. The main aim is to investigate what linguistic norms are conveyed to the students through the teachers’ comments on the students’ texts and through various forms of writing instructions. The material consists of student texts with teacher comments and documentation on various phases of instruction with a focus on writing. Teacher comments on texts written during the first year of the law programme are analyzed and categorized. The analysis stems from two models. The first model is based on different text levels, like formal conventions of writing, sentence construction, text structure, word choice and style, and content. The second model distinguishes different linguistic norms based on three layers: The first layer consists of written language norms in general language practice, the second of academic language norms and the third of norms that are specific to the use of legal language. The results show that word choice and style is the most common category for the teachers’ comments in the first term of the law programme and content is the most common in the second term (with word choice and style the second most common). Formal conventions of writing, sentence structure and different types of grammatical constructions are some of the things the teachers criticize. Surprisingly few of the teachers’ comments concern more overarching aspects such as text structure or the aim and genre of the text. Comments are made on local features in the text, but rarely on more global features. The teaching practice that the writing of law students belongs to entails, among other things, that the students’ texts are assessed anonymously for the sake of fairness. This means that there is not much opportunity for a student to discuss the text with the teacher who commented on and assessed it. The construction of the teachers’ text comments is particularly important when dialogue between student and teacher on the text draft and final version is not an integral part of instruction. The teachers’ written comments are usually brief and do not allow much space for a consideration of linguistic norms and text patterns, which reduces the opportunities for the teachers and the law programme to contribute to a deeper linguistic awareness in the law students.
In this paper, we propose two different language modeling approaches, namely skip trigram and across sentence boundary, to capture the long range dependencies. The skip trigram model is able to cover more predecessor words of the present word compared to the normal trigram while the same memory space is required. The across sentence boundary model uses the word distribution of the previous sentences to calculate the unigram probability which is applied as the emission probability in the word and the class model frameworks. Our experiments on the Penn Treebank [1] show that each of our proposed models and also their combination significantly outperform the baseline for both the word and the class models and their linear interpolation. The linear interpolation of the word and the class models with the proposed skip trigram and across sentence boundary models achieves 118.4 perplexity while the best state-of-the-art language model has a perplexity of 137.2 on the same dataset. 1.
OBJECTIVE: To explore the possible differences in subjective analysis of the emotional stimuli from the International Affective Picture System between elderly and young samples. METHOD: 187 elderly subjects ranked the International Affective Picture System images according to the directions from the Manual of Affective Ratings. Their scores were compared to those obtained from International Affective Picture System studies with young people. RESULT: There is an age-related difference in arousal and valence in the International Affective Picture System rating. The correlation between affective valence and arousal is strong, and negative for the elderly. The expected versus the observed frequency of International Affective Picture System images between elderly and young samples show a statistical difference. CONCLUSION: This study shows an inter-age statistical dichotomy in how elderly and young people subjectively evaluate International Affective Picture System images.
There often exist multiple corpora for the same natural language processing (NLP) tasks. However, such corpora are generally used independently due to distinctions in annotation standards. For the purpose of full use of readily available human annotations, it is significant to simultaneously utilize multiple corpora of different annotation standards. In this paper, we focus on the challenge of constituent syntactic parsing with treebanks of different annotations and propose a collaborative decoding (or co-decoding) approach to improve parsing accuracy by leveraging bracket structure consensus between multiple parsing decoders trained on individual treebanks. Experimental results show the effectiveness of the proposed approach, which outperforms stateof-the-art baselines, especially on long sentences. 1
We present how the conceptually and numerically simple concept of a fuzzy linguistic database summary can be a very powerful tool for gaining much insight into the very essence of data. The use of linguistic summaries provides tools for the verbalisation of data analysis (mining) results which, in addition to the more commonly used visualisation, e.g. via a graphical user interface, can contribute to an increased human consistency and ease of use, notably for supporting decision makers via the data-driven decision support system paradigm. Two new relevant aspects of the analysis are also outlined which were first initiated by the authors. First, following Kacprzyk and Zadrożny, it is further considered how linguistic data summarisation is closely related to some types of solutions used in natural language generation (NLG). This can make it possible to use more and more effective and efficient tools and techniques developed in NLG. Second, similar remarks are given on relations to systemic functional linguistics. Moreover, following Kacprzyk and Zadrożny, comments are given on an extremely relevant aspect of scalability of linguistic summarisation of data, using a new concept of a conceptual scalability.
Proceedings of the Ninth International Workshop \non Treebanks and Linguistic Theories. \nEditors: Markus Dickinson, Kaili Müürisep and Marco Passarotti. \nNEALT Proceedings Series, Vol. 9 (2010), 103-113. \n© 2010 The editors and contributors. \nPublished by \nNorthern European Association for Language \nTechnology (NEALT) \nhttp://omilia.uio.no/nealt. \nElectronically published at \nTartu University Library (Estonia) \nhttp://hdl.handle.net/10062/15891.
Lexicon-Grammar tables are a very rich syntactic lexicon for the French language. This linguistic database is nevertheless not directly suitable for use by computer programs, as it is incomplete and lacks consistency. Tables are defined on the basis of features which are not explicitly recorded in the lexicon. These features are only described in literature. Our aim is to define for each tables these essential properties to make them usable in various Natural Language Processing (NLP) applications, such as parsing.
We present a probabilistic model extension to the Tesnière Dependency Structure (TDS) framework formulated in (Sangati and Mazza, 2009). This representation incorporates aspects from both constituency and dependency theory. In addition, it makes use of junction structures to handle coordination constructions. We test our model on parsing the English Penn WSJ treebank using a re-ranking framework. This technique allows us to efficiently test our model without needing a specialized parser, and to use the standard evaluation metric on the original Phrase Structure version of the treebank. We obtain encouraging results: we achieve a small improvement over state-of-the-art results when re-ranking a small number of candidate structures, on all the evaluation metrics except for chunking.
This article investigates a relatively underdeveloped subject in natural language processing---the generation of punctuation marks. From a theoretical perspective, we study 16 Chinese punctuation marks as defined in the Chinese national standard of punctuation usage, and categorize these punctuation marks into three different types according to their syntactic properties. We implement a three-tier maximum entropy model incorporating linguistically-motivated features for generating the commonly used Chinese punctuation marks in unpunctuated sentences output by a surface realizer. Furthermore, we present a method to automatically extract cue words indicating sentence-final punctuation marks as a specialized feature to construct a more precise model. Evaluating on the Penn Chinese Treebank data, the MaxEnt model achieves an f -score of 79.83% for punctuation insertion and 74.61% for punctuation restoration using gold data input, 79.50% for insertion and 73.32% for restoration using parser-based imperfect input. The experiments show that the MaxEnt model significantly outperforms a baseline 5-gram language model that scores 54.99% for punctuation insertion and 52.01% for restoration. We show that our results are not far from human performance on the same task with human insertion f -scores in the range of 81-87% and human restoration in the range of 71-82%. Finally, a manual error analysis of the generation output shows that close to 40% of the mismatched punctuation marks do in fact result in acceptable choices, a fact obscured in the automatic string-matching based evaluation scores.
In order to parse Arabic texts, we have chosen to use a machine learning approach. It learns from an Arabic Treebank. The knowledge enclosed in this Treebank is structured as patterns of syntactic trees. These patterns are representative models of syntactic components of the Arabic language. They are not only layered but also both structurally and contextually rich. They serve as an informational source for guiding the parsing process. Our parser is progressive given that it proceeds by treating a sentence into a number of stages, equal to the number of its words. At each step, the parser affects the target word with the most likely patterns to represent it in the context where it is put. Then, it joins the selected patterns with those collected in the previous steps so as to construct the representative syntactic tree(s) of the whole sentence. Preliminary tests have yielded to obtain accuracy and f-score which are respectively equal to 84.78% and 77.52%.
OBJECTIVE: To investigate whether interviewer personality, sex or being of the same sex as the interviewee, and training account for variance between interviewers' ratings in a medical student selection interview. DESIGN, SETTING AND PARTICIPANTS: In 2006 and 2007, data were collected from cohorts of each year's interviewers (by survey) and interviewees (by interview) participating in a multiple mini-interview (MMI) process to select students for an undergraduate medical degree in Australia. MMI scores were analysed and, to account for the nested nature of the data, multilevel modelling was used. MAIN OUTCOME MEASURES: Interviewer ratings; variance in interviewee scores. RESULTS: In 2006, 153 interviewers (94% response rate) and 268 interviewees (78%) participated in the study. In 2007, 139 interviewers (86%) and 238 interviewees (74%) participated. Interviewers with high levels of agreeableness gave higher interview ratings (correlation coefficient [r] = 0.26 in 2006; r = 0.24 in 2007) and, in 2007, those with high levels of neuroticism gave lower ratings (r = -0.25). In 2006 but not 2007, female interviewers gave higher overall ratings to male and female interviewees (t = 2.99, P = 0.003 in 2006; t = 2.16, P = 0.03 in 2007) but interviewer and interviewee being of the same sex did not affect ratings in either year. The amount of variance in interviewee scores attributable to differences between interviewers ranged from 3.1% to 24.8%, with the mean variance reducing after skills-based training (20.2% to 7.0%; t = 4.42, P = 0.004). CONCLUSION: This study indicates that rating leniency is associated with personality and sex of interviewers, but the effect is small. Random allocation of interviewers, similar proportions of male and female interviewers across applicant interview groups, use of the MMI format, and skills-based interviewer training are all likely to reduce the effect of variance between interviewers.
Developing large-scale deep grammars in a constraint-based framework such as Lexical Functional Grammar (LFG) is time-consuming and requires significant linguistic insight. Recently, treebank-based constraint-grammar acquisition approaches have been developed as an alternative to hand-crafting such resources. While treebank-based approaches are wide coverage and robust and achieve competitive evaluation results for many languages, the granularity of the linguistic analyses provided by treebank-based resources tends to be less fine-grained than what is offered by state-of-the-art handcrafted grammars. This paper presents an approach to extend the English DCU LFG annotation algorithm with more detailed f-structure information to provide probabilistic treebank-based LFG grammars with rich feature information comparable to that implemented by the hand-crafted English XLE grammar, while maintaining the robustness and the coverage of treebankbased stochastic grammars.
Electronic dictionaries offer many possibilities unavailable in paper dictionaries to view, display or access information. However, even these resources fall short when it comes to access words sharing semantic features and certain aspects of form: few applications offer the possibility to access a word via a morphologically or semantically related word. In this paper, we present such an application, POLYMOTS, a lexical database for contemporary French containing 20.000 words grouped in 2.000 families. The purpose of this resource is to group words into families on the basis of shared morpho-phonological and semantic information. Words with a common stem form a family; words in a family also share a set of common conceptual fragments (in some families there is a continuity of meaning, in others meaning is distributed). With this approach, we capitalize on the bidirectional link between semantics and morpho-phonology: the user can thus access words not only on the basis of ideas, but also on the basis of formal characteristics of the word, i.e. its morphological features. The resulting lexical database should help people learn French vocabulary and assist them to find words they are looking for, going thus beyond other existing lexical resources. 1.
We describe a model for the lexical analy-sis of Arabic text, using the lists of alterna-tives supplied by a broad-coverage morpho-logical analyzer, SAMA, which include sta-ble lemma IDs that correspond to combina-tions of broad word sense categories and POS tags. We break down each of the hundreds of thousands of possible lexical labels into its constituent elements, including lemma ID and part-of-speech. Features are computed for each lexical token based on its local and document-level context and used in a novel, simple, and highly efficient two-stage super-vised machine learning algorithm that over-comes the extreme sparsity of label distribu-tion in the training data. The resulting system achieves accuracy of 90.6 % for its first choice, and 96.2 % for its top two choices, in selecting among the alternatives provided by the SAMA lexical analyzer. We have successfully used this system in applications such as an online reading helper for intermediate learners of the Arabic language, and a tool for improving the productivity of Arabic Treebank annotators. 1
Understanding the syntactic structure of a sentence is a necessary preliminary to understanding its semantics and therefore for many practical applications. The field of natural language processing has achieved a high degree of accuracy in parsing, at least in English. However, the syntactic structures produced by the most commonly used parsers are less detailed than those structures found in the treebanks the parsers were trained on. In particular, these parsers typically lack the null elements used to indicate wh-movement, control, and other phenomena. This thesis presents a system for inserting these null elements into parse trees in English. It then examines the problem in Arabic, which motivates a second, joint-inference system which has improved performance on English as well. Finally, it examines the application of information derived from the Google Web 1T corpus as a way of reducing certain data sparsity issues related to wh-movement.
In this paper, we propose automatic categorization and summarization of documentaries using subtitles of videos. We propose two methods for video categorization. The first makes unsupervised categorization by applying natural language processing techniques on video subtitles and uses the WordNet lexical database and WordNet domains. The second has the same extraction steps but uses a learning module to categorize. Experiments with documentary videos give promising results in discovering the correct categories of videos. We also propose a video summarization method using the subtitles of videos and text summarization techniques. Significant sentences in the subtitles of a video are identified using these techniques and a video summary is then composed by finding the video parts corresponding to these summary sentences.
Proceedings of the Ninth International Workshop \non Treebanks and Linguistic Theories. \nEditors: Markus Dickinson, Kaili Müürisep and Marco Passarotti. \nNEALT Proceedings Series, Vol. 9 (2010), 19-30. \n© 2010 The editors and contributors. \nPublished by \nNorthern European Association for Language \nTechnology (NEALT) \nhttp://omilia.uio.no/nealt. \nElectronically published at \nTartu University Library (Estonia) \nhttp://hdl.handle.net/10062/15891.
This paper investigates whether high-quality annotations for tasks involving semantic disambiguation can be obtained without a major investment in time or expense. We examine the use of untrained human volunteers from Amazon’s Mechanical Turk in disambiguating prepositional phrase (PP) attachment over sentences drawn from the Wall Street Journal corpus. Our goal is to compare the performance of these crowdsourced judgments to the annotations supplied by trained linguists for the Penn Treebank project in order to indicate the viability of this approach for annotation projects that involve contextual disambiguation. The results of our experiments show that invoking majority agreement between multiple human workers can yield PP attachments with fairly high precision, confirming that this crowdsourcing approach to syntactic annotation holds promise for the generation of training corpora in new domains and genres.
Proceedings of the Ninth International Workshop \non Treebanks and Linguistic Theories. \nEditors: Markus Dickinson, Kaili Müürisep and Marco Passarotti. \nNEALT Proceedings Series, Vol. 9 (2010), 67-78. \n© 2010 The editors and contributors. \nPublished by \nNorthern European Association for Language \nTechnology (NEALT) \nhttp://omilia.uio.no/nealt. \nElectronically published at \nTartu University Library (Estonia) \nhttp://hdl.handle.net/10062/15891.
Auditory displays have been used in both human-machine and computer interfaces. However, the use of non-speech audio in assistive communication for people with language disabilities, or in other applications that employ visual representations, is still under-investigated. In this paper, we introduce SoundNet, a linguistic database that associates natural environmental sounds with words and concepts. A sound labeling study was carried out to verify SoundNet associations and to investigate how well the sounds evoke concepts. A second study was conducted using the verified SoundNet data to explore the power of environmental sounds to convey concepts in sentence contexts, compared with conventional icons and animations. Our results show that sounds can effectively illustrate (especially concrete) concepts and can be applied to assistive interfaces.