Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Consistency of corpus annotation is an essential property for the many uses of annotated corpora in computational and theoretical linguistics. While some research addresses the detection of inconsistencies in part-of-speech and other positional annotation (van Halteren, 2000; Eskin, 2000; Dickinson and Meurers, 2003a), more recently work has also started to address errors in syntactic and other structural annotation (Dickinson and Meurers, 2003b, 2005; Ule and Simov, 2004; Dickinson, 2005). Spoken language differs in many respects from written language, but to the best of our knowledge the issue of detecting errors in the annotation of spoken language corpora has not yet been systematically addressed. This is significant since spoken data is increasingly relevant for linguistic and computational research—and such corpora are starting to become more readily available, as illustrated by the holdings of the Linguistic Data Consortium (http://www.ldc.upenn.edu). This paper addresses the issue, based on the variation n-gram error detection approach developed in Dickinson and Meurers (2003a). We use the German Verbmobil treebank (Hinrichs et al., 2000) as an exemplar of a spoken language corpus and discuss properties of such corpora which are relevant when adapting the variation n-gram approach for detecting errors in syntactic annotation of spoken language corpora.
Since ancient linguistics, the studies of Indo-European word order work with the conception of universal natural word order (ordo naturalis) – an order of the verb-dependent constituents in the linear organization of a clause. The description of the natural word order is usually based on occasional (and in some degree random) observations of clauses in a certain language. In Czech linguistics, the idea of the natural word order was formulated in a more precise way as the hypothesis of the systemic ordering (Sgall, Hajicova and Buraňova, 1980). According to the authors, the contextually non-bound participants and adverbials are ordered as follows (o. c., page 77):
The present paper proposes a method by which to translate outputs of a robust HPSG parser into semantic representations of Typed Dynamic Logic (TDL), a dynamic plural semantics defined in typed lambda calculus. With its higher-order representations of contexts, TDL analyzes and describes the inherently inter-sentential nature of quantification and anaphora in a strictly lexicalized and compositional manner. The present study shows that the proposed translation method successfully combines robustness and descriptive adequacy of contemporary semantics. The present implementation achieves high coverage, approximately 90%, for the real text of the Penn Treebank corpus.
Group identifications and intergroup relations among Turkish Dutch respondents What determines group identification processes among ethnic minority groups and how are these processes related to in-group and out-group evaluations? This article focuses on Turkish and Dutch identification among Turkish Dutch respondents and their feelings towards different ethnic and religious groups. The results show that Turkish identification is strong and Dutch identification rather weak, and that both group identifications are not strongly associated. Perceived socio-structural characteristics of intergroup relations (stability, legitimacy, permeability, and discrimination) affected both Turkish and Dutch identification. Group identification was positively related to (ethnic and religious) in-group evaluation, but there were few relationships with out-group evaluations. The affective ratings of Moroccans, Antilleans, Jews and non-believers were quite negative.
Many natural language processing tasks make use of a lexicon – typically the words collected from some annotated training data along with their associated properties. We demonstrate here the utility of corpora-independent lexicons derived from machine readable dictionaries. Lexical information is encoded in the form of features in a Conditional Random Field tagger providing improved performance in cases where: i) limited training data is made available ii) the data is case-less and iii) the test data genre or domain is different than that of the training data. We show substantial error reductions, especially on unknown words, for the tasks of part-of-speech tagging and shallow parsing, achieving up to 20 % error reduction on Penn TreeBank part-of-speech tagging and up to a 15.7 % error reduction for shallow parsing using the CoNLL 2000 data. Our results here point towards a simple, but effective methodology for increasing the adaptability of text processing systems by training models with annotated data in one genre augmented with general lexical information or lexical information pertinent to the target genre (or domain). 1.
In this paper we describe the current state of a new Japanese lexical resource: the Hinoki treebank. The treebank is built from dictionary definition sentences, and uses an HPSG based Japanese grammar to encode both syntactic and semantic information. It is combined with an ontology based on the definition sentences to give a detailed sense level description of the most familiar 28,000 words of Japanese.
This paper presents an approach to dependency parsing which can utilize any standard machine learning (classification) algorithm. A decision list learner was used in this work. The training data provided in the form of a treebank is converted to a format in which each instance represents information about one word pair, and the classification indicates the existence, direction, and type of the link between the words of the pair. Several distinct models are built to identify the links between word pairs at different distances. These models are applied sequentially to give the dependency parse of a sentence, favoring shorter links. An analysis of the errors, attribute selection, and comparison of different languages is presented.
The suitability of computer‐based instruction (CBI) for workers with limited education was evaluated in an Hispanic orchard workforce that reported little computer experience and 5.6 mean years of formal education. Ladder safety training was completed by employees who rated the training highly (effect size [d_gain] = 5.68), and their knowledge of ladder safety improved (d_gain = 1.45). There was a significant increase (p < 0.01) in safe work practices immediately after training (d_gain = 0.70), at 40 days post training (d_gain = 0.87) and at 60 days (d_gain = 1.40), indicating durability. As in mainstream populations, reaction or affective ratings correlated well with utility ratings, but not with behavior change. This demonstrates that an agricultural workforce with limited formal education can learn job safety from CBI and translate the knowledge to work practice changes, and those changes are durable.
Objective:This study aimed to examine the effects of haloperidol and amphetamine on human startle response modulated by emotionally-toned film clips. Method: Sixty participants, in two groups (one receiving haloperidol and the other receiving amphetamine) were tested using electromyography (EMG) to measure eye-blink muscle (orbicular oculi) while different emotions were induced by six 2-minute film clips. Results: An affective rating shows the negative and positive effects of the two drugs on emotional reactivity, neither amphetamine nor haloperidol had any impact on the modulation of the startle response. Conclusion: The methodological and theoretical aspects of the study and findings will be discussed.
We present an automatic approach to tree annotation in which basic nonterminal symbols are alternately split and merged to maximize the likelihood of a training treebank. Starting with a simple X-bar grammar, we learn a new grammar whose nonterminals are subsymbols of the original nonterminals. In contrast with previous work, we are able to split various terminals to different degrees, as appropriate to the actual complexity in the data. Our grammars automatically learn the kinds of linguistic distinctions exhibited in previous work on manual tree annotation. On the other hand, our grammars are much more compact and substantially more accurate than previous work on automatic annotation. Despite its simplicity, our best grammar achieves an F1 of 90.2% on the Penn Treebank, higher than fully lexicalized systems.
Data-driven grammatical function tag assignment has been studied for English using the Penn-II Treebank data. In this paper we address the question of whether such methods can be applied successfully to other languages and treebank resources. In addition to tag assignment accuracy and f-scores we also present results of a task-based evaluation. We use three machine-learning methods to assign Cast3LB function tags to sentences parsed with Bikel's parser trained on the Cast3LB treebank. The best performing method, SVM, achieves an f-score of 86.87% on gold-standard trees and 66.67% on parser output - a statistically significant improvement of 6.74% over the baseline. In a task-based evaluation we generate LFG functional-structures from the function-tag-enriched trees. On this task we achive an f-score of 75.67%, a statistically significant 3.4% improvement over the baseline.
The Hamburg implementation of the Weighted Constraint Dependency Grammar formalism (WCDG) includes an example grammar with comprehensive coverage for written German. This manual is the annotation guideline that was used to define the goals of the grammar and to create the Hamburg Dependency Treebank also published in the course of this project.
Recently proposed deterministic classifier-based parsers (Nivre and Scholz, 2004; Sagae and Lavie, 2005; Yamada and Mat-sumoto, 2003) offer attractive alternatives to generative statistical parsers. Deterministic parsers are fast, efficient, and simple to implement, but generally less accurate than optimal (or nearly optimal) statistical parsers. We present a statistical shift-reduce parser that bridges the gap between deterministic and probabilistic parsers. The parsing model is essentially the same as one previously used for deterministic parsing, but the parser performs a best-first search instead of a greedy search. Using the standard sections of the WSJ corpus of the Penn Treebank for training and testing, our parser has 88.1% precision and 87.8% recall (using automatically assigned part-of-speech tags). Perhaps more interestingly, the parsing model is significantly different from the generative models used by other well-known accurate parsers, allowing for a simple combination that produces precision and recall of 90.9% and 90.7%, respectively.
We study unsupervised methods for learning refinements of the nonterminals in a treebank. Following Matsuzaki et al. (2005) and Prescher (2005), we may for example split NP without supervision into NP[0] and NP[1], which behave differently. We first propose to learn a PCFG that adds such features to nonterminals in such a way that they respect patterns of linguistic feature passing: each node's nonterminal features are either identical to, or independent of, those of its parent. This linguistic constraint reduces runtime and the number of parameters to be learned. However, it did not yield improvements when training on the Penn Treebank. An orthogonal strategy was more successful: to improve the performance of the EM learner by treebank preprocessing and by annealing methods that split nonterminals selectively. Using these methods, we can maintain high parsing accuracy while dramatically reducing the model size.
Abstract Facial masculinity may be used as a cue in female mate choice, as it reflects the success of the male genotype in its developmental environment. Women may maximize reproductive success by using a conditional strategy favoring highly masculine facial features for short‐term relationships and feminized facial features in men for long‐term relationships. Three studies examine reactions to masculinized and feminized male facial composites. Properties of the original composite image affect ratings of critical attributes and the magnitude of the differences in ratings between versions undergoing identical processes of geometric manipulation (Study 1). Both men and women attribute personality, behavior, and mating strategies consistent with predictions derived from the good genes and mating trade‐off hypotheses (Study 2). Participants accurately grouped behavioral tendencies related to high mating effort/risky strategies and high parenting effort/risk adverse strategies and associated mating effort more so with masculinized faces and parenting effort more so with feminized faces (Study 3). These results indicate that male facial masculinity serves as a visual cue for inferring personality and reproductive strategy.
In this paper we describe the structure and development of the Brandeis Semantic Ontology (BSO), a large generative lexicon ontology and lexical database. The BSO has been designed to allow for more widespread access to Generative Lexicon-based lexical resources and help researchers in a variety of computational tasks. The specification of the type system used in the BSO largely follows that proposed by the SIMPLE specification (Busa et al., 2001), which was adopted by the EU-sponsored SIMPLE project (Lenci et al., 2000). 1.
CT resulted in variable functional and structural changes in dementia, and conclusions are limited by heterogeneity and study quality. Larger, more robust studies are required to correlate these findings with clinical benefits from CT.
The PropBank primarily adds semantic role labels to the syntactic constituents in the parsed trees of the Treebank. The goal is for automatic semantic role labeling to be able to use the domain of locality of a predicate in order to find its arguments. In principle, this is exactly what is wanted, but in practice the PropBank annotators often make choices that do not actually conform to the Treebank parses. As a result, the syntactic features extracted by automatic semantic role labeling systems are often inconsistent and contradictory. This paper discusses in detail the types of mismatches between the syntactic bracketing and the semantic role labeling that can be found, and our plans for reconciling them.
BACKGROUND: Differential responses in terms of gender and antisocial behaviour in emotional reactivity to affective pictures using the International Affective Picture System (IAPS) have been demonstrated in adult and adolescent samples. Moreover, a quadratic relationship between the arousal (intensity) and valence (degree of unpleasantness) has been suggested. The picture perception methodology has rarely been applied to middle school-aged children. We examined the subjective ratings of emotional reactivity in children for: i) the relationship between arousal and valence, ii) gender differences, and iii) its association with measures of antisocial behaviour. METHOD: Twenty-seven IAPS pictures were selected to cover a wide range of affective content and were individually administered to a non-referred community sample of 659 7-11-year-old children using a paper-and-pencil version. Concurrent symptoms of conduct disorder, oppositional defiance and psychopathy were collected from multiple sources (teacher-, parent- and self-report). RESULTS: A quadratic relationship between arousal and valence, similar to that previously reported in adults, was demonstrated. A gender difference was found for valence ratings, with girls rating aversive pictures more unpleasant than boys. No gender differences for arousal ratings were found. A significant difference was found between groups scoring above and below cut-off scores on measures of antisocial behaviour. Children above cut-off reported lower arousal to unpleasant pictures, but higher arousal to pleasant pictures. CONCLUSIONS: We confirmed that a paper-and-pencil version of the IAPS for evaluating emotion response to affectively valent and arousing stimuli can be used in school settings and that comparable gender differences in emotional reactivity can be found in children. The differential emotional reactivity of children above cut-off on measures of antisocial behaviour suggested these symptoms to be associated with a combination of increased reward and decreased punishment sensitivity.
The Prague Dependency Treebank 2.0 (PDT 2.0) contains a large amount of Czech texts with complex and interlinked morphological (two million words), syntactic (1.5 MW) and complex semantic annotation (0.8 MW); in addition, certain properties of sentence information structure and coreference relations are annotated at the semantic level. PDT 2.0 is based on the long-standing Praguian linguistic tradition, adapted for the current Computational Linguistics research needs. The corpus itself uses the latest annotation technology. Software tools for corpus search, annotation and language analysis are included. Extensive documentation (in English) is provided as well.
Semantic Web applications require robust and accurate annotation tools that are capable of automating the assignment of ontological classes to words in naturally occurring text (ontological annotation). Most current ontologies do not include rich lexical databases and are therefore not easily integrated with word sense disambiguation algorithms that are needed to automate ontological annotation. WordNet provides a potentially ideal solution to this problem as it offers a highly structured lexical conceptual representation that has been extensively used to develop word sense disambiguation algorithms. However, WordNet has not been designed as an ontology, and while it can be easily turned into one, the result of doing this would present users with serious practical limitations due to the great number of concepts (synonym sets) it contains. Moreover, mapping WordNet to an existing ontology may be difficult and requires substantial labor. We propose to overcome these limitations by developing an analytical platform that (1) provides a WordNet-based ontology offering a manageable and yet comprehensive set of concept classes, (2) leverages the lexical richness of WordNet to give an extensive characterization of concept class in terms of lexical instances, and (3) integrates a class recognition algorithm that automates the assignment of concept classes to words in naturally occurring text. The ensuing framework makes available an ontological annotation platform that can be effectively integrated with intelligence analysis systems to facilitate evidence marshaling and sustain the creation and validation of inference models.
This paper describes one of the ways how to overcome some of the major limitations of current fulltext search engines. It deals with synonymy of the web search engine results by clustering them into rele- vant synonym category of given word. It employs WordNet lexical database and several linguistic approaches to classify results in search engine re- sult page (SERP) in appropriate synonym category according to Word- Net synsets. Some methods to refine the classification are proposed and some initial experiments and results are described and discussed.
This paper aims at developing methods of extracting figurative expressions from a domain-specific corpus and extending the multi-lingual lexical database entries for those expressions to be included. The nominal entries which can be used figuratively are selected from the annotated corpus, using the hierarchical semantic class of Sejong Electronic Dictionary. And from the raw corpus, the partially-parsed corpus are constructed automatically. To extract candidates for figurative expression, we fixed some syntactic patterns in advance, for example [NP+sbj-marker NP+obj-marker V], and extract those patterns from partially-parsed corpus. After investigating the extracted expressions and classifying them whether figurative or literal, we concluded that some nominal entries are used more figuratively in specific syntactic patterns.
In this article, we describe the Naturalistic University of Alberta Nonlinear Correlation Explorer (NUANCE), a computer program for data exploration and analysis. NUANCE is specialized for finding nonlinear relations between any number of predictors and a dependent value to be predicted. It searches the space of possible relations between the predictors and the dependent value by using natural selection to evolve equations that maximize the correlation between their output and the dependent value. In this article, we introduce the program, describe how to use it, and provide illustrative examples. NUANCE is written in Java, which runs on most computer platforms. We have contributed NUANCE to the archival Web site of the Psychonomic Society (www.psychonomic.org/archive), from which it may be freely downloaded.
Studies on attribution in the moral domain often involve the use of specific behavior examples. To make valid comparisons across trait dimensions (such as honesty and friendliness), it is important to equate the intensities of the specific behaviors used. Pretesting specific behaviors can be a costly effort, but it is often necessary for research in social psychology. Our study provides a rich source of such pretested behaviors. Positive and negative examples of behaviors in the categories of honesty, loyalty, friendliness, charitableness, and cooperativeness were solicited from participants and then rated on the relevant trait dimension by an independent group. The result is data representing rankings, raw scores, andz-scores in an index of 500 behaviors across 10 trait categories that can be used by researchers to study moral and immoral behaviors. The full index of behaviors is available at www .psychonomic.org/archive/.
Previously, we introduced a new computational tool for nonlinear curve fitting and data set exploration: the Naturalistic University of Alberta Nonlinear Correlation Explorer (NUANCE) (Hollis & Westbury, 2006). We demonstrated that NUANCE was capable of providing useful descriptions of data for two toy problems. Since then, we have extended the functionality of NUANCE in a new release (NUANCE 3.0) and fruitfully applied the tool to real psychological problems. Here, we discuss the results of two studies carried out with the aid of NUANCE 3.0. We demonstrate that NUANCE can be a useful tool to aid research in psychology in at least two ways: It can be harnessed to simplify complex models of human behavior, and it is capable of highlighting useful knowledge that might be overlooked by more traditional analytical and factorial approaches. NUANCE 3.0 can be downloaded from the Psychonomic Society Archive of Norms, Stimuli, and Data at www.psychonomic.org/archive.
Detecting idioms in a sentence is important to sentence understanding. This paper discusses the linguistic knowledge for idiom detection. The challenges are that idioms can be ambiguous between literal and idiomatic meanings, and that they can be “transformed” when expressed in a sentence. However, there has been little research on Japanese idiom detection with its ambiguity and transformations taken into account. We propose a set of linguistic knowledge for idiom detection that is implemented in an idiom dictionary. We evaluated the linguistic knowledge by measuring the performance of an idiom detector that exploits the dictionary. As a result, more than 90% of the idioms are detected with 90% accuracy.
An HMM-based single character recovery (SCR) model is proposed in this paper to extract a large set of atomic abbreviations and their full forms from a text corpus. By an “atomic abbreviation,” it refers to an abbreviated word consisting of a single Chinese character. This task is important since Chinese abbreviations cannot be enumerated exhaustively but the abbreviation process for compound words seems to be compositional. One can often decode an abbreviated word character by character to its full form. With a large atomic abbreviation dictionary, one may be able to handle multiple character abbreviation problems more easily based on the compositional property of abbreviations.
The last decade has seen a large increase in the number of available corpus query systems. Some of these are optimized for a particular kind of linguistic annotation (e.g., time-aligned, treebank, word-oriented, etc.). In this paper, we report on our own corpus query system, called Emdros. Emdros is very generic, and can be applied to almost any kind of linguistic annotation using almost any linguistic theory. We describe Emdros and its query language, showing some of the benefits that linguists can derive from using Emdros for their corpora. We then describe the underlying database model of Emdros, and show how two corpora can be imported into the system. One of the two is a parallel corpus of Hungarian and English (the Hunglish corpus), while the other is a treebank of German (the TIGER Corpus). In order to evaluate the performance of Emdros, we then run some performance tests. It is shown that Emdros has extremely good performance on “small ” corpora (less than 1 million words), and that it scales well to corpora of many millions of words. 1.
This paper focuses on the electronic literacy practices of two Korean-American heritage language learners who manage Korean weblogs.Online users deliberately alter standard forms of written language and play with symbols, characters, and words to economize typing effort, mimic oral language, or convey qualities of their linguistic identity such as gender, age, and emotional states.However, little is known about the impact of computer-mediated nonstandard language use on heritage learners' linguistic development.Through in-depth case studies of two siblings, the study examines the linguistic and pragmatic practices of these learners online and the perceived effects of non-standard forms of computer-mediated language on their heritage language development and maintenance.The data show that electronic literacy practices provide authentic opportunities to use the language and support the development of a social network of Korean speakers, which results in greater sociopsychological attachment to the Korean language and culture.The informants report that the deviant language forms found in e-texts enable them to engage in online interactions without the pressures of having to spell the words correctly.However, they express frustrations in not being able to distinguish between correct and non-standard forms of the language, which appear to be affecting their offline language use. THE KOREAN CONTEXTThe Republic of Korea has one of the fastest-growing cybercommunities in the world.According to the Korea Network Information Center, over 63% of the entire South Korean population are Internet users, and 95% of individuals in the 6-29 age bracket report using it on a daily basis.Internet sites that enable users to create "personal spaces" to share and document their changing lives and keep connected with people they know are immensely popular among Koreans.A case in point is "Cyworld," an upgraded blog that features chatting, commentaries, pictures, music, a guest book, avatars and links to other homepages prompting users to network with their friends, family, and colleagues.As of August 2005, there are over 11 million Cyworld registered users.Participation in online forums such as Cyworld engages its members in a social process of learning through shared practices, internally constructed membership, and the formation of personal and group identities (Holmes & Jin Sook Lee Electronic literacy and heritage language maintenance Language Learning & Technology 94Myerhoff, 1999).Members are involved in a community of practice, where a group of people who come together around a joint enterprise develop common beliefs, values, and ways of doing things, which all influence the ways in which members communicate with one another (Eckert, 2000;Wenger, 1998).New forms of expression are constantly being negotiated and shared among online users, making it difficult to keep current with the changing face of electronic text.Computer-mediated communication is unique in that, despite its similarities to oral speech, it invites substantial deregulation effects on communication, which can foster the use of creative, non-standard language play (Sproull & Kiesler, 1986).Studies have documented non-standard 1 uses of language in online interactions (a) to mark certain individual characteristics such as provincial dialects, social class, gender, age, and/or personality traits, (b) to economize typing efforts, and/or (c) to mimic spoken language (Barnes, 2003;Herring, 2001;Song, 2002;Sproull & Kiesler, 1986).For example, Su ( 2004) found an emergent mock Taiwanese accent among Internet users as a form of language play to jointly construct "a young, lively, congenial, and witty presence" (p.61).Androutsopoulos ( 2000) also revealed that non-standard orthography in online fan media texts was representative of spoken language and purely graphemic modifications, which are used to serve as contextualization cues and cues of subcultural positioning.Although all natural languages inevitably change over time, drastic deviances from standard language ranging from non-standard orthography and incorrect grammar to unfamiliar lexical items and symbols have brought forth great concern about the preservation of standard orthography, grammar, and pragmatic uses of the Korean language (Choi, 2003;Kim, 2005;Park, 1989).Educators across grade levels in Korea are reporting that students display electronic textual features in their school work: they have difficulty with spelling and with the proper word spacing used to delineate word boundaries due to non-standard ways of Internet language use, which flout conventional norms of literacy practices (Ahn, 2000;Choi, 2003;Kim, 2005;Noh, 2000).For young children and Korean as foreign/second language learners who have not fully acquired literacy in the language, exposure to electronic texts may have adverse effects on their language development.However, Meskill, Mossop, and Bates (1999) state that "children in the age of electronic text are developing unique skills and strategies for inventing novel forms of understanding these texts that are quite often independent of formal instructional ('school') literacy training" (p.4), thus, highlighting the positive ways in which the development of electronic texts can benefit students' cognitive flexibility and skills.
Web searchers reformulate their queries, as they adapt to search engine behavior, learn more about a topic, or simply correct typing errors. Automatic query rewriting can help user web search, by augmenting a user’s query, or replacing the query with one likely to retrieve better results. One example of query-rewriting is spell-correction. We may also be interested in changing words to synonyms or other related terms. For Japanese, the opportunities for improving results are greater than for languages with a single character set, since documents may be written in multiple character sets, and a user may express the same meaning using different character sets. We give a description of the characteristics of Japanese search query logs and manual query reformulations carried out by Japanese web searchers. We use characteristics of Japanese query reformulations to extend previous work on automatic query rewriting in English, taking into account the Japanese writing system. We introduce several new features for building models resulting from this difference and discuss their impact on automatic query rewriting. We also examine enhancements in the form of rules which block conversion between some character sets, to address Japanese homophones. The precision/recall curves show significant improvement with the new feature set and blocking rules, and are often better than the English counterpart.
This special issue on Data Resources, Evaluation, and Dialogue Interaction is based on five thoroughly revised and extended papers from the sixth SIGdial Workshop held in Lisbon, Portugal, in September 2005. SIGdial is a special interest group on discourse and dialogue whose parent organisations are the Association for Computational Linguistics (ACL) and the International Speech Communication Association (ISCA). SIGdial workshops accommodate a broad range of topics related to discourse and dialogue. Among these topics are data resources, evaluation, and dialogue interaction. The papers selected for this special issue have in common that they all deal with aspects of these topics and each paper has its focus on at least one of them.
This paper describes a corpus of syntactic structures and associated sentences. However, it is not a traditional treebank. The syntactic structures are created first and are then associated with sentences in a human language. We therefore call it a reverse treebank (RTB).1 The RTB has been created for elicitation of sentences in low resource languages. First, a corpus of feature structures is created using a tool suite built by the authors. The second step is to add sentences in a widely spoken language like English or Spanish that express the meanings of each feature structure. We will call this language the Elicitation Language. The third step is to have a bilingual informant translate the sentences into a low resource language. Using an elicitation tool, the informant can also graphically align the words of the Elicitation Language to the words of the low resource language. The result is a high quality parallel, word aligned corpus annotated with feature structures, which we will call a parallel RTB. RTB sentences may have multiple clauses, but they are generally short in comparison to naturally occurring sentences in treebanks. The reason is that parallel RTBs provide small, but highly structured corpora for machine learning with small amounts of resources. Corpora such as these have been used for automatic learning of transfer rules for machine translation [8].
This paper describes a methodology aimed at grouping Catalan verbs according to their syntactic behavior. Our goal is to acquire a small number of basic classes with a high level of accuracy, using minimal resources. Information on syntactic class, expensive and slow to compile by hand, is useful for any NLP task requiring specific lexical information. We show that it is possible to acquire this kind of information using only a POS-tagged corpus. We perform two clustering experiments. The first one aims at classifying verbs into transitive, intransitive and verbs alternating with a se-construction. Our system achieves an average 0.84 F-score, for a task with a 0.33 baseline. The second experiment aims at further distinguishing among pure intransitives and verbs bearing a prepositional object. The baseline for the task is 0.51 and the upperbound 0.98. The system achieves an average 0.88 F-score.
In this paper, we address the issue of generating in-domain language model training data when little or no real user data are available. The two-stage approach taken begins with a data induction phase whereby linguistic constructs from out-of-domain sentences are harvested and integrated with artificially constructed in-domain phrases. After some syntactic and semantic filtering, a large corpus of synthetically assembled user utterances is induced. In the second stage, two sampling methods are explored to filter the synthetic corpus to achieve a desired probability distribution of the semantic content, both on the sentence level and on the class level. The first method utilizes user simulation technology, which obtains the probability model via an interplay between a probabilistic user model and the dialogue system. The second method synthesizes novel dialogue interactions from the raw data by modelling after a small set of dialogues produced by the developers during the course of system refinement. Evaluation is conducted on recognition performance in a restaurant information domain. We show that a partial match to usage-appropriate semantic content distribution can be achieved via user simulations. Furthermore, word error rate can be reduced when limited amounts of in-domain training data are augmented with synthetic data derived by our methods.
In this paper, we present a semiautomatic approach for annotating semantic information in biomedical texts. The information is used to construct a biomedical proposition bank called BioProp. Like PropBank in the newswire domain, BioProp contains annotations of predicate argument structures and semantic roles in a treebank schema. To construct BioProp, a semantic role labeling (SRL) system trained on PropBank is used to annotate BioProp. Incorrect tagging results are then corrected by human annotators. To suit the needs in the biomedical domain, we modify the Prop-Bank annotation guidelines and characterize semantic roles as components of biological events. The method can substantially reduce annotation efforts, and we introduce a measure of an upper bound for the saving of annotation efforts. Thus far, the method has been applied experimentally to a 4,389-sentence treebank corpus for the construction of Bio-Prop. Inter-annotator agreement measured by kappa statistic reaches.95 for combined decision of role identification and classification when all argument labels are considered. In addition, we show that, when trained on BioProp, our biomedical SRL system called BIOSMILE achieves an F-score of 87%.