Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Coordinations in noun phrases often pose the problem that elliptified parts have to be reconstructed for proper semantic interpretation. Unfortunately, the detection of coordinated heads and identification of elliptified elements notoriously lead to ambiguous reconstruction alternatives. While linguistic intuition suggests that semantic criteria might play an important, if not superior, role in disambiguating resolution alternatives, our experiments on the reannotated WSJ part of the Penn Treebank indicate that solely morpho-syntactic criteria are more predictive than solely lexico-semantic ones. We also found that the combination of both criteria does not yield any substantial improvement.
The effects of emotional content and emotional voice on speech intelligibility in younger and older adults was investigated. Twenty-eight younger adults with good health and clinically normal hearing thresholds in the speech range were tested. The stimuli used were the 200 sentences from the NU6 lists. The stimuli were presented to one group visually as text on paper, and to two groups auditorally, through two loudspeakers in a sound-attenuating booth. Means were obtained for both valence and arousal ratings for all three groups. The SNR threshold data were collected on young adults with normal hearing by Richard Wilson and colleagues using the female voice. Analyses revealed a significant positive correlation between valence and arousal for participants in the visual condition. The results indicated that the emotional arousal of listeners to a particular word can affect intelligibility, depending on the modality of presentation.
Nicotine, like several other abused drugs, is known to act on the reward system in the brain. Smoking-associated cues produce smoking urges and cravings accompanied by autonomic dysfunction to these cues in smokers. The present study was aimed at investigating whether cues related to smoking elicit the autonomic response in smokers. The subjective and physiological reactivity of 7 smokers and 12 nonsmokers in a supine position to smoking-related visual cues was assessed under indirect dim light using a self-assessment manikin and a specially designed pupillometer. The experimental procedure consisted of the elicitation and measurement of pupil size (PS) while the subjects viewed a smoking image and images from three valence-defined categories (i.e., pleasant, unpleasant, and neutral), based on normative affective ratings selected from the International Affective Picture System. Both groups produced significantly larger PS increases in response to pleasant or unpleasant images compared to neutral images. Smokers, viewing smoking-related visual cues but no other affective images, produced significantly larger PS's compared to nonsmokers. Moreover, smokers rated the smoking image with more pleasure and arousal than nonsmokers. These findings suggest that cues related to smoking induce not only a subjective emotional alteration, but also sympathetic activation, measured by the time-series PS data in smokers.
Abstract Despite recent calls for establishing a paremiological minimum for the United States, there has been little systematic empirical attention devoted to the issue of proverb familiarity in this country or to the possibility that geographic differences might preclude establishing a single paremiological minimum that would be equally representative of proverb familiarity in diverse regions around the country. This article reviews the call for establishing a paremiological minimum for the United States and the relevant efforts to date. The article then summarizes the results of a study of proverb familiarity among college students in four different regions of the United States, presenting results from both a proverb familiarity rating task and a proverb generation task. These data make it possible to determine the relative familiarity of several hundred proverbs and to assess the possibility that familiarity with these proverbs varies across regions. The success of this approach is evidence that this type of nomothetic and quantitative analysis of folkloric material can be a useful complement to performance-oriented research. For example, the results suggest some limits to the verisimilitude of people’s intuitions as to the currency and familiarity of proverbial texts, as proverbs often listed as "common" were sometimes shown to be relatively unfamiliar (at least to contemporary college students). Further, although statistical analyses of the data did reveal that there may be some difference in absolute levels of familiarity across regions, it was also clear that the relative familiarity of various proverbs nonetheless shows strong stability from one region to another. In general, then, the data presented here provide clear evidence that proverbs familiar in one region of the United States can generally be expected to be familiar in other regions as well; this finding suggests that a truly national paremiological minimum may well be achievable.
The normative speech associated with Japanese women today has been identified as a product of the Meiji government’s modernization project in the early 20th Century. In this article, we examine the speech style in question, which is found in novels, magazines, and other print media during the Meiji period (1868--1912), in conjunction with other notable expressions of the time (e.g. foreign borrowings). We also examine other cultural expressions of femininity, for example, female students’ clothing and hairstyle. The analysis reveals that female students’ speech style that is now categorized as ‘feminine’ was part of the vernacular, rather than emanating from the context of the school, as is generally asserted. It was criticized by older linguistic norm holders (e.g. educators, novelists) as being coarse, crude and unladylike, in contrast to upper-class women’s speech in the preceding Edo period. Drawing comparisons with the linguistic innovation of current young Japanese women, we suggest that young female speakers of the Meiji period can be viewed as the trendsetters of the era, not simply as passive targets of ideological conditioning.
In this paper, we investigate which information is useful for the detection of rhetorical (RST) relations between (Multi-) Sentential Discourse Units ((M-)SDUs)–text spans consisting of one or more sentences–within the same paragraph. In order to do so, we simplified the task of discourse parsing to a decision problem in which we decided whether an (M-)SDU is either rhetorically related to a preceding or a following (M-)SDU. Employing the RST Treebank (Carlson et al. 2003), we offered this choice to machine learning algorithms together with syntactic, lexical, referential, discourse and surface features. Next, the features were ranked on the basis of (1) models established by the classification algorithms and (2) feature selection metrics. Highly ranked features that predict the presence of a rhetorical relation are syntactic similarity, word overlap, word similarity, continuous punctuation and many reference features. Other features are used to introduce new topics or arguments: time references, proper nouns, definite articles and the word further. 1
This article considers three periods of the theory of predication in the Prague Linguistic Circle. The first belongs to the classical period, the second to the 1960s and the final and current one begins in the 1990s. The work of three particular authors, Mathesius, DaneÅ¡, and Sgall, is discussed. Four questions arise: 1) Was Mathesius an inspirational source for the second 1929 Thesis? 2) What were the historical and epistemological roots of the rather strong link between philosophy of language and linguistics during the first period? 3) Why is indexicality not recognized by DaneÅ¡ as the semiotic device which associates predication with those aspects of the sentence which potentially place it in relation with a situation? 4)Why is the relationship between the Praguian idea of âfunctionâ and the Fregean one, i.e. the logico-semantic concept which gives foundation to the philosophical truth-functional approach in semantics, and which explains the so-called âpredicative-argumental structureâ, not recognized as noteworthy? What, in fact, is\na function? In conclusion, this article sketches the project of a Latin Dependency Treebank developed on the basis of the PDT, which bears witness to the far-reaching implications of the Circleâs work and to an international, intercontinental prosecution of a tâche abordée.
The increasing use of ontologies in recent years has led to much of the information represented in them appear in differently expressed ways. When integration processes of ontologies are required within the same domain, problems arise in the correspondence between the terms that appear in ontologies involved in the process. To solve this problem, this article presents a method, which provides the user with a value of similarity between the terms of two ontologies to assist in the integration process. The method is based on the use of the lexical database defined by WordNet and the application of semantic similarity algorithms. It also introduces a software tool that carries out the implementation of this method.
The process of developing, implementing, and refining a registry data validation system is integral to optimal trauma registry operations. Describing registrar skill and proficiency in a manner that was once subjective can be replaced with objective assessment through the use of concrete rating guidelines and examples. The ability to standardize the evaluation of each registry abstract becomes the foundation for analyzing the overall accuracy of registry data. Key to the process is incorporating the validation rating tool as part of the data abstract. If properly implemented, the methodology described becomes a practical means for accuracy reporting, peer benchmarking, orientation and training, and performance management.
The aim of this presentation is to describe our on-going work on the building of a written Hindi treebank, referred to here as the Uppsala Hindi
A number of researchers have recently conducted experiments comparing “deep” hand-crafted wide-coverage with “shallow” treebank- and machine-learning-based parsers at the level of dependencies, using simple and automatic methods to convert tree output generated by the shallow parsers into dependencies. In this article, we revisit such experiments, this time using sophisticated automatic LFG f-structure annotation methodologies with surprising results. We compare various PCFG and history-based parsers to find a baseline parsing system that fits best into our automatic dependency structure annotation technique. This combined system of syntactic parser and dependency structure annotation is compared to two hand-crafted, deep constraint-based parsers, RASP and XLE. We evaluate using dependency-based gold standards and use the Approximate Randomization Test to test the statistical significance of the results. Our experiments show that machine-learning-based shallow grammars augmented with sophisticated automatic dependency annotation technology outperform hand-crafted, deep, wide-coverage constraint grammars. Currently our best system achieves an f-score of 82.73% against the PARC 700 Dependency Bank, a statistically significant improvement of 2.18% over the most recent results of 80.55% for the hand-crafted LFG grammar and XLE parsing system and an f-score of 80.23% against the CBS 500 Dependency Bank, a statistically significant 3.66% improvement over the 76.57% achieved by the hand-crafted RASP grammar and parsing system.
This paper presents a corpus study of parenthetical constructions in two different corpora: the Penn Discourse Treebank (PDTB, (PDTB-Group, 2008)) and the RST Discourse Treebank (Carlson et al., 2001). The motivation for the study is to gain a better understanding of the rhetorical properties of parentheticals in order to enable a natural language generation system to produce parentheticals as part of a rhetorically well-formed output. We argue that there is a correlation between syntactic and rhetorical types of parentheticals and establish two main categories: ELABORATION/EXPANSION-type NP-modifier parentheticals and NON-ELABORATION/EXPANSION-type VP- or S-modifier parentheticals. We show several strategies for extracting these from the two corpora and discuss how the seemingly contradictory results obtained can be reconciled in light of the rhetorical and syntactic properties of parentheticals as well as the decisions taken in the annotation guidelines. 1. Definition Parentheticals are constructions that typically occur embed-ded in the middle of a clause. They are not part of the main predicate-argument structure of the sentence and are marked by special punctuation (e.g. parentheses, dashes, commas) in written texts, or by special intonation in speech.
The suitability of different parsing methods for different languages is an important topic in syntactic parsing. Especially lesser-studied languages, typologically different from the languages for which methods have originally been developed, pose interesting challenges in this respect. This article presents an investigation of data-driven dependency parsing of Turkish, an agglutinative, free constituent order language that can be seen as the representative of a wider class of languages of similar type. Our investigations show that morphological structure plays an essential role in finding syntactic relations in such a language. In particular, we show that employing sublexical units called inflectional groups, rather than word forms, as the basic parsing units improves parsing accuracy. We test our claim on two different parsing methods, one based on a probabilistic model with beam search and the other based on discriminative classifiers and a deterministic parsing strategy, and show that the usefulness of sublexical units holds regardless of the parsing method. We examine the impact of morphological and lexical information in detail and show that, properly used, this kind of information can improve parsing accuracy substantially. Applying the techniques presented in this article, we achieve the highest reported accuracy for parsing the Turkish Treebank.
In this paper, we give a description of the machine translation (MT) system developed at DCU that was used for our third participation in the evaluation campaign of the International Workshop on Spoken Language Translation (IWSLT 2008). In this participation, we focus on various techniques for word and phrase alignment to improve system quality. Specifically, we try out our word packing and syntax-enhanced word alignment techniques for the Chinese–English task and for the English–Chinese task for the first time. For all translation tasks except Arabic–English, we exploit linguistically motivated bilingual phrase pairs extracted from parallel treebanks. We smooth our translation tables with out-of-domain word translations for the Arabic–English and Chinese–English tasks in order to solve the problem of the high number of out of vocabulary items. We also carried out experiments combining both in-domain and out-of-domain data to improve system performance and, finally, we deploy a majority voting procedure combining a language model based method and a translation-based method for case and punctuation restoration. We participated in all the translation \ntasks and translated both the single-best ASR hypotheses and \nthe correct recognition results. The translation results confirm that our new word and phrase alignment techniques are often helpful in improving translation quality, and the data combination method we proposed can significantly improve system performance.
This contribution article focuses on German-language collocation research and lexicographic practice from a corpus linguistic perspective. Although there is no dictionary called “Deutsches Kollokationswörterbuch” (German collocation dictionary), the collocation perspective acquires increasing popularity in linguistic research and dictionary work in the German-speaking area. On the one hand, this tendency is due to the growing number of studies dealing with German as a contrast language and works on foreign language didactics. On the other hand, powerful electronic resources such as large corpora and lexical databases, which are nowadays available for the German language, are recognised as a valuable empirical basis. Nevertheless, the application of novel corpus linguistic methods in lexicographic practice is still unsatisfactory. Therefore, this article concludes by discussing innovative aspects of corpus linguistic empirical research on the basis of collocations. These ideas are presented as an incentive for further research as well as practical application.
We present a new model of language learning which is based on the following idea: if a language learner does not know which phrase-structure trees should be assigned to initial sentences, s/he allows (implicitly) for all possible trees and lets linguistic experience decide which is the ‘best’ tree for each sentence. The best tree is obtained by maximizing ‘structural analogy ’ between a sentence and previous sentences, which is formalized by the most probable shortest combination of subtrees from all trees of previous sentences. Corpus-based experiments with this model on the Penn Treebank and the Childes database indicate that it can learn both exemplar-based and rulebased aspects of language, ranging from phrasal verbs to auxiliary fronting. By having learned the syntactic structures of sentences, we have also learned the grammar implicit in these structures, which can in turn be used to produce new sentences. We show that our model mimicks children’s language development from item-based constructions to abstract constructions, and that the model can simulate some of the errors made by children in producing complex questions.
We describe three PCFG-based models for Chinese sentence realisation from Lexical-Functional Grammar (LFG) f-structures. Both the lexicalised model and the history-based model improve on the accuracy of a simple wide-coverage PCFG model by adding lexical and contextual information to weaken inappropriate independence assumptions implicit in the PCFG models. In addition, we provide techniques for lexical smoothing and rule smoothing to increase the generation coverage. Trained on 15,663 automatically LFG f-structure annotated sentences of the Penn Chinese treebank and tested on 500 sentences randomly selected from the treebank test set, the lexicalised model achieves a BLEU score of 0.7265 at 100% coverage, while the history-based model achieves a BLEU score of 0.7245 also at 100% coverage.
In this paper, we introduce a novel method based on context awareness, semantic similarities and customized weights for different categories to improve keyword matching. The algorithm is able to weight terms by using category information and semantic relationships with WordNet as a lexical database. To demonstrate the usefulness of the approach, several tests are run comparing against other existing methods such as Salton, Glasgow and Balanced Term Weighting schemes.
We propose two general and robust methods for enriching resources annotated in the Frame Semantic paradigm with syntactic dependency graphs, which can provide useful additional information for applications such as semantic role labeling methods. One method incorporates information of a dependency parser, while the other one assumes the resource to be based on a treebank and uses dependency graphs converted from phrase structure trees. Coverage and accuracy of the methods are evaluated on the English FrameNet and German SALSA corpora. It is shown that large proportions of those resources can be accurately enriched by mapping their annotations onto dependency graphs. Failures to do so are found to be largely due to parser errors and can therefore be seen as an indicator of incorrect parses, which helps to improve parse selection. The remaining failures are analyzed and an outlook on ways of improving the results by adaptation to specific resources is given. 1.
We report on some recent parse selection experiments carried out with GG, a large-scale HPSG grammar for German. Using a manually disambiguated treebank derived from the Verbmobil corpus, we achieve over 81% exact match accuracy compared to a 21.4% random baseline, corresponding to an error reduction rate of 3.8.
The interest in dependency grammar has clearly been on the rise in the last 10– 15 years. The parsing community has seen a number of benefits in the use of and parsing with dependency-based representations of syntactic phenomena. On the one hand, dependency trees are much simpler than phrase-structure trees – they contain exactly the same number of nodes as there are tokens in the sentence. The words in the sentence are linked by binary, asymmetric relations, the so called dependencies. In addition, every link is often labelled with the syntactic or semantic relation which holds between the two words. A tree labelled with such relations encodes the predicate-argument structure of a clause in a much straightforward way than phrase-structure trees. Finally, a number of parsing algorithms have been proposed whose running time is linear with respect to the number of tokens in a sentence. The availability of syntactically annotated corpora has facilitated the use of supervised machine learning methods in the field of parsing. Data-driven parsers, which rely on corpora for the training stage can be acquired for any language for which such an annotated corpus exists. In addition, they are robust and efficient for the given task. Therefore in the present work we have used data-driven parsing with syntactic dependency representations for Bulgarian. This thesis deals with three topics, with each part building on the findings of the previous one. First, we have chosen to adopt the framework of dependency grammar and apply it to syntactic phenomena in Bulgarian. We have identified dependency structures and relations which cover a number of phenomena in the language. However, the central issue is the proper encoding of the predicateargument structure within the chosen framework. No other studies of Bulgarian have dealt with this issue from a dependency grammar point of view. Secondly, once the crucial structures and relations have been identified, we have used those priciniples in the conversion of a constituency-based treebank of Bulgarian. Again we argue for the need of the proper transfer of the predicateargument structure from a constituent tree into the dependency graph. This need reflects the nature of the underlying structures in the original treebank, where complex verb-phrases have a flat annotation. In addition, the principles we adopt lead to increase of non-projective structures compared to previous conversions of the treebank. Finally, the performance of two data-driven parsers have been tested on the acquired treebank. The aim is not only to achieve state-of-the-art results for parsing Bulgarian, but what is more important – to test the performance with respect to a higher number of non-projective structures and a verb argument structure which is more complex than previously created.
This paper describes a study of the levels at which different rhetorical relations occur in rhetorical structure trees. In a previous empirical study (Williams and Reiter, 2003) of the RST-DT (Rhetorical Structure Theory Discourse Treebank) Corpus (Carlson et al., 2003), we noticed that certain rhetorical relations tended to occur more frequently at higher levels in a rhetorical structure tree, whereas others seemed to occur more often at lower levels. The present study takes a closer look at the data, partly to test this observation, and partly to investigate related issues such as the relative complexity of satellite and nucleus for each type of relation. One practical application of this investigation would be to guide discourse planning in Natural Language Generation (NLG), so that it reflects more accurately the structures found in documents written by human authors. We present our preliminary findings and discuss their relevance for discourse planning.
Stereotypes regarding social status lead to the categorization of individuals as belonging to high or low social-status groups, based on little information, such as looks or possession of certain traits. The present study examined the relative effect of looks and musical preference on the inference of other traits relating to high and low social status. Seventy participants were asked to rate photos of eight individuals (four males and four females). Compatible and incompatible pairing of high- and low-status looks and liking for high- and low-status music were created. Findings show that more positive traits were attributed to females, high-status looking individuals and individuals with a preference for high-status music. An interaction between looks and music status was found in which liking for low-status music lowered evaluations in high-status looking individuals, but liking for high-status music did not affect evaluations of low-status looking individuals. Participants' own musical preference did not consistently affect ratings of photographed individuals.
We present a dependency-driven parser that parses both dependency structures and constituent structures. Constituency representations are automatically transformed into dependency representations with complex arc labels, which makes it possible to recover the constituent structure with both constituent labels and grammatical functions. We report a labeled attachment score close to 90% for dependency versions of the TIGER and TúBa-D/Z treebanks. Moreover, the parser is able to recover both constituent labels and grammatical functions with an F-Score over 75% for TüBa-D/Z and over 65% for TIGER.
Computer simulated avatars and humanoid robots have an increasingly prominent place in today's world. Acceptance of these synthetic characters depends on their ability to properly and recognizably convey basic emotion states to a user population. This study presents an analysis of audio-visual features that can be used to predict user evaluations of synthetic character emotion displays. These features include prosodic, spectral, and semantic properties of audio signals in addition to FACS-inspired video features. The goal of this paper is to identify the audio-visual features that explain the variance in the emotional evaluations of naive listeners through the utilization of information gain feature selection in conjunction with support vector machines. These results suggest that there exists an emotionally salient subset of the audio-visual feature space. The features that contribute most to the explanation of evaluator variance are the prior knowledge audio statistics (e.g., average valence rating), the high energy band spectral components, and the quartile pitch range. This feature subset should be correctly modeled and implemented in the design of synthetic expressive displays to convey the desired emotions.
Recently, most of the research in NLP has concentrated on the creation of applications coping with textual entailment. However, there still exist very few resourses for the evaluation of such applications. We argue that the reason for this resides not only in the novelty of the research field but also and mainly in the difficulty of defining the linguistic phenomena which are responsible for inference. As the TSNLP project has shown test suites provide an optimal diagnostic and evaluation tools for NLP applications, as contrary to text corpora they provide a deep insight in the linguistic phenomena allowing control over the data. Thus in this paper, we present a test suite specifically developed for studying inference problems shown by English adjectives. The construction of the test suite is based on the deep linguistic analysis and following classification of entailment patterns of adjectives and follows the TSNLP guidelines on linguistic databases providing a clear coverage, systematic annotation of inference tasks, large reusability and simple maintenance. With the design of this test suite we aim at creating a resource supporting the evaluation of computational systems handling natural language inference and in particular at providing a benchmark against which to evaluate and compare existing semantic analysers.
In this paper, we investigate the use of a machine-learning based approach to the specific problem of scientific term detection in patient information. Lacking lexical databases which differentiate between the scientific and popular nature of medical terms, we used local context, morphosyntactic, morphological and statistical information to design a learner which accurately detects scientific medical terms. This study is the first step towards the automatic replacement of a scientific term by its popular counterpart, which should have a beneficial effect on readability. We show an F-score of 84 % for the prediction of scientific terms in an English and Dutch EPAR corpus. Since recasting the term extraction problem as a classification problem leads to a large skewedness of the resulting data set, we rebalanced the data set through the application of some simple TF-IDF-based and Log-likelihood-based filters. We show that filtering indeed has a beneficial effect on the learner’s performance. However, the results of the filtering approach combined with the learning-based approach remain below those of the learning-based approach. 1.
The query language in TIGERSearch is limited due to its lack of universal quantification.This restriction makes it impossible to make simple queries like "Find sentences that do not include a certain word".We propose an easy way to formulate such queries.We have implemented this extension to the query language in a tool that allows querying parallel treebanks, while including their alignment constraints.Our implementation of universal quantification relies on the view of node sets rather than single node unification.Our query tool is freely available.
This paper reports on the work carried out developing MedLex+, a medical corpuslexicon workbench for Swedish. This project, which is still under active development, has been going on for some years now within the Department of Swedish language at Goteborg University. At the moment, the workbench incorporates: - an annotated collection of medical texts-including 20 million tokens and 45,000 documents, - a number of language processing software programs, including tools for collocation extraction, compound segmentation and thesaurus-based semantic annotation, and - a lexical database of medical terms-containing 5,000 medical entries. MedLex+ is a multifunctional lexical resource due to a structural design and content which can be easily queried. The medical workbench is intended to support lexicographers compiling lexicons and also lexicon users more or less initiated in the medical domain. MedLex+ can also assist researchers working on either lexical semantics or natural language processing (NLP) applications with focus on medical language. The linguistically and semantically annotated medical texts in combination with a set of smart queries turn the corpora into a rich repository of semasiological and onomasiological knowledge about medical terms and their linguistic, lexical and pragmatic properties. These properties are recorded in the lexical database with a cognitive profile. The MedLex+ workbench seems to offer a constructive help in many different lexical tasks.
Abstract. The paper addresses a problem of extraction of semantic information from Czech texts from the Web. The method described in this paper exploits existing linguistic tools created originally for a syntactically annotated corpus, Prague Dependency Treebank (PDT 2.0). We are working on development of a system which captures text of web-pages, annotates it linguistically by linguistic tools, extracts data and interprets the extracted data semantically in terms of web ontologies. The proposed extraction method is based on extraction rules – tree queries, which are adopted from the Netgraph application. Semantic interpretation of these rules provides semantics of the extracted data. We present some initial experiments in the domain of reports of traffic accidents.
Abstract. This paper describes the implementation and system details of Klex, a finite-state transducer lexicon for the Korean language, developed using XRCE’s Xerox Finite State Tool (XFST). Klex is essentially a transducer network representing the lexicon of the Korean language with the lexical string on the upper side and the inflected surface string on the lower side. Two major applications for Klex are morphological analysis and generation: given a well-formed inflected lower string, a languageindependent algorithm derives the upper lexical string from the network and vice versa. Klex was written to conform to the part-of-speech tagging standards of the Korean Treebank Project, and is currently operating as the morphological analysis engine for the project. 1
What kinds of lexical resources are helpful for extracting useful information from domain-specific documents? Although domain-specific documents contain much useful knowledge, it is not obvious how to extract such knowledge efficiently from the documents. We need to develop techniques for extracting hidden information from such domain-specific documents. These techniques do not necessarily use state-of-the-art technologies and achieve deep and accurate language understanding, but are based on huge amounts of linguistic resources, such as domain-specific lexical databases. In this paper, we introduce two techniques for extracting informative expressions from documents: the extraction of related words that are not only taxonomically related but also thematically related, and the acquisition of salient terms and phrases. With these techniques we then attempt to automatically and statistically extract domain-specific informative expressions in aviation documents as an example and evaluate the results. 1.
This paper describes a method of accurately projecting Propbank roles onto constituents in the CCGbank with near perfect accuracy and automatically annotating verbal categories with the semantic roles of their arguments. The current version of the CCGbank annotates arguments and adjuncts in a suboptimal way – it relies heavily on the Penn Treebank CLR tag, which is widely considered unreliable. By incorporating Propbank roles we are able to modify the derivation to better reflect linguistic reality. Tagging of nodes in the CCG derivation also permits us to annotate verbal categories with semantic roles corresponding to their syntactic arguments, which has strong implications for many NLP tasks.
This paper proposes an novel approach to annotate function tags for unparsed text. What distinguishes our work from other attempts in such task is that we assign function tags directly basing on lexical information other than on parsed trees. In order to demonstrate the effectiveness and versatility of our method, we investigate two statistical models for automatic annotation, one is log-linear maximum entropy model and the other is margin maximum based support vector machine model, which achieve the best F-score of 82.8 and 86.4 respectively when tested on the text from Penn Chinese Treebank. We also quantity the effect of POS tagger accuracy on system performance. Our results indicate that the function tag types could be determined via flexible and powerful feature representations from words, POS tags and word position indicators, and that, similarly to syntactic parsing, the main difficulty lies in complex constituents with long-distance dependency.
In this paper we present LXGram, a general purpose grammar for the deep linguistic processing of Portuguese that aims at delivering detailed and high precision meaning representations. LXGram is grounded on the linguistic framework of Head-Driven Phrase Structure Grammar (HPSG). HPSG is a declarative formalism resorting to unification and a type system with multiple inheritance. The semantic representations that LXGram associates with linguistic expressions use the Minimal Recursion Semantics (MRS) format, which allows for the underspecification of scope effects. LXGram is developed in the Linguistic Knowledge Builder (LKB) system, a grammar development environment that provides debugging tools and efficient algorithms for parsing and generation. The implementation of LXGram has focused on the structure of Noun Phrases, and LXGram accounts for many NP related phenomena. Its coverage continues to be increased with new phenomena, and there is active work on extending the grammar's lexicon. We have already integrated, or plan to integrate, LXGram in a few applications, namely paraphrasing, treebanking and language variant detection. Grammar coverage has been tested on newspaper text.
To determine how differences in emotion representation and/or inhibitory ability affect adolescents’ responses to emotion words, 13-yr and 16-yr olds, as well as adults, were compared on the processing of emotion-laden and neutral words. Word ratings revealed that 16-yr olds tended towards perceiving all words as more arousing than did adults, irrespective of valence. Also, they rated words more negatively than 13-yr olds. Performance on an Affective Simon task revealed a marked incongruency effect only for 13-yr olds (and then only for negative words) but not for 16-yr olds (who responded fastest) or adults. Performance on a sustained attention task confirmed the expected age-related increase in inhibitory ability and a concomitant increase in response latencies. Our conclusions are two-fold. First, there are age-related differences in lexical representation which appear more marked for 16-yr olds. Second, 16-yr olds are more reactive, irrespective of the emotional content they are processing, yet appear to control its impact as efficiently as adults.
We present an initial ontology for tactical behaviors conducted by unmanned ground vehicles (UGVs). We focus on activities, which are the denotations of verbs, notably 'move' but also 'look (for)' and several others. These take collective subjects, allowing activities to be attributed to units at various hierarchical levels. The semantics of verbs must consider the denotations of their grammatical complements; that is, we must consider entire verb frames. The thematic relations of the noun-phrase complements are critical, but prepositions also play an important role. FrameNet is an online lexical database of frames derived from text corpora. Our other major resource is Levin's classification of verbs according to how changes in their frames affect their meanings. Although natural languages have a large variety of words for aspects of tactical behaviors, there is motivation to get by with as few basic verbs as possible. A variety of meanings can often be associated with a verb by altering its frame, and we can impose co-reference constraints on combinations of frames to generate structures denoting more complex activities. A simple grammar is developed for the verbs of interest. Protege-Frames ontologies include classes that inherit from linguistically inspired classes but capture domain-specific notions.
Text-to-phoneme (TTP) mapping, also called grapheme-to-phoneme (GTP) conversion, defines the process of transforming a written text into its corresponding phonetic transcription. Text-to-phoneme mapping is a necessary step in any state-of-the-art automatic speech recognition (ASR) and text-to-speech (TTS) system, where the textual information changes dynamically (i.e., new contact entries for name dialing, or new short messages or emails to be read out by a device). There are significant differences between the implementation requirements of a text-to-phoneme mapping module embedded into the automatic speech recognition and into the text-to-speech systems: in automatic speech recognition systems the errors of the text-to-phoneme mapping module are tolerated better (leading to occasional recognition errors) than in the text-to-speech applications, where the effect is immediately and in all cases audible. Automatic speech recognition systems typically use text-to-phoneme mapping to lower the footprint (to avoid storing the lexicon), while maintaining quality. The use of text-to-phoneme mapping in the text-to-speech systems is different. In addition to the phonetic information, the text-to-speech systems also need prosodic information to be able to produce high quality speech, which cannot be predicted by text-to-phoneme mapping. Most state-of-the-art text-to-speech systems use explicit pronunciation lexicon, which is aimed at providing the widest possible coverage, in the order of 100K words, with high quality pronunciation information. Because of this reason, text-to-phoneme mapping is typically used as a fall-back strategy, when the system encounters very rare or non-native words and the quality of a ext-to-speech system is indirectly affected by the quality of the grapheme-to-phoneme conversion. Another important issue is the question of training the text-to-phoneme mapping module. The problem of grapheme-to-phoneme conversion is a static one and such a system is trained off-line. The correspondence between the written and spoken form of a language is usually unchanged in the lifetime of an application. So the complexity/speed of the model training is of secondary importance compared to e.g., the speed of convergence or model size.\n\nIn this thesis, the problem of text-to-phoneme mapping using neural networks is studied. One of the main goals of the thesis is to provide a comprehensive analysis of different neural network structures which can be implemented to convert a written text into its corresponding phonetic transcription. Another important target, of this work, is to provide new solutions that improve the performance of the existing algorithms, in terms of convergence speed and phoneme accuracy. Three main neural network classes are studied in this thesis: the multilayer perceptron (MLP) neural network, the recurrent neural network (RNN) and the bidirectional recurrent neural network (BRNN).\n\nDue to their ability of self adaptation, neural networks have been shown to be a viable solution in applications that require modeling abilities. Such an application is the text-tophoneme mapping where the correspondence between letters of a written text and their corresponding phonetic transcription must be modeled.\n\nOne of the main concerns in all practical implementations, where neural networks are used, is to develop algorithms which provide fast convergence of the synaptic weights and in the same time good mapping performances. When a neural network is trained for text-to-phoneme mapping, at every iteration, a letter-phoneme pair is presented to the network such that, the number of letters and the number of training iterations are equal. As a result, fast convergence of the neural network means smaller size of the training dictionary since fast convergence is in fact similar to less necessary training letters1. A fast convergence speed is important in applications where only a small linguistic database is available. Of course, one solution could be to use a small dictionary (with very few words) which is presented at the input of the neural network many times until the convergence of the synaptic weights is reached. In this case the time of training becomes more important. Taking into account these two sides of the convergence speed (the size of the training dictionary and the processing time during training) one can understand the importance of having algorithms that ensure fast convergence of the neural network.\n\nIt is well known that the error back-propagation algorithm which is used to train the MLP neural network, possess sometimes a quite slow convergence (a very large number of iterations required to reach the stability point). In order to increase the convergence speed two novel alternative solutions are proposed in this thesis: one using an adaptive learning rate in the training process and another which is a transform domain implementation of the multilayer perceptron neural network. The computational complexity of the two proposed training algorithms is slightly higher than the computational complexity of the error back-propagation algorithm but the number of training iterations is highly reduced. Due to this fact, although the three algorithms might have the same training time, the novel algorithms necessitate smaller training dictionary.\n\nDue to the limitations of the processing power that usually are encountered in real devices, another very important requirement for a text-to-phoneme mapping system is to have low computational and memory costs. In the case of text-to-phoneme mapping systems based in neural networks, the computational complexity is mainly linked to the mathematical complexity of the training algorithm as well as to the number of the synaptic weights of the neural network. Memory load is due to the number of synaptic weights of the neural network which must be stored.\n\nTaking into account all these limitations and implementation requirements, in this thesis, several neural network structures with different number of synaptic weights and trained with various training algorithms, are studied. The modeling capability of the neural networks is addressed, which is translated in the text-to-phoneme mapping case into the phoneme accuracy. Different neural network structures, training algorithms and network complexities are analyzed also from this point of view. As a remark here, we mention that input letter encoding plays a very important role in the phoneme accuracy of the grapheme-to-phoneme conversion system. This is why special attention has been paid to the comparative analysis of the performances (in terms of phoneme accuracy) obtained with several orthogonal and non-orthogonal encoding of the input letters.\n\nThe thesis is structured into four main parts. Chapter 1 brings the reader into the world of text-to-phoneme mapping. In Chapter 2 several different neural network structures and their corresponding training algorithms are described and two new training algorithms are introduced and analyzed. In Chapter 3 the experimental results, for the problem of monolingual text-to-phoneme mapping, obtained with the neural networks described in Chapter 2 are shown. Chapter 4 is dedicated to the problem of bilingual grapheme-to-phoneme conversion and Chapter 5 concludes the thesis.
This presentation will describe Floresta Sintactica, a syntactic Treebank for Portuguese. Some new linguistic features will be presented, as well as examples of how Floresta can be used to explore aspects of Portuguese syntax. The new interface of Floresta, Milhafre, work in progress, will also be shown.
Abstract. Many state-of-the-art statistical parsers for English can be viewed as Probabilistic Context-Free Grammars (PCFGs) acquired from treebanks consisting of phrase-structure trees enriched with a variety of contextual, derivational (e.g., markovization) and lexical information. In this paper we empirically investigate the applicability and adequacy of the unlexicalized variety of such parsing models to Modern Hebrew, a Semitic language that differs in structure and characteristics from English. We show that contrary to experience with parsing the WSJ, the markovized, head-driven unlexicalized variety does not necessarily outperform plain PCFGs for Semitic languages. We demonstrate that enriching unlexicalized PCFGs with morphologically marked agreement features percolated up the parse tree (e.g., definiteness) outperforms plain PCFGs as well as a simple head-driven variation on the MH treebank. We further show that an (unlexicalized) head-driven variety enriched with the same features achieves even better performance. We conclude
The paper analyzes various possible linguistic norms that could govern the feminine forms, which slowly appear in the Polish language, and which correspond to the masculine names of professions. Adopting a basically feminist standpoint leads one to reject those proposals, which would legislate that the masculine forms ought to be applied to men while the feminine forms ought to be applied to women. The article considers in particular the inferential roles of concepts to argue for a gender-neutral rendition of the historically masculine forms.
In sound perception the focus often lies in the cognitive aspect of the sound. We argue that the emotional aspect has to be added to get a fuller picture of sound perception. By using emotions as parameter in design of auditory alerts, one can reach a more accurate reaction to the alert. In this paper we studied the emotional connection to some attributes, common in music psychology, that are possible to describe by simple parameters. Short stimuli were created from these parameters in a factorial test design. The sounds were presented over headphones, with same signal fed to both ears, to 30 participants. The participants were asked to rate level of valence and activation, using a pictorial scale (SAM). Statistical differences was mostly found in ratings of activation, but differences were also shown in valence ratings. Results will be discussed in relation to theories of sound perception as well as music psychology.
The Berkeley FrameNet Project (BFN) is making an English lexical database called FrameNet, which describes syntactic and semantic properties of an English lexicon extracted from large electronic text corpora (Baker et al., 1998). Other projects dealing with Spanish, German and Japanese follow a similar approach and annotate large corpora. FrameSQL is a web-based application developed by the author, and it allows the user to search the BFN database in a variety of ways (Sato, 2003). FrameSQL shows a clear view of the headword’s grammar and combinatorial properties offered by the FrameNet database. FrameSQL has been developing and new functions were implemented for processing the Spanish FrameNet data (Subirats and Sato, 2004). FrameSQL is also in the process of incorporating the data of the Japanese FrameNet Project (Ohara et al., 2003) and that of the Saarbrücken Lexical Semantics Acquisition Project (Erk et al., 2003) into the database and will offer the same user-interface for searching these lexical data. This paper describes new functions of FrameSQL, showing how FrameSQL deals with the lexical data of English, Spanish, Japanese and German seamlessly. 1.
An improved chart parser based on active edges sharing the leftmost common elements is presented. After analyzing the mechanism of avoiding redundant work in the traditional chart parsing algorithm, the inefficient treatment on active edges possessing the same leftmost common elements was discovered. Then a new presentation of the active edge was proposed to share the same leftmost elements, and thus an improved chart parser was realized with great decrease of generated active edges. The experimental results on Chinese Treebank show that the improved chart parser significantly outperforms the packed chart parser in terms of both speed (about 10 times faster) and space consumption.
Considering the popularity of the Internet, an automatic interactive feedback system for Elearning websites is becoming increasingly desirable. However, computers still have problems understanding natural languages, especially the Chinese language, firstly because the Chinese language has no space to segment lexical entries (its segmentation method is more difficult than that of English) and secondly because of the lack of a complete grammar in the Chinese language, making parsing more difficult and complicated. Building an automated Chinese feedback system for special application domains could solve these problems. This paper proposes an interactive feedback mechanism in a virtual campus that can parse, understand and respond to Chinese sentences. This mechanism utilizes a specific lexical database according to the particular application. In this way, a virtual campus website can implement a special application domain that chooses the proper response in a user friendly, accurate and timely manner.
We previously observed robust activation in the hippocampal region in response to novel valenced stimuli during an fMRI recognition memory paradigm in healthy individuals across the lifespan. In this study, we compared activation during the memory task in elderly controls (EC) and individuals with mild cognitive impairment (MCI). In 23 right-handed participants (15 EC, 8 MCI; 16M/7F; mean age=71.6), neural activity was compared for novel and previously learned (familiar) items. During encoding, participants viewed 10 B/W pictures of baby and elderly faces with happy/sad expressions repeated in a block 6x, alternating with blocks of 10 circles. Participants indicated whether each face was happy or sad. After a 20-minute consolidation period, the 10 encoded faces were presented 4x each intermixed with 40 new faces (half happy, half sad). Participants identified “new” or previously learned (“old”) faces. Random effects group analysis was performed using SPM5 to compare activity in response to novel vs familiar faces. Outside the scanner, participants viewed 40 new faces (half sad, half happy), 10 faces from the encoding scan, and 80 novel faces from the recognition scan. They indicated whether each face was “new” or “old” and rated valence and arousal of each face. EC showed significant activity in right fusiform and hippocampal regions in response to novel faces compared to previously learned faces (MNI coordinates FF: 40, -58, -14, T=8.07; pcorrected=.0001; HC: 22, -12, -16; T=5.60, puncorrected=.048). Significant activations were also observed in occipital and frontal cortices. A similar pattern was found in response to baby faces but did not hold for elderly faces, suggesting that arousal is important. MCI did not show increased activity in response to novel versus familiar faces in predicted regions. Groups did not significantly differ in overall reaction time/accuracy during encoding, valence/arousal ratings or accuracy during the post-scan task, or on the Florida Affect Battery, suggesting MCIs' affective perception was intact. MCIs were slower and less accurate during recognition scans. Lack of encoding-associated activity in MCIs suggests that continued studies of the functional correlates of the emotional-memory enhancement effect in the early stage of Alzheimer's disease are warranted.