Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
Previous research found that the duration of segments decreases as children grow older. The development of suprasegmental duration, however, has not been explored. The present study investigated developmental changes in duration of the four Mandarin tones. 5-, 8-, and 12-year-old monolingual Mandarin-speaking children and young adults participated in the study. Tone durations were measured in participants’ production of monosyllabic target words elicited by picture identification tasks. The results were as follows (1) For each tone category, tone duration and variability decreased with age: 5- and 8-year-old children showed significantly longer durations than adults. Tone durations in 12-year-old children approximated adult values. (2) Despite longer durations, adultlike duration patterns across tone categories existed in all children: dipping tones were the longest, followed by rising and level tones, with falling tones being the shortest. (3) Duration differences between the rising and dipping tones became larger as children grew older. The results may be indicative of the general maturation of laryngeal control over age. Although 5- and 8-year-old children have already established lexical contrasts of tone, adultlike phonetic norms are still in the process of development. The developmental data also provide support for a hybrid account of speech production from a suprasegmental perspective.
The complexity of Korean numeral classifiers demands semantic as well as computational approaches that employ natural language processing (NLP) techniques. The classifier is a universal linguistic device, having the two functions of quantifying and classifying nouns in noun phrase constructions. Many linguistic studies have focused on the fact that numeral classifiers afford decisive clues to categorizing nouns. However, few studies have dealt with the semantic categorization of classifiers and their semantic relations to the nouns they quantify and categorize in building ontologies. In this article, we propose the semantic recategorization of the Korean numeral classifiers in the context of classifier ontology based on large corpora and KorLex Noun 1.5 (Korean wordnet; Korean Lexical Semantic Network), considering its high applicability in the NLP domain. In particular, the classifier can be effectively used to predict the semantic characteristics of nouns and to process them appropriately in NLP. The major challenge is to make such semantic classification and the attendant NLP techniques efficient. Accordingly, a Korean numeral classifier ontology (KorLexClas 1.0), including semantic hierarchies and relations to nouns, was constructed.
Visual psychophysicists, who study object, color, and light perception, have a demand for software that produces complex but, at the same time, physically accurate stimuli for their experiments. The number of computer graphic packages that simulate the physical interaction of light and surfaces is limited, and mostly they require the purchase of a license. RADIANCE (Ward, 1994), however, is freely available and popular in the visual perception community, making it a prime candidate. We have shown previously that RADIANCE’S simulation accuracy is greatly improved when color is coded by spectra, rather than by the originally envisaged RGB triplets (Ruppertsberg & Bloj, 2006). Here, we present a method for spectral rendering with RADIANCE to generate hyperspectral images that can be converted to XYZ images (CIE 1931 system) and then to machine-dependent RGB images. Generating XYZ stimuli has the added advantage of making stimulus images independent of display devices and, thereby, facilitating the process of reproducing results across different labs. Materials associated with this article may be downloaded from www.psychonomic.org.
The reflection of the category of the comic in literary speech is under investigation in the article. The basis of the taxonomic description of morphological means in the Russian language, offered by the author, consists of the following ways of creation of the comic: the usage of homonymy and contiguous phenomena to it, play up of the meanings of the same linguistic unit, repetition of the word in different grammatical forms (polyptot), divergence from the linguistic norms.
This paper presents a methodology for automatic learning of ontologies from Thai text corpora, by extraction of terms and relations. A shallow parser is used to chunk texts on which we identify taxonomic relations with the help of cues: lexico-syntactic patterns and item lists. The main advantage of the approach is that it simplify the task of concept and relation labeling since cues help for identifying the ontological concept and hinting their relation. However, these techniques pose certain problems, i.e. cue word ambiguity, item list identification, and numerous candidate terms. We also propose the methodology to solve these problems by using lexicon and co-occurrence features and weighting them with information gain. The precision, recall and F-measure of the system are 0.74, 0.78 and 0.76, respectively.
This study investigated cued odor identification performance with a set of 64 natural common odors (half of edible and half of nonedible stimuli) in three groups of participants: one group of 30 young adults (mean age 25.3 years, range 18–30, SD 3.1) and two groups of older adults—20 young-old (mean age 64.4 years, range 60–69, SD 2.8) and 21 old-old (mean age 74.6 years, range 70–79, SD 2.5). The results showed that 49 of the 64 odors were correctly identified by over 70% of the participants in all groups. The odor identification performance of the young-old adults did not differ from that of the young adults. However, the oldest group showed a significant loss of performance in the task. Women in the young-old group performed better than men, whereas no gender differences were found in the other two age groups. The data obtained in this study will be useful for further perceptual and memory studies conducted in the olfactory modality with young as well as with older participants.
Semantic intrusions are inappropriate responses frequently observed in patients with Alzheimer's disease. They belong to the same category as the words to be remembered, but their prototypic value remains largely unexplored. The prototype is the most representative word in a particular lexical category. The prototypic value is measured according to different criteria: written and oral lexical frequency, frequency of use, degree of typicality, degree of familiarity and rank of quotation. The objective of the study was to evaluate the prototypic value of intrusions produced by 17 Alzheimer's patients with mild to severe dementia, during the cued recall of the Grober & Buschke procedure (RL/RI 16 items). The prototypic value was compared to the categorial norms provided by 1) 17 control subjects and 2) the lexical database 'Lexique 3'. The results show that intrusions had a significantly higher prototypic value than targeted items. The prototypic value increased with the progression of the disease, and according to the evaluation criteria used. Thus with the criteria 'frequency of use', 'degree of typicality' and 'degree of familiarity,' the prototypic value increased exponentially with the severity of dementia. In contrast, in spite of the development of the pathology, the prototypic value decreased when assessed by the criteria of 'rank of quotation', and 'lexical frequency' (oral and written). In conclusion, the qualitative analysis of the prototypic value of intrusion errors in Alzheimers opens up new clinical and methodological considerations. (PsycINFO Database Record (c) 2016 APA, all rights reserved)
Statistical parsing of noun phrase (NP) structure has been hampered by a lack of goldstandard data. This is a significant problem for CCGbank, where binary branching NP derivations are often incorrect, a result of the automatic conversion from the Penn Treebank. We correct these errors in CCGbank using a gold-standard corpus of NP structure, resulting in a much more accurate corpus. We also implement novel NER features that generalise the lexical information needed to parse NPs and provide important semantic information. Finally, evaluating against DepBank demonstrates the effectiveness of our modified corpus and novel features, with an increase in parser performance of 1.51%. 1
We address corpus building situations, where complete annotations to the whole corpus is time consuming and unrealistic. Thus, annotation is done only on crucial part of sentences, or contains unresolved label ambiguities. We propose a parameter estimation method for Conditional Random Fields (CRFs), which enables us to use such incomplete annotations. We show promising results of our method as applied to two types of NLP tasks: a domain adaptation task of a Japanese word segmentation using partial annotations, and a part-of-speech tagging task using ambiguous tags in the Penn treebank corpus.
There are many expressive and structural differences between product names and general named entities such as person names, location names and organization names. To date, there has been little research on product named entity recognition (NER), which is crucial and valuable for information extraction in the field of market intelligence. This paper focuses on product NER (PRO NER) in Chinese text. First, we describe our efforts on data annotation, including well-defined specifications, data analysis and development of a corpus with annotated product named entities. Second, a hierarchical hidden Markov model-based approach to PRO NER is proposed and evaluated. Extensive experiments show that the proposed method outperforms the cascaded maximum entropy model and obtains promising results on the data sets of two different electronic product domains (digital and cell phone).
In my thesis I have attempted to develop an integrated translation approach materialized in the form of a Dynamic Translation Model (DTM). This endeavour can be justified to the extent that Translation Studies is perceived so far as a fragmentary discipline with implicitly and explicitly opposed and apparently irreconcilable points of view: linguistics-oriented approaches and culture-and-literature-oriented approaches. The main problem arising from this lack of common ground for further developing Translation Studies is that the disciplinary boundaries are not well-established and therefore the discipline itself cannot be developed coherently. Besides, Translation Studies is still to be constructed as an autonomous and an independent discipline that has a common core of theoretical and practical problems. This lack of coherent development of the discipline is due, I think, to an epistemological mistake: to believe that one single approach can account for (that is, describe and explain) all the translational reality. I propose to distinguish a two-phase epistemological move: 1. each translation approach works on its own research interests and acknowledges that its approach deals only with one part of the whole subject matter of Translation Studies; and 2. the results obtained by each translation approach are incorporated into a holistic integrative model like the Dynamic Translation Model I propose. In order to achieve this goal I have attempted to show the key tenets of modern translation approaches, both linguistics-oriented and culture-and-literature-oriented, by quoting the main theses of the representatives of these approaches. I have then presented the most important criticisms that have been raised in relation to these diverse translation approaches, together with my own criticisms (chapters 1 and 2). Also, I have introduced the theoretical basis for an integrated approach taking Holmes’ differentiation between theoretical (product-, process-, and function-oriented) and practical approaches as a point of departure. Likewise, I have discussed the problems of integrating Translation Studies, as well as Snell-Hornby’s integrated proposal and some key aspects of literary translation relevant for my integrative endeavour (chapter 3). Finally, I have developed my proposal for a Dynamic Translation Model (chapter 4). As to the conclusions of my thesis, I can say that my holistic DTM was able to integrate functionally aspects from both linguistics-oriented and culture-and-literature-oriented approaches: historico-cultural context (Leipzig School and postcolonial studies); norms, ideology and power (Descriptive Translation Studies; G. Toury and A. Lefevere); translation commisioner (Skopos theory); sender’s communicative purpose (linguistic and pragmatic approaches: W. Koller, J. House, H. Gerzymisch-Arbogast, etc); importance of source language text (linguistic and textlinguistic approaches; stylistic approaches; B. Spillner, B. Sandig); translator’s comprehension process (hermeneutic, deconstructive, and poststructural approaches), target language receiver in the target language historico-cultural context (Descriptive Translation Studies; postcolonial and gender studies). On the other hand, the three levels of the Dynamic Translation Model help to explain the flux of translational proceses and the variables that are activated or neutralized therein. They also incorporate concepts from other disciplines such as text linguistics, pragmatics, stylistics, and the communication theory. In my integrative endeavour I also proposed new concepts and, accordingly, coined new terms: Compulsory Translational Forces (CTF) (which include both Initiator’s Translational Instructions (ITI) and Target Language Valid Translational Norms (TL-VTN), Default Equivalence Position (DEP). In the pragmatic dimension of the model special attention is paid to what I call Text Illocutionary Indicators (TII) as well as the strengthening (upgraders) and weakening (downgraders) illocutionary mechanisms in relation to the Source Language Text (SLT) and the Target Language Text (TLT). Semantic/lexical fields play a crucial role in the establishment of equivalences between SLT and TLT in the text semantic dimension, as well as what I have called Fictionalizing Stylistic Shifts in the text stylistic dimension. As to the future developments of translation research within the framework of the Dynamic Translation Model I would say that some modificationbs may be called for so that interpretation can also be accounted for. This proposal can be used profitably in the field of translation criticism. As is the case with any other integrative approach, DTM should be widely discussed and criticized in order to validate its theoretical soundness and its application in Translation Studies. This thesis is an attempt to contribute in this research direction.
Recent neuroimaging studies have identified a set of brain regions that are metabolically active during wakeful rest and consistently deactivate in a variety the performance of demanding tasks. This ''default network'' has been functionally linked to the stream of thoughts occurring automatically in the absence of goal-directed activity and which constitutes an aspect of mental behavior specifically addressed by many meditative practices. Zen meditation, in particular, is traditionally associated with a mental state of full awareness but reduced conceptual content, to be attained via a disciplined regulation of attention and bodily posture. Using fMRI and a simplified meditative condition interspersed with a lexical decision task, we investigated the neural correlates of conceptual processing during meditation in regular Zen practitioners and matched control subjects. While behavioral performance did not differ between groups, Zen practitioners displayed a reduced duration of the neur)
The semantic annotation of texts with senses from a computational lexicon is a complex and often subjective task. As a matter of fact, the fine granularity of the WordNet sense inventory [Fellbaum, Christiane (ed.). 1998. WordNet: An Electronic Lexical Database MIT Press], a de facto standard within the research community, is one of the main causes of a low inter-tagger agreement ranging between 70% and 80% and the disappointing performance of automated fine-grained disambiguation systems (around 65% state of the art in the Senseval-3 English all-words task). In order to improve the performance of both manual and automated sense taggers, either we change the sense inventory (e.g. adopting a new dictionary or clustering WordNet senses) or we aim at resolving the disagreements between annotators by dealing with the fineness of sense distinctions. The former approach is not viable in the short term, as wide-coverage resources are not publicly available and no large-scale reliable clustering of WordNet senses has been released to date. The latter approach requires the ability to distinguish between subtle or misleading sense distinctions. In this paper, we propose the use of structural semantic interconnections—a specific kind of lexical chains—for the adjudication of disagreed sense assignments to words in context. The approach relies on the exploitation of the lexicon structure as a support to smooth possible divergencies between sense annotators and foster coherent choices. We perform a twofold experimental evaluation of the approach applied to manual annotations from the SemCor corpus, and automatic annotations from the Senseval-3 English all-words competition. Both sets of experiments and results are entirely novel: structural adjudication allows to improve the state-of-the-art performance in all-words disambiguation by 3.3 points (achieving a 68.5% Fl-score) and attains figures around 80% precision and 60% recall in the adjudication of disagreements from human annotators. (PsycINFO Database Record (c) 2016 APA, all rights reserved)
Previous work suggests that phonological neighborhood density is a key factor in shaping early lexical acquisition. Such studies have, however, have not considered how semantic neighborhoods may influence word-learning. We studied how phonological and semantic densities affect both comprehension and production of nouns from the Macarthur-Bates Communicative Development Inventory (MCDI). New measures of semantic and phonological densities, along with child-directed word frequency counts were used to predict the percentage of children who know each word at different ages (8 - 30 months) as indicated in MCDI lexical norms. Production was predicted by frequency and phonological density at all time points, replicating previous research. Semantic density predicted production only at 30 months. Comprehension norms were predicted by frequency and semantic density, and never by phonological density. Two- and three-way interactions reveal that semantic density may moderate effects in production, while sound density may moderate effects in comprehension.
Critical to vision research is the generation of visual displays with precise control over stimulus metrics. Generating stimuli often requires adapting commercial software or developing specialized software for specific research applications. In order to facilitate this process, we give here an overview that allows nonexpert users to generate and customize stimuli for vision research. We first give a review of relevant hardware and software considerations, to allow the selection of display hardware, operating system, programming language, and graphics packages most appropriate for specific research applications. We then describe the framework of a generic computer program that can be adapted for use with a broad range of experimental applications. Stimuli are generated in the context of trial events, allowing the display of text messages, the monitoring of subject responses and reaction times, and the inclusion of contingency algorithms. This approach allows direct control and management of computer-generated visual stimuli while utilizing the full capabilities of modern hardware and software systems. The flowchart and source code for the stimulus-generating program may be downloaded from www.psychonomic.org/archive.
Mediation analysis is widely used in the social sciences. Despite the popularity of mediation models, few researchers have used graphical methods, other than structural path diagrams, to represent their models. Plots of the mediated effect can help a researcher better understand the results of the analysis and convey these results to others. This article presents a method for creating and interpreting plots of the mediated effect for a variety of mediation models, including models with (1) a dichotomous independent variable, (2) a continuous independent variable, and (3) an interaction between an independent variable and the mediating variable. An empirical example is then presented to illustrate these plots. Sample code for creating plots of the mediated effect in R and SAS is also included, and may be downloaded from www.psychonomic.org/archive.
Rixte, Jean-Claude, éd. 2007 Louis Monder, Dictionnaire des dialectes dauphinois anciens et modernes. Préface Jean-Claude Bouvier. Montclimar: IFO-Drôme & ELLUG ISBN 2-95135186 -0. 899 pp. Hailed by Walther von Wartburg as "l'un des ouvrages les plus remarquables qu'il y ait dans ce genre." the posthumous Dictionnaire des chuléeles dauphinois anciens el modernes of l'abbé Louis Montier ( 1831-1903) is now available in a splendid first-time edition, thanks to Jean-Claude Rixte and the I EO-Drome. Moutier's work covers the North Occitan and Franco-Provençal varieties ofdauphinois from several départements (Drôme. Isère. Hautes-Alpes), with their "diversité étonnante"—the word is JeanClaude Bouvier's, in his preface to the new edition. Frédéric Mistral used an early version in the preparation of his own monumental Tresor dein Felibrige. The dictionary opens with an extremely useful two-page map of Dauphiné, localizing all points covered by Moutier (the accuracy of his localizationsjust one ofthe remarkable aspects of the abbé's work). Cleanly delimited are bas. moyen and haut dauphinois and their borders with provençal alpin (along the valley of the Durance, to the east) and Franco-Provençal (haut dauphinois) to the north. An annex (737-881 ) contains all Occitan words, transcribed into IEO norm but lemmatized according to Moutier's graphy, e.g., "eifoursou" esforçon 'petit effort," "mignoutiso" minliotisa 'cajolerie, caresse,' "poutrassá" poltrussur/pollrussat 'gâter, dissiper/dissipé.' Jean-Claude Rixte has reconstituted a valuable listing of Moutier's sources and added a rich secondary bibliography (883897 ). The edited work includes more than 25,000 entries (some 37,600 lexical items), which will be available online by late April 2008 to purchasers of the print version, and through other subscription opportunities as well (). Kathryn Klingebiel University of Hawai 7 Meinoa 147...
Research on insight—the phenomenon of suddenly solving an apparently intransigent problem—has been hampered because stimulus problems have been few, ad hoc, heterogeneous, and difficult to solve. Responding to the need for a larger pool of problems of a similar type and of varying level of difficulty, we report an experiment testing the validity of rebuses as insight problems. A rebus combines verbal and visual clues to a common phrase, such as PAINS (“growing pains”). Solving a rebus requires breaking implicit assumptions of normal reading, similar to the restructuring required in insight. We hypothesized that, the more implicit assumptions are involved, the more difficult the solution. The results of a two-part experiment supported the hypothesis, with participants solving more problems involving one assumption than they did problems involving two or more. Also, rebus performance correlated significantly with self-rated insight and with scores on remote associates, but not with general verbal ability. The findings suggest that rebus puzzles may be a useful source of theoretically grounded insight problems.
It is generally acknowledged nowadays that and are inseparable and there exists a direct connection between a and the used by its members. Culture occupies a prominent position on the foreign teaching agenda for the time being and the role of cultural learning has become one of the essential issues in foreign teaching theory today. There are a lot of definitions of culture suggested by different authors from various perspectives. For us as teachers of English as a foreign one of the most useful approaches in this context might be the definition provided by G. Hofstede who sees as collective programming of the mind which distinguishes the members of one group or category of people from another. M. Seidl proposes to consider a concept of that links it to a oriented analysis that in turn defines in terms of the norms and values shared by the members of a social group. The author states that language proficiency, … is a matter of familiarity with commonly held norms and values which constitute hidden meaning encoded in discourse structures. She believes that when someone learns a foreign and wants to understand another it is not enough to come to terms with another lexical or grammatical code. One has to view the world from a different perspective since speaking another means adopting another point of view.
This article examines the phenomenon of code switching in The Map of Love (1999) by the Egyptian—British writer Ahdaf Soueif. Though she chooses English as a medium for her creative expression, Soueif deploys Arabic in her narrative to represent different aspects of the linguistic and cultural norms of Egyptian society. The article's methodology is informed by Kachru's framework on contact literature and his categorization of the occurrence of literary code switching or bilingual creativity into different strategies that encompass cultural and linguistic processes. The results indicate the predominance in The Map of Love of the discourse strategies of employing lexical borrowing, culture-bound references and translational transfer. Finally, the article analyzes the functional motivation of code switching in the postcolonial context of the novel and how the use of certain creative strategies might enhance or diminish the narrative's effectiveness and readability.
ABSTRACT Using lexical items from Martin Durrell's classification of register variation as a sample, the study investigates how the current advanced monolingual learners' dictionaries of German as an additional language treat such variation and indicate to their users what they consider to be standard usage: how do they set the standard? Abbreviated usage labels as conventionally found in dictionaries for first‐language users are the primary indications, and the dictionaries seldom go beyond such labels. German Standard German is the norm. At its core are unmarked or unlabelled items, while its range extends to include less formal items from everyday use, especially spoken, which are typically labelled umg. or gespr., and more formal items, more particularly found in written usage, which are labelled geh. or geschr. Non‐standard items, if entered as headwords, may be labelled derb or vulgär, veraltet or lit. No one dictionary stands out from the others as setting the standard in terms of treating register variation, and it must be questioned whether learners of German as an additional language would not be better served by more detailed, discursive information on different contexts of use and stylistic levels.
Reviewed by: Lexicalization and language change Jesús Fernández-Domínguez Laurel J. BrintonElizabeth Closs Traugott. 2005. Lexicalization and language change. In the series Research Surveys in Linguistics. Cambridge: Cambridge University Press. Pp. xii + 207. US $34.99 (softcover). Lexicalization has been customarily defined as “a gradual historical process, involving graphemic, phonological and semantic changes and the loss of motivation” (Lipka 2005:40), and can affect a word in its phonology, morphology, semantics, or syntax. Because it can affect the makeup of virtually any item, it stands as a central phenomenon in language change, and as such it has gathered the attention of scholars for decades. The aim of [End Page 104] Brinton and Traugott’s work is to provide a wide coverage for what has been traditionally considered under lexicalization, as well as to discuss related concepts necessary for its understanding. Lexicalization and language change develops along six chapters and progressively introduces the various conceptualizations given to the processes of language modification. Chapter 1 (pp. 1–31) sets the theoretical context of the book and introduces some basic notions, and Chapter 2 (pp. 32–61) provides a background in terms of definitions and viewpoints for lexicalization. The authors discuss next the relationship between lexicalization and grammaticalization, first in a general fashion in Chapter 3 (pp. 62–88) and then in further detail in Chapter 4 (pp. 89–110). The most relevant contents of the work are exemplified in Chapter 5 (pp. 111–140), and some conclusions and research questions are offered in Chapter 6 (pp. 141–160). Among the concepts introduced in Chapter 1, the notion of lexicon bears a special significance, as there exist various senses to it which must be clarified before attempting a definition of lexicalization (see Aronoff 1989, not mentioned by the authors). To this end, Brinton and Traugott devote several pages to outline holistic vs. componential approaches to the lexicon, to the categories of the lexicon, and to the lexicon viewed as a continuum of productivity, thus laying the conceptual background required for a proper comprehension of the book. This overview is a suitable introduction to the subject also because it is contrasted with concepts like grammar, language change, or productivity, all of which have a bearing on lexicalization and are seen by the authors as a matter of gradation. Brinton and Traugott also offer a summary of the remainder of contents, and set a number of assumptions for a study of language change “from a historical, functionalist perspective” (p. 31). Chapter 2 immerses into lexicalization proper. After a brief introduction, a central section is “Ordinary processes of word formation” (pp. 33–45), a summary of the major devices of contemporary English: compounding, derivation, conversion, back-formation, initialism, etc. Here, Brinton and Traugott rightly note that lexicalization is to be distinguished from word-formation as far as only the latter has the capacity to produce new items in a regular and predictable manner, a discussion picked up later in Chapter 4. Their review proves valuable because it offers the reader the general features of present-day word-formation in a concise and satisfactory manner, even if one can hardly agree with the inclusion of loan translation, root creation, or coinage under word-formation (see Štekauer 2005:214). A subsequent logical step is the indispensable though brief explanation of institutionalization, that is, “the spread of a usage to a community and its establishment as the norm” (p. 45), usually taken as a stage following word-formation and preceding lexicalization (see Bauer 1983:45–48; Hohenhaus 2005). A number of opinions are explained and illustrated here before turning to the core of the chapter: lexicalization as fusion (pp. 47–57) and as increase in autonomy (pp. 57–60). The authors complain of the very little attention that lexicalization as fusion has received from a historical point of view, and define it as “the development of a form from a more complex to a simpler sequence” (p. 47). The present chapter truly represents a deep and up-to-date review of the typology of the phenomenon, given that it covers lexicalization as affecting phrasal and syntactic constructions (p. 48–50), word-formation (p. 50–52), phonological...
Data-driven learning based on shift reduce parsing algorithms has emerged dependency parsing and shown excellent performance to many Treebanks. In this paper, we investigate the extension of those methods while considerably improved the runtime and training time efficiency via L2-SVMs. We also present several properties and constraints to enhance the parser completeness in runtime. We further integrate root-level and bottom-level syntactic information by using sequential taggers. The experimental results show the positive effect of the root-level and bottom-level features that improve our parser from 81.17 % to 81.41 % and 81.16 % to 81.57 % labeled attachment scores with modified Yamada’s and Nivre’s method, respectively on the Chinese Treebank. In comparison to well-known parsers, such as Malt-Parser (80.74%) and MSTParser (78.08%), our methods produce not only better accuracy, but also drastically reduced testing time in 0.07 and 0.11, respectively. 1
Unlike previous emotional studies using functional neuroimaging that have focused on either locating discrete emotions in the brain or linking emotional response to an external behavior, this study investigated brain regions in order to validate a three-dimensional construct--namely pleasure, arousal, and dominance (PAD) of emotion induced by marketing communication. Emotional responses to five television commercials were measured with Advertisement Self-Assessment Manikins (AdSAM) for PAD and with functional magnetic resonance imaging (fMRI) to identify corresponding patterns of brain activation. We found significant differences in the AdSAM scores on the pleasure and arousal rating scales among the stimuli. Using the AdSAM response as a model for the fMRI image analysis, we showed bilateral activations in the inferior frontal gyri and middle temporal gyri associated with the difference on the pleasure dimension, and activations in the right superior temporal gyrus and right middle frontal gyrus associated with the difference on the arousal dimension. These findings suggest a dimensional approach of constructing emotional changes in the brain and provide a better understanding of human behavior in response to advertising stimuli.
We describe a parsing approach that makes use of the perceptron algorithm, in conjunction with dynamic programming methods, to recover full constituent-based parse trees. The formalism allows a rich set of parse-tree features, including PCFG-based features, bigram and trigram dependency features, and surface features. A severe challenge in applying such an approach to full syntactic parsing is the efficiency of the parsing algorithms involved. We show that efficient training is feasible, using a Tree Adjoining Grammar (TAG) based parsing formalism. A lower-order dependency parsing model is used to restrict the search space of the full model, thereby making it efficient. Experiments on the Penn WSJ treebank show that the model achieves state-of-the-art performance, for both constituent and dependency accuracy.
This study examined the effects of appraisal of sexual stimuli on sexual arousal in women with superficial dyspareunia (n = 50) and sexually functional women (n = 25). To elicit different appraisals of an erotic film fragment, participants received an instruction prior to viewing it, with a focus on genital pain or on sexual enjoyment. A neutral instruction served as a control condition. Assignment to instruction condition was randomized. Genital arousal (vaginal pulse amplitude) and self-report ratings of affect and genital sensations were obtained in response to the erotic stimulus. As predicted, appraisal of the erotic stimulus affected genital responding, albeit marginally significant. Follow-up tests indicated that women who received the genital pain instruction responded with marginally significant lower genital arousal levels than women who received the sexual enjoyment instruction (d = 0.67). A significant instruction effect for negative affect was found, signifying that negative affect ratings were highest after the genital pain instruction and lowest after the sexual enjoyment instruction (d = 0.80). A marginally significant group by instruction interaction effect was observed for positive affect, indicating that women with dyspareunia reported significantly less positive affect than controls after the sexual enjoyment instruction (d = 1.48). Whereas women with dyspareunia reported overall marginally significant more negative affect than controls (d = 0.48), there were no differences in genital responsiveness between groups. These results provided preliminary evidence for the modulatory effects of appraisal of sexual stimuli on subsequent genital responding and affect in women with and without sexual complaints.
Broad-coverage parsing has come to a point where distinct approaches can offer (seemingly) comparable performance: statistical parsers acquired from the Penn Treebank (PTB); data-driven dependency parsers; deep parsers trained off enriched treebanks (in linguistic frameworks like CCG, HPSG, or LFG); and hybrid deep parsers, employing hand-built grammars in, for example, HPSG, LFG, or LTAG. Evaluation against trees in the Wall Street Journal (WSJ) section of the PTB has helped advance parsing research over the course of the past decade. Despite some skepticism, the crisp and, over time, stable task of maximizing ParsEval metrics (i.e. constituent labeling precision and recall) over PTB trees has served as a dominating benchmark. However, modern treebank parsers still restrict themselves to only a subset of PTB annotation; there is reason to worry about the idiosyncrasies of this particular corpus; it remains unknown how much the ParsEval metric (or any intrinsic evaluation) can inform NLP application developers; and PTB-style analyses leave a lot to be desired in terms of linguistic information.
To date, parsers have made limited use of semantic information, but there is evidence to suggest that semantic features can enhance parse disambiguation. This paper shows that semantic classes help to obtain significant improvement in both parsing and PP attachment tasks. We devise a gold-standard sense- and parse tree-annotated dataset based on the intersection of the Penn Treebank and SemCor, and experiment with different approaches to both semantic representation and disambiguation. For the Bikel parser, we achieved a maximal error reduction rate over the baseline parser of 6.9% and 20.5%, for parsing and PP-attachment respectively, using an unsupervised WSD strategy. This demonstrates that word sense information can indeed enhance the performance of syntactic disambiguation. © 2008 Association for Computational Linguistics.
Grammar induction is one of attractive research areas of natural language processing. Since both supervised and to some extent semi-supervised grammar induction methods require large treebanks, and for many languages, such treebanks do not currently exist, we focused our attention on unsupervised approaches. Constituent Context Model (CCM) seems to be the state of the art in unsupervised grammar induction. In this paper, we show that the performance of CCM in free word order languages (FWOLs) such as Persian is inferior to that of fixed order languages such as English. We also introduce a novel approach, called parent-based constituent context model (PCCM), and show that by using some history notion of context and constituent information of each span's parent, the performance of CCM, especially in dealing with FWOLs, can be significantly improved.
Modern statistical parsers are trained on large annotated corpora (treebanks). These treebanks usually consist of sentences addressing different subdomains (e.g. sports, politics, music), which implies that the statistics gathered by current statistical parsers are mixtures of subdomains of language use. In this paper we present a method that exploits raw subdomain corpora gathered from the web to introduce subdomain sensitivity into a given parser. We employ statistical techniques for creating an ensemble of domain sensitive parsers, and explore methods for amalgamating their predictions. Our experiments show that introducing domain sensitivity by exploiting raw corpora can improve over a tough, state-of-the-art baseline. 1.
The query language in TIGERSearch is limited due to its lack of universal quantification. This restriction makes it impossible to ask simple queries like „Find sentences that do not include a certain word”. We propose an easy way to formulate such queries. We have implemented this extension to the query language in a tool that allows querying parallel treebanks including their alignment constraints. Our implementation of universal quantification relies on the view of node sets rather than single node unification. Our query tool is freely available.
This paper investigates transforms of split dependency grammars into unlexicalised context-free grammars annotated with hidden symbols. Our best unlexicalised grammar achieves an accuracy of 88% on the Penn Treebank data set, that represents a 50% reduction in error over previously published results on unlexicalised dependency parsing.
Combining Statistical and Rule-Based Approaches to Morphological Tagging of Czech Texts This article is an extract of the PhD thesis (Spoustová, 2007) and it extends the article (Spoustová et al., 2007). Several hybrid disambiguation methods are described which combine the strength of hand-written disambiguation rules and statistical taggers. Three different statistical taggers (HMM, Maximum-Entropy and Averaged Perceptron) and a large set of hand-written rules are used in a tagging experiment using Prague Dependency Treebank. The results of the hybrid system are better than any other method tried for Czech tagging so far.
Many problems in Natural Language Processing (NLP) involves an efficient search for the best derivation over (exponentially) many candidates. For example, a parser aims to find the best syntactic tree for a given sentence among all derivations under a grammar, and a machine translation (MT) decoder explores the space of all possible translations of the source-language sentence. In these cases, the concept of packed forest provides a compact representation of huge search spaces by sharing common sub-derivations, where efficient algorithms based on Dynamic Programming (DP) are possible. Building upon the hypergraph formulation of forests and well-known 1-best DP algorithms, this dissertation develops fast and exact k-best DP algorithms on forests, which are orders of magnitudes faster than previously used methods on state-of-the-art parsers. We also show empirically how the improved output of our algorithms has the potential to improve results from parse reranking systems and other applications. We then extend these algorithms to approximate search when the forests are too big for exact inference. We discuss two particular instances of this new method, forest rescoring for MT decoding, and forest reranking for parsing. In both cases, our methods perform orders of magnitudes faster than conventional approaches. In the latter, faster search also leads to better learning, where our approximate decoding makes whole-Treebank discriminative training practical and results in an accuracy better than any previously reported systems trained on the Treebank. Finally, we apply the above materials to the problem of syntax-based translation and propose a new paradigm, forest-based translation. This scheme translates a packed forest of the source sentence into a target sentence, rather than just using 1-best or k -best parses as in usual practice. By considering exponentially many alternatives, it alleviates the propagation of parsing errors into translation, yet only comes with fractional overhead in running time. We also push this direction further to extract translation rules from packed forests. The combined results of forest-based decoding and rule extraction show significant improvements in translation quality with large-scale experiments, and consistently outperform the hierarchical system Hiero, one of the best performing systems to date.
This paper provides a frame-based account of the inclusion of non-prototypical members in lexical categorization and semantic classification. It proposes that sense extensions based on metaphorical mappings can be viewed as frame-to-frame transfer, constituting an essential part of our lexical knowledge. Adopting the perspective of frame semantics (Fillmore and Atkins 1991), the study helps delimit and anchors the broad notion of 'domain,' a key concept in defining metaphors, into lexically-attested 'semantic frames' in a principled and systematic manner. Most metaphorical or extended meanings, such as the use of mo 'to touch' in wo mo bu qing ta de yong yi 'I don't understand his intension,' are often neglected by the existing databases. Literally, as a verb of touching, mo is classified as a contact verb, but the above usage of mo clearly indicates its affiliation with cognition verbs. By exploring a number of non-prototypical cognition verbs, kan 'to see', xiu 'to smell', mo, and chi 'to eat', this paper aims to propose a mechanism that allows an effective categorization of non-core members into the appropriate verb class by incorporating and redefining metaphorical extensions in a frame-based approach within the framework of frame semantics. In previous studies of lexical semantics, extended meanings are often excluded from the lexical database. For example, both Levin (1993) and Fillmore, Wooters & Baker (2001) fail to grasp the cognition-related meaning of buy 'to believe' as in I don't buy your story and that of catch 'to understand' in I just can't catch the idea of the book. However, the frequent association of the cognition sense with the two verbs still requires an explanation. In an attempt to account for the usages of various non-prototypical Mandarin cognition verbs, we propose that the usages can be viewed as metaphorical in nature and semantically, the conceptual transfer (cf. Lakoff and Johnson 1980) can be redefined as a partial transfer of frame elements from a source domain/frame to a target domain/frame. For instance, the meaning of kan in [wo/Cognizer] kan kan [yao bu yao bang ta /Issue] 'I am considering whether to help him or not,' belongs to the Cogitating Frame (Target Domain) through an extension from the meaning of kan in [wo/Perceiver] kan kan [zhe zhang zi tiao /Phenomenon] 'I take a look at this note,' in the Perception_active Frame (Source Domain). Through a convergence of frame elements, the key frame element in the Cogitating frame, the Issue, can be viewed as the corresponding frame element in the Perception frame, the Phenomenon. Syntactically, the non-prototypical cognition verb (e.g. kan) displays a similar range of lexical and grammatical collocations in its source domain (as a perception verb) as well as in the target domain (as a cognition verb). Based on lexical aspectual properties, it is shown that Mandarin cognition verbs may encode three different event types, viz. activity achievement and state (cf. Hu 2007). Examples are sorted by the three event types and operated in this paper. With clear operational mechanisms under the framework of frame semantics, metaphorically extended meanings of non-prototypical members can be readily included and well represented in lexical databases and verbal classification. The study ultimately provides a unified account of extended verb members in a grammatically-relevant and semantically-motivated way.
While the effect of domain variation on Penn-treebank- \ntrained probabilistic parsers has been investigated in previous work, we study its effect on a Penn-Treebank-trained probabilistic generator. We show that applying the generator to data from the British National Corpus \nresults in a performance drop (from a BLEU score of 0.66 on the standard WSJ test set to a BLEU score of 0.54 on our BNC test set). We develop a generator retraining method where the domain-specific training data is automatically \nproduced using state-of-the-art parser output. The retraining method recovers a substantial portion of the performance drop, resulting in a generator which achieves a BLEU score of 0.61 on our BNC test data.
Socio-economic decisions are commonly explained by rational cost versus benefit considerations, whereas person variables have not much been considered. The present study aimed at investigating the degree to which dispositional power motivation and affective states predict socio-economic decisions. The power motive was assessed both indirectly and directly using a TAT-like picture test and a power motive self-report, respectively. After 9 months, 62 students completed an affect rating and performed on a money allocation task (social values questionnaire). We hypothesized and confirmed that dispositional power should be associated with a tendency to maximize one’s profit but to care less about another party’s profit. Additionally, positive affect showed effects in the same direction. The results are discussed with respect to a motivational approach explaining socio-economic behaviour.
Traditionally, parsers are evaluated against gold standard test data. This can cause problems if there is a mismatch between the data structures and representations used by the parser and the gold standard. A particular case in point is German, for which two treebanks (TiGer and TüBa-D/Z) are available with highly different
Coordinations in noun phrases often pose the problem that elliptified parts have to be reconstructed for proper semantic interpretation. Unfortunately, the detection of coordinated heads and identification of elliptified elements notoriously lead to ambiguous reconstruction alternatives. While linguistic intuition suggests that semantic criteria might play an important, if not superior, role in disambiguating resolution alternatives, our experiments on the reannotated WSJ part of the Penn Treebank indicate that solely morpho-syntactic criteria are more predictive than solely lexico-semantic ones. We also found that the combination of both criteria does not yield any substantial improvement.
The effects of emotional content and emotional voice on speech intelligibility in younger and older adults was investigated. Twenty-eight younger adults with good health and clinically normal hearing thresholds in the speech range were tested. The stimuli used were the 200 sentences from the NU6 lists. The stimuli were presented to one group visually as text on paper, and to two groups auditorally, through two loudspeakers in a sound-attenuating booth. Means were obtained for both valence and arousal ratings for all three groups. The SNR threshold data were collected on young adults with normal hearing by Richard Wilson and colleagues using the female voice. Analyses revealed a significant positive correlation between valence and arousal for participants in the visual condition. The results indicated that the emotional arousal of listeners to a particular word can affect intelligibility, depending on the modality of presentation.
Nicotine, like several other abused drugs, is known to act on the reward system in the brain. Smoking-associated cues produce smoking urges and cravings accompanied by autonomic dysfunction to these cues in smokers. The present study was aimed at investigating whether cues related to smoking elicit the autonomic response in smokers. The subjective and physiological reactivity of 7 smokers and 12 nonsmokers in a supine position to smoking-related visual cues was assessed under indirect dim light using a self-assessment manikin and a specially designed pupillometer. The experimental procedure consisted of the elicitation and measurement of pupil size (PS) while the subjects viewed a smoking image and images from three valence-defined categories (i.e., pleasant, unpleasant, and neutral), based on normative affective ratings selected from the International Affective Picture System. Both groups produced significantly larger PS increases in response to pleasant or unpleasant images compared to neutral images. Smokers, viewing smoking-related visual cues but no other affective images, produced significantly larger PS's compared to nonsmokers. Moreover, smokers rated the smoking image with more pleasure and arousal than nonsmokers. These findings suggest that cues related to smoking induce not only a subjective emotional alteration, but also sympathetic activation, measured by the time-series PS data in smokers.
One problem facing the extraction of treebank grammars is that of ad hoc rules, rules used for constructions specific to one data set and unlikely to be used on new data (Dickinson, 2008). These rules can be erroneous, cover ungrammatical text, or reveal issues with the treebank’s annotation scheme.
In this paper, we investigate which information is useful for the detection of rhetorical (RST) relations between (Multi-) Sentential Discourse Units ((M-)SDUs)–text spans consisting of one or more sentences–within the same paragraph. In order to do so, we simplified the task of discourse parsing to a decision problem in which we decided whether an (M-)SDU is either rhetorically related to a preceding or a following (M-)SDU. Employing the RST Treebank (Carlson et al. 2003), we offered this choice to machine learning algorithms together with syntactic, lexical, referential, discourse and surface features. Next, the features were ranked on the basis of (1) models established by the classification algorithms and (2) feature selection metrics. Highly ranked features that predict the presence of a rhetorical relation are syntactic similarity, word overlap, word similarity, continuous punctuation and many reference features. Other features are used to introduce new topics or arguments: time references, proper nouns, definite articles and the word further. 1
A treebank is a text corpus in which each sentence has been annotated with its syntactic structure. Although the construction of a treebank is an expensive task, we believe that it is indispensable for the development of real applications in the field of Natural Language Processing (NLP) and also for the development of the
Expression of the serotonin transporter is affected by the genotype of the 5-HTTLPR (short and long forms) as well as the genotype of the SNP rs25531 within this region. Based on the combined genotypes for these polymorphisms, we designated each allele as a high or low expressing allele according to established expression levels-resulting in HiHi, HiLo, & LoLo genotype groups for analysis. We evaluated effects of gender and the promoter genotype on induction of negative affect by intravenous infusion of L: -tryptophan (TRP). The protocol consisted of a day-1 sham saline infusion and a day-2 active TRP infusion. Models assessed 5-HTTLPR composite genotype and gender as predictors of change in ratings of negative emotion during TRP infusion. During sham infusion there were no significant changes from baseline in mood ratings. During TRP infusion all negative affect ratings increased significantly from baseline (P's <.02). The genotype x gender interaction was a significant predictor of depression-dejection (P =.013), and trended towards predicting anger-hostility (P =.084). Males in the HiHi group had greater increases in negative affect during infusion, compared to all groups except LoLo females, who also showed increased negative affect.