Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
We outline the problem of ad hoc rules in treebanks, rules used for specific constructions in one data set and unlikely to be used again. These include ungeneralizable rules, erroneous rules, rules for ungrammatical text, and rules which are not consistent with the rest of the annotation scheme. Based on a simple notion of rule equivalence and on the idea of finding rules unlike any others, we develop two methods for detecting ad hoc rules in flat treebanks and show they are successful in detecting such rules. This is done by examining evidence across the grammar and without making any reference to context. 1
OBJECTIVE: We aimed to study the neural processing of emotion-denoting words based on a circumplex model of affect, which posits that all emotions can be described as a linear combination of two neurophysiological dimensions, valence and arousal. Based on the circumplex model, we predicted a linear relationship between neural activity and incremental changes in these two affective dimensions. METHODS: Using functional magnetic resonance imaging, we assessed in 10 subjects the correlations of BOLD (blood oxygen level dependent) signal with ratings of valence and arousal during the presentation of emotion-denoting words. RESULTS: Valence ratings correlated positively with neural activity in the left insular cortex and inversely with neural activity in the right dorsolateral prefrontal and precuneus cortices. The absolute value of valence ratings (reflecting the positive and negative extremes of valence) correlated positively with neural activity in the left dorsolateral and medial prefrontal cortex (PFC), dorsal anterior cingulate cortex, posterior cingulate cortex, and right dorsal PFC, and inversely with neural activity in the left medial temporal cortex and right amygdala. Arousal ratings and neural activity correlated positively in the left parahippocampus and dorsal anterior cingulate cortex, and inversely in the left dorsolateral PFC and dorsal cerebellum. CONCLUSION: We found evidence for two neural networks subserving the affective dimensions of valence and arousal. These findings clarify inconsistencies from prior imaging studies of affect by suggesting that two underlying neurophysiological systems, valence and arousal, may subserve the processing of affective stimuli, consistent with the circumplex model of affect.
As a result of the widespread use of English in science and scholarship, there is an increasing need of reference tools which provide accurate information to non-nativeespecially junior-researchers on the correct use of lexico-grammatical patterns of nontechnical words when writing their scientific papers in English and on the conventionalized phraseological characteristics of the genre. Our aim is to present SciE-Lex, a lexical database which provides information to help Spanish researchers to write research papers in English accurately. Whereas there are specialized monolingual and bilingual dictionaries with specific terminological information, there is a shortage of reference tools supplying information on the correct use of syntactic and collocational patterns of nontechnical words in the scientific register and on the conventionalized phraseological characteristics of the genre. Based on the analysis of a 3+ million word corpus of scientific English, in its first stage, SciE-Lex displays information on: word class, morphological variants, equivalent(s) in Spanish, patterns of occurrence, list of collocations, examples of real use, and notes to clarify usage. In a second stage we plan to include lexical bundles, compositional recurrent sequences of words, since several studies have confirmed the difficulties that learners have with them. Further research will provide SciE-Lex with information about the distribution of lexical bundles across the different sections and/or moves of the academic research article as well as their function in discourse.
In this study, 120 males (60 sexual offenders and 60 non-sexual offenders) in psychiatric treatment while in prison were evaluated using neuropsychological, psychological, and sociological/demographic measures. All sexual offenders (N = 60) would be evaluated for potential civil commitment as sexually violent predators before prison release. Non-sexual offenders (N = 60) had not been convicted of a sexual offense. Sexual offenders demonstrated significantly more overall neuropsychological impairment suggesting diffuse brain differences, with dysfunction primarily associated with temporal and frontal brain cortexes; higher Psychopathy Checklist-Revised Factor 1 (Interpersonal/Affective) ratings and Rorschach responses indicated disordered attachment, disordered self-perception, and impulsive emotionality. Sexual offenders also were more likely to be younger and unmarried. Stepwise logistic regression analysis resulted in 80.20% accuracy of prediction of sexual offenders. Potential application of this empirically derived multidimensional description to treatment of sexual offenders is discussed. Potential limitations to generalization of this information are also discussed.
In this paper we present a corpus representation format which unifies the representation of a wide range of dependency treebanks within a single model. This approach provides interoperability and reusability of annotated syntactic data which in turn extends its applicability within various research contexts. We demonstrate our approach by means of dependency treebanks of 11 languages. Further, we perform a comparative quantitative analysis of these treebanks in order to demonstrate the interoperability of our approach.
Hungarian Academy of ScienceEötvös Loránd UniversityThis paper examines the Afro-Asiatic etymologies of Chadic lexical roots discussed by Olga V. Stolbova in her Chadic Lexical Database, Issue I (2005). The analysis is arranged according to the following sections: (1) Common Chadic reconstructions, (2) Isolated Chadic roots that nevertheless have Afro-Asiatic cognates. The paper represents the third part of my longer series of papers on addenda et corrigenda to Chadic lexical roots.
We present a dependency-driven parser that parses both dependency structures and constituent structures. Constituency representations are automatically transformed into dependency representations with complex arc labels, which makes it possible to recover the constituent structure with both constituent labels and grammatical functions. We report a labeled attachment score close to 90% for dependency versions of the TIGER and TúBa-D/Z treebanks. Moreover, the parser is able to recover both constituent labels and grammatical functions with an F-Score over 75% for TüBa-D/Z and over 65% for TIGER.
We previously observed robust activation in the hippocampal region in response to novel valenced stimuli during an fMRI recognition memory paradigm in healthy individuals across the lifespan. In this study, we compared activation during the memory task in elderly controls (EC) and individuals with mild cognitive impairment (MCI). In 23 right-handed participants (15 EC, 8 MCI; 16M/7F; mean age=71.6), neural activity was compared for novel and previously learned (familiar) items. During encoding, participants viewed 10 B/W pictures of baby and elderly faces with happy/sad expressions repeated in a block 6x, alternating with blocks of 10 circles. Participants indicated whether each face was happy or sad. After a 20-minute consolidation period, the 10 encoded faces were presented 4x each intermixed with 40 new faces (half happy, half sad). Participants identified “new” or previously learned (“old”) faces. Random effects group analysis was performed using SPM5 to compare activity in response to novel vs familiar faces. Outside the scanner, participants viewed 40 new faces (half sad, half happy), 10 faces from the encoding scan, and 80 novel faces from the recognition scan. They indicated whether each face was “new” or “old” and rated valence and arousal of each face. EC showed significant activity in right fusiform and hippocampal regions in response to novel faces compared to previously learned faces (MNI coordinates FF: 40, -58, -14, T=8.07; pcorrected=.0001; HC: 22, -12, -16; T=5.60, puncorrected=.048). Significant activations were also observed in occipital and frontal cortices. A similar pattern was found in response to baby faces but did not hold for elderly faces, suggesting that arousal is important. MCI did not show increased activity in response to novel versus familiar faces in predicted regions. Groups did not significantly differ in overall reaction time/accuracy during encoding, valence/arousal ratings or accuracy during the post-scan task, or on the Florida Affect Battery, suggesting MCIs' affective perception was intact. MCIs were slower and less accurate during recognition scans. Lack of encoding-associated activity in MCIs suggests that continued studies of the functional correlates of the emotional-memory enhancement effect in the early stage of Alzheimer's disease are warranted.
This paper presented an experiment on semantic role labeling by using SVM.This experiment was based on Chinese PropBank 5.0,which consisted of 1 652 sentences.The role-labeling set of this experiment included subject,object,indirect object,time and location.It used two-phase classification method with eight features,including path,phrase type,etc.For the small scaled training set,the experiment on testing set could reach the accuracy of 89.73% and the recall of 91.26% for semantic role labeling.Results highlight the effectiveness and efficiency of proposed approach for shallow semantic parsing of Chinese.
We previously observed robust activation in the hippocampal region in response to novel valenced stimuli during an fMRI recognition memory paradigm in healthy individuals across the lifespan. In this study, we compared activation during the memory task in elderly controls (EC) and individuals with mild cognitive impairment (MCI). In 23 right-handed participants (15 EC, 8 MCI; 16M/7F; mean age=71.6), neural activity was compared for novel and previously learned (familiar) items. During encoding, participants viewed 10 B/W pictures of baby and elderly faces with happy/sad expressions repeated in a block 6x, alternating with blocks of 10 circles. Participants indicated whether each face was happy or sad. After a 20-minute consolidation period, the 10 encoded faces were presented 4x each intermixed with 40 new faces (half happy, half sad). Participants identified “new” or previously learned (“old”) faces. Random effects group analysis was performed using SPM5 to compare activity in response to novel vs familiar faces. Outside the scanner, participants viewed 40 new faces (half sad, half happy), 10 faces from the encoding scan, and 80 novel faces from the recognition scan. They indicated whether each face was “new” or “old” and rated valence and arousal of each face. EC showed significant activity in right fusiform and hippocampal regions in response to novel faces compared to previously learned faces (MNI coordinates FF: 40, -58, -14, T=8.07; pcorrected=.0001; HC: 22, -12, -16; T=5.60, puncorrected=.048). Significant activations were also observed in occipital and frontal cortices. A similar pattern was found in response to baby faces but did not hold for elderly faces, suggesting that arousal is important. MCI did not show increased activity in response to novel versus familiar faces in predicted regions. Groups did not significantly differ in overall reaction time/accuracy during encoding, valence/arousal ratings or accuracy during the post-scan task, or on the Florida Affect Battery, suggesting MCIs' affective perception was intact. MCIs were slower and less accurate during recognition scans. Lack of encoding-associated activity in MCIs suggests that continued studies of the functional correlates of the emotional-memory enhancement effect in the early stage of Alzheimer's disease are warranted.
This paper investigates transforms of split dependency grammars into unlexicalised context-free grammars annotated with hidden symbols. Our best unlexicalised grammar achieves an accuracy of 88% on the Penn Treebank data set, that represents a 50% reduction in error over previously published results on unlexicalised dependency parsing.
An algorithm based on information retrieval that applies the lexical database WordNet together with a linear discriminant function is proposed. It calculates the degree of similarity between words and their relative importance to support the development of distributed applications based on web services. The algorithm uses the semantic information contained in the Web Service Description Language specifications and ranks web services based on their similarity to the one the developer is searching for. It is applied to a set of 48 real web services in five categories, then compared them to four other algorithms based on information retrieval, showing an averaged improvement over all data between 0.6% and 1.9% in precision and 0.7% and 3.1% in recall for the top 15 ranked web services. The objective was to reduce the burden and time spent searching web services during the development of distributed applications, and it can be used as an alternative to current web service discovery systems such as brokers in the Universal Description, Discovery, and Integration (UDDI) platform.
Problematic types of prepositions, conjuctions, and particles and possible ways of their treatment in the lexical database.
German genitive attributes are usually tagged as such in treebanks. However, it is well known that this information is not sufficient for determining the type of relation between head nouns and attributes, as genitive attributes can express many different semantic relations. Various linguistic classifications have been worked out, but to my knowledge, nobody has so far proposed to apply this linguistic knowledge to a corpus. The challenge here is to come up with a classification that is both easy to verify and sufficiently fine-grained. Using earlier linguistic approaches as guidelines, I propose in this paper a detailed annotation scheme for German genitive attributes based on readily identifiable noun features. First insights from its application to the Smultron Treebank show that it is easy to distinguish between the proposed classes and that my classification of genitive attributes can be related to a more general semantic annotation level.
Social relations between humans critically depend on our affective experiences of others. Oxytocin enhances prosocial behavior, but its effect on humans' affective experience of others is not known. We tested whether oxytocin influences affective ratings, and underlying brain activity, of faces that have been aversively conditioned. Using a standard conditioning procedure, we induced differential negative affective ratings in faces exposed to an aversive conditioning compared with nonconditioning manipulation. This differential negative evaluative effect was abolished by treatment with oxytocin, an effect associated with an attenuation of activity in anterior medial temporal and anterior cingulate cortices. In amygdala and fusiform gyrus, this modulation was stronger for faces with direct gaze, relative to averted gaze, consistent with a relative specificity for socially relevant cues. The data suggest that oxytocin modulates the expression of evaluative conditioning for socially relevant faces via influences on amygdala and fusiform gyrus, an effect that may explain its prosocial effects.
Treebanks have become crucial for the development of data-driven approaches to natural language processing, human language technologies, grammar extraction, and linguistic research in general. Manifold projects aim at compiling representative treebanks for specific languages. Other projects focus on the development of tools for exploration of annotated treebanks, or explore annotation beyond syntactic structure and beyond single languages. The Seventh International Workshop on Treebanks and Linguistic Theories (TLT7) provides a forum for researchers in the field of Computational Linguistics who are experts in the design, creation and exploitation of treebanks and their relation to linguistic theories. A selection of 16 workshop papers is published in these proceedings. Together, they cover a wide range of topics, including the building, querying, exploring, exploiting and evaluating of treebanks.
While the effect of domain variation on Penn-treebank- \ntrained probabilistic parsers has been investigated in previous work, we study its effect on a Penn-Treebank-trained probabilistic generator. We show that applying the generator to data from the British National Corpus \nresults in a performance drop (from a BLEU score of 0.66 on the standard WSJ test set to a BLEU score of 0.54 on our BNC test set). We develop a generator retraining method where the domain-specific training data is automatically \nproduced using state-of-the-art parser output. The retraining method recovers a substantial portion of the performance drop, resulting in a generator which achieves a BLEU score of 0.61 on our BNC test data.
This paper describes the use of two machine learning techniques, naive Bayes and decision trees, to address the task of assigning function tags to nodes in a syntactic parse tree. Function tags are extra functional information, such as logical subject or predicate, that can be added to certain nodes in syntactic parse trees. We model the function tags assignment problem as a classification problem. Each function tag is regarded as a class and the task is to find what class/tag a given node in a parse tree belongs to from a set of predefined classes/tags. The paper offers the first systematic comparison of the two techniques, naive Bayes and decision trees, for the task of function tags assignment. The comparison is based on a standardized data set, the Penn Treebank, a collection of sentences annotated with syntactic information including function tags. We found out that decision trees generally outperform naive Bayes for the task of function tagging. Furthermore, this is the first large scale evaluation of decision trees based solutions to the task of functional tagging.
Two Languages - One Annotation Scenario? Experience from the Prague Dependency Treebank This paper compares the two FGD-based annotation scenarios for Czech and for English, with the Czech as the basis. We discuss the secondary predication expressed by infinitive and its functions in Czech and English, respectively. We give a few examples of English constructions that do not have direct counterparts in Czech (e.g., tough movement and causative constructions with make, get, and have ), as well as some phenomena central in English but much less employed in Czech (object raising or control in adjectives as nominal predicates), and, last, structures more or less parallel both in their function and distribution, whose respective annotation differs due to significant differences in the respective linguistic traditions (verbs of perception).
The paper describes an approach to automati-cally annotate a Hindi Treebank using Pan-inian dependency framework. The annotator is a rule based system and the rules use certain syntactic cues available in a sentence. This automated annotation scheme aims at facilitat-ing manual annotation by reducing time and effort of manual annotators. Also, the aim of automatic annotation, among other things, is to increase the efficiency of a broad coverage constraint based Hindi parser. We also evalu-ate this tool and show its accuracy and cover-age. 1
In this paper, we present the Web-based resource sharing system KnownStyleNoLife, which allows users explicitly annotate fashion-related images. KnownStyleNoLife harnesses human power and collects image metadata such as object locations, object labels, image rating and semantic relationships among images. Metadata of that type are invaluable as it is difficult for ordinary computer software to extract equivalent metadata from the Web. Acquiring image metadata such as the type and location of objects in images requires a computer to use machine learning techniques, needing training with sample images and Web search techniques, with the outcome that only limited image metadata are extracted. Furthermore, the semantic image relationship metadata is based upon peoplepsilas image perception, hence user feedback is required to acquire it. KnownStyleNoLife provides a social value to users in return for their explicit annotations. It is a novel approach to harness human power to acquire relevant image metadata.
To determine how differences in emotion representation and/or inhibitory ability affect adolescents’ responses to emotion words, 13-yr and 16-yr olds, as well as adults, were compared on the processing of emotion-laden and neutral words. Word ratings revealed that 16-yr olds tended towards perceiving all words as more arousing than did adults, irrespective of valence. Also, they rated words more negatively than 13-yr olds. Performance on an Affective Simon task revealed a marked incongruency effect only for 13-yr olds (and then only for negative words) but not for 16-yr olds (who responded fastest) or adults. Performance on a sustained attention task confirmed the expected age-related increase in inhibitory ability and a concomitant increase in response latencies. Our conclusions are two-fold. First, there are age-related differences in lexical representation which appear more marked for 16-yr olds. Second, 16-yr olds are more reactive, irrespective of the emotional content they are processing, yet appear to control its impact as efficiently as adults.
In this paper we present LXGram, a general purpose grammar for the deep linguistic processing of Portuguese that aims at delivering detailed and high precision meaning representations. LXGram is grounded on the linguistic framework of Head-Driven Phrase Structure Grammar (HPSG). HPSG is a declarative formalism resorting to unification and a type system with multiple inheritance. The semantic representations that LXGram associates with linguistic expressions use the Minimal Recursion Semantics (MRS) format, which allows for the underspecification of scope effects. LXGram is developed in the Linguistic Knowledge Builder (LKB) system, a grammar development environment that provides debugging tools and efficient algorithms for parsing and generation. The implementation of LXGram has focused on the structure of Noun Phrases, and LXGram accounts for many NP related phenomena. Its coverage continues to be increased with new phenomena, and there is active work on extending the grammar's lexicon. We have already integrated, or plan to integrate, LXGram in a few applications, namely paraphrasing, treebanking and language variant detection. Grammar coverage has been tested on newspaper text.
In this paper, we will give an overview of the reconstruction process of the Swedish treebank Talbanken, created in the first half of the 70’s. Talbanken contains both written and spoken material, both encoded in the MAMBA-format. The goal has been to construct two new versions of the original data, one based on phrase structure and one on dependency structure. The outcome of the reconstruction, i.e. different versions of Talbanken, is available for non-commercial research and educational purposes. 1
Stereotypes regarding social status lead to the categorization of individuals as belonging to high or low social-status groups, based on little information, such as looks or possession of certain traits. The present study examined the relative effect of looks and musical preference on the inference of other traits relating to high and low social status. Seventy participants were asked to rate photos of eight individuals (four males and four females). Compatible and incompatible pairing of high- and low-status looks and liking for high- and low-status music were created. Findings show that more positive traits were attributed to females, high-status looking individuals and individuals with a preference for high-status music. An interaction between looks and music status was found in which liking for low-status music lowered evaluations in high-status looking individuals, but liking for high-status music did not affect evaluations of low-status looking individuals. Participants' own musical preference did not consistently affect ratings of photographed individuals.
This article examines the role of subjective familiarity in the implicit and explicit learning of artificial grammars. Experiment 1 found that objective measures of similarity (including fragment frequency and repetition structure) predicted ratings of familiarity, that familiarity ratings predicted grammaticality judgments, and that the extremity of familiarity ratings predicted confidence. Familiarity was further shown to predict judgments in the absence of confidence, hence contributing to above-chance guessing. Experiment 2 found that confidence developed as participants refined their knowledge of the distribution of familiarity and that differences in familiarity could be exploited prior to confidence developing. Experiment 3 found that familiarity was consciously exploited to make grammaticality judgments including those made without confidence and that familiarity could in some instances influence participants' grammaticality judgments apparently without their awareness. All 3 experiments found that knowledge distinct from familiarity was derived only under deliberate learning conditions. The results provide decisive evidence that familiarity is the essential source of knowledge in artificial grammar learning while also supporting a dual-process model of implicit and explicit learning.
Studies examining factors that influence when words are learned typically investigate one lexical category or a small set of words. We provide the first evaluation of the relation between input frequency and age of acquisition for a large sample of words. The MacArthur-Bates Communicative Development Inventory provides norming data on age of acquisition for 562 individual words collected from the parents of children aged 0; 8 to 2; 6. The CHILDES database provides estimates of frequency with which parents use these words with their children (age: 0; 7-7; 5; mean age: 36 months). For production, across all words higher parental frequency is associated with later acquisition. Within lexical categories, however, higher frequency is related to earlier acquisition. For comprehension, parental frequency correlates significantly with the age of acquisition only for common nouns. Frequency effects change with development. Thus, frequency impacts vocabulary acquisition in a complex interaction with category, modality and developmental stage.
In my thesis I have attempted to develop an integrated translation approach materialized in the form of a Dynamic Translation Model (DTM). This endeavour can be justified to the extent that Translation Studies is perceived so far as a fragmentary discipline with implicitly and explicitly opposed and apparently irreconcilable points of view: linguistics-oriented approaches and culture-and-literature-oriented approaches. The main problem arising from this lack of common ground for further developing Translation Studies is that the disciplinary boundaries are not well-established and therefore the discipline itself cannot be developed coherently. Besides, Translation Studies is still to be constructed as an autonomous and an independent discipline that has a common core of theoretical and practical problems. This lack of coherent development of the discipline is due, I think, to an epistemological mistake: to believe that one single approach can account for (that is, describe and explain) all the translational reality. I propose to distinguish a two-phase epistemological move: 1. each translation approach works on its own research interests and acknowledges that its approach deals only with one part of the whole subject matter of Translation Studies; and 2. the results obtained by each translation approach are incorporated into a holistic integrative model like the Dynamic Translation Model I propose. In order to achieve this goal I have attempted to show the key tenets of modern translation approaches, both linguistics-oriented and culture-and-literature-oriented, by quoting the main theses of the representatives of these approaches. I have then presented the most important criticisms that have been raised in relation to these diverse translation approaches, together with my own criticisms (chapters 1 and 2). Also, I have introduced the theoretical basis for an integrated approach taking Holmes’ differentiation between theoretical (product-, process-, and function-oriented) and practical approaches as a point of departure. Likewise, I have discussed the problems of integrating Translation Studies, as well as Snell-Hornby’s integrated proposal and some key aspects of literary translation relevant for my integrative endeavour (chapter 3). Finally, I have developed my proposal for a Dynamic Translation Model (chapter 4). As to the conclusions of my thesis, I can say that my holistic DTM was able to integrate functionally aspects from both linguistics-oriented and culture-and-literature-oriented approaches: historico-cultural context (Leipzig School and postcolonial studies); norms, ideology and power (Descriptive Translation Studies; G. Toury and A. Lefevere); translation commisioner (Skopos theory); sender’s communicative purpose (linguistic and pragmatic approaches: W. Koller, J. House, H. Gerzymisch-Arbogast, etc); importance of source language text (linguistic and textlinguistic approaches; stylistic approaches; B. Spillner, B. Sandig); translator’s comprehension process (hermeneutic, deconstructive, and poststructural approaches), target language receiver in the target language historico-cultural context (Descriptive Translation Studies; postcolonial and gender studies). On the other hand, the three levels of the Dynamic Translation Model help to explain the flux of translational proceses and the variables that are activated or neutralized therein. They also incorporate concepts from other disciplines such as text linguistics, pragmatics, stylistics, and the communication theory. In my integrative endeavour I also proposed new concepts and, accordingly, coined new terms: Compulsory Translational Forces (CTF) (which include both Initiator’s Translational Instructions (ITI) and Target Language Valid Translational Norms (TL-VTN), Default Equivalence Position (DEP). In the pragmatic dimension of the model special attention is paid to what I call Text Illocutionary Indicators (TII) as well as the strengthening (upgraders) and weakening (downgraders) illocutionary mechanisms in relation to the Source Language Text (SLT) and the Target Language Text (TLT). Semantic/lexical fields play a crucial role in the establishment of equivalences between SLT and TLT in the text semantic dimension, as well as what I have called Fictionalizing Stylistic Shifts in the text stylistic dimension. As to the future developments of translation research within the framework of the Dynamic Translation Model I would say that some modificationbs may be called for so that interpretation can also be accounted for. This proposal can be used profitably in the field of translation criticism. As is the case with any other integrative approach, DTM should be widely discussed and criticized in order to validate its theoretical soundness and its application in Translation Studies. This thesis is an attempt to contribute in this research direction.
Previous work suggests that phonological neighborhood density is a key factor in shaping early lexical acquisition. Such studies have, however, have not considered how semantic neighborhoods may influence word-learning. We studied how phonological and semantic densities affect both comprehension and production of nouns from the Macarthur-Bates Communicative Development Inventory (MCDI). New measures of semantic and phonological densities, along with child-directed word frequency counts were used to predict the percentage of children who know each word at different ages (8 - 30 months) as indicated in MCDI lexical norms. Production was predicted by frequency and phonological density at all time points, replicating previous research. Semantic density predicted production only at 30 months. Comprehension norms were predicted by frequency and semantic density, and never by phonological density. Two- and three-way interactions reveal that semantic density may moderate effects in production, while sound density may moderate effects in comprehension.
Despite the importance of spoken vocabulary use to improve spoken skills, little has been conducted to describe spoken features of Korean learners. The purpose of this study is to investigate spoken vocabulary use of Korean learners and find out how far they deviate from native speaker norms. For this purpose, 40 Korean college students` spoken interaction data are transcribed and analysed. Using Wordsmith tool with 17,436 words, frequent single words category (modal items, delexical verbs, interactive words, and discourse markers) and frequent multi-words clusters (discourse markers, vagueness & approximation, politeness & face, and hedgeing) are compared with the spoken British National Corpus (BNC). Frequency analysis revealed that among 56 lexical items investigated, 34 items were underused and 17 items are not represented in KLC. It is evident that Korean learners used limited variety and range in those spoken vocabulary use. Based on this result, some suggestions are made to improve Korean learners` spoken vocabulary teaching and learning.
Critical to vision research is the generation of visual displays with precise control over stimulus metrics. Generating stimuli often requires adapting commercial software or developing specialized software for specific research applications. In order to facilitate this process, we give here an overview that allows nonexpert users to generate and customize stimuli for vision research. We first give a review of relevant hardware and software considerations, to allow the selection of display hardware, operating system, programming language, and graphics packages most appropriate for specific research applications. We then describe the framework of a generic computer program that can be adapted for use with a broad range of experimental applications. Stimuli are generated in the context of trial events, allowing the display of text messages, the monitoring of subject responses and reaction times, and the inclusion of contingency algorithms. This approach allows direct control and management of computer-generated visual stimuli while utilizing the full capabilities of modern hardware and software systems. The flowchart and source code for the stimulus-generating program may be downloaded from www.psychonomic.org/archive.
Correct identification of word meaning is a long-standing problem for lexicography, language teaching, linguistic theory, and computer processing of text. Traditional approaches typically proceed word by word, relying on evidence from introspection – and have failed. A new theory of meaning is needed. In this prototype-based approach, called the Theory of Norms and Exploitations (TNE), the first step is identifying the phraseological patterns with which each word is associated. Meanings are then associated with patterns, rather than with isolated words. Words are highly ambiguous, but patterns are mostly unambiguous. \n Patterns cannot be identified by valency alone, but require statistical analysis and semantic typing of collocates. For example, (1) blowing up a bridge and (2) blowing up a balloon activate different meanings of blow up. But how many other contexts have the same effect on the meaning of the phrasal verb? Relevant members of the lexical set for (1) include building, factory, house, hotel, etc. Such lexical sets provide a basis for machine learning and text processing. \n Authentic uses of words are classified either as normal components of a pattern or as exploitations of norms. For example, “blowing up a condom” is not normal, but exploits (2). Creative metaphors are also exploitations.
Graph-based and transition-based approaches to dependency parsing adopt very different views of the problem, each view having its own strengths and limitations. We study both approaches under the framework of beam-search. By developing a graph-based and a transition-based dependency parser, we show that a beam-search decoder is a competitive choice for both methods. More importantly, we propose a beam-search-based parser that combines both graph-based and transition-based parsing into a single system for training and decoding, showing that it outperforms both the pure graph-based and the pure transition-based parsers. Testing on the English and Chinese Penn Treebank data, the combined system gave state-of-the-art accuracies of 92.1% and 86.2%, respectively.
Speech monitoring encompasses detection and self-repair of errors.This paper first reviews types of errors and self-repairs,then focuses on three theoretical accounts of how the monitoring mechanism works to detect and correct errors.Product-based theory assumes that there is a monitor which is equipped with phonological,lexical and syntactical rules and pragmatic norms and whose sole function is to monitor errors at varying levels when language is produced.Perception-based theory posits that a central monitor within the conceptualizer functions to accomplish the monitoring job.Node structure theory accounts for monitoring from the node activation hypothesis,i.e.,detection and correction of errors is tied to the activation strength or level of node committed or uncommitted.
There is an increasing interest in multimodal communication as suggested by several national and international projects (ISLE, HUMAINE, SIMILAR, CHIL, AMI, CALO, VACE, CALLAS), the attention devoted to the topic by well-known institutions and organizations (the National Institute of Standards and Technology, the Linguistic Data Consortium), and the success of conferences related to multimodal communication (ICMI, IVA, Gesture, Measuring Behavior, Nordic Symposium on Multimodal Communication, LREC Workshops on Multimodal Corpora).
The reflection of the category of the comic in literary speech is under investigation in the article. The basis of the taxonomic description of morphological means in the Russian language, offered by the author, consists of the following ways of creation of the comic: the usage of homonymy and contiguous phenomena to it, play up of the meanings of the same linguistic unit, repetition of the word in different grammatical forms (polyptot), divergence from the linguistic norms.
We address corpus building situations, where complete annotations to the whole corpus is time consuming and unrealistic. Thus, annotation is done only on crucial part of sentences, or contains unresolved label ambiguities. We propose a parameter estimation method for Conditional Random Fields (CRFs), which enables us to use such incomplete annotations. We show promising results of our method as applied to two types of NLP tasks: a domain adaptation task of a Japanese word segmentation using partial annotations, and a part-of-speech tagging task using ambiguous tags in the Penn treebank corpus.
Semantic intrusions are inappropriate responses frequently observed in patients with Alzheimer's disease. They belong to the same category as the words to be remembered, but their prototypic value remains largely unexplored. The prototype is the most representative word in a particular lexical category. The prototypic value is measured according to different criteria: written and oral lexical frequency, frequency of use, degree of typicality, degree of familiarity and rank of quotation. The objective of the study was to evaluate the prototypic value of intrusions produced by 17 Alzheimer's patients with mild to severe dementia, during the cued recall of the Grober & Buschke procedure (RL/RI 16 items). The prototypic value was compared to the categorial norms provided by 1) 17 control subjects and 2) the lexical database "Lexique 3". The results show that intrusions had a significantly higher prototypic value than targeted items. The prototypic value increased with the progression of the disease, and according to the evaluation criteria used. Thus with the criteria "frequency of use", "degree of typicality" and "degree of familiarity," the prototypic value increased exponentially with the severity of dementia. In contrast, in spite of the development of the pathology, the prototypic value decreased when assessed by the criteria of "rank of quotation", and "lexical frequency" (oral and written). In conclusion, the qualitative analysis of the prototypic value of intrusion errors in Alzheimers opens up new clinical and methodological considerations.
Clinical interviews are a powerful method for assessing students’ knowledge and conceptual development. However, the analysis of the resulting data is time-consuming and can create a “bottleneck” in large-scale studies. This article demonstrates the utility of computational methods in supporting such an analysis. Thirty-four 7th-grade student explanations of the causes of Earth’s seasons were assessed using latent semantic analysis (LSA). Analyses were performed on transcriptions of student responses during interviews administered, prior to (n = 21) and after (n = 13) receiving earth science instruction. An instrument that uses LSA technology was developed to identify misconceptions and assess conceptual change in students’ thinking. Its accuracy, as determined by comparing its classifications to the independent coding performed by four human raters, reached 90%. Techniques for adapting LSA technology to support the analysis of interview data, as well as some limitations, are discussed.
Statistical parsing of noun phrase (NP) structure has been hampered by a lack of goldstandard data. This is a significant problem for CCGbank, where binary branching NP derivations are often incorrect, a result of the automatic conversion from the Penn Treebank. We correct these errors in CCGbank using a gold-standard corpus of NP structure, resulting in a much more accurate corpus. We also implement novel NER features that generalise the lexical information needed to parse NPs and provide important semantic information. Finally, evaluating against DepBank demonstrates the effectiveness of our modified corpus and novel features, with an increase in parser performance of 1.51%. 1
In certain areas of the behavioral sciences, such as cognitive and perceptual psychology, researchers may choose to have their experiments partially or completely driven by software programs, which instruct and guide subjects through the sequence of tasks. Despite distinct advantages of unattended trial execution, on frequent occasions, experimenters may desire to keep track of the progress or to be notified of certain events, such as when the subject has completed a task. The Tracer software library presented here is a lightweight Windows programming interface that provides experimenters the ability to trace events and status notifications on one or more remote computers and log files. With only a few lines of additional code or script code, researchers can monitor the real-time progress of one or more unattended experiments running on remote computers of the local area network or the Internet. This article describes the functionality and usage of the Tracer library. The Tracer binaries, include files, sample code, and documentation files may be downloaded from the Psychonomic Society Archive of Norms, Stimuli, and Data at www.psychonomic.org/archive.
Visual psychophysicists, who study object, color, and light perception, have a demand for software that produces complex but, at the same time, physically accurate stimuli for their experiments. The number of computer graphic packages that simulate the physical interaction of light and surfaces is limited, and mostly they require the purchase of a license. RADIANCE (Ward, 1994), however, is freely available and popular in the visual perception community, making it a prime candidate. We have shown previously that RADIANCE’S simulation accuracy is greatly improved when color is coded by spectra, rather than by the originally envisaged RGB triplets (Ruppertsberg & Bloj, 2006). Here, we present a method for spectral rendering with RADIANCE to generate hyperspectral images that can be converted to XYZ images (CIE 1931 system) and then to machine-dependent RGB images. Generating XYZ stimuli has the added advantage of making stimulus images independent of display devices and, thereby, facilitating the process of reproducing results across different labs. Materials associated with this article may be downloaded from www.psychonomic.org.
Since norms for vocabulary acquisition in Maltese children do not yet exist, documentation of productive vocabulary acquisition may contribute to establishing a baseline of lexical development. Clinical implications may thus be derived. The current study is a small-scale investigation of the proportions of Maltese and English lexemes in the vocabularies of ten normally-developing Maltese children aged between 12 and 30 months. The participants were primarily exposed to Maltese within their immediate environments, while receiving indirect exposure to English. Outcomes of parental report and language sampling were analysed for evidence of a bilingual dimension in these children's productive vocabularies. Translation equivalents were reported on by parents, but negligible evidence of equivalents emerged in conversational language use. In contrast, lexical borrowings were both reported and sampled. A substantial proportion of English lexemes were reported by the parents in the absence of Maltese equivalents.
ABSTRACT Using lexical items from Martin Durrell's classification of register variation as a sample, the study investigates how the current advanced monolingual learners' dictionaries of German as an additional language treat such variation and indicate to their users what they consider to be standard usage: how do they set the standard? Abbreviated usage labels as conventionally found in dictionaries for first‐language users are the primary indications, and the dictionaries seldom go beyond such labels. German Standard German is the norm. At its core are unmarked or unlabelled items, while its range extends to include less formal items from everyday use, especially spoken, which are typically labelled umg. or gespr., and more formal items, more particularly found in written usage, which are labelled geh. or geschr. Non‐standard items, if entered as headwords, may be labelled derb or vulgär, veraltet or lit. No one dictionary stands out from the others as setting the standard in terms of treating register variation, and it must be questioned whether learners of German as an additional language would not be better served by more detailed, discursive information on different contexts of use and stylistic levels.
Abstract This paper presents an analysis of a sample of intentional deviations from the typical stress pattern of German words. These deviations are described as stress shifts in which the main stress is in a different position to the norm. This process is optional, mainly found in media speech and used for emphatic purposes. All stress shifts involve an interchange of primary and secondary stress, thereby demonstrating their sensitivity to a prosodic-similarity constraint. Stress retractions by far outnumber stress advancements, which can be jointly explained by a probabilistic association of the main stress and the word-initial position in language structure and an anticipatory bias in the language production system. Stress shifts show a strong overrepresentation of adjectives because this word class codes evaluative aspects most naturally and it is evaluations that speakers prefer to emphasize. From a social-psychological perspective, stress shifts are claimed to be a means by which speakers may boast their knowledge and, from a rhetorical perspective, a strategy of making the event being talked about more spectacular. Stress shifts are minority patterns in the sense that the constraints on them are so strong that only relatively few lexical items are eligible. This raises the issue of what speakers do with those items which they wish to emphasize but which do not lend themselves readily to stress shifting. Whether they turn to alternative means of expression or whether they leave their intentions unexpressed remains to be determined.
The purpose of this paper is to describe the way some of the Spanish general monolingual dictionaries published during the last twelve years have dealt with lexical collocations, that is, those combinations of words that present certain combinatorial restrictions in the norm, basically semantic restrictions, imposed by usage (Corpas 1996). These have been the analyzed dictionaries: Diccionario Salamanca de la lengua espanola, directed by Juan Gutierrez (1996); Diccionario del espanol actual, by Manuel Seco, Olimpia Andres y Gabino Ramos (1999); RAE's Diccionario de la Lengua Espanola (2001); and Gran diccionario de uso del espanol actual. Basado en el Corpus Cumbre, directed by Aquilino Sanchez (2001). We have based our research on a corpus of 52 lexical collocations, which has been built on the analysis of the subentries starting with b in the chosen dictionaries. After that, we have looked up the entries corresponding to each element that constitutes the collocation, in order to know if these dictionaries account for those same combinations in other parts of the lexicographical article. The analysis of the lexicographic information has focused on our aspects: a) the preliminary pages of each dictionary; b) the position of collocations in the lexicographic article; c) the inclusion of these units in a given article; and d) the grammatical category.
The complexity of Korean numeral classifiers demands semantic as well as computational approaches that employ natural language processing (NLP) techniques. The classifier is a universal linguistic device, having the two functions of quantifying and classifying nouns in noun phrase constructions. Many linguistic studies have focused on the fact that numeral classifiers afford decisive clues to categorizing nouns. However, few studies have dealt with the semantic categorization of classifiers and their semantic relations to the nouns they quantify and categorize in building ontologies. In this article, we propose the semantic recategorization of the Korean numeral classifiers in the context of classifier ontology based on large corpora and KorLex Noun 1.5 (Korean wordnet; Korean Lexical Semantic Network), considering its high applicability in the NLP domain. In particular, the classifier can be effectively used to predict the semantic characteristics of nouns and to process them appropriately in NLP. The major challenge is to make such semantic classification and the attendant NLP techniques efficient. Accordingly, a Korean numeral classifier ontology (KorLexClas 1.0), including semantic hierarchies and relations to nouns, was constructed.