Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
Partial cognates are pairs of words in two languages that have the same meaning in some, but not all contexts. Detecting the actual meaning of a partial cognate in context can be useful for Machine Translation tools and for Computer-Assisted Language Learning tools. We propose a supervised and a semi-supervised method to disambiguate partial cognates between two languages: French and English. The methods use only automatically-labeled data; therefore they can be applied to other pairs of languages as well. The aim of our work is to automatically detect the meaning of a French partial cognate word in a specific context.
This special issue of Language Resources and Evaluation, entitled “New Frontiers in Asian Language Resources”, complements the earlier special double issue on Asian Language Processing: State of the Art Resources and Processing (Huang et al. 2006) by presenting eight papers describing specific Asian language resources. As Bird and Simons (2003) explain, research on language resources must deal with how the resources can be acquired and documented as well as how the resources can be accessed and used. Among the eight papers in this issue, the first four papers focus on resources, while the latter four target specific application tasks and describe resource building in the contexts of these applications.
Young (n = 24) and old (n = 24) participants rated 160 faces of young and old individuals taken from the CAL/PAL Face Database (Minear & Park, 2004) with regard to attractiveness, likeability, distinctiveness, goal orientation, energy, mood, and age. Ratings are reported for each face separately. Further analyses showed that the age groups differed in their ratings of young and old faces. On average, old participants evaluated the faces as more positive (i.e., more attractive, more energetic) than did young participants. In line with research on a negative aging stereotype, old faces were judged as less positive than young faces. They were, for instance, seen as less attractive, less likeable, less distinctive, less growth-oriented, and less energetic. The findings of the present study can serve as a basis for the selection of appropriate facial stimuli in age-comparative studies of face perception, face processing, or memory for faces. All face-specific data are archived at www.psychonomic.org/archive.
Critical to vision research is the generation of visual displays with precise control over stimulus metrics. Generating stimuli often requires adapting commercial software or developing specialized software for specific research applications. In order to facilitate this process, we give here an overview that allows nonexpert users to generate and customize stimuli for vision research. We first give a review of relevant hardware and software considerations, to allow the selection of display hardware, operating system, programming language, and graphics packages most appropriate for specific research applications. We then describe the framework of a generic computer program that can be adapted for use with a broad range of experimental applications. Stimuli are generated in the context of trial events, allowing the display of text messages, the monitoring of subject responses and reaction times, and the inclusion of contingency algorithms. This approach allows direct control and management of computer-generated visual stimuli while utilizing the full capabilities of modern hardware and software systems. The flowchart and source code for the stimulus-generating program may be downloaded from www.psychonomic.org/archive.
While large-scale corpora and various corpus query tools have long been recognized as essential language resources, the value of word association norms as language resources has been largely overlooked. This paper conducts some initial comparisons of the lexical relationships observed within Japanese collocation data extracted from a large corpus using the Japanese language version of the Sketch Engine (SkE) tool (Srdanović et al., 2008) and the relationships found within Japanese word association sets taken from the large-scale Japanese Word Association Database (JWAD) under ongoing construction by Joyce (2005, 2007). The comparison results indicate that while some relationships are common to both linguistic resources, many lexical relationships are only observed in one resource. These findings suggest that both resources are necessary in order to more adequately cover the diverse range of lexical relationships. Finally, the paper reflects briefly on the implementation of association-based word-search strategies into electronic dictionaries proposed by Zock and Bilac (2004) and Zock (2006).
Background: Understanding the time course of how listeners reconstruct a missing fundamental component in an auditory stimulus remains elusive. We report MEG evidence that the missing fundamental component of a complex auditory stimulus is recovered in auditory cortex within 100 ms post stimulus onset. Methodology: Two outside tones of four-tone complex stimuli were held constant (1200 Hz and 2400 Hz), while two inside tones were systematically modulated (between 1300 Hz and 2300 Hz), such that the restored fundamental (also knows as ''virtual pitch'') changed from 100 Hz to 600 Hz. Constructing the auditory stimuli in this manner controls for a number of spectral properties known to modulate the neuromagnetic signal. The tone complex stimuli only diverged on the value of the missing fundamental component. Principal Findings: We compared the M100 latencies of these tone complexes to the M100 latencies elicited by their respective pure tone (spectral pitch) counterparts. The M100 laten)
Mediation analysis is widely used in the social sciences. Despite the popularity of mediation models, few researchers have used graphical methods, other than structural path diagrams, to represent their models. Plots of the mediated effect can help a researcher better understand the results of the analysis and convey these results to others. This article presents a method for creating and interpreting plots of the mediated effect for a variety of mediation models, including models with (1) a dichotomous independent variable, (2) a continuous independent variable, and (3) an interaction between an independent variable and the mediating variable. An empirical example is then presented to illustrate these plots. Sample code for creating plots of the mediated effect in R and SAS is also included, and may be downloaded from www.psychonomic.org/archive.
There is an increasing interest in multimodal communication as suggested by several national and international projects (ISLE, HUMAINE, SIMILAR, CHIL, AMI, CALO, VACE, CALLAS), the attention devoted to the topic by well-known institutions and organizations (the National Institute of Standards and Technology, the Linguistic Data Consortium), and the success of conferences related to multimodal communication (ICMI, IVA, Gesture, Measuring Behavior, Nordic Symposium on Multimodal Communication, LREC Workshops on Multimodal Corpora).
Abstract This paper presents an analysis of a sample of intentional deviations from the typical stress pattern of German words. These deviations are described as stress shifts in which the main stress is in a different position to the norm. This process is optional, mainly found in media speech and used for emphatic purposes. All stress shifts involve an interchange of primary and secondary stress, thereby demonstrating their sensitivity to a prosodic-similarity constraint. Stress retractions by far outnumber stress advancements, which can be jointly explained by a probabilistic association of the main stress and the word-initial position in language structure and an anticipatory bias in the language production system. Stress shifts show a strong overrepresentation of adjectives because this word class codes evaluative aspects most naturally and it is evaluations that speakers prefer to emphasize. From a social-psychological perspective, stress shifts are claimed to be a means by which speakers may boast their knowledge and, from a rhetorical perspective, a strategy of making the event being talked about more spectacular. Stress shifts are minority patterns in the sense that the constraints on them are so strong that only relatively few lexical items are eligible. This raises the issue of what speakers do with those items which they wish to emphasize but which do not lend themselves readily to stress shifting. Whether they turn to alternative means of expression or whether they leave their intentions unexpressed remains to be determined.
Predicting Word-Naming and Lexical Decision Times from a Semantic Space Model Brendan T. Johns (johns4@indiana.edu) Department of Psychological and Brain Sciences, 1101 E. Tenth St. Bloomington, In 47405 USA Michael N. Jones (jonesmn@indiana.edu) Department of Psychological and Brain Sciences, 1101 E. Tenth St. Bloomington, In 47405 USA Abstract organization of semantic memory. In a lexical decision task, a letter string is presented and the participant provides a speeded response of whether the string is a word or not. In a naming task, the participant’s task is to name the presented word aloud as quickly as possible. Both measures produce an index of a word’s identification latency. Orthographic and phonological factors are certainly large components of both LDT and NT, but semantics plays a significant role as well, and co-occurrence models have yet to be extended to predicting reaction time variance for these single-word identification tasks. Modeling of retrieval times is usually done by looking for the best environmental correlates of LDT and NT (Adelman & Brown, 2008). Some of the most influential models of retrieval times are based upon word frequency. Word frequency (WF) has been used to drive many different types of models, including serial-searched rank frequency models (Murray & Forster, 2004), threshold activation models (Coltheart, et al., 2001), and connectionist models (Seidenberg & McClelland, 1989). However, recent evidence suggests that word frequency may not drive retrieval times but, rather, the causal factor is a word’s contextual diversity (Adelman, Brown, & Quesada, 2006; Adelman & Brown, 2008). Contextual diversity (CD) is the number of different contexts that a word appears in, and is based on the rational analysis of memory (Anderson & Milson, 1989), particularly the principle of likely need (PLN). PLN states that the more unique contexts a word appears in, the more likely the word will be needed in any future context. Hence, a word with a high CD should be faster to retrieve under this principle. A word’s CD value is typically computed by simply counting the number of different documents in which it appears across a text corpus. This measure has been shown to be a better predictor of LDT and NT than WF (Adelman, et al., 2006). However, operationalizing CD as the number of documents in which a word occurs may not be a fair instantiation of PLN. A word that appears in many documents may have a high WF, but it should have a low CD if those documents are highly redundant, as is the case with words that belong to a popular discourse topic for which many documents exist. It is the number of different contexts and the uniqueness of contexts that determines a word’s likely need. This calls for a measure of CD that considers the semantic uniqueness of documents that a word appears in. Based on PLN, it is reasonable to assume that if a word appears in a context it has never before occurred in, We propose a method to derive predictions for single-word retrieval times from a semantic space model trained on text corpora. In Experiment 1 we present a large corpus analysis demonstrating that it is the number of unique semantic contexts a word appears in across language, rather than simply the number of contexts or the frequency of the word, that is the most salient predictor of lexical decision and naming times. In Experiment 2, we develop a co-occurrence learning model that weights new contextual uses of a word based on fit to what currently exists in the word’s memory representation, and demonstrate this model’s superiority in fitting the human data compared to models built using information about the word’s frequency or number of contexts. Finally, in Experiment 3 we find that building lexical representations using semantic distinctiveness naturally produces a better-organized semantic space to make predictions for semantic similarity between words. Keywords: Co-occurrence model; Lexical-decision; LSA; Contextual distinctiveness Introduction The last decade has seen remarkable progress with co- occurrence models of lexical semantics (e.g., Lund & Burgess, 1996; Landauer & Dumais, 1997). These models learn semantic representations for words by observing lexical co-occurrence patterns across a large text corpus, typically representing the words in a high-dimensional semantic space. This approach provides both an account of the semantic representation for words and an account of the learning mechanisms humans use to build and organize semantic memory. Co-occurrence models have seen considerable success at accounting for data in a wide variety of semantic tasks, including TOEFL synonyms (Landauer & Dumais, 1997), semantic similarity ratings and exemplar categorization (Jones & Mewhort, 2007), and free association norms (Griffiths, Steyvers, & Tenenbaum, To date, all applications of co-occurrence models have been to semantic similarity between two words or two documents. The standard prediction of semantic similarity in these models is some measure of the angle between two vectors. However, co-occurrence models should, in theory, contain sufficient information in the magnitude of their representations to make predictions about single word retrieval as well. Lexical decision time (LDT) and word naming time (NT) are both important variables that offer insight into the
Recent neuroimaging studies have identified a set of brain regions that are metabolically active during wakeful rest and consistently deactivate in a variety the performance of demanding tasks. This ''default network'' has been functionally linked to the stream of thoughts occurring automatically in the absence of goal-directed activity and which constitutes an aspect of mental behavior specifically addressed by many meditative practices. Zen meditation, in particular, is traditionally associated with a mental state of full awareness but reduced conceptual content, to be attained via a disciplined regulation of attention and bodily posture. Using fMRI and a simplified meditative condition interspersed with a lexical decision task, we investigated the neural correlates of conceptual processing during meditation in regular Zen practitioners and matched control subjects. While behavioral performance did not differ between groups, Zen practitioners displayed a reduced duration of the neur)
Background: The alcohol dehydrogenases (ADH) are widely studied enzymes and the evolution of the mammalian gene cluster encoding these enzymes is also well studied. Previous studies have shown that the ADH1B*47His allele at one of the seven genes in humans is associated with a decrease in the risk of alcoholism and the core molecular region with this allele has been selected for in some East Asian populations. As the frequency of ADH1B*47His is highest in East Asia, and very low in most of the rest of the world, we have undertaken more detailed investigation in this geographic region. Methodology/Principal Findings: Here we report new data on 30 SNPs in the ADH7 and Class I ADH region in samples of 24 populations from China and Laos. These populations cover a wide geographic region and diverse ethnicities. Combined with our previously published East Asian data for these SNPs in 8 populations, we have typed populations from all of the 6 major linguistic phyla (Altaic including Korean-J)
Background: Genome-wide data provide a powerful tool for inferring patterns of genetic variation and structure of human populations. Principal Findings: In this study, we analysed almost 250,000 SNPs from a total of 945 samples from Eastern and Western Finland, Sweden, Northern Germany and Great Britain complemented with HapMap data. Small but statistically significant differences were observed between the European populations (FST = 0.0040, p<10-4), also between Eastern and Western Finland (FST = 0.0032, p<10-3). The latter indicated the existence of a relatively strong autosomal substructure within the country, similar to that observed earlier with smaller numbers of markers. The Germans and British were less differentiated than the Swedes, Western Finns and especially the Eastern Finns who also showed other signs of genetic drift. This is likely caused by the later founding of the northern populations, together with subsequent founder and bottleneck effects, and a smaller populatio)
Since norms for vocabulary acquisition in Maltese children do not yet exist, documentation of productive vocabulary acquisition may contribute to establishing a baseline of lexical development. Clinical implications may thus be derived. The current study is a small-scale investigation of the proportions of Maltese and English lexemes in the vocabularies of ten normally-developing Maltese children aged between 12 and 30 months. The participants were primarily exposed to Maltese within their immediate environments, while receiving indirect exposure to English. Outcomes of parental report and language sampling were analysed for evidence of a bilingual dimension in these children's productive vocabularies. Translation equivalents were reported on by parents, but negligible evidence of equivalents emerged in conversational language use. In contrast, lexical borrowings were both reported and sampled. A substantial proportion of English lexemes were reported by the parents in the absence of Maltese equivalents.
The reflection of the category of the comic in literary speech is under investigation in the article. The basis of the taxonomic description of morphological means in the Russian language, offered by the author, consists of the following ways of creation of the comic: the usage of homonymy and contiguous phenomena to it, play up of the meanings of the same linguistic unit, repetition of the word in different grammatical forms (polyptot), divergence from the linguistic norms.
This article presents a Web-based tool for the creation of divergent-thinking and open-ended creativity tasks. A Java program generates HTML forms with PHP scripting that run an Alternate Uses Task and/or open-ended response items. Researchers may specify their own instructions, objects, and time limits, or use default settings. Participants can also be prompted to select their best responses to the Alternate Uses Task (Silvia et al., 2008). Minimal programming knowledge is required. The program runs on any server, and responses are recorded in a standard MySQL database. Responses can be scored using the consensual assessment technique (Amabile, 1996) or Torrance’s (1998) traditional scoring method. Adoption of this Web-based tool should facilitate creativity research across cultures and access to eminent creators. The Creative Task Creator may be downloaded from the Psychonomic Society’s Archive of Norms, Stimuli, and Data, www.psychonomic.org/archive.
THE VOICE TEACHER IS REGULARLY BESET WITH CHALLENGES in the studio regarding consonant clusters in sung German, as is the singer who approaches any vocal work in the German language. The reputation of the German language as being consonant rather than vowel oriented is commonly appreciated and justifiable. Statistical studies show that the burden of text intelligibility is carried principally by the consonants, to a greater extent than most languages. A language that can produce lexical items such as entsturzt [ent'∫tYrtst] and kraftstrotzend ['kraft∫trctsent] adopts a strongly marked position among the world's languages with respect to the involvement of consonants in its sound system. These words contain ten and fourteen phonemes respectively, of which only two or three are vowels. The remaining clusters of consonants are samples of the subject of this article. The consonant clusters normally encountered in German phonology will be inventoried and contrasted with English. The material is likely to be familiar to many readers, albeit presented in a different, perhaps more systematic perspective than is normally encountered. The subject of German consonant clusters is best dealt with in terms of phonetic, not orthographic consonants. A firm grasp of the relationship between spelling and pronunciation is naturally also essential. Two or three successive letters may represent a single phoneme, as in [arrow right] /c/ or /x/ [arrow right] /∫/ [arrow right] /k/ [arrow right] /t/ Conversely, a single written consonant may serve to indicate more than one phoneme, as in [arrow right] /ts/ This situation is familiar because it is even more pronounced in English. The word scythe contains two consonant digraphs and two letter-vowels, but phonetically only one diphthong and no clusters at all. Since the greatest challenge in consonant clusters is visual (i.e., orthographic), thinking in terms of phonetic consonants should serve to simplify the matter for a student. German, more than most other languages, has absorbed lexical items from other languages into its own vocabulary, particularly from English, French, and Italian. Thus Duden, the principal lexicographic publisher in modern Germany, devotes an entire book to Fremdworter in its series of dictionaries. The process of lexical transfer is a complex aspect of German linguistics, particularly regarding pronunciation norms. Some words, such as Situation, have been subsumed into the phonological patterning of German, while others have retained the pronunciation of the word in the language from whence it came, or have struck a middle ground, such as Orange and Weekend. This diversity of phonetic transfer gives modern spoken German a particular flavor, and reflects the country's central geographic position in Europe. There are similar examples in English, such as cul-de-sac (where the French pronunciation has been distorted) and naive (which retains the original, although English idiosyncratically employs only the feminine form). This article will confine itself to the consonant clusters that occur regularly in the standard lexis, referring to combinations resulting from foreign influences only when appropriate. It is useful to consider consonant clusters in two quite distinct groups: syllable-interior, and across syllable or word boundaries. Part I of the article will concern itself with the former; Part II (to appear in the March/April 2008 issue), with the latter. Recognition of which group an example belongs to is the first step toward establishing correct pronunciation, and in some cases is necessary to discriminate between two potentially correct pronunciations. Before outlining in tabular form the cluster environments of German and English, it will be useful to consider all the consonantal combinations that are admissible in each language. A detailed theoretical account of the phonotactic rules and constraints of each language will not be necessary for our purposes. …
Reviews the book, Le Francais en Amérique du Nord: État présent by Albert Valdman, Julie Auger, and Deborah Piston-Hatlen (eds.) (2005). Poirier, Boivin, Trepanier & Verreault 1994 gave us the first general overview of the French linguistic legacy in North America. Valdman and his team have updated that work, incorporating the findings of sociolinguistic research carried out over the past decade, notably in language obsolescence. The result is comprehensive treatment of all the main areas of North America where French is spoken. It is well organized, with an introductory chapter providing a succinct overview of the four sections to follow: The first describes where, how, and to what extent French is spoken in North America; the second examines language variation in each of these areas and the effects of language contact; the third looks at linguistic norms and language planning; and a fourth is devoted to more general comparative and historical issues. It is possible to point to improvements that could have been incorporated into the volume. There are occasional production blemishes. However, this volume offers an invaluable tool not only for students of French but also for sociolinguists concerned with the effects of language contact, language change, and language obsolescence. (PsycINFO Database Record (c) 2016 APA, all rights reserved)
Statistical parsing of noun phrase (NP) structure has been hampered by a lack of goldstandard data. This is a significant problem for CCGbank, where binary branching NP derivations are often incorrect, a result of the automatic conversion from the Penn Treebank. We correct these errors in CCGbank using a gold-standard corpus of NP structure, resulting in a much more accurate corpus. We also implement novel NER features that generalise the lexical information needed to parse NPs and provide important semantic information. Finally, evaluating against DepBank demonstrates the effectiveness of our modified corpus and novel features, with an increase in parser performance of 1.51%. 1
This study investigated cued odor identification performance with a set of 64 natural common odors (half of edible and half of nonedible stimuli) in three groups of participants: one group of 30 young adults (mean age 25.3 years, range 18–30, SD 3.1) and two groups of older adults—20 young-old (mean age 64.4 years, range 60–69, SD 2.8) and 21 old-old (mean age 74.6 years, range 70–79, SD 2.5). The results showed that 49 of the 64 odors were correctly identified by over 70% of the participants in all groups. The odor identification performance of the young-old adults did not differ from that of the young adults. However, the oldest group showed a significant loss of performance in the task. Women in the young-old group performed better than men, whereas no gender differences were found in the other two age groups. The data obtained in this study will be useful for further perceptual and memory studies conducted in the olfactory modality with young as well as with older participants.
As a linguistic phenomenon language play is directly connected with norms or anomaly. The aim of this paper is to present the essence of dictemes with abnormal constituents such as zero formal representation or reduced semantic form and circumstances in which they appear on the one hand, and changes in syntactical as well as in lexical semantics of these phenomena on the other hand.
In my thesis I have attempted to develop an integrated translation approach materialized in the form of a Dynamic Translation Model (DTM). This endeavour can be justified to the extent that Translation Studies is perceived so far as a fragmentary discipline with implicitly and explicitly opposed and apparently irreconcilable points of view: linguistics-oriented approaches and culture-and-literature-oriented approaches. The main problem arising from this lack of common ground for further developing Translation Studies is that the disciplinary boundaries are not well-established and therefore the discipline itself cannot be developed coherently. Besides, Translation Studies is still to be constructed as an autonomous and an independent discipline that has a common core of theoretical and practical problems. This lack of coherent development of the discipline is due, I think, to an epistemological mistake: to believe that one single approach can account for (that is, describe and explain) all the translational reality. I propose to distinguish a two-phase epistemological move: 1. each translation approach works on its own research interests and acknowledges that its approach deals only with one part of the whole subject matter of Translation Studies; and 2. the results obtained by each translation approach are incorporated into a holistic integrative model like the Dynamic Translation Model I propose. In order to achieve this goal I have attempted to show the key tenets of modern translation approaches, both linguistics-oriented and culture-and-literature-oriented, by quoting the main theses of the representatives of these approaches. I have then presented the most important criticisms that have been raised in relation to these diverse translation approaches, together with my own criticisms (chapters 1 and 2). Also, I have introduced the theoretical basis for an integrated approach taking Holmes’ differentiation between theoretical (product-, process-, and function-oriented) and practical approaches as a point of departure. Likewise, I have discussed the problems of integrating Translation Studies, as well as Snell-Hornby’s integrated proposal and some key aspects of literary translation relevant for my integrative endeavour (chapter 3). Finally, I have developed my proposal for a Dynamic Translation Model (chapter 4). As to the conclusions of my thesis, I can say that my holistic DTM was able to integrate functionally aspects from both linguistics-oriented and culture-and-literature-oriented approaches: historico-cultural context (Leipzig School and postcolonial studies); norms, ideology and power (Descriptive Translation Studies; G. Toury and A. Lefevere); translation commisioner (Skopos theory); sender’s communicative purpose (linguistic and pragmatic approaches: W. Koller, J. House, H. Gerzymisch-Arbogast, etc); importance of source language text (linguistic and textlinguistic approaches; stylistic approaches; B. Spillner, B. Sandig); translator’s comprehension process (hermeneutic, deconstructive, and poststructural approaches), target language receiver in the target language historico-cultural context (Descriptive Translation Studies; postcolonial and gender studies). On the other hand, the three levels of the Dynamic Translation Model help to explain the flux of translational proceses and the variables that are activated or neutralized therein. They also incorporate concepts from other disciplines such as text linguistics, pragmatics, stylistics, and the communication theory. In my integrative endeavour I also proposed new concepts and, accordingly, coined new terms: Compulsory Translational Forces (CTF) (which include both Initiator’s Translational Instructions (ITI) and Target Language Valid Translational Norms (TL-VTN), Default Equivalence Position (DEP). In the pragmatic dimension of the model special attention is paid to what I call Text Illocutionary Indicators (TII) as well as the strengthening (upgraders) and weakening (downgraders) illocutionary mechanisms in relation to the Source Language Text (SLT) and the Target Language Text (TLT). Semantic/lexical fields play a crucial role in the establishment of equivalences between SLT and TLT in the text semantic dimension, as well as what I have called Fictionalizing Stylistic Shifts in the text stylistic dimension. As to the future developments of translation research within the framework of the Dynamic Translation Model I would say that some modificationbs may be called for so that interpretation can also be accounted for. This proposal can be used profitably in the field of translation criticism. As is the case with any other integrative approach, DTM should be widely discussed and criticized in order to validate its theoretical soundness and its application in Translation Studies. This thesis is an attempt to contribute in this research direction.
The semantic annotation of texts with senses from a computational lexicon is a complex and often subjective task. As a matter of fact, the fine granularity of the WordNet sense inventory [Fellbaum, Christiane (ed.). 1998. WordNet: An Electronic Lexical Database MIT Press], a de facto standard within the research community, is one of the main causes of a low inter-tagger agreement ranging between 70% and 80% and the disappointing performance of automated fine-grained disambiguation systems (around 65% state of the art in the Senseval-3 English all-words task). In order to improve the performance of both manual and automated sense taggers, either we change the sense inventory (e.g. adopting a new dictionary or clustering WordNet senses) or we aim at resolving the disagreements between annotators by dealing with the fineness of sense distinctions. The former approach is not viable in the short term, as wide-coverage resources are not publicly available and no large-scale reliable clustering of WordNet senses has been released to date. The latter approach requires the ability to distinguish between subtle or misleading sense distinctions. In this paper, we propose the use of structural semantic interconnections—a specific kind of lexical chains—for the adjudication of disagreed sense assignments to words in context. The approach relies on the exploitation of the lexicon structure as a support to smooth possible divergencies between sense annotators and foster coherent choices. We perform a twofold experimental evaluation of the approach applied to manual annotations from the SemCor corpus, and automatic annotations from the Senseval-3 English all-words competition. Both sets of experiments and results are entirely novel: structural adjudication allows to improve the state-of-the-art performance in all-words disambiguation by 3.3 points (achieving a 68.5% Fl-score) and attains figures around 80% precision and 60% recall in the adjudication of disagreements from human annotators. (PsycINFO Database Record (c) 2016 APA, all rights reserved)
Real-time communication platforms such as ICQ, MSN and online chat rooms are getting more popular than ever on the Internet. There are, however, real risks where criminals and terrorists can perpetrate illegal and criminal abuses. This highlights the security significance of accurate detection and translation of the chat language to its stand language counterpart. The language used on these platforms differs significantly from the standard language. This language, referred to as chat language, is comparatively informal, anomalous and dynamic. Such features render conventional language resources such as dictionaries, and processing tools such as parsers ineffective. In this paper, we present the NIL corpus, a chat language text collection annotated to facilitate training and testing of chat language processing algorithms. We analyse the NIL corpus to study the linguistic characteristics and contextual behaviour of a chat language. First we observe that majority of the chat terms, i.e. informal words in a chat text, is formed by phonetic mapping. We then propose the eXtended Source Channel Model (XSCM) for the normalization of the chat language, which is a process to convert messages expressed in a chat language to its standard language counterpart. Experimental results indicate that the performance of XSCM in terms of chat term recognition and normalization accuracy is superior to its Source Channel Model (SCM) counterparts, and is also more consistent over time.
Semantic intrusions are inappropriate responses frequently observed in patients with Alzheimer's disease. They belong to the same category as the words to be remembered, but their prototypic value remains largely unexplored. The prototype is the most representative word in a particular lexical category. The prototypic value is measured according to different criteria: written and oral lexical frequency, frequency of use, degree of typicality, degree of familiarity and rank of quotation. The objective of the study was to evaluate the prototypic value of intrusions produced by 17 Alzheimer's patients with mild to severe dementia, during the cued recall of the Grober & Buschke procedure (RL/RI 16 items). The prototypic value was compared to the categorial norms provided by 1) 17 control subjects and 2) the lexical database 'Lexique 3'. The results show that intrusions had a significantly higher prototypic value than targeted items. The prototypic value increased with the progression of the disease, and according to the evaluation criteria used. Thus with the criteria 'frequency of use', 'degree of typicality' and 'degree of familiarity,' the prototypic value increased exponentially with the severity of dementia. In contrast, in spite of the development of the pathology, the prototypic value decreased when assessed by the criteria of 'rank of quotation', and 'lexical frequency' (oral and written). In conclusion, the qualitative analysis of the prototypic value of intrusion errors in Alzheimers opens up new clinical and methodological considerations. (PsycINFO Database Record (c) 2016 APA, all rights reserved)
We survey the evaluation methodology adopted in information extraction (IE), as defined in a few different efforts applying machine learning (ML) to IE. We identify a number of critical issues that hamper comparison of the results obtained by different researchers. Some of these issues are common to other NLP-related tasks: e.g., the difficulty of exactly identifying the effects on performance of the data (sample selection and sample size), of the domain theory (features selected), and of algorithm parameter settings. Some issues are specific to IE: how leniently to assess inexact identification of filler boundaries, the possibility of multiple fillers for a slot, and how the counting is performed. We argue that, when specifying an IE task, these issues should be explicitly addressed, and a number of methodological characteristics should be clearly defined. To empirically verify the practical impact of the issues mentioned above, we perform a survey of the results of different algorithms when applied to a few standard datasets. The survey shows a serious lack of consensus on these issues, which makes it difficult to draw firm conclusions on a comparative evaluation of the algorithms. Our aim is to elaborate a clear and detailed experimental methodology and propose it to the IE community. Widespread agreement on this proposal should lead to future IE comparative evaluations that are fair and reliable. To demonstrate the way the methodology is to be applied we have organized and run a comparative evaluation of ML-based IE systems (the Pascal Challenge on ML-based IE) where the principles described in this article are put into practice. In this article we describe the proposed methodology and its motivations. The Pascal evaluation is then described and its results presented.
Reply by the current authors to the review by Constant Leung (see record [rid]2008-09991-006[/rid]) on the original book, Language testing: The social dimension (2006). Leung has drawn attention to possibly the most obvious gap in our treatment: a properly elaborated discussion of the assessment of English as a lingua franca. While consideration of this issue has begun in the work of authors he cites in his review, and elsewhere, it is, as Leung points out, a multiply complex issue, in which the social dimension is the crux of the problem, ‘in terms of speaker subject positions, lexicogrammatical norms, transcultural pragmatic conventions and so on’. It is clear that language testing—particularly significant here as a site of authority about linguistic norms, not unlike a dictionary—has an important role to play in authorizing or de-authorizing English as a lingua franca communication as a proper target for language learning; a further demonstration, if one were needed, of the power of tests. But beyond this political and institutional dimension, the psychometric problems inherent in designing tests based on a construct that is local, situated, and fluid pose a difficult but productive challenge to testers. (PsycINFO Database Record (c) 2016 APA, all rights reserved)
This article describes a project that aimed to uncover the effects of different forms of conflict on team performance during the important feasibility, requirements analysis, and design phases of software engineering (SE) projects. The research subjects were master of science students who were working to produce software commissioned by real-world clients. A template was developed that allowed researchers to record details of any conflicts that occurred. It was found that some forms of conflict were more damaging than others and that the frequency and intensity of specific conflicts are important factors to consider. The experience of the researchers when using the final template suggests that it is a valuable weapon to have in one’s arsenal if one is interested in observing and recording the details of conflict in either SE teams or teams in different contexts.
Visual psychophysicists, who study object, color, and light perception, have a demand for software that produces complex but, at the same time, physically accurate stimuli for their experiments. The number of computer graphic packages that simulate the physical interaction of light and surfaces is limited, and mostly they require the purchase of a license. RADIANCE (Ward, 1994), however, is freely available and popular in the visual perception community, making it a prime candidate. We have shown previously that RADIANCE’S simulation accuracy is greatly improved when color is coded by spectra, rather than by the originally envisaged RGB triplets (Ruppertsberg & Bloj, 2006). Here, we present a method for spectral rendering with RADIANCE to generate hyperspectral images that can be converted to XYZ images (CIE 1931 system) and then to machine-dependent RGB images. Generating XYZ stimuli has the added advantage of making stimulus images independent of display devices and, thereby, facilitating the process of reproducing results across different labs. Materials associated with this article may be downloaded from www.psychonomic.org.
Previous research found that the duration of segments decreases as children grow older. The development of suprasegmental duration, however, has not been explored. The present study investigated developmental changes in duration of the four Mandarin tones. 5-, 8-, and 12-year-old monolingual Mandarin-speaking children and young adults participated in the study. Tone durations were measured in participants’ production of monosyllabic target words elicited by picture identification tasks. The results were as follows (1) For each tone category, tone duration and variability decreased with age: 5- and 8-year-old children showed significantly longer durations than adults. Tone durations in 12-year-old children approximated adult values. (2) Despite longer durations, adultlike duration patterns across tone categories existed in all children: dipping tones were the longest, followed by rising and level tones, with falling tones being the shortest. (3) Duration differences between the rising and dipping tones became larger as children grew older. The results may be indicative of the general maturation of laryngeal control over age. Although 5- and 8-year-old children have already established lexical contrasts of tone, adultlike phonetic norms are still in the process of development. The developmental data also provide support for a hybrid account of speech production from a suprasegmental perspective.
In our research on using information extraction to help populate semantic web resources, we have encountered significant obstacles to interoperability between the technologies. We believe these obstacles to be endemic to the basic paradigms and not quirks of the specific implementations we have worked with. In particular, we identify five dimensions of interoperability that must be addressed to successfully employ information extraction systems to populate semantic web resources that are suitable for reasoning. We call the task of transforming IE data into knowledge-based resources knowledge integration and we report results of experiments in which the knowledge integration process uses the deeper semantics of OWL ontologies to improve by between 8% and 13% the precision of relation extraction from text.
Hungarian Academy of ScienceEötvös Loránd UniversityThis paper examines the Afro-Asiatic etymologies of Chadic lexical roots discussed by Olga V. Stolbova in her Chadic Lexical Database, Issue I (2005). The analysis is arranged according to the following sections: (1) Common Chadic reconstructions, (2) Isolated Chadic roots that nevertheless have Afro-Asiatic cognates. The paper represents the third part of my longer series of papers on addenda et corrigenda to Chadic lexical roots.
Data-driven learning based on shift reduce parsing algorithms has emerged dependency parsing and shown excellent performance to many Treebanks. In this paper, we investigate the extension of those methods while considerably improved the runtime and training time efficiency via L2-SVMs. We also present several properties and constraints to enhance the parser completeness in runtime. We further integrate root-level and bottom-level syntactic information by using sequential taggers. The experimental results show the positive effect of the root-level and bottom-level features that improve our parser from 81.17 % to 81.41 % and 81.16 % to 81.57 % labeled attachment scores with modified Yamada’s and Nivre’s method, respectively on the Chinese Treebank. In comparison to well-known parsers, such as Malt-Parser (80.74%) and MSTParser (78.08%), our methods produce not only better accuracy, but also drastically reduced testing time in 0.07 and 0.11, respectively. 1
In this paper we describe some technical and theoretical aspects related to a manually aligned bilingual treebank Italian (ITA) – Italian Sign Language (LIS) provided with both constituency and dependency annotation (Siena University Treebank, SUT). We briefly discuss the linguistic rationale behind the feature set and the dependency/constituency structure we adopted. Moreover we discuss the tool we used to annotate, semi-automatically, the treebank that, in the end, will be evaluated qualitatively with respect to a specific Transfer-Based Machine Translation (TB-MT) task.
The process of developing, implementing, and refining a registry data validation system is integral to optimal trauma registry operations. Describing registrar skill and proficiency in a manner that was once subjective can be replaced with objective assessment through the use of concrete rating guidelines and examples. The ability to standardize the evaluation of each registry abstract becomes the foundation for analyzing the overall accuracy of registry data. Key to the process is incorporating the validation rating tool as part of the data abstract. If properly implemented, the methodology described becomes a practical means for accuracy reporting, peer benchmarking, orientation and training, and performance management.
The aim of this presentation is to describe our on-going work on the building of a written Hindi treebank, referred to here as the Uppsala Hindi
The increasing use of ontologies in recent years has led to much of the information represented in them appear in differently expressed ways. When integration processes of ontologies are required within the same domain, problems arise in the correspondence between the terms that appear in ontologies involved in the process. To solve this problem, this article presents a method, which provides the user with a value of similarity between the terms of two ontologies to assist in the integration process. The method is based on the use of the lexical database defined by WordNet and the application of semantic similarity algorithms. It also introduces a software tool that carries out the implementation of this method.
We introduce LTAG-spinal, a novel variant of traditional Lexicalized Tree Adjoining Grammar (LTAG) with desirable linguistic, computational and statistical properties. Unlike in traditional LTAG, subcategorization frames and the argument-adjunct distinction are left underspecified in LTAG-spinal. LTAG-spinal with adjunction constraints is weakly equivalent to LTAG. The LTAG-spinal for- malism is used to extract an LTAG-spinal Treebank from the Penn Treebank with Propbank annotation. Based on Propbank annotation, predicate coordination and LTAG adjunction structures are successfully extracted. The LTAG-spinal Treebank makes explicit semantic relations that are implicit or absent from the original PTB. LTAG-spinal provides a very desirable resource for statistical LTAG parsing, incremental parsing, dependency parsing, and semantic parsing. This treebank has been successfully used to train an incremental LTAG-spinal parser and a bidirec- tional LTAG dependency parser.
Text-to-phoneme (TTP) mapping, also called grapheme-to-phoneme (GTP) conversion, defines the process of transforming a written text into its corresponding phonetic transcription. Text-to-phoneme mapping is a necessary step in any state-of-the-art automatic speech recognition (ASR) and text-to-speech (TTS) system, where the textual information changes dynamically (i.e., new contact entries for name dialing, or new short messages or emails to be read out by a device). There are significant differences between the implementation requirements of a text-to-phoneme mapping module embedded into the automatic speech recognition and into the text-to-speech systems: in automatic speech recognition systems the errors of the text-to-phoneme mapping module are tolerated better (leading to occasional recognition errors) than in the text-to-speech applications, where the effect is immediately and in all cases audible. Automatic speech recognition systems typically use text-to-phoneme mapping to lower the footprint (to avoid storing the lexicon), while maintaining quality. The use of text-to-phoneme mapping in the text-to-speech systems is different. In addition to the phonetic information, the text-to-speech systems also need prosodic information to be able to produce high quality speech, which cannot be predicted by text-to-phoneme mapping. Most state-of-the-art text-to-speech systems use explicit pronunciation lexicon, which is aimed at providing the widest possible coverage, in the order of 100K words, with high quality pronunciation information. Because of this reason, text-to-phoneme mapping is typically used as a fall-back strategy, when the system encounters very rare or non-native words and the quality of a ext-to-speech system is indirectly affected by the quality of the grapheme-to-phoneme conversion. Another important issue is the question of training the text-to-phoneme mapping module. The problem of grapheme-to-phoneme conversion is a static one and such a system is trained off-line. The correspondence between the written and spoken form of a language is usually unchanged in the lifetime of an application. So the complexity/speed of the model training is of secondary importance compared to e.g., the speed of convergence or model size.\n\nIn this thesis, the problem of text-to-phoneme mapping using neural networks is studied. One of the main goals of the thesis is to provide a comprehensive analysis of different neural network structures which can be implemented to convert a written text into its corresponding phonetic transcription. Another important target, of this work, is to provide new solutions that improve the performance of the existing algorithms, in terms of convergence speed and phoneme accuracy. Three main neural network classes are studied in this thesis: the multilayer perceptron (MLP) neural network, the recurrent neural network (RNN) and the bidirectional recurrent neural network (BRNN).\n\nDue to their ability of self adaptation, neural networks have been shown to be a viable solution in applications that require modeling abilities. Such an application is the text-tophoneme mapping where the correspondence between letters of a written text and their corresponding phonetic transcription must be modeled.\n\nOne of the main concerns in all practical implementations, where neural networks are used, is to develop algorithms which provide fast convergence of the synaptic weights and in the same time good mapping performances. When a neural network is trained for text-to-phoneme mapping, at every iteration, a letter-phoneme pair is presented to the network such that, the number of letters and the number of training iterations are equal. As a result, fast convergence of the neural network means smaller size of the training dictionary since fast convergence is in fact similar to less necessary training letters1. A fast convergence speed is important in applications where only a small linguistic database is available. Of course, one solution could be to use a small dictionary (with very few words) which is presented at the input of the neural network many times until the convergence of the synaptic weights is reached. In this case the time of training becomes more important. Taking into account these two sides of the convergence speed (the size of the training dictionary and the processing time during training) one can understand the importance of having algorithms that ensure fast convergence of the neural network.\n\nIt is well known that the error back-propagation algorithm which is used to train the MLP neural network, possess sometimes a quite slow convergence (a very large number of iterations required to reach the stability point). In order to increase the convergence speed two novel alternative solutions are proposed in this thesis: one using an adaptive learning rate in the training process and another which is a transform domain implementation of the multilayer perceptron neural network. The computational complexity of the two proposed training algorithms is slightly higher than the computational complexity of the error back-propagation algorithm but the number of training iterations is highly reduced. Due to this fact, although the three algorithms might have the same training time, the novel algorithms necessitate smaller training dictionary.\n\nDue to the limitations of the processing power that usually are encountered in real devices, another very important requirement for a text-to-phoneme mapping system is to have low computational and memory costs. In the case of text-to-phoneme mapping systems based in neural networks, the computational complexity is mainly linked to the mathematical complexity of the training algorithm as well as to the number of the synaptic weights of the neural network. Memory load is due to the number of synaptic weights of the neural network which must be stored.\n\nTaking into account all these limitations and implementation requirements, in this thesis, several neural network structures with different number of synaptic weights and trained with various training algorithms, are studied. The modeling capability of the neural networks is addressed, which is translated in the text-to-phoneme mapping case into the phoneme accuracy. Different neural network structures, training algorithms and network complexities are analyzed also from this point of view. As a remark here, we mention that input letter encoding plays a very important role in the phoneme accuracy of the grapheme-to-phoneme conversion system. This is why special attention has been paid to the comparative analysis of the performances (in terms of phoneme accuracy) obtained with several orthogonal and non-orthogonal encoding of the input letters.\n\nThe thesis is structured into four main parts. Chapter 1 brings the reader into the world of text-to-phoneme mapping. In Chapter 2 several different neural network structures and their corresponding training algorithms are described and two new training algorithms are introduced and analyzed. In Chapter 3 the experimental results, for the problem of monolingual text-to-phoneme mapping, obtained with the neural networks described in Chapter 2 are shown. Chapter 4 is dedicated to the problem of bilingual grapheme-to-phoneme conversion and Chapter 5 concludes the thesis.
A treebank is a text corpus in which each sentence has been annotated with its syntactic structure. Although the construction of a treebank is an expensive task, we believe that it is indispensable for the development of real applications in the field of Natural Language Processing (NLP) and also for the development of the
Grammar induction is one of attractive research areas of natural language processing. Since both supervised and to some extent semi-supervised grammar induction methods require large treebanks, and for many languages, such treebanks do not currently exist, we focused our attention on unsupervised approaches. Constituent Context Model (CCM) seems to be the state of the art in unsupervised grammar induction. In this paper, we show that the performance of CCM in free word order languages (FWOLs) such as Persian is inferior to that of fixed order languages such as English. We also introduce a novel approach, called parent-based constituent context model (PCCM), and show that by using some history notion of context and constituent information of each span's parent, the performance of CCM, especially in dealing with FWOLs, can be significantly improved.
This research shows a new approach and development of a design methodology, based on the perspective of meanings. In this study the design process is explored as a development of the structure of meanings. The processes of search and evaluation of meanings form the foundations of developing this structure. In order to facilitate the use and operation of the meanings, the WordNet lexical database and an existing visualization of WordNet — Visuwords — is used for the process of meaning search. The basic tool used for evaluation process is the WordNet::Similarity software, measuring the relatedness of meanings in the database. In this way it is measuring the degree of interconnections between different meanings. This kind of search and evaluation techniques are later on incorporated into our methodology of the structure of meanings to support the design process. The measures of relatedness of meanings are developed as convergence criteria for application in the processes of evaluation. Further on, the methodology for the structure of meanings developed here is used to construct meanings in a verification of product design. The steps of the design methodology, including the search and evaluation processes involved in developing the structure of the meanings, are elucidated. The choices, made by the designer in terms of meanings are supported by consequent searches and evaluations of meanings to be implemented in the designed product. In conclusion, the paper presents directions for developing and further extensions of the proposed design methodology.
The Portal da Lingua Portuguesa is a website containing information about the Portuguese language oriented towards the general public. The largest part of the information on the Portal is lexical information concerning formal characteristics of words, such as orthography, derivations, loanwords and gentiles. The lexical information comes from a lexical database called MorDebe-or more precisely, a network of lexical databases called the Open Source Lexical Information Network (OSLIN). This abstract shows the general set-up of and major functions of MorDebe Admin, which is the lexicon management system for OSLIN. MorDebe Admin provides an easy and secure way of updating and editing the content of the different databases of OSLIN. Furthermore, much of the data on the Portal are organised as mini-dictionaries and MorDebe Admin provides an integrated collection of tools dedicated to the maintenance of these mini-dictionaries, as well as a built-in neologism tracking system. The software demonstration will illustrate these functions from a user perspective, and how easy ir is to maintain the data behind the Portal.
We present an initial ontology for tactical behaviors conducted by unmanned ground vehicles (UGVs). We focus on activities, which are the denotations of verbs, notably 'move' but also 'look (for)' and several others. These take collective subjects, allowing activities to be attributed to units at various hierarchical levels. The semantics of verbs must consider the denotations of their grammatical complements; that is, we must consider entire verb frames. The thematic relations of the noun-phrase complements are critical, but prepositions also play an important role. FrameNet is an online lexical database of frames derived from text corpora. Our other major resource is Levin's classification of verbs according to how changes in their frames affect their meanings. Although natural languages have a large variety of words for aspects of tactical behaviors, there is motivation to get by with as few basic verbs as possible. A variety of meanings can often be associated with a verb by altering its frame, and we can impose co-reference constraints on combinations of frames to generate structures denoting more complex activities. A simple grammar is developed for the verbs of interest. Protege-Frames ontologies include classes that inherit from linguistically inspired classes but capture domain-specific notions.
This article considers three periods of the theory of predication in the Prague Linguistic Circle. The first belongs to the classical period, the second to the 1960s and the final and current one begins in the 1990s. The work of three particular authors, Mathesius, DaneÅ¡, and Sgall, is discussed. Four questions arise: 1) Was Mathesius an inspirational source for the second 1929 Thesis? 2) What were the historical and epistemological roots of the rather strong link between philosophy of language and linguistics during the first period? 3) Why is indexicality not recognized by DaneÅ¡ as the semiotic device which associates predication with those aspects of the sentence which potentially place it in relation with a situation? 4)Why is the relationship between the Praguian idea of âfunctionâ and the Fregean one, i.e. the logico-semantic concept which gives foundation to the philosophical truth-functional approach in semantics, and which explains the so-called âpredicative-argumental structureâ, not recognized as noteworthy? What, in fact, is\na function? In conclusion, this article sketches the project of a Latin Dependency Treebank developed on the basis of the PDT, which bears witness to the far-reaching implications of the Circleâs work and to an international, intercontinental prosecution of a tâche abordée.
Socio-economic decisions are commonly explained by rational cost vs. benefit considerations, whereas person variables have not usually been considered. The present study aims at investigating the degree to which dispositional power motivation and affective states predict socio-economic decisions. The power motive was assessed both indirectly and directly using a TAT-like picture test and a power motive self-report, respectively. After nine months, 62 students completed an affect rating and performed on a money allocation task (Social Values Questionnaire). We hypothesized and confirmed that dispositional power should be associated with a tendency to maximize one's profit but to care less about another party's profit. Additionally, positive affect showed effects in the same direction. The results are discussed with respect to a motivational approach explaining socio-economic behaviour.