Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Psycholinguistic research shows that key properties of the human sentence processor are incrementality, connectedness (partial structures contain no unattached nodes), and prediction (upcoming syntactic structure is anticipated). There is currently no broad-coverage parsing model with these properties, however. In this article, we present the first broad-coverage probabilistic parser for PLTAG, a variant of TAG that supports all three requirements. We train our parser on a TAG-transformed version of the Penn Treebank and show that it achieves performance comparable to existing TAG parsers that are incremental but not predictive. We also use our PLTAG model to predict human reading times, demonstrating a better fit on the Dundee eye-tracking corpus than a standard surprisal model.
We present a novel method ("waste") for the segmentation of text into tokens and sentences. Our approach makes use of a Hidden Markov Model for the detection of segment boundaries. Model parameters can be estimated from pre-segmented text which is widely available in the form of treebanks or aligned multi-lingual corpora. We formally define the waste boundary detection model and evaluate the system's performance on corpora from various languages as well as a small corpus of computer-mediated communication.
We present an affective text analysis model that can directly estimate and combine affective ratings of multi-word terms, with application to the problem of sentence polarity/semantic orientation detection. Starting from a hierarchical compositional method for generating sentence ratings, we expand the model by adding multi-word terms that can capture non-compositional semantics. The method operates similarly to a bigram language model, using bigram terms or backing off to unigrams based on a (degree of) compositionality criterion. The affective ratings for n-gram terms of different orders are estimated via a corpus-based method using distributional semantic similarity metrics between unseen words and a set of seed words. N-gram ratings are then combined into sentence ratings via simple algebraic formulas. The proposed framework produces state-of-the-art results for word-level tasks in English and German and the sentence-level news headlines classification SemEval'07-Task14 task. The inclusion of bigram terms to the model provides significant performance improvement, even if no term selection is applied.
We present a reformulation of the word pair features typically used for the task of disambiguating implicit relations in the Penn Discourse Treebank. Our word pair features achieve significantly higher performance than the previous formulation when evaluated without additional features. In addition, we present results for a full system using additional features which achieves close to state of the art performance without resorting to gold syntactic parses or to context outside the relation.
People believe that women are more emotionally intense than men, but the scientific evidence is equivocal. In this study, we tested the novel hypothesis that men and women differ in the neural correlates of affective experience, rather than in the intensity of neural activity, with women being more internally (interoceptively) focused and men being more externally (visually) focused. Adult men (n = 17) and women (n = 17) completed a functional magnetic resonance imaging study while viewing affectively potent images and rating their moment-to-moment feelings of subjective arousal. We found that men and women do not differ overall in their intensity of moment-to-moment affective experiences when viewing evocative images, but instead, as predicted, women showed a greater association between the momentary arousal ratings and neural responses in the anterior insula cortex, which represents bodily sensations, whereas men showed stronger correlations between their momentary arousal ratings and neural responses in the visual cortex. Men also showed enhanced functional connectivity between the dorsal anterior insula cortex and the dorsal anterior cingulate cortex, which constitutes the circuitry involved with regulating shifts of attention to the world. These results demonstrate that the same affective experience is realized differently in different people, such that women's feelings are relatively more self-focused, whereas men's feelings are relatively more world-focused.
This paper investigates the effect of the label bias problem of maximum entropy Markov models for part-of-speech tagging, a typical sequence prediction task in natural language processing. This problem has been underexploited and underappreciated. The investigation reveals useful information about the entropy of local transition probability distributions of the tagging model which enables us to exploit and quantify the label bias effect of part-of-speech tagging. Experiments on a Vietnamese treebank and on a French treebank show a significant effect of the label bias problem in both of the languages.
This study examines the impact of online word of mouth (WOM) and expert reviews on movies' box office revenues, both in the U.S. domestic market and in the international markets. Using a sample of 169 movies released in 2008, the study discovered that the frequency of online WOM and the valence rating of expert reviews were significant factors for box office outcomes in the domestic market. The study also found that only the frequency of online WOM was a significant factor in the international markets. The findings suggest that online WOM and expert reviews play a critical role in moviegoers' consumption behavior in the age of the Internet and social media.
Large-scale linguistically annotated cor-pora have played a crucial role in advanc-ing the state of the art of key natural lan-guage technologies such as syntactic, se-mantic and discourse analyzers, and they serve as training data as well as evaluation benchmarks. Up till now, however, most of the evaluation has been done on mono-lithic corpora such as the Penn Treebank, the Proposition Bank. As a result, it is still unclear how the state-of-the-art analyzers perform in general on data from a vari-ety of genres or domains. The completion of the OntoNotes corpus, a large-scale, multi-genre, multilingual corpus manually annotated with syntactic, semantic and discourse information, makes it possible to perform such an evaluation. This paper presents an analysis of the performance of publicly available, state-of-the-art tools on all layers and languages in the OntoNotes v5.0 corpus. This should set the bench-mark for future development of various NLP components in syntax and semantics, and possibly encourage research towards an integrated system that makes use of the various layers jointly to improve overall performance. 1
ABSTRACT Online travel reviews are emerging as a powerful source of information affecting tourists' pre-purchase evaluation of a hotel organization. This trend has highlighted the need for a greater understanding of the impact of online reviews on consumer attitudes and behaviors. In view of this need, we investigate the influence of online hotel reviews on consumers' attributions of service quality and firms' ability to control service delivery. An experimental design was used to examine the effects of four independent variables: framing; valence; ratings; and target. The results suggest that in reviews evaluating a hotel, remarks related to core services are more likely to induce positive service quality attributions. Recent reviews affect customers' attributions of controllability for service delivery, with negative reviews exerting an unfavorable influence on consumers' perceptions. The findings highlight the importance of managing the core service and the need for managers to act promptly in addressing customer service problems.
peer reviewed
The present paper focuses on ways in which the pragmatic (functional) meaning that arises from various contextual features, known in corpus linguistics as semantic prosody, can become an integral part of lexicographical descriptions as they are represented in the Slovene Lexical Database (SLD). This is particularly important for the treatment of phraseology and idiomatics. First, the theoretical background is provided, with the focus on the prototype theory and its practical implications for monolingual lexicography. A parallel is drawn with the model of meaning analysis in the SLD. The second part begins with a brief introduction to semantic prosody and continues with an analysis of monolingual meaning descriptions in the SLD against a number of authentic corpus examples, investigating how their pragmatic components have been identified. The analysis of corpus data shows that pragmatics is an important contributor to the process of sense discrimination in works of lexical and lexicographic relevance.
With the growing interest in statistical parsing, special attention has recently been devoted to the problem of comparing different treebanks to assess which languages or domains are more difficult to parse relative to a given model. A common methodology for comparing parsing difficulty across treebanks is based on the use of the standard labeled precision and recall measures. As an alternative, in this article we propose an information-theoretic measure, called the expected conditional cross-entropy (ECC). One important advantage with respect to standard performance measures is that ECC can be directly expressed as a function of the parameters of the model. We evaluate ECC across several treebanks for English, French, German, and Italian, and show that ECC is an effective measure of parsing difficulty, with an increase in ECC always accompanied by a degradation in parsing accuracy.
National audience
This paper discusses the extension of a sys-tem developed for automatic discovery of tree-bank annotation inconsistencies over an entire corpus to the particular case of evaluation of inter-annotator agreement. This system makes for a more informative IAA evaluation than other systems because it pinpoints the incon-sistencies and groups them by their structural types. We evaluate the system on two corpora- (1) a corpus of English web text, and (2) a corpus of Modern British English. 1
This paper presents our preliminary conclusions as part of an ongoing effort to construct a new dependency representation framework for Turkish.We aim for this new framework to accommodate the highly agglutinative morphology of Turkish as well as to allow the annotation of unedited web data, and shape our decisions around these considerations.In this paper, we firstly describe a novel syntactic representation for morphosyntactic sub-word units (namely inflectional groups (IGs) in Turkish) which allows inter-IG relations to be discerned with perfect accuracy without having to hide lexical information.Secondly, we investigate alternative annotation schemes for coordination structures and present a better scheme (nearly 11% increase in recall scores) than the one in Turkish Treebank (Oflazer et al., 2003) for both parsing accuracies and compatibility for colloquial language.
Although several syntactically annotated corpora (or treebanks) exist for Dutch, they are seldomly used for descriptive linguistic research because there are no easy-to-use exploitation tools available.This demonstration paper describes GrETEL, a linguistic search engine (http:// nederbooms.ccl.kuleuven.be/eng/gretel)that enables non-technical users to consult treebanks in a user-friendly way.Instead of a formal search expression, a natural language example is used as input to the system, allowing users to search for similar constructions as the example they provide.In the first version of GrETEL, only written Dutch (LASSY) was included.Based on user requests we have now included the Spoken Dutch Corpus (CGN) as well.
We present, here, our analysis of systematic divergences in parallel English-Hindi dependency treebanks based on the Computational Paninian Grammar (CPG) framework. Study of structural divergences in parallel treebanks not only helps in developing larger treebanks automatically, but can also be useful for many NLP applications such as data-driven machine translation (MT) systems. Given that the two treebanks are based on the same grammatical model, a study of divergences in them could be of advantage to such tasks, along with making it more interesting to study how and where they diverge. We consider two parallel trees divergent based on differences in constructions, relations marked, frequency of annotation labels and tree depth. Some interesting instances of structural divergences in the treebanks have been discussed in the course of this paper. We also present our task of alignment of the two treebanks, wherein we talk about our extraction of divergent structures in the trees, and discuss the results of this exercise. 1
This paper introduces an advanced, efficient approach for rule based English to Bengali (E2B) machine translation (MT), where Penn-Treebank parts of speech (PoS) tags, HMM (Hidden Markov Model) Tagger is used.Fuzzy-If-Then-Rule approach is used to select the lemma from rule-based-knowledge. The proposed E2B-MT has been tested through F-Score measurement, and the accuracy is more than eighty percent.
This paper investigates the appropriateness of using lexical cohesion analysis to assess Chinese readability. In addition to term frequency features, we derive features from the result of lexical chaining to capture the lexical cohesive information, where E-HowNet lexical database is used to compute semantic similarity between nouns with high word frequency. Classification models for assessing readability of Chinese text are learned from the features using support vector machines. We select articles from textbooks of elementary schools to train and test the classification models. The experiments compare the prediction results of different sets of features.
В статье рассматривается проблема нормы и нормативного подхода к языку в диахроническом плане.Определяется специфика нормативного похода к языковым средствам в различных лингвистических традициях и выявляются основные характеристики лингвистической нормы.В статье указывается, что на каждом этапе развития языка складываются свои нормы как резуль
Aiming at the area of machine translation applications,this paper conduct research on the construction of Chinese Sentence-Category Dependency Treebank(CSCDT) based on the theory of hierarchical network of concepts.Conceptual category tagset and sentence-category relation tagset for the treebank are presented also with the example tree of CSCDT.
In this paper, we provide a quantitative analysis of non-projective constructions attested in the Ancient Greek Dependency Treebank (AGDT). We consider the different types of formal constraints and metrics that have become standardized in the literature on non-projectivity (planarity, wellnestedness, gap-degree, edge-degree). We also discuss some of the linguistic factors that cause non-projective edges in Ancient Greek. Our results confirm the remarkable extension of non-projectivity in the AGDT, both in terms of quantitative incidence of non-projective nodes and for their complexity, which is not paralleled by the corpora of modern languages considered in the literature. At the same time, the usefulness of other constraint (especially well-nestedness) is confirmed by our researches. 1
In this paper, we propose a method for au-tomatic clause boundary annotation in the Hindi Dependency Treebank. We show that the clausal information implicitly encoded in a dependency structure can be made explicit with no or less human interven-tion. We exercised the proposed approach on 16,000 sentences of Hindi Dependency Treebank. Our approach gives an accuracy of 94.44 % for clause boundary identifica-tion evaluated over 238 clauses. The resul-tant corpus has varied usages and can be utilized for developing a statistical clause boundary identifier. 1
We investigate statistical dependency parsing of two closely related languages, Croatian and Serbian.As these two morphologically complex languages of relaxed word order are generally under-resourced -with the topic of dependency parsing still largely unaddressed, especially for Serbian -we make use of the two available dependency treebanks of Croatian to produce state-of-the-art parsing models for both languages.We observe parsing accuracy on four test sets from two domains.We give insight into overall parser performance for Croatian and Serbian, impact of preprocessing for lemmas and morphosyntactic tags and influence of selected morphosyntactic features on parsing accuracy.
To facilitate future research in unsupervised induction of syntactic structure and to standardize best-practices, we propose a tagset that consists of twelve universal part-of-speech categories. In addition to the tagset, we develop a mapping from 25 different treebank tagsets to this universal set. As a result, when combined with the original treebank data, this universal tagset and mapping produce a dataset consisting of common parts-of-speech for 22 different languages. We highlight the use of this resource via two experiments, including one that reports competitive accuracies for unsupervised grammar induction without gold standard part-of-speech tags.
We present results of a study investigating evaluative learning in dementia patients with a classic evaluative conditioning paradigm. Picture pairs of three unfamiliar faces with liked, disliked, or neutral faces, that were rated prior to the presentation, were presented 10 times each to a group of dementia patients (N = 15) and healthy controls (N = 14) in random order. Valence ratings of all faces were assessed before and after presentation. In contrast to controls, dementia patients changed their valence ratings of unfamiliar faces according to their pairing with either a liked or disliked face, although they were not able to explicitly assign the picture pairs after the presentation. Our finding suggests preserved evaluative conditioning in dementia patients. However, the result has to be considered preliminary, as it is unclear which factors prevented the predicted rating changes in the expected direction in the control group.
Chinese word structure annotation is potentially useful for many NLP tasks, especially for Chinese word segmentation. Li and Zhou (2012) have presented an annotation for word structures in the Penn Chinese Treebank. But they only consider words that have productive affixes, which covers 35% of word types in that corpus. In this paper, we propose a linguistically inspired annotation that covers various morphological derivations of Chinese in a more general way, such that almost all multiple-character words can be structurally analyzed. As manual annotation is expensive, we propose a semi-supervised approach to automatic annotation, which combines the maximum entropy learning and the EM iteration for the Gaussian mixture model. The proposed method has achieved an accuracy of 90% on the testing set. © 2021 CLP 2012 - 2nd CIPS-SIGHAN Joint Conference on Chinese Language Processing. All Rights Reserved.
Cet article décrit les étapes qui composent notre analyse du discours, en partant du texte brut, et pour en produire une représentation sémantique dans le cadre de la Discourse Representation Theory, désormais DRT (Kamp and Reyle, 1993). Une chaîne complète de traitement est proposée et testée sur le corpus Itipy, "Itinéraires Pyrénéens", lequel a été proposé par la médiathèque de Pau. Le premier but applicatif consiste à attacher un lieu aux portions de texte narrant une action dans ce lieu. Nous exploitons alors ce corpus de récits de voyage du XIXème siècle dans l'objectif d'extraire automatiquement les itinéraires décrits et afin d'indexer les portions de texte prenant effectivement pour décors les lieux géographiques en question. Notre outil, Grail est un parser pour grammaire logique de types avec un ensemble restreint de règles fixes et utilisant un lexique riche. Tout d'abord, la première phase a consisté en l'acquisition de la grammaire sur un corpus annoté (Paris 7 Treebank). Ce corpus nous a permis d'obtenir les informations grammaticales propres aux unités du lexique de la langue française présentes dans le corpus, le lexique produit ne contient donc pas la totalité des mots du français et contient plusieurs catégories pour les entrées les plus fréquentes. Dans la chaine de traitement, la méthode d'attribution de la catégorie intègre une approche statistique: lors- qu'un mot est absent du lexique, l'analyse propose une catégorie ou lorsqu'il présente plusieurs catégories possibles, elle sélectionne la plus appropriée. Chaque mot du texte est taggé, puis supertaggé en fonction des autres unités se trouvant dans son contexte proche (la phrase). Le supertagger propose plusieurs formules qui correspondent à une analyse syntaxique partielle pour chaque phrase du texte dans le cadre des grammaires catégorielles, et plus précisément du calcul de Lambek. S'ensuit une étape de combinaison de toutes les analyses partielles pour donner l'analyse globale. La structure obtenant la meilleure probabilité étant sélectionnée, on garde cette structure comme organisation du calcul de la représentation sémantique en fonction des unités qui la composent. On associe alors à chaque mot son λ-terme à partir du lexique sémantique cette fois et dont la formule correspond à celle présente dans le lexique grammatical pour cette même entrée (Moot, 2010). Le λ -terme pour chaque unité sémantique est saisi à la main dans le style de la λ -DRT. La représentation sémantique étant produite automatiquement à partir de l'analyse syntaxique, nous obtenons une représentation logique sémantique bien formée. La dimension pragmatique quant à elle ne peut être reléguée à un plan inférieur dans l'interprétation du discours. En effet, une analyse du discours impose de fait une interaction entre la sémantique des unités de langue dont on doit interpréter le sens en discours et la prise en compte de la dimension pragmatique de ce qui est dit. Notre approche s'inspire de l'approche de Busquets et al. (2001), "une théorie de l'interprétation des discours doit être aussi en fait une théorie de la sémantique, de la pragmatique, et de leur interaction, c'est-à-dire une théorie de l'interface pragmatique-sémantique". Certains phénomènes sémantiques restent cependant difficiles à traiter, certains cas de glissement de sens montrent qu'une flexibilité dans le typage doit être permise, alors que dans les cas les plus courants le typage doit être rigide pour éviter une repré- sentation inappropriée. Nous donnerons quelques exemples à propos et proposons donc afin d'améliorer les résultats de notre chaîne traitement de traiter ces phénomènes par l'affinement des λ -termes du lexique dans le cadre du système F, λ -calcul d'ordre supérieur. Nous détaillerons ici notre corpus et nos objectifs applicatifs quant à celui-ci, nous présenterons les étapes de traitement du discours, commençant par l'acquisition de la grammaire du français sur corpus annoté, puis l'analyse syntaxique dans le cadre des grammaires catégorielles. Nous expliquerons plus amplement l'interface syntaxe-sémantique dans la théorie des types logiques permettant la construction de nos repré- sentations sémantiques en λ-DRT. Nous présenterons le système F et notre traitement des phénomènes discursifs mettant en jeu l'interaction sémantique-pragmatique puis nous présenterons les perspectives de ce travail.
Evidence on the impact of nature images has been found in research with hospital patients (Ulrich, 2008, Nanda, Hathorn & Neumann, 2007). The use of art in healthcare environments has become increasingly common (Nanda, Eisen & Baladandayuthapani, 2008). Art is viewed as a positive distraction from stress of the hospital among patients and possibly staff (Ulrich et. al. 1991; Ulrich, Zimring, Quan, & Joseph, 2006). In a previous study art preference study (Nanda, Eisen & Baladandayuthapani, 2008) showed significant difference in the ratings of design students and patients. Findings showed that there was a significant difference in the ratings of the two groups. Furthermore, the emotional rating scale (how does the art picture make you feel) was highly correlated to the selection scale (would you put this art picture in your room) for hospital patients- while this was not the case with the design students. What is the role of culture in the above questions and in how does it impact healthcare design?A total of more than 600 design and non-design students from National University of Mexico, National University of Singapore and University of Texas San Antonio rated images of visual art included abstract, representational and nature images from Mexico, Singapore and Texas representative of the unique cultural contexts, in addition to images that strictly adhere to the evidence-based guidelines for healthcare art laid down by Ulrich & Gilpin (2003) and examples of classic high art.At the end of the survey students re-rated the images again as if they were hospitalized and lying in a patient room. An analysis of preferences across cultures, design disciplines and emotion and selection was undertaken.Results show a surprising amount of agreement across cultures on image rating for hospital rooms. Level of agreement for art selection for personal rooms is significantly lower. This is true in both design and non-design students. Landscapes with a high depth of field, bright colors and verdant foliage were rated consistently high across all cultures, regardless of indigenous elements, with few exceptions, that suggests that there is a certain universal appeal for restorative images of nature that go beyond cultural and educational boundaries. The study showed that empathy (how this art would make you feel) is a stronger determinant of selection than culture, or education, when it comes to art selection for hospitals.
Semantic verbal fluency (SVF) often shows early and disproportionate decline in AD relative to other language, attention, and executive abilities. Successful performance on SVF depends on the ability to organize conceptual information into related clusters and efficiently access these clusters. Current methods for clustering and switching assessment are labor-intensive and subjective. We developed an automated computational linguistic approach to quantify the semantic content of SVF responses. Neuropsychological and resting state fMRI data were obtained from the work-up of 52 patients presenting to the Minneapolis VAMC GRECC Memory Loss Clinic. Participants included had a clinical diagnosis of MCI or AD, a completed MRI protocol with good quality data, and neuropsychological evaluation including the SVF task (animals). Imaging data were collected on a Philips 1.5T system at the Minneapolis VAMC. Semantic indices based on pairs of words on the SVF task were quantified in two ways: based on the length of their hierarchical relations in WordNet, an electronic lexical database of English (similarity), or calculated using a computerized algorithm based on a variant of principal components analysis (relatedness). Mean cumulative and sequential indices were produced for each method: cumulative similarity and relatedness were computed between all possible pairs of words produced regardless of order; and sequential similarity and relatedness were computed only between pairs of adjacent words. Higher scores reflect larger clusters and reduced switching. Several resting state fMRI network measures were related to the four automated semantic fluency indices. Nodal diversity, local efficiency, and the mean clustering coefficient, were all significantly correlated with cumulative and sequential measures of semantic similarity and relatedness. Pearson r values ranged from.323-.407 with corresponding p-values of.012-.003. All correlations survived multiple comparison correction. The traditional SVF score was not significantly related to imaging indices. We found that computational linguistic measurements of similarity and relatedness were significantly related to network measures obtained from resting state fMRI. These results suggest automated assessment of SVF has correlates with brain function in MCI and AD, and may outperform the traditional SVF score. This approach provides an easy way to standardize clustering and switching assessment without adding burden.
Medical discharge documents are summaries written by a physician about the patient’s condition and aim at transferring information to other health care personnel but also to the patient. According to the legislation, the patient should be able to understand the document. In practice, however, this has been shown to be problematic. This paper studies discharge documents from the patients’ perspective and examines how they fulfil the legislation’s demands on understandability. Concentrating on the vocabulary of the texts, we analyse the frequency of domainadapted terms, abbreviations and foreign words. The material consists of 23 528 heart patients’ discharge documents (5 747 126 words). The analysis is performed with the morphological analyser FinTWOL (http://www2.lingsoft.fi/cgi-bin/fintwol). Altogether, FinTWOL analyses 24% of the corpus as unknown or foreign words, abbreviations or medical terms. The most common category, unknown words, includes misspellings and medical terms, such as l.dex. Of these, 100 most common cover for 43% of the total. These terms thus seem to be relatively fixed. Of the words analysed as abbreviations, some are common also in standard language, but others are still very domain-specific, such as I.V. (intravenous). Also the used abbreviations are very fixed: the 100 most common ones cover for 94% of the total. This, however, does not help the patient who probably reads only one document. Similarly, even though misspellings are globally infrequent, they still occur more than once per document. In order to place the obtained results in a context, we performed a similar analysis on general Finnish university newspaper text from Turku Dependency Treebank. In comparison with the 24% obtained with the discharge documents, from the total of 10 687 words, 8,6% were given a special tag. The results show that that terms and abbreviations are considerably more used in discharge documents than in general newspaper text. It is clear that a text with such a vocabulary is domain-specific and distinct from the language that the patient is used to. Also e.g. the varying use of upper and lower case letters (dg and DG for diagnosis) emphasize the particularity of the language. In standard language texts such writing would not be acceptable. Standard writing would, however, help the patients to better understand the texts.
We often use tactile-input in order to recognize familiar objects and to acquire information about unfamiliar ones. We also use our hands to manipulate objects and utilize them as tools. However, research on object affordances has mainly been focused on visual-input and, thus, limiting the level of detail one can get about object features and uses. In addition to the limited multisensory-input, data on object affordances has also been hindered by limited participant input (e.g., naming task). In order to address the above mention limitations, we aimed at identifying a new methodology for obtaining undirected, rich information regarding people’s perception of a given object and the uses it can afford without necessarily viewing the particular object. Specifically, 40 participants were video-recorded in a three-block experiment. During the experiment, participants were exposed to pictures of objects, pictures of someone holding the objects, and the actual objects and they were allowed to provide unconstrained verbal responses on the description and possible uses of the stimuli presented. The stimuli presented were lithic tools given the: novelty, man-made design, design for specific use/action, and absence of functional knowledge and movement associations. The experiment resulted in a large linguistic database, which was linguistically analyzed following a response-based specification. Analysis of the data revealed significant contribution of visual- and tactile-input in naming and definition of object-attributes (color/condition/shape/size/texture/weight), while no significant tactile-information was obtained for object-features of material, visual-pattern, and volume. Overall, this new approach highlights the importance of multisensory-input in the study of object affordances.
The task of automatic machine translation (MT) is the focus of a huge variety of active research efforts, both because of the intrinsic utility of this difficult task, and the theoretical and linguistic insights that arise from modeling relationships between natural languages. However, MT systems that leverage syntactic information are only recently becoming practical, and in a typical system of this sort, syntactic information is generated by monolingual parsers; the task of explicitly modeling syntactic relationships between target and source languages is yet to be fully explored. This thesis investigates the problem of finding syntactic parse trees of target and/or source sentences that are more appropriate for use in a syntactic MT system. Two basic methodologies are explored. First, we present a sequence of two statistical models that leverage bilingual information to improve the linguistic quality of syntactic parses, as measured by their ability to replicate human-generated gold-standard annotations. The first model uses word to word alignments as an external source of information, while the second models the alignments jointly. These models are both quite effective at improving the intrinsic quality of the parse trees, and the second model additionally improves word alignment performance. However, while the two models achieve similar parsing improvements, we find that improving parses in conjunction with word alignments is much more helpful for the downstream machine translation task. In the next part of the thesis, we explore this finding further by investigating the effects on MT performance of agreement between parse trees and word alignments. We present a simple method for transforming input trees in a way that ignores gold-standard annotations, concentrating instead on improving syntactic agreement directly. In experiments, we find that though we obviously lose fidelity to more linguistically informed treebank annotation guidelines, this transformation-based approach yields the strongest improvements in syntactic machine translation.
This dissertation explores the nature and extent of retroflex consonant harmony in South Asia. Using statistics calculated over lexical databases from a broad sample of languages, the study demonstrates that retroflex consonant harmony is an areal trait affecting most languages in the northern half of the South Asian subcontinent, including languages from at least three of the four major families in the region: Dravidian, Indo-Aryan and Munda (but not Tibeto-Burman). Dravidian and Indo-Aryan languages in the southern half of the subcontinent do not exhibit retroflex consonant harmony. In South Asia, retroflex consonant harmony is manifested primarily as a static co-occurrence restriction on coronal consonants in roots/words. Historical-comparative evidence reveals that this pattern is the result of retroflex assimilation that is non-local, regressive and conditioned by the similarity of interacting segments. These typological properties stand in contrast to those of other retroflex assimilation patterns, which are local, primarily progressive, and not conditioned by similarity. This is argued to support the hypothesis that local feature spreading and long-distance feature agreement constitute two independent mechanisms of assimilation, each with its own set of typological properties, and that retroflex consonant harmony is the product of agreement, not spreading. Building on this hypothesis, the study offers a formal account of retroflex consonant harmony within the Agreement by Correspondence (ABC) model of Rose & Walker (2004) and Hansson (2001; 2010). Two Indo-Aryan languages, Kalasha and Indus Kohistani, figure prominently throughout the dissertation. These languages exhibit similarity effects that have not been clearly observed in other retroflex consonant harmony systems; retroflexion is contrastive in both non-sibilant (i.e., plosive) and sibilant obstruents (i.e., affricates and fricatives), but harmony applies only within each manner class, not between them. At the same time, harmony is not sensitive to laryngeal features. Theoretical implications of these and other similarity effects are discussed.
New Irish speakers in Belfast play a crucial, complex part in the revitalization and change of both the city and Irish within Northern Ireland. This paper examines the role of new Irish speakers in transforming Belfast, whose emergence from a post-conflict period involves a reassessment of communal cultural expressions. Markers of ethno-national identity are bitterly contentious locally, and yet increasingly celebrated, in line with international trends, as high status cultural forms and potentially profitable tourist attractions. Irish in Belfast currently occupies an ambiguous position: divisive enough for a sign reading ‘Happy Christmas’ in Irish to be experienced as an insult by some city councillors, yet a secure enough part of the establishment for a neighbourhood to be officially rebranded as the Gaeltacht Quarter. <br/>When, how and where new Irish speakers use the language in Belfast has implications for the relationship of Irishness to the Northern Irish state and for the place of Belfast within regional frameworks across the UK, Ireland and Europe. Adult learners and young people exiting Irish medium education have an impact on life in Belfast beyond its small population of Irish speakers. Urbanisation fuelled by new speakers, which shifts the balance of Irish language resources and speakers away from traditional rural Gaeltacht areas and towards cities, also has implications for the language itself. Recent increase in new Irish speakers in Belfast is due to expansion in the Irish-medium sector as well as to adult learners, whose decisions contribute to the school expansion. <br/>Urbanisation, multilingualism and intergenerational shift combine in Belfast to produce new linguistic norms. Moreover, in a minority language community where hierarchies of ‘authenticity’ are weighted towards the rural and the native speaker, where the rural and the native have traditionally been conflated, and where indigeneity is a central concept to contested nationalisms, the emergence of a self-confident, youthful Irish speaking community in Northern Ireland’s biggest city involves a recalibration of the qualities signifying ‘gaelicness’. As students, professionals, hobbyists and activists, new Irish speakers in Belfast occupy a vital position at the crux of changing ideas about place, language and identity.<br/>
Early-latency theories of emotional processing state that at least coarse monitoring of the emotional valence (a pleasure-displeasure continuum) of facial expressions should be both rapid and highly automated (LeDoux, 1995; Russell, 1980). Research has largely substantiated early-latency differential processing of emotional versus non-emotional facial expressions; however, the effect of valence on early-latency processing of emotional facial expression remains unclear. In an effort to delineate the effects of valence on early-latency emotional facial expression processing, the current investigation compared ERP responses to positive (happy and surprise), neutral, and negative (afraid and sad) basic facial expression photographs as well as to positive (happy-surprise), neutral (afraid-surprise, happy-afraid, happy-sad, sad-surprise), and negative (sad-afraid) morph facial expression photographs during a valence-rating task. Morphing manipulations have been shown to decrease the familiarity of facial patterns and thus preclude any overlearned responses to specific facial codes. Accordingly, it was proposed that morph stimuli would disrupt more detailed emotional identification to reveal a valence response independent of a specific identifiable emotion (Balconi & Lucchiari, 2005; Schweinberger, Burton & Kelly, 1999). ERP results revealed early-latency differentiation between positive, neutral, and negative morph facial expressions approximately 108 milliseconds post-stimulus (P1) within the right electrode cluster; negative morph facial expressions continued to elicit significantly smaller ERP amplitudes than other valence categories approximately 164 milliseconds post-stimulus (N170). Consistent with previous imaging research on emotional facial expression processing, source localization revealed substantial dipole activation within regions of the mesolimbic dopamine system. Thus, these findings confirm rapid valence processing of facial expressions and suggest that negative valence processing may continue to modulate subsequent structural facial processing.
Individuals who effectively regulate or mildly increase their systolic blood pressure (SBP) in response to an orthostatic challenge exhibit healthier affective status, cognitive functioning, and better quality of life. Thus, increased SBP in response to an orthostatic challenge serves as a proxy for several underlying changes. This study examined the relationship between SBP regulation and self-esteem in children. Data were collected from 92 boys and girls, aged 8–11 years. Systolic, diastolic, and pulse measurements were obtained after 5 minutes of remaining supine and again after 1 minute of standing. Children also provided affective ratings on the Children's Depression Inventory. The Negative Self-Esteem subscale was examined for this study. A multiple regression analysis revealed that poorer orthostatic regulation was associated with higher levels of negative self-esteem among children aged 8–11 years. Thus, orthostatic BP regulation may serve as a biological marker for poor self-esteem in children. This may have further implications for children's emotional functioning as low self-esteem may serve as a risk factor for future negative affective states.
Alexithymia is a personality trait characterised by difficulties in identifying and describing one’s emotions, constricted imaginal processing, and an externally oriented cognitive style. Alexithymia is associated with psychopathology and interpersonal problems. The aim of the current study was to evaluate the psychometric properties of the most frequently used measure of alexithymia, the self-report 20-item Toronto Alexithymia Scale (TAS-20). Specifically, the study aimed to (1) cross-validate the hypothesised three-factor structure of the TAS-20, (2) determine whether the measure indirectly assesses the constricted imaginal thinking component of alexithymia, despite its absence of imagination items, (3) examine the overlap between the TAS-20 and measures of psychopathology, and (4) determine whether the TAS-20 assesses actual, rather than merely perceived, emotional understanding. Participants were 194 (138 female) university students and community members who completed an online survey. Confirmatory factor analyses showed mixed support for the hypothesised three-factor model of the TAS-20; however, this model provided a better fit to the data than either a one- or two-factor model. Inverse relationships were found between the TAS-20 and measures of perspective-taking and fantasy (although this relationship was only marginally significant for fantasy). There were moderate to large positive associations between the TAS-20 and measures of depression, anxiety, stress, and negative affectivity. Inverse relationships were found between the TAS-20 and objective measures of emotional ability; however, these relationships were no longer significant after the effects of negative emotions and affectivity were partialled out. Higher TAS-20 scores were also associated with more moderate affective valence ratings of emotion-evoking stimuli, and thus lower self-reported arousal. Together, these findings suggest problems with the TAS-20’s construct validity. Theoretical and practical implications are discussed.
This paper focuses on the links between contemporary literature and the various positions choosenchosenby authors facing the problematics of translation. Beginning with the observation that translation studies should develop from a theoretical point of view in Japan--an emblematic country for translations--this paper shows that currently, translation in Japan has to be considered as a cultural exportation trend and not only as the importation trend that dominated the cultural scene during the 20th century. For example, data on published translations in France show that since 2007, Japanese is the second most frequently translated language after American-English--due to the popularity of mangas in France. In the literary field, new phenomenons can also be observed in Japan. In this paper, four case studies are presented. The most remarkable case concerns Murakami Haruki's strategy, in which he, being an important translator of the Great American Novel, crosses the boundaries between countries and languages in order to represent a new kind of nationless writer, i.e. a global writer appreciated all over the world. On the other hand, Mizumura Minae mixes English and Japanese in her I novel from left to right, making it untranslatable into English. This for her represents the resistance of a minor language, Japanese, to the domination of English. Tawada Yôko, for her part, writes in two languages, Japanese and German, and in doing so tries to deconstruct both cultural and linguistic norms, enhancing translation as an impossible tool. Finally, the American-born Hideo Levy's three-piece band features Japanese, English and Chinese members, interconnected by the belief in translation as an ideal vector of communication. All these new streams contribute to the reshifting of Japanese literature in the world and induce a necessary renewal of the critical approaches.
The explosion of information in the World Wide Web is overwhelming readers with limitless information. Large internet articles or journals are often cumbersome to read as well as comprehend. More often than not, readers are immersed in a pool of information with limited time to assimilate all of the articles. It leads to information overload whereby readers are trying to deal with more information than they can process. Hence, there is an apparent need for an automatic text summarizer as to produce summaries quicker than humans. The text summarization research on mobile platform has been inspired by the new paradigm shift in accessing information ubiquitously at anytime and anywhere on Smartphones or smart devices. In this research, a semantic and syntactic based summarization is implemented in a text summarizer to solve the overload problem whilst providing a more coherent summary. Additionally, WordNet is used as the lexical database to semantically extract the text document which provides a more efficient and accurate algorithm than the existing summary system. The objective of the paper is to integrate WordNet into the proposed system called TextSumIt which condenses lengthy documents into shorter summarized text that gives a higher readability to Android mobile users. The experimental results are done using recall, precision and F-Score to evaluate on the summary output, in comparison with the existing automated summarizer. Human-generated summaries from Document Understanding Conference (DUC) are taken as the reference summaries for the evaluation. The evaluation of experimental results shows satisfactory results.
Statistické jazykové modely jsou důležitou součástí mnoha úspěšných aplikací, mezi něž patří například automatické rozpoznávání řeči a strojový překlad (příkladem je známá aplikace Google Translate). Tradiční techniky pro odhad těchto modelů jsou založeny na tzv. N-gramech. Navzdory známým nedostatkům těchto technik a obrovskému úsilí výzkumných skupin napříč mnoha oblastmi (rozpoznávání řeči, automatický překlad, neuroscience, umělá inteligence, zpracování přirozeného jazyka, komprese dat, psychologie atd.), N-gramy v podstatě zůstaly nejúspěšnější technikou. Cílem této práce je prezentace několika architektur jazykových modelůzaložených na neuronových sítích. Ačkoliv jsou tyto modely výpočetně náročnější než N-gramové modely, s technikami vyvinutými v této práci je možné jejich efektivní použití v reálných aplikacích. Dosažené snížení počtu chyb při rozpoznávání řeči oproti nejlepším N-gramovým modelům dosahuje 20%. Model založený na rekurentní neurovové síti dosahuje nejlepších publikovaných výsledků na velmi známé datové sadě (Penn Treebank).