Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
We present “CDG Lab”, an integrated environment for development of dependency grammars and treebanks. It uses the Categorial Dependency Grammars (CDG) as a formal model of dependency grammars. CDG are very expressive. They generate unlimited dependency structures, are analyzed in polynomial time and are conservatively extendable by regular type expressions without loss of parsing efficiency. Due to these features, they are well adapted to definition of large scale grammars. CDG Lab supports the analysis of correctness of treebanks developed in parallel with evolving grammars.
Data-driven research in linguistics typically involves the processes of data annotation, data visualization and identification of relevant patterns. We describe our experience in incorporating these processes at an undergraduate course on language information technology. Students collectively annotated the syntactic structures of a set of Classical Chinese poems; the resulting treebank was put on a platform for corpus search and visualization; finally, using this platform, students investigated research questions about the text of the treebank. 1
This paper investigates the impact on French dependency parsing of lexical generalization methods beyond lemmatization and morphological analysis. A distributional thesaurus is created from a large text corpus and used for distributional clustering and WordNet automatic sense ranking. The standard approach for lexical generalization in parsing is to map a word to a single generalized class, either replacing the word with the class or adding a new feature for the class. We use a richer framework that allows for probabilistic generalization, with a word represented as a probability distribution over a space of generalized classes: lemmas, clusters, or synsets. Probabilistic lexical information is introduced into parser feature vectors by modifying the weights of lexical features. We obtain improvements in parsing accuracy with some lexical generalization configurations in experiments run on the French Treebank and two out-of-domain treebanks, with slightly better performance for the probabilistic lexical generalization approach compared to the standard single-mapping approach. 1
This study investigates whether age and/or hearing loss influence the perception of the emotion dimensions arousal (calm vs. aroused) and valence (positive vs. negative attitude) in conversational speech fragments. Specifically, this study focuses on the relationship between participants' ratings of affective speech and acoustic parameters known to be associated with arousal and valence (mean F0, intensity, and articulation rate). Ten normal-hearing younger and ten older adults with varying hearing loss were tested on two rating tasks. Stimuli consisted of short sentences taken from a corpus of conversational affective speech. In both rating tasks, participants estimated the value of the emotion dimension at hand using a 5-point scale. For arousal, higher intensity was generally associated with higher arousal in both age groups. Compared to younger participants, older participants rated the utterances as less aroused, and showed a smaller effect of intensity on their arousal ratings. For valence, higher mean F0 was associated with more negative ratings in both age groups. Generally, age group differences in rating affective utterances may not relate to age group differences in hearing loss, but rather to other differences between the age groups, as older participants' rating patterns were not associated with their individual hearing loss.
BACKGROUND: In Alzheimer's disease (AD), some patients present with cognitive impairment other than episodic memory disturbances. We evaluated whether occurrence of posterior atrophy (PA) and medial temporal lobe atrophy (MTA) could account for differences in cognitive domains affected. METHODS: In 329 patients with AD, we assessed five cognitive domains: memory, language, visuospatial functioning, executive functioning, and attention. Magnetic resonance imaging (MRI) was rated visually for the presence of MTA and PA. Two-way analyses of variance were performed with MTA and PA as independent variables, and cognitive domains as dependent variables. Gender, age, and education were covariates. As PA is often encountered in younger patients, analyses were repeated after stratification for age of onset (early onset, ≤65 years). RESULTS: The mean age of the participants was 67 years, 175 (53%) were female, and the mean Mini-Mental State Examination (score±standard deviation) was 20±5 points. Based on dichotomized magnetic resonance imaging ratings, 84 patients (26%) had MTA and PA, 98 (30%) had MTA, 57 (17%) had PA, and 90 (27%) had neither. MTA was associated with worse performance on memory, language, and attention (all, P<.05), and PA was associated with worse performance on visuospatial and executive functioning (both, P<.05). Stratification for age showed in patients with late-onset AD (n=173) associations between MTA and impairment on memory, language, visuospatial functioning, and attention (all, P<.05); in early-onset AD (n=156), patients with PA tended to perform worse on visuospatial functioning. CONCLUSIONS: Regional atrophy is related to impairment in specific cognitive domains in AD. The prevalence of PA in a large set of patients with AD and its association with cognitive functioning provides support for the usefulness of this visual rating scale in the diagnostic evaluation of AD.
Based on Mongolian dependency grammar,this paper develops the semantic role classification method and designs the tag-sets of Mongolian language from the perspective of Mongolian information processing.A manual annotation is carried on certain amount of Mongolian Syntactic Dependency Treebank by focusing on semantic role of Mongolian with a reference to the semantic role labeling theory and methods in other languages.
Please note: This article is in Greek. Compiling a dialectal dictionary: the “Syntychies” lexical database: This paper introduces\nthe reader to the issues of making an online dialectal dictionary, presenting some\nof the matters that have arisen while producing a lexical database of the Cypriot Greek\ndialect. Most problems related to the selection of data and compilation of lemmas were\ncaused by the great variation in orthographic and/or morphological representation of\nCypriot word forms. The database which has been created as part of the “Syntychies”\nresearch program is available on the website http://lexcy.library.ucy.ac.cy. The choices\nthat have been adopted in this database after a lexical analysis of a large amount of data\noutline a framework for compiling other dialectical dictionaries of Greek.
The fundamental goal of this dissertation is to establish that deep, efficient, accurate parsing models can be acquired for Chinese, through parsers founded on Combinatory Categorial Grammar (CCG), a grammar formalism which has already enabled the creation of rich parsing models for English. We harness these CCG analyses of cross-linguistic syntax, harmonising them with modern accounts from Chinese generative syntax, contributing the first analysis of Chinese syntax through CCG in the literature. Supervised statistical parsing approaches rely on the availability of large annotated corpora. To avoid the cost of manual annotation, we adopt the corpus conversion methodology, in which an automatic corpus conversion algorithm projects annotations from a source corpus into the target formalism. The central contribution of this thesis is Chinese CCGbank, a corpus of 750,000 words automatically extracted from the Penn Chinese Treebank, reifying the abstract analysis through corpus conversion. We then take three state-of-the-art CCG parsers from the literature — the split-merge PCFG parser of Petrov and Klein, the transition-based CCG parser of Zhang et al., and the maximum entropy parser of Clark and Curran — and train and evaluate all three on Chinese CCGbank, achieving the first Chinese CCG parsing models in the literature. We demonstrate that while the three parsers are only separated by a small margin trained on English CCGbank, a substantial gulf of 4.8% separates the same parsers trained on Chinese CCGbank. We also confirm that the gap between the states-of-the-art in English and Chinese PSG parsing can be observed in CCG parsing. Our parsing experiments establish Chinese CCG parsing as a new and substantial challenge, a line of empirical investigation directly enabled by Chinese CCGbank.
African Languages WordNet is an ongoing project which is based on the English WordNet.WordNet is an electronic lexical database that groups words in synonym sets (synsets).In this project, words are translated from English to African Languages.As such, this paper comments on the semantic aspects of verbs in one of the African Languages, namely Northern Sotho, that are translated from English.Verbs are generally understood as expressions of action or state of being.These two languages are typologically dissimilar, with different cultural-historical backgrounds.Northern Sotho is a Bantu 1 language, agglutinating with extensive and productive use of affixes while English is not.In the first place, structural differences between these two languages pose equivalence challenges, both linguistically and computationally.Secondly, verb equivalents may not be affected by the same collocational restrictions in the source and target languages.Another issue is that a verb in one language may invoke certain connotations, which may not apply to its equivalent in another language.Finally, some concepts may be foreign and others culture-specific to one language and not the other, thus resulting in omission of some target language concepts.Attention to these equivalence challenges may enhance technological development of the target language lexicon.
Thresholded two-tone ("Mooney") images are of interest for vision science because the object hidden within the image can be hard to recognize, with recognition times in the second to minute range. However, once a subject has seen the original grayscale image from which the Mooney is generated, recognition is much accelerated. Typically, suitable "Mooney" images need to be painstakingly generated by hand. Here, we present an approach for automatically generating a two-tone image database. This is based on large number of images collected from the internet. We first selected concrete words from a linguistic database. Using these words as search words, we automatically downloaded images from an online image database (www.flickr.com). Subsequently, the images were preprocessed and thresholded using a histogram based thresholding algorithm to generate the two-tone images. We provide an image set with 330 Mooney images and psychophysical results obtained from six subjects. With a presentation time of 20 s, the average recognition time was 9.36 s ± 7.40 s. Additionally, subjective ratings (confidence, Aha, and difficulty ratings) were obtained and are presented for each subject and image. This image set is, to our knowledge, the largest two-tone image set available to the vision and cognitive science research community (https://sites.google.com/site/hayneslab/links). We provide a Matlab toolbox that makes the extension of the image database possible. Using this toolbox, the researcher can add new object names as search words and create new two-tone images easily. Furthermore, we will present possibilities to extent this toolbox using another image database called ImageNet and introduce the use of Amazon Mechanical Turk to select useful images for a particular experiment. This image database can be useful for studying the mechanisms of rapid learning of visual image recognition, and has applications in research on conscious vision, learning, priming, reward and insight. Meeting abstract presented at VSS 2013
В статье рассматривается проблема нормы и нормативного подхода к языку в диахроническом плане.Определяется специфика нормативного похода к языковым средствам в различных лингвистических традициях и выявляются основные характеристики лингвистической нормы.В статье указывается, что на каждом этапе развития языка складываются свои нормы как резуль
We describe Abstract Meaning Representation (AMR), a semantic representation language in which we are writing down the meanings of thousands of English sentences. We hope that a sembank of simple, whole-sentence semantic structures will spur new work in statistical natural language understanding and generation, like the Penn Treebank encouraged work on statistical parsing. This paper gives an overview of AMR and tools associated with it.
Paratactic syntactic structures are notoriously difficult to represent in dependency formalisms. This has painful consequences such as high frequency of parsing errors related to coordination. In other words, coordination is a pending problem in dependency analysis of natural languages. This paper tries to shed some light on this area by bringing a systematizing view of various formal means developed for encoding coordination structures. We introduce a novel taxonomy of such approaches and apply it to treebanks across a typologically diverse range of 26 languages. In addition, empirical observations on convertibility between selected styles of representations are shown too. 1
A time-windowing feature extraction approach based on time-frequency (TF) analysis is adopted here to investigate the time-course of the discrimination between musical appraisal electroencephalogram (EEG) responses, under the parameter of familiarity. An EEG data set, formed by the responses of nine subjects during music listening, along with self-reported ratings of liking and familiarity, is used. Features are extracted from the beta (13-30 Hz) and gamma (30-49 Hz) EEG bands in time windows of various lengths, by employing three TF distributions (spectrogram, Hilbert-Huang spectrum, and Zhao-Atlas-Marks transform). Subsequently, two classifiers (k-NN and SVM) are used to classify feature vectors in two categories, i.e., "likeâ and "dislike,â under three cases of familiarity, i.e., regardless of familiarity (LD), familiar music (LDF), and unfamiliar music (LDUF). Key findings show that best classification accuracy (CA) is higher and it is achieved earlier in the LDF case {91.02 ± 1.45% (7.5-10.5 s)} as compared to the LDUF case {87.10 ± 1.84% (10-15 s)}. Additionally, best CAs in LDF and LDUF cases are higher as compared to the general LD case {85.28 ± 0.77%}. The latter results, along with neurophysiological correlates, are further discussed in the context of the existing literature on the time-course of music-induced affective responses and the role of familiarity.
We have introduced here a new type of corpus annotation which we call Etymological Annotation (EA). We propose this new type because although, over the years, scientists have proposed corpus annotation of various types (Atkins, Clear and Ostler 1992, Biber 1993, Leech 2005), nobody has ever suggested that words included within corpora need to be annotated at their etymological level so that one can retrieve necessary linguistic information relating to antiquity of words and terms used in corpora. The applicational relevance of etymologically annotated corpora may be visualized in language description, language planning, language education, lexicology, language technology as well as in compilation of general, historical, learner and special dictionaries. In case of those languages, where one comes across large number of words borrowed from neighbouring and foreign languages, the proper identification of source of origin of words carries tremendous referential relevance in cross-lingual lexical database generation, morphological processing, part-of-speech tagging, e-learning, digital lexical profile generation, information retrieval, machine learning, and language documentation. Thus, etymologically annotated corpora become an essential resource of applied linguistics and language technology. We propose here to define this new event with necessary direction and guidance to develop etymologically tagged language corpora for all natural languages.
We propose a method of affective text analysis and modeling that is capable of generating continuous valence ratings at the sentence level starting from word and multi-word term valence ratings. Motivated from the language modeling literature, a back-off algorithm is employed to efficiently fuse the valence of single-word and multi-word terms. Specifically, a term detection criterion is used to select the appropriate n-gram terms, starting with bigrams and potentially backing off to unigrams. Term affective ratings are generated by a lexicon expansion method, using semantic similarity estimates computed on a large web corpus. The proposed framework provides state-of-the art results in the sentence level SemEval'07 task of news headline polarity detection, reaching an accuracy of 75%.
This experiment investigated whether affective information from unfamiliar people can influence the affective ratings for unfamiliar foods, before and after participants have ate the foods. The participants rated a food product’s appearance in terms of how palatable it looked on a 7-point Likert scale before they had eaten it (affective expectation). They also rated how palatable the food actually was after they had eaten it (affective evaluation). Results showed that there was a significant interaction between when participants provided the ratings and whether they had been informed of other people’s affective evaluation. Exposure to affective information did not influence affective expectation, but it did increase affective evaluation. These results suggest that affective information that is presented simultaneously with a visual experience might only be influential when perceivers are able to test the information through their perceptual experience.
Much of the research exploring the relationship between taste quality and affective state suggests that sweet-tasting foods are associated with pleasant feelings, and sour- and spicy-tasting foods are associated with unpleasant feelings. The findings of arousal response as a component of overall affective state are less clear with respect to taste quality. The present study investigated the relationship between taste quality and affective state by comparing arousal and pleasantness ratings of neutral images from the International Affective Picture System (IAPS). Participants (N = 55) recorded these ratings during consumption of sprays which varied in taste quality (sweet, sour, or spicy). As hypothesized, sweet sprays elicited significantly higher ratings of pleasantness than sour or spicy sprays ( η 2 p =.14) on the neutral images. However, arousal ratings did not differ among the three taste quality conditions. Implications of the findings in a broader framework and suggestions for future research are discussed.
Prior research has suggested that configural resemblance between a current scene and a previously experienced but forgotten one may trigger déjà vu experiences. The present study examined whether there is a relationship between the frequency of actual déjé vu experiences, measured by questionnaires, and sensitivity to a configural resemblance between past and present events, measured by questionnaires, and between two scenes presented simultaneously in the laboratory. We measured familiarity ratings and remember–know judgements of several scenes. Some scenes had been previously presented, some were similar to previously presented scenes and the others were dissimilar. Déjà vu tendencies were significantly correlated with sensitivity to similarity in the measured questionnaires and in the laboratory, as well as to a feeling of familiarity for similar scenes. In this study, we found for the first time that people who more frequently experience déjé vu states were also more likely to regard themselves as sensitive to similarity and more likely to notice the similarity between two scenes in the laboratory.
AIM: A group-based multisensory activity program (Sensory Day) for residents with dementia was developed, to address the challenge of providing personalised activities within tight operational constraints in residential aged care facilities. METHOD: Fourteen participants with severe and very severe dementia were observed before, during and after participation in one of four Sensory Day sessions. The Menorah Park Rating Scale was used to yield four levels of engagement. The Philadelphia Geriatric Affect Rating Scale was used to identify four affect states. Dementia severity was ascertained by PAS-CIS scores mapped onto the Global Deterioration Scale. RESULTS: Increased levels of constructive engagement and positive affect were observed during participation in the Sensory Day sessions, relative to measures taken before the session. CONCLUSIONS: This novel approach to activity programming demonstrates that it is possible to provide group-based activities for residents with severe and very severe dementia which result in increased engagement and positive mood.
Patrick Hanks is well known as a lexicographer and as the author of several remarkable articles on phraseology, collocations and co-occurrences and on the description of meaning. He was the chief editor, or one of the editors, of several dictionaries, some of which are highly original, particularly in their treatment of polysemy. Many people were hoping that he would eventually develop his views in a book, and this had been 'announced as forthcoming for many years'. 'Some people... had given up hope that it would ever appear' (xv), but now, at last, after a period of preparation of sixteen years (215), what began as a 'disjointed collection of short essays and other fragments' has become 'a coherent text' (xv) of almost 500 pages.
Sprkbanken at the National Library of Norway is currently building up gold-standard Dependency Grammar treebanks for Norwegian Bokml and Nynorsk.The treebanks are manually annotated for morphological features, syntactic functions and dependency relations.This paper explains the choice of texts and format of the treebanks, some key aspects of the morphological and syntactic annotation, and it is illustrated how the treebanks can be used.
Methods are proposed for measuring affective valence and arousal in speech. The methods apply support vector regression to prosodic and text features to predict human valence and arousal ratings of three stimulus types: speech, delexicalized speech, and text transcripts. Text features are extracted from transcripts via a lookup table listing per-word valence and arousal values and computing per-utterance statistics from the per-word values. Prediction of arousal ratings of delexicalized speech and of speech from prosodic features was successful, with accuracy levels not far from limits set by the reliability of the human ratings. Prediction of valence for these stimulus types as well as prediction of both dimensions for text stimuli proved more difficult, even though the corresponding human ratings were as reliable. Text based features did add, however, to the accuracy of prediction of valence for speech stimuli. We conclude that arousal of speech can be measured reliably, but not valence, and that improving the latter requires better lexical features.
El objetivo del trabajo consiste en reutilizar el Treebank de dependencias EPEC-DEP (BDT) para construir el gold standard de la sintaxis superficial del euskera. El paso basico consiste en el estudio comparativo de los dos formalismos aplicados sobre el mismo corpus: el formalismo de la Gramatica de Restricciones (Constraint Grammar, CG) y la Gramatica de Dependencias (Dependency Grammar, DP). Como resultado de dicho estudio hemos establecido los criterios linguisticos necesarios para derivar la funciones sintacticas en estilo CG. Dichos criterios han sido implementados y evaluados, asi en el 75% de los casos somos capaces de derivar automaticamente las funciones sintacticas para construir el gold standard.
This article presents a sociolinguistic lexical-grammar representation in the media by means of an analysis of the polemics on the textbook for young and adult education delivered by MEC, in 2011. The analysis was based on a theoretical basis which refutes the idea of a pure and homogeneous language. The theoretical background used in this analysis is based on the Functional-systemic linguistics (HALLIDAY; MATTHIESSEN, 2004), more specifically, on the ideational meta function that is responsible for the expression of the experience of an inunciative interior material world. Results point to a representation founded in the dichotomy between norm and language use and in the canonical conception of science.
We can determine whether two texts are paraphrases of each other by finding out the extent to which the texts are similar. The typical lexical matching technique works by matching the sequence of tokens between the texts to recognize paraphrases, and fails when different words are used to convey the same meaning. We can improve this simple method by combining lexical with syntactic or semantic representations of the input texts. The present work makes use of syntactical information in the texts and computes the similarity between them using word similarity measures based on WordNet and lexical databases. The texts are converted into a unified semantic structural model through which the semantic similarity of the texts is obtained. An approach is presented to assess the semantic similarity and the results of applying this approach is evaluated using the Microsoft Research Paraphrase (MSRP) Corpus.
Abstract This article is a critical survey of contemporary Arabic lexicography, covering a period beginning roughly with the first edition of the Hans Wehr dictionary and extending to present-day corpus-based machine-readable dictionaries and online lexical databases. It covers both monolingual and bilingual Arabic lexicography relating to the standard written language, Modern Standard Arabic, and to the dialects; however, only bilingual publications in which Arabic is the source language are discussed. The article reviews the different solutions that both Standard Arabic and dialect dictionaries have found to the issue of lemma representation (i.e., the choice of phonetic script) and the issue of lemma organization (i.e., whether to follow a root-based arrangement or one that is purely phonetic and alphabetical).
This paper addresses semantic search of Web services using natural language processing. First we survey various existing approaches, focusing on the fact that the expensive costs of current semantic annotation frameworks result in limited use of semantic search for large scale applications. We then propose a service search framework based on the vector space model to combine the traditional frequency weighted term-document matrix, the syntactical information extracted from a lexical database and a dependency grammar parser. In particular, instead of using terms as the rows in a term-document matrix, we propose using synsets from WordNet to distinguish different meanings of a word under different contexts as well as clustering different words with similar meanings. Also based on the characteristics of Web services descriptions, we propose an approach to identifying semantically important terms to adjust weightings. Our experiments show that our approach achieves its goal well.
The paper addresses the challenge of converting MIDT, an existing dependency‐ based Italian treebank resulting from the harmonization and merging of smaller resources, into the Stanford Dependencies annotation formalism, with the final aim of constructing a standard‐compliant resource for the Italian language. Achieved results include a methodology for converting treebank annotations belonging to the same dependency‐based family, the Italian Stanford Dependency Treebank (ISDT), and an Italian localization of the Stanford Dependency scheme.
The paper describes a broadly applicable method of designing multilingual semantics-syntactic analyzers of recommender systems. The user inputs may include the questions of many kinds formed with the help of interrogative words (or without interrogative words), verbs, nouns, attributes, prepositions, the designations of the digital values of various parameters. For the queries in English and German, the developed algorithm of semantic-syntactic analysis processes the questions of many kinds, the commands, and the statements from a restricted sublanguage of NL. For the queries in Russian, the algorithm is additionally able to process the requests with participle constructions and attributive clauses. As a semantic intermediary language, the algorithm uses the SK-language determined by the considered linguistic database. The class of SK-languages is introduced by the theory of K-representations (knowledge representations), its current version is mainly stated in a monograph of the author published by Springer in 2010. The developed algorithm is implemented by means of the programming language PYTHON.
The concept of emotion and how to regulate it is a central aspect of modern psychology. Within the process model of emotion regulation (Gross, 1998), one issue is how attentional deployment affects emotion regulation and how this can be measured. In task 1, pictures of positive or negative valence were showed in two conditions, either attend or decrease emotional reaction, while participants’ eye movements were followed with an eye tracker. Ratings of arousal and valence were significantly affected by instruction, but dwell times were only significant for positive pictures. In task 2, participants were directed either to emotional or non-emotional parts of emotional pictures while skin conductance was recorded. Arousal and valence ratings decreased significantly in non-emotional areas, but no effect could be found for skin conductance data. Results were generally weak in regards to the effectiveness of measuring gaze to indicate emotion regulation in the form of attentional deployment. For future studies, research of individual differences in habitual usage of attentional deployment for emotion regulation was suggested.
Solution Ranking for a Symbolic Parser by Tatiana EkeinhorRanking is one of the methods used to improve data-driven parsers.The task of the ranker is to define a function that will assign a score to each parse candidate tree obtained in the parsing process.We want to apply such models to resolve ambiguity problems in LEOPAR, a grammar-driven parser.LEOPAR produces a high number of parse solutions especially when the size of the input sentence grows.We want to choose the most suitable solutions among the parse candidates proposed by LEOPAR.We have to notice that a scoring system exists already in LEOPAR.It is based on some handcrafted rules.We present in this document a ranking solution for LEOPAR based on statistical techniques.The candidate parses provided by LEOPAR are used as input for our ranker.We test our approach on the Sequoia TreeBank and obtain an improvement of the system compared to the handcrafted rules.
We investigate acoustic modeling, feature extraction and feature selection for the problem of affective content recognition of generic, non-speech, non-music sounds. We annotate and analyze a database of generic sounds containing a subset of the BBC sound effects library. We use regression models, longterm features and wrapper-based feature selection to model affect in the continuous 3-D (arousal, valence, dominance) emotional space. The frame-level features for modeling are extracted from each audio clip and combined with functionals to estimate long term temporal patterns over the duration of the clip. Experimental results show that the regression models provide similar categorical performance as the more popular Gaussian Mixture Models. They are also capable of predicting accurate affective ratings on continuous scales, achieving 62-67% 3-class accuracy and 0.69-0.75 correlation with human ratings, higher than comparable numbers in literature.
The global prevalence of English language is interpreted through two conflicting views: one being a Global English paradigm which, it is said, oppresses non-native English speakers (Phillipson, 1992); and the other, the World Englishes paradigm, in which speakers liberate themselves from binding linguistic norms by adhering to their own culture and mother language creating a version of heteroglossic, or pluralized English language (Kachru, 1992: 11). A boon in settling the arguments between former and the latter is a description of the context in which English language teaching and English language learning is carried out. The particular context of EFL from which I will describe my observations, takes place in secondary education. I am convinced that in the context of my work I have found some evidence showing that, although Japanese curriculum focuses on acrolectal forms, that is learning English for external communication using external standards of formal language (Yano, 2001: 123); in an instance of curriculum change, when the focus changes to basilectal forms, meaning less prestigious language, the indigenization of English takes place. Given this, I will show how formal (EFL) education can display characteristics of the two paradigms affecting one context simultaneously. The interplay between Global English and World Englishes paradigms results in a merger, creating a dual-existence view in the classroom. As an aid to this description, an English language relation with Japanese society, as well as current issues in EFL will be provided. A discussion about what can be done in this intricate context will follow. In addition, this paper intends to underpin the notion of the English language having the ability to carry and maintain a different culture (Mahboob, 2009: 183), thereby undermining the claims of the existence of linguistic imperialism in the expanding circle. Finally, a short description of the curriculum and a small sample of students' creative work will be provided to support my conclusions.
Introduction: The Adult Attachment Interview (AAI) is considered the gold standard of attachment assessment, but is expensive and time consuming. The Relationship Scale Questionaire (RSQ) is a self-report assessment of attachment, but measures different constructs. Objectives: We investigated how each measure correlates with brain activity in attachment-related tasks: conscious valence and salience appraisal of mother's face. Methods: 28 female subjects ages 18–30 were given the AAI and RSQ. Using fMRI subjects viewed pictures of their mother and were asked to rate how good the image made them feel (valence rating) and how related they felt (salience rating). Brain activity correlating with AAI and RSQ measures of attachment security and dismissingness was determined by linear regression. Results: Salience processing and valence processing were associated with increased thalamo-striatal, posterior cingulate and visual cortex activity. Salience processing was associated with bilateral decrease in PFC activity. In salience processing AAI secure subjects had attenuated visual cortex response and increased lingual gyrus activity while RSQ secure subjects had decreased right temporal lobe activity and AAI dismissing subjects had enhanced left PFC deactiviation while RSQ dismissing subjects had increased left temporal pole and visual cortex activity. In valence processing RSQ secure subjects demonstrated increased left insula activity and AAI dismissing subjects demonstrated bilateral increase in thalamic and posterior cingulate activation. Conclusions: AAI and RSQ measures of attachment measure different constructs with divergent patterns of associated brain activity. AAI measures of attachment tap midline structures (pre-reflective/unconscious attachment processing), while RSQ measures tap lateral structures (cognitive/conscious attachment processing).
This magister degree project is a quantitative, real-time study concerning Swedes’ pronunciation of English, their choice of English accent and the degree of mixing of accents by individual speakers. The informants of the study are Swedish television journalists who speak English on television, in various interview situations. In order to determine which accent/s the journalists adopt, the classical RP/GA differences have been observed. For the purpose of the study a corpus of television clips was created, using The Swedish Media Database (Svensk Mediedatabas). The time span of the gathered material stretches from 1970 until 2009, covering four full decades. The speech of TV journalists is particularly interesting from a sociolinguistic point of view, as it can be argued that it is a form of performed speech where the concern for linguistic norm or context appropriateness is higher than in normal speech. The accent that the journalists adopt could therefore be particularly indicative of which English accent is considered most prestigious or most appropriate, among Swedish speakers. British English was the exclusive educational norm in Sweden until 1994 when American English was accepted as an alternative. Students have since been encouraged to choose one of these accents and to avoid mixing of accents. At the same time Swedish speakers are increasingly exposed to American English through media. The hypothesis underlying this study was therefore that we should see a growing tendency in favour of American English in the journalists’ speech and that the tendency to mix accents would be less frequent in earlier years and more common today. Results of the study show a very modest increase of American accent, which peaks in the 1990s and seems to abate by 2000. The data indicates a surprisingly stable situation in favour of British English over the four decades, with a general 30-40 percent mix of American English features. All the informants mix accents, typically up to 30 percent, already in the 1970s. The data cannot fully confirm an increasing American English influence on Swedes’ choice of English accent. However, the study indicates that mixing of accents is, and has been, a common and probably inevitable phenomenon.
Ces dernières années ont vu un regain d’intérêt dans l’utilisation de données semi-structurées, grâce à la standardisation de formats d’échange de données sur le Web tels que XML et RDF. On notera en particulier le Linking Open Data Project qui comptait plus de 31 milliard de triplets RDF à la fin de l’année 2011. XML reste, pour sa part, l’un des formats de données privilégié de nombreuses bases de données de grandes tailles dont Uniprot, Open Government Initiative et Penn Treebank. Cet accroissement du volume de données semi-structurées a suscité un intérêt croissant pour le développement de bases de données adaptées. Parmi les différentes approches proposées, on peut distinguer les approches relationnelles et les approches graphes, comme détaillé au Chapitre 3. Les premières visent à exploiter les moteurs de bases de données relationnelles existants, en y intégrant des techniques spécialisées. Les secondes voient les données semistructurées comme des graphes, c’est-à-dire un ensemble de noeuds liés entre eux par des arêtes étiquetées, dont elles exploitent la structure. L’une des techniques de ce domaine, connue sous le nom d’indexation structurelle, vise à résumer les graphes de données, de sorte à pouvoir identifier rapidement les données utiles au traitement d’une requête. Les index structurels classiques sont construits sur base des notions de simulation et de bisimulation sur des graphes. Ces notions, qui sont d’usage dans de nombreux domaines tels que la vérification, la sécurité, et le stockage de données, sont des relations sur les noeuds des graphes. Fondamentalement, ces notions caractérisent le fait que deux noeuds partagent certaines caractéristiques telles qu’un même voisinage. Bien que les approches graphes soient efficaces en pratique, elles présentent des limitations dans le cadre de RDF et son langage de requêtes SPARQL. Les étiquettes sont, dans cette optique, distinctes des noeuds du graphe.Dans le modèle décrit par RDF et supporté par SPARQL, les étiquettes et noeuds font néanmoins partie du même ensemble. C’est pourquoi, les approches graphes ne supportent qu’un sous-ensemble des requêtes SPARQL. Au contraire, les approches relationnelles sont fidèles au modèle RDF, et peuvent répondre au différentes requêtes SPARQL. La question à laquelle nous souhaitons répondre dans cette thèse est de savoir si les approches relationnelles et graphes sont incompatible, ou s’il est possible de les combiner de manière avantageuse. En particulier, il serait souhaitable de pouvoir conserver la performance des approches graphe, et la généralité des approches relationnelles. Dans ce cadre, nous réalisons un index structurel adapté aux données relationnelles. Nous nous basons sur une méthodologie décrite par Fletcher et ses coauteurs pour la conception d’index structurels. Cette méthodologie repose sur trois composants principaux. Un premier composant est une caractérisation dite structurelle du langage de requêtes à supporter. Il s’agit ici de pouvoir identifier les données qui sont retournées en même temps par n’importe quelle requête du langage aussi précisément que possible. Un second composant est un algorithme qui doit permettre de grouper efficacement les données qui sont retournées en même temps, d’après la caractérisation structurelle. Le troisième composant est l’index en tant que tel. Il s’agit d’une structure de données qui doit permettre d’identifier les groupes de données, générés par l’algorithme précédent pour répondre aux requêtes. Dans un premier temps, il faut remarquer que le langage SPARQL pris dans sa totalité ne se prête pas à la réalisation d’index structurels efficaces. En effet, le fondement des requêtes SPARQL se situe dans l’expression de requêtes conjonctives. La caractérisation structurelle des requêtes conjonctives est connue, mais ne se prête pas à la construction d’algorithmes efficaces pour le groupement. Néanmoins, l’étude empirique des requêtes SPARQL posées en pratique que nous réalisons au Chapitre 5 montre que celles-ci sont principalement des requêtes conjonctives acycliques. Les requêtes conjonctives acycliques sont connues dans la littérature pour admettre des algorithmes d’évaluation efficaces. Le premier composant de notre index structurel, introduit au Chapitre 6, est une caractérisation des requêtes conjonctives acycliques. Cette caractérisation est faite en termes de guarded simulation. Pour les graphes la notion de simulation est une version restreinte de la notion de bisimulation. Similairement, nous introduisons la notion de guarded simulation comme une restriction de la notion de guarded bisimulation, une extension connue de la notion de bisimulation aux données relationelles. Le Chapitre 7 offre un second composant de notre index structurel. Ce composant est une structure de données appelée guarded structural index qui supporte le traitement de requêtes conjonctives quelconques. Nous montrons que, couplé à la caractérisation structurelle précédente, cet index permet d’identifier de manière optimale les données utiles au traitement de requêtes conjonctives acycliques. Le Chapitre 8 constitue le troisième composant de notre index structurel et propose des méthodes efficaces pour calculer la notion de guarded simulation. Notre algorithme consiste essentiellement en une transformation d’une base de données en un graphe particulier, sur lequel les notions de simulation et guarded simulation correspondent. Il devient alors possible de réutiliser les algorithmes existants pour calculer des relations de simulation. Si les chapitres précédents définissent une base nécessaire pour un index structurel visant les données relationnelles, ils n’intègrent pas encore cet index dans le contexte d’un moteur de bases de données relationnelles. C’est ce que propose le Chapitre 9, en développant des méthodes qui permettent de prendre en compte l’index durant le traitement d’une requête SPARQL. Des résultats expérimentaux probants complètent cette étude. Ce travail apporte donc une première réponse positive à la question de savoir s’il est possible de combiner de manière avantageuse les approches relationnelles et graphes de stockage de données RDF.
The paper presents our work on the annotation of intra-chunk dependencies on an English treebank that was previously annotated with Inter-chunk dependencies, and for which there exists a fully expanded parallel Hindi dependency treebank. This provides fully parsed dependency trees for the English treebank. We also report an analysis of the inter-annotator agreement for this chunk expansion task. Further, these fully expanded parallel Hindi and English treebanks were word aligned and an analysis for the task has been given. Issues related to intra-chunk expansion and alignment for the language pair HindiEnglish are discussed and guidelines for these tasks have been prepared and released.
The objective of the present contribution is to give a survey of the annotation of information structure in the Czech part of the Prague Czech-English Dependency Treebank. We report on this first step in the process of building a parallel annotation of information structure in this corpus, and elaborate on the automatic pre-annotation procedure for the Czech part. The results of the pre-annotation are evaluated, based on the comparison of the automatic and manual annotation.
The paper introduces a dependency annotation effort which aims to fully annotate an Uyghur corpus. It is the first attempt of its kind to develop a large scale tree-bank for Uyghur. In this paper, we provide the motivation for following the dependency theory as the annotation scheme and argue that the dependency grammar is better suited to model the various linguistic phenomena in Uyghur. In our solution, the syntactic relations are encoded as labeled dependency relations among segments of lexical items and sequence of inflectional groups separated by derivational boundaries. We present the basic annotation scheme including morphological and syntactically dependency relation. We also show how the scheme handles some phenomenon such as omissions in copula sentences, punctuations and coordinations, etc.
While working on valency lexicons for Czech and English, it was necessary to define treatment of multiword entities (MWEs) with the verb as the central lexical unit. Morphological, syntactic and semantic properties of such MWEs had to be formally specified in order to create lexicon entries and use them in treebank annotation. Such a formal specification has also been used for automated quality control of the annotation vs. the lexicon entries. We present a corpus-based study, concentrating on multilayer specification of verbal MWEs, their properties in Czech and English, and a comparison between the two languages using the parallel Czech-English Dependency Treebank (PCEDT). This comparison revealed interesting differences in the use of verbal MWEs in translation (discovering that such MWEs are actually rarely translated as MWEs, at least between Czech and English) as well as some inconsistencies in their annotation. Adding MWE-based checks should thus result in better quality control of future treebank/lexicon annotation. Since Czech and English are typologically different languages, we believe that our findings will also contribute to a better understanding of verbal MWEs and possibly their more unified treatment across languages. This work has been supported by the Grant No.
This paper presents unpublished materials of Ivan Pankevitch’s dictionary of Southern Carpathian dialects that are now handled in the form of an electronic lexical database in the Slavonic Institute of the Academy of Sciences of the Czech Republic. Analysis of the selected materials showed representation of individual sources in the lexical database, allowed a preliminary determination of the literary sources and in the case of direct field records provided an opportunity to specify their geographical distribution.
OBJECTIVE: To create annotated clinical narratives with layers of syntactic and semantic labels to facilitate advances in clinical natural language processing (NLP). To develop NLP algorithms and open source components. METHODS: Manual annotation of a clinical narrative corpus of 127 606 tokens following the Treebank schema for syntactic information, PropBank schema for predicate-argument structures, and the Unified Medical Language System (UMLS) schema for semantic information. NLP components were developed. RESULTS: The final corpus consists of 13 091 sentences containing 1772 distinct predicate lemmas. Of the 766 newly created PropBank frames, 74 are verbs. There are 28 539 named entity (NE) annotations spread over 15 UMLS semantic groups, one UMLS semantic type, and the Person semantic category. The most frequent annotations belong to the UMLS semantic groups of Procedures (15.71%), Disorders (14.74%), Concepts and Ideas (15.10%), Anatomy (12.80%), Chemicals and Drugs (7.49%), and the UMLS semantic type of Sign or Symptom (12.46%). Inter-annotator agreement results: Treebank (0.926), PropBank (0.891-0.931), NE (0.697-0.750). The part-of-speech tagger, constituency parser, dependency parser, and semantic role labeler are built from the corpus and released open source. A significant limitation uncovered by this project is the need for the NLP community to develop a widely agreed-upon schema for the annotation of clinical concepts and their relations. CONCLUSIONS: This project takes a foundational step towards bringing the field of clinical NLP up to par with NLP in the general domain. The corpus creation and NLP components provide a resource for research and application development that would have been previously impossible.