Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Parsing requires the quantitative information of grammatical functions of part of speec h.This paper studied grammatical functions of part of speech by using the Chinese Dependency Treebank based on the Probabilistic Valency Pattern Theor y.According to the frequency,we divided grammatical functions of verbs into principal function,secondary function and part functio n.From the aspect of quantitative analysis,we validated and complemented formers’ conclusion,which helped us to have a clearer understanding of grammatical functions of verb s.This paper is also a development of the Probabilistic Valency Pattern Theor y.
Theories proposing that how one thinks and feels is influenced by feedback from the body remain controversial. A central but untested prediction of many of these proposals is that how well individuals can perceive subtle bodily changes (interoception) determines the strength of the relationship between bodily reactions and cognitive-affective processing. In Study 1, we demonstrated that the more accurately participants could track their heartbeat, the stronger the observed link between their heart rate reactions and their subjective arousal (but not valence) ratings of emotional images. In Study 2, we found that increasing interoception ability either helped or hindered adaptive intuitive decision making, depending on whether the anticipatory bodily signals generated favored advantageous or disadvantageous choices. These findings identify both the generation and the perception of bodily responses as pivotal sources of variability in emotion experience and intuition, and offer strong supporting evidence for bodily feedback theories, suggesting that cognitive-affective processing does in significant part relate to "following the heart."
Complications arise for standoff annotation when the annotation is not on the source text itself, but on a more abstract representation. This is particularly the case in a language such as Arabic with morphological and orthographic challenges, and we discuss various aspects of these issues in the context of the Arabic Treebank. The Standard Arabic Morphological Analyzer (SAMA) is closely integrated into the annotation workflow, as the basis for the abstraction between the explicit source text and the more abstract token representation. However, this integration with SAMA gives rise to various problems for the annotation workflow and for maintaining the link between the Treebank and SAMA. In this paper we discuss how we have overcome these problems with consistent and more precise categorization of all of the tokens for their relationship with SAMA. We also discuss how we have improved the creation of several distinct alternative forms of the tokens used in the syntactic trees. As a result, the Treebank provides a resource relating the different forms of the same underlying token with varying degrees of vocalization, in terms of how they relate (1) to each other, (2) to the syntactic structure, and (3) to the morphological analyzer. 1.
In this paper, we introduce our recent work on Chinese HPSG grammar development through treebank conversion. By manually defining grammatical constraints and anno-tation rules, we convert the bracketing trees in the Penn Chinese Treebank (CTB) to be an HPSG treebank. Then, a large-scale lexi-con is automatically extracted from the HPSG treebank. Experimental results on the CTB 6.0 show that a HPSG lexicon was successfully extracted with 97.24 % accu-racy; furthermore, the obtained lexicon achieved 98.51 % lexical coverage and 76.51 % sentential coverage for unseen text, which are comparable to the state-of-the-art works for English. 1
The Varro toolkit is a system for identifying and counting a major class of regularity in treebanks and annotated natural language data in the form of treestructures: frequently recurring unordered subtrees. This software has been designed for use in linguistics to be maximally applicable to actually existing treebanks and other stores of tree-structurable natural language data. It minimizes memory use so that moderately large treebanks are tractable on commonly available computer hardware. This article introduces condensed canonically ordered trees as a data structure for efficiently discovering frequently recurring unordered subtrees.
In this paper we describe the development of a schema for the annotation of attribution relations and present the first findings and some relevant issues concerning this phenomenon. Following the D-LTAG approach to discourse, we have developed a lexically anchored description of attribution, considering this relation, contrary to the approach in the PDTB, independently from other discourse relations. This approach has allowed us to deal with the phenomenon in a broader perspective than previous studies, reaching therefore a more accurate description of it and making it possible to raise some still unaddressed issues. Following this analysis, we propose an annotation schema and discuss the first results concerning its applicability. The schema has been applied to a pilot portion of the ISST corpus of Italian and represents the initial phase of a project aiming at the creation of an Italian Discourse Treebank. We believe this work will raise some awareness concerning the fundamental importance of attribution relations. The identification of the source has in fact strong implications for the attributed material. Moreover, it will make overt the complexity of a phenomenon for long underestimated. 1.
Language users are increasingly turning to electronic resources to address their lexical information needs, due to their convenience and their ability to simultaneously capture different facets of lexical knowledge in a single interface. In this paper, we discuss techniques to respond to a user’s lexical queries by providing multilingual and multimodal information, and facilitating navigating along different types of links. To this end, structured information from sources like WordNet, Wikipedia, Wiktionary, as well as Web services is linked and integrated to provide a multi-faceted yet consistent response to user queries. The meanings of words in many different languages are characterized by mapping them to appropriate WordNet sense identifiers and adding multilingual gloss descriptions as well as example sentences. Relationships are derived from WordNet and Wiktionary to allow users to discover semantically related words, etymologically related words, alternative spellings, as well as misspellings. Last but not least, images, audio recordings, and geographical maps extracted from Wikipedia and Wiktionary allow for a multimodal experience. 1.
We have manually curated a polarity lexicon for German, comprising word polarities and polarity strength values of about 8,000 words: nouns, verbs and adjectives. The decisions were primarily carried out using the synsets from GermaNet, a WordNet-like lexical database. In an evaluation on German novels, it turned out that the stock of adjectives was too small. We carried out experiments to automatically learn new subjective adjectives together with their polarity orientation and polarity strength. For this purpose, we applied a corpus-based approach that works with pairs of coordinated adjectives extracted from a large German newspaper corpus. In the context of this work, we evaluated two subtasks in detail. First, how good are we at reproducing the polarity classification – including our three- level strength measure – contained in our initial lexicon by machine learning methods. Second, because adding of training material did not improve the results at the expected rate, we evaluated the human intercoder agreement on polarity classifications in an experiment. The results show that judgements about the strength of polarity do vary considerably between different persons. Given these problems related to the design and automatic augmentation of polarity lexicons, we have successfully experimented with a semi-automatically approach where a list of reliable candidate words (here: adjectives) is generated to ease the manual annotation process.
We are in the process of creating a multi-representational and multi-layered treebank for Hindi/Urdu (Palmer et al., 2009), which has three main layers: dependency structure, predicate-argument structure (PropBank), and phrase structure. This paper discusses an important issue in treebank design which is often neglected: the use of empty categories (ECs). All three levels of representation make use of ECs. We make a high-level distinction between two types of ECs, trace and silent, on the basis of whether they are postulated to mark displacement or not. Each type is further refined into several subtypes based on the underlying linguistic phenomena which the ECs are introduced to handle. This paper discusses the stages at which we add ECs to the Hindi/Urdu treebank and why. We investigate methodically the different types of ECs and their role in our syntactic and semantic representations. We also examine our decisions whether or not to coindex each type of ECs with other elements in the representation. 1.
Corpora of sentences annotated with grammatical information have been deployed by extending the basic lexical and morphological data with increasingly complex information, such as phrase constituency, syntactic functions, semantic roles, etc. As these corpora grow in size and the linguistic information to be encoded reaches higher levels of sophistication, the utilization of annotation tools and, above all, supporting computational grammars appear no longer as a matter of convenience but of necessity. In this paper, we report on the design features, the development conditions and the methodological options of a deep linguistic databank, the CINTIL DeepGramBank. In this corpus, sentences are annotated with fully fledged linguistically informed grammatical representations that are produced by a deep linguistic processing grammar, thus consistently integrating morphological, syntactic and semantic information. We also report on how such corpus permits to straightforwardly obtain a whole range of past generation annotated corpora (POS, NER and morphology), current generation treebanks (constituency treebanks, dependency banks, propbanks) and next generation databanks (logical form banks) simply by means of a very residual selection/extraction effort to get the appropriate “views ” exposing the relevant layers of information. 1.
The paper presents a system for querying treebanks in a uniform way. The system is able to work with both dependency and constituency\nbased treebanks in any language. We demonstrate its abilities on 11 different treebanks. The query language used by the system\nprovides many features not available in other existing systems while still keeping the performance efficient. The paper also describes\nthe conversion of ten treebanks into a common XML-based format used by the system, touching the question of standards and formats.\nThe paper then shows several examples of linguistically interesting questions that the system is able to answer, for example browsing\nverbal clauses without subjects or extraposed relative clauses, generating the underlying grammar in a constituency treebank, searching\nfor non-projective edges in a dependency treebank, or word-order typology of a language based on the treebank. The performance of\nseveral implementations of the system is also discussed by measuring the time requirements of some of the queries.
The Quran is a significant religious text, followed by the 1.5 billion believers of the Islamic faith worldwide. The text dates to 610–632 CE and is written in Quranic Arabic, the direct ancestor language of modern standard Arabic in use today. This paper presents the Quranic Arabic Dependency Treebank (QADT) and reports on the approaches and solutions used to apply Natural Language Processing to the unique and challenging language of the Quran. This project differs from other Arabic treebanks by providing a deep computational linguistic model based on historical traditional Arabic grammar($$$$). The treebank is part of the Quranic Arabic Corpus (http://corpus.quran.com), a popular free Arabic resource developed at the University of Leeds. Motivated by the importance of the Quran as a central religious text, we also report on how online collaborative annotation was used to bring together Quranic scholars and Arabic language experts to ensure a high level of accuracy for grammatical analysis of the entire Quran.
We first describe the automatic conversion of the French Treebank (Abeillé and Barrier, 2004), a constituency treebank, into typed projective dependency trees. In order to evaluate the overall quality of the resulting dependency treebank, and to quantify the cases where the projectivity constraint leads to wrong dependencies, we compare a subset of the converted treebank to manually validated dependency trees. We then compare the performance of two treebank-trained parsers that output typed dependency parses. The first parser is the MST parser (Mcdonald et al., 2006), which we directly train on dependency trees. The second parser is a combination of the Berkeley parser (Petrov et al., 2006) and a functional role labeler: trained on the original constituency treebank, the Berkeley parser first outputs constituency trees, which are then labeled with functional roles, and then converted into dependency trees. We found that used in combination with a high-accuracy French POS tagger, the MST parser performs a little better for unlabeled dependencies (UAS=90.3 % versus 89.6%), and better for labeled dependencies (LAS=87.6 % versus 85.6%). 1.
Excessive or addictive Internet use can be linked to different online activities, such as Internet gaming or cybersex. The usage of Internet pornography sites is one important facet of online sexual activity. The aim of the present work was to examine potential predictors of a tendency toward cybersex addiction in terms of subjective complaints in everyday life due to online sexual activities. We focused on the subjective evaluation of Internet pornographic material with respect to sexual arousal and emotional valence, as well as on psychological symptoms as potential predictors. We examined 89 heterosexual, male participants with an experimental task assessing subjective sexual arousal and emotional valence of Internet pornographic pictures. The Internet Addiction Test (IAT) and a modified version of the IAT for online sexual activities (IATsex), as well as several further questionnaires measuring psychological symptoms and facets of personality were also administered to the participants. Results indicate that self-reported problems in daily life linked to online sexual activities were predicted by subjective sexual arousal ratings of the pornographic material, global severity of psychological symptoms, and the number of sex applications used when being on Internet sex sites in daily life, while the time spent on Internet sex sites (minutes per day) did not significantly contribute to explanation of variance in IATsex score. Personality facets were not significantly correlated with the IATsex score. The study demonstrates the important role of subjective arousal and psychological symptoms as potential correlates of development or maintenance of excessive online sexual activity.
"En este art ́ıculo presentamos el desarrollo de un nuevo recurso de c ́odigo abierto para el espa ̃ nol: el treebank Tibidabo. La anotaci ́on se est ́a llevando a cabo de forma semi–autom ́atica en la que, en primer lugar, el corpus es analizado au- tom ́ aticamente con una gram ́ atica simb ́ olica del espa ̃ nol basada en HPSG e im- plementada en el sistema Linguistic Knowledge Builder, y, en segundo lugar, los resultados del proceso de an ́alisis se desambiguan manualmente. La existencia del treebank Tibidabo nos permitir ́a futuros trabajos de investigaci ́on para el desar- rollo y evaluaci ́on de una arquitectura h ́ıbrida que combine metodos simb ́olicos y estad ́ısticos para el PLN, as ́ı como investigaciones orientadas a la hibridizaci ́on de t ́ecnicas de bajo y alto nivel para el PLN."
Treebank annotation is a labor-intensive and time-consuming task. In this paper, we show that a simple statistical ranking model can significantly improve treebanking efficiency by prompting human annotators, well-trained in disambiguation tasks for treebanking but not necessarily grammar experts, to the most relevant linguistic disambiguation decisions. Experiments were carried out to evaluate the impact of such techniques on annotation efficiency and quality. The detailed analysis of outputs from the ranking model shows strong correlation to the human annotator behavior. When integrated into the treebanking environment, the model brings a significant annotation speed-up with improved inter-annotator agreement. †
This paper describes the application of probabilistic part of speech taggers to the Dzongkha language. A tag set containing 66 tags is designed, which is based on the Penn Treebank. A training corpus of 40,247 tokens is utilized to train the model. Using the lexicon extracted from the training corpus and lexicon from the available word list, we used two statistical taggers for comparison reasons. The best result achieved was 93.1% accuracy in a 10-fold cross validation on the training set. The winning tagger was thereafter applied to annotate a 570,247 token corpus.
Several studies have investigated the neural responses triggered by emotional pictures, but the specificity of the involved structures such as the amygdala or the ventral striatum is still under debate. Furthermore, only few studies examined the association of stimuli's valence and arousal and the underlying brain responses. Therefore, we investigated brain responses with functional magnetic resonance imaging of 17 healthy participants to pleasant and unpleasant affective pictures and afterwards assessed ratings of valence and arousal. As expected, unpleasant pictures strongly activated the right and left amygdala, the right hippocampus, and the medial occipital lobe, whereas pleasant pictures elicited significant activations in left occipital regions, and in parts of the medial temporal lobe. The direct comparison of unpleasant and pleasant pictures, which were comparable in arousal clearly indicated stronger amygdala activation in response to the unpleasant pictures. Most important, correlational analyses revealed on the one hand that the arousal of unpleasant pictures was significantly associated with activations in the right amygdala and the left caudate body. On the other hand, valence of pleasant pictures was significantly correlated with activations in the right caudate head, extending to the nucleus accumbens (NAcc) and the left dorsolateral prefrontal cortex. These findings support the notion that the amygdala is primarily involved in processing of unpleasant stimuli, particularly to more arousing unpleasant stimuli. Reward-related structures like the caudate and NAcc primarily respond to pleasant stimuli, the stronger the more positive the valence of these stimuli is.
In this paper, we focus on the challenge of automatically converting a constituency treebank (source treebank) to fit the standard of another constituency treebank (target treebank). We formalize the conversion problem as an informed decoding procedure: information from original annotations in a source treebank is incorporated into the decoding phase of a parser trained on a target treebank during the parser assigning parse trees to sentences in the source treebank. Experiments on two Chinese treebanks show significant improvements in conversion accuracy over baseline systems, especially when training data used for building the parser is small in size. 1
In 2003, the special issue of “Computational linguistics” (September, 29, 3) dedicated to the Web as corpus and edited by Adam Kilgarriff and Gregory Grefenstette was a landmark event for a promising field of study. Today, this book makes for a fine update, even if it is more limited in scope than its predecessor and less recent in its content than its date of publication would lead to believe. The articles included are in fact partially “based on papers presented at the symposium Corpus linguistics—Perspectives for the Future held (…) in Heidelberg in October 2004” (p. 4). However, the editors state that some of the articles were commissioned later, and many of the texts have in fact been brought up to date to take recent developments into account.
Due to idiosyncrasies in their syntax, semantics or frequency, Multiword Expressions (MWEs) have received special attention from the NLP community, as the methods and techniques developed for the treatment of simplex words are not necessarily suitable for them. This is certainly the case for the automatic acquisition of MWEs from corpora. A lot of effort has been directed to the task of automatically identifying them, with considerable success. In this paper, we propose an approach for the identification of MWEs in a multilingual context, as a by-product of a word alignment process, that not only deals with the identification of possible MWE candidates, but also associates some multiword expressions with semantics. The results obtained indicate the feasibility and low costs in terms of tools and resources demanded by this approach, which could, for example, facilitate and speed up lexicographic work.
We present an extensive empirical evaluation of collocation extraction methods based on lexical association measures and their combination. The experiments are performed on three sets of collocation candidates extracted from the Prague Dependency Treebank with manual morphosyntactic annotation and from the Czech National Corpus with automatically assigned lemmas and part-of-speech tags. The collocation candidates were manually labeled as collocational or non-collocational. The evaluation is based on measuring the quality of ranking the candidates according to their chance to form collocations. Performance of the methods is compared by precision-recall curves and mean average precision scores. The work is focused on two-word (bigram) collocations only. We experiment with bigrams extracted from sentence dependency structure as well as from surface word order. Further, we study the effect of corpus size on the performance of the individual methods and their combination.
We propose a method for automatically identifying individual instances of English verb-particle constructions (VPCs) in raw text. Our method employs the RASP parser and analysis of the sentential context of each VPC candidate to differentiate VPCs from simple combinations of a verb and prepositional phrase. We show that our proposed method has an F-score of 0.974 at VPC identification over the Brown Corpus and Wall Street Journal.
Over the past two decades or so, Multi-Word Expressions (MWEs; also called Multi-word Units) have been an increasingly important concern for Computational Linguistics and Natural Language Processing (NLP). The term MWE has been used to refer to various types of linguistic units and expressions, including idioms, noun compounds, phrasal verbs, light verbs and other habitual collocations. However, while there is no universally agreed definition for MWE as yet, most researchers use the term to refer to those frequently occurring phrasal units which are subject to certain level of semantic opaqueness, or non-compositionality. Non-compositional MWEs pose tough challenges for automatic analysis because their interpretation cannot be achieved by directly combining the semantics of their constituents, thereby causing the “pain in the neck of NLP” (Sag et al. 2001).
Background: The geographical position of Maharashtra state makes it rather essential to study the dispersal of modern humans in South Asia. Several hypotheses have been proposed to explain the cultural, linguistic and geographical affinity of the populations living in Maharashtra state with other South Asian populations. The genetic origin of populations living in this state is poorly understood and hitherto been described at low molecular resolution level. Methodology/Principal Findings: To address this issue, we have analyzed the mitochondrial DNA (mtDNA) of 185 individuals and NRY (non-recombining region of Y chromosome) of 98 individuals belonging to two major tribal populations of Maharashtra, and compared their molecular variations with that of 54 South Asian contemporary populations of adjacent states. Inter and intra population comparisons reveal that the maternal gene pool of Maharashtra state populations is composed of mainly South Asian haplogroups with traces of east and w)
Studies of human memory often generate data on the sequence and timing of recalled items, but scoring such data using conventional methods is difficult or impossible. We describe a Python-based semiautomated system that greatly simplifies this task. This software, called PyParse, can easily be used in conjunction with many common experiment authoring systems. Scored data is output in a simple ASCII format and can be accessed with the programming language of choice, allowing for the identification of features such as correct responses, prior-list intrusions, extra-list intrusions, and repetitions.
The purpose of the article is to describe the limits of phonetic variance within consonant groups listed in the article’s title and evidenced in Polish letters from the years 1525−1550, additionally taking into account the degree of standardization of individual variants, and also their chronological, phonetic, lexical, geographic and textual (idiolectic) conditioning. It stems from the presented analysis that within the discussed questions, the textual norm of the letters has a rather conservative character, which is corroborated by sole occurrence of traditional forms in the group of norm-creating variants. Despite this, the Polish language of the letters is characterised by significant openness – owing to a considerable share of innovative forms in the group of variants remaining outside the norm, although the majority of the analysed phonetic representations (13 forms) are the variants which were not accepted in the future by usage and the norm; some of them are the forms with limited, often dialectal (and even slang) scope. It must be also noted that within the analysed consonant groups, the Polish language used in the letters represents the situation basically identical with the situation presented in the printed texts from that period. There is just one significant difference concerning the simplification of the group (-)xv- ≥ (-)f-, with quite numerous evidence in the discussed letters (44%), and unrecorded in the studies on the Polish language of the printed material of the first half of the 16th century.
This paper presents a rational argument based on examples of real language to make the case that lay definitions of parts-of-speech are more complex than commercial language pedagogy appreciates. Put simply, school grammars are misleading. They tend to pick the most convenient words for explanation and categorize them as if there were few or no variants within that category, when in reality, however, variation is the norm. Word class categories as presented in typical textbook illustration function as a handicap to future learning. I thus argue two points in this paper. First, that the definitions of lexical categories ought to be made in the form of respecting distinct linguistic dimensions and not in oversimplified and misleading one- dimensional categories which must be unlearned in order for learners to begin actually learning about how languages function. Secondly, a proper theory that radically separates the representation of linguistic expressions in the various grammatical components must be adopted for pedagogy to develop. I illustrate these points with examples drawn from English and Japanese.
Background and Aim: Understanding and defining developmental norms of auditory comprehension is a necessity for detecting auditory-verbal comprehension impairments in children. We hereby investigated lexical auditory development of Persian (Farsi) speaking children.Methods: In this cross-sectional study, auditory comprehension of four 2-5 year old normal children of adult’s child-directed utterance at available nurseries was observed by researchers primarily to gain a great number of comprehendible words for the children of the same age. The words were classified into nouns, verbs and adjectives. Auditory-verbal comprehension task items were also considered in 2 sections of subordinates and superordinates auditory comprehension. Colored pictures were provided for each item. Thirty 2-5 year old normal children were randomly selected from nurseries all over Tehran. Children were tested by this task and subsequently, mean of their correct response were analyzed. Results: The findings revealed that there is a high positive correlation between auditory-verbal comprehension and age (r=0.804, p=0.001). Comparing children in 3 age groups of 2-3, 3-4 and 4-5 year old, showed that subordinate and superordinate auditory comprehension of the former group is significantly lower (p<0.05) than the others. Intra-group comparisons revealed no significant difference between nouns, verbs and adjectives (p>0.05), while the difference between subordinate and superordinate auditory comprehension was significant in all age groups (p<0.05).Conclusion: Auditory-verbal comprehension develop much faster at lower than older ages and there is no prominent difference between word linguistic classes including nouns, verbs and adjectives. Slower development of superordinate auditory comprehension implies semantic hierarchical evolution of words.
Especially since the mid 20th century, Newfoundland English has experienced considerable change, much of which appears to involve weakening or even loss of local speech features, and greater alignment with supralocal (typically, North American) norms. This chapter begins by contextualising language change relative to (largely negative) insider and outsider attitudes to Newfoundland dialects. Using illustrative examples, the chapter documents the social and stylistic patterns associated with ongoing phonetic and grammatical change. Despite fairly rapid intergenerational decline in the use of some local features, others are shown to be more robust: they display obvious style shifting, in that they tend to be avoided by younger speakers in formal, though not in casual, speech styles. Rapid change is also in evidence at the levels of vocabulary and discourse. Loss of traditional lexicon is countered by the borrowing of lexical innovations from outside the province, along with such “trendy” discourse features as quotative be like, and the prosodic features of creaky voice and high rising intonation in statements.
IntroductionA common experience in education is lack of synthetic view upon data - not only among students but also teachers. Obviously, students are taught to a certain extent to recognise and understand dependencies, interrelations within single fields of study, and they encounter methods of both distinction between analysis and synthesis but they are hardly ever capable of carrying out similar activities on their own, not to mention their serious scarcities in recognising and interpreting connections among different fields of study such as geography and literature, or physics and biology.Contemporary approaches to foreign language teaching often stress importance of using literature in language classroom as it provides a wide range of topics for students. Graded readers are becoming extremely popular with those preparing for state and international language examinations, but also with learners out of institutional framework - even these works are regarded as authentic. Although graded readers are undoubtedly useful for this type of approach, they are limited to a finite number of lexical items and a definite level of grammar, and as such, they are capable of transmitting a small number of cultural characteristics.J. Thompson defines culture as the pattern of meanings embodied in symbolic forms, including actions, utterances and meaningful objects of various kinds, by virtue of which individuals communicate with one another and share their experiences, conceptions and beliefs. (Thompson 1990:132) His definition includes significant constituents: pattern, which is syntax in a broad sense, meanings, which are studied in semantics, whereas symbolic forms are signs, use of which - communication - is dealt with by pragmatics; his definition is, therefore, another semiotic definition of culture, a little more detailed, thus applicable to education. According to his point of view, we can assume that authenticity of literary pieces in English refers to true reflection of Anglophone pattern of meanings. By 'Anglophone' is meant a multicultural, multinational and multilingual vortex, as English language is incessantly pushing its boundaries outwards by taking in new grammatical and lexical elements, thus broadening its register and improving its grammatical flexibility or tolerance in order to meet needs of various cultures employing it as a lingua franca. Its permanent relationship with other languages offers a great variety of unfamiliar items, with unusual characteristics that are welcomed or refused by English language, depending on its relative acceptability on receiving side.As for case of language teaching and learning, broadening set of devices employed by a language means immeasurable challenge for both teachers and learners, therefore, it is a must to consider observation of Claire Kramsch thatnative speakers of a language speak not only with their own individual voices, but through them speak also established knowledge of their native community and society, stock of metaphors this community lives by, and categories they use to represent their experience.(Kramsch 1993:43)Non-native speakers, learners of foreign languages usually do not share above elements, simply because underlying patterns of their mother tongue, even among members of one language family, differ from those in target language, and so structuring of information and art of expression have very little in common, and acquisition of this kind of linguistic experience requires incredible effort. Obviously, task of meeting needs and expectations of target language community is always very difficult, and for this reason, use of literature in language classroom proves to be a considerable contribution to intercultural education.Foreign language learning is always a process of getting to know another experience of existence, meeting another culture, people, and standards, norms and values of living. …
Abstract: The paper aims to make a comparison study between word association of native speakers and that of Chinese English learners (CELs). Through data analysis of the word association results, the nature of the second language (L2) mental lexicon is explored. A continuous free word association test (WAT) was conducted to 150 students from Dalian University of Technology (DUT). And the Minnesota word association norms are selected as a native speakers' word association test for the comparison. The results of WATs are classified and analyzed with respect to response type and part of speech. The major findings in the paper are as follows: (1) The words in L2 mental lexicon are essentially semantically-related, just like the mental lexicon of L1 speakers. But phonological relation plays a more important role in L2 mental lexicon than in L1 mental lexicon. (2) Nouns are easy to be activated for both native speakers and L2 learners. And responses of the same part of speech as the stimulus word are easier to be activated. (3) Difference in culture and limitation of language competence may cause the different word association of natives and L2 learners. And L2 learners' native language is likely to have influence on their L2 mental lexicon. Keywords: mental lexicon; word association; Chinese English learners Resume: Le document vise a faire une etude comparative entre l'association de mots entre les locuteurs de langue maternelle anglaise et les apprenants chinois de l'anglais (ACA). Grâce a l'analyse des donnees des resultats d'association de mots, la nature de lexique mental de la deuxieme langue (L2) est exploree. Un test continu de l'association de mots libre (TAM) a ete realisee chez 150 etudiants de l'Universite de Technologie de Dalian (UTD). Et les normes d'association de mots de Minnesota sont selectionnee comme un test d'association de mots chez les locuteurs natifs pour faire la comparaison. Les resultats de TAM sont classes et analyses en fonction du type de reponse et de la partie du discours. Les conclusions principales de cet article sont les suivantes: (1) Les mots dans le lexique mental L2 sont semantiquement lies, tout comme le lexique mental des locuteurs de L1. Mais les relations phonologiques jouent un role plus important dans le lexique mental L2 que dans le lexique mental L1. (2) Les noms sont faciles a etre actives pour les locuteurs natifs et les apprenants de L2. Et les reponses de la meme partie du discours en tant que le mot de stimulus sont plus faciles a activer. (3) La difference de culture et la limitation de la competence linguistique peuvent causer une association de mots differente des autochtones et des apprenants de L2. Et la langue maternelle des apprenants de L2 est susceptible d'avoir une influence sur leur lexique mental L2. Mots-cles: lexique mental; association de mots; apprenants chinois de l'anglais INTRODUCTION For any language, vocabulary plays a significant role. Without vocabulary, communication cannot happen in a meaningful way. Thus lexical researches have aroused more and more interest among linguists. And the study of mental lexicon has drawn special attention from researchers. In the past thirty years, there has been great development in lexical research. Researchers make great efforts to try revealing the organization of mental lexicon which contains an extremely large amount of information. By now, agreement has been reached on the organization of the fust language (Ll) mental lexicon. Researchers commonly agree that words in L1 mental lexicon are connected with each other semantically and are stored in mind around a semantic network. However, there is still disagreement among researchers on the organization of the second language (L2) mental lexicon. Three kinds of viewpoints have been advanced, namely phonological view, semantic view and syntactic view. With the application of word association test (WAT) to linguistic study, more and more researchers have started to use this efficient method in the study of L2 mental lexicon to try to find answers to this unsettled issue. …
This article focuses on the variability of one of the subtypes of multi-word expressions, namely those consisting of a verb and a particle or a verb and its complement(s). We build on evidence from Estonian, an agglutinative language with free word order, analysing the behaviour of verbal multi-word expressions (opaque and transparent idioms, support verb constructions and particle verbs). Using this data we analyse such phenomena as the order of the components of a multi-word expression, lexical substitution and morphosyntactic flexibility.
We investigate the performance of an easyfirst, non-directional dependency parser on the Hebrew Dependency treebank. We show that with a basic feature set the greedy parser’s accuracy is on a par with that of a first-order globally optimized MST parser. The addition of morphological-agreement feature improves the parsing accuracy, making it on-par with a second-order globally optimized MST parser. The improvement due to the morphological agreement information is persistent both when gold-standard and automatically-induced morphological information is used. 1
In this paper, we argue for and demonstrate the use of Prolog as a tool to query annotated corpora. We present a case study based on the German TüBa-D/Z Treebank to show that flexible and efficient corpus querying can be started with a minimal amount of effort. We end this paper with a brief discussion of performance, that suggests that the approach is both fast enough and scalable. 1
All in-text\treferences\tunderlined\tin\tblue\tare\tlinked\tto\tpublications\ton\tResearchGate, letting you\taccess\tand\tread\tthem\timmediately.
Corpus-based techniques have proved to be very beneficial in the development of efficient and accurate approaches to word sense disambiguation (WSD) despite the fact that they generally represent relatively shallow knowledge. It has always been thought, however, that WSD could also benefit from deeper knowledge sources. We describe a novel approach to WSD using inductive logic programming to learn theories from first-order logic representations that allows corpus-based evidence to be combined with any kind of background knowledge. This approach has been shown to be effective over several disambiguation tasks using a combination of deep and shallow knowledge sources. Is it important to understand the contribution of the various knowledge sources used in such a system. This paper investigates the contribution of nine knowledge sources to the performance of the disambiguation models produced for the SemEval-2007 English lexical sample task. The outcome of this analysis will assist future work on WSD in concentrating on the most useful knowledge sources.
Empty categories represent an important source of information in syntactic parses annotated in the generative linguistic tradition, but empty category recovery has only started to receive serious attention until very recently, after substantial progress in statistical parsing. This paper describes a unified framework in recovering empty categories in the Chinese Treebank. Our results show that given skeletal gold standard parses, the empty categories can be detected with very high accuracy. We report very promising results for empty category recovery for automatic parses as well. 1
Natural Language Processing (NLP) is experiencing rapid growth as its theories and methods are more and more deployed in a wide range of different fields. In the humanities, the work on corpora is gaining increasing prominence. Within industry, people need NLP for market analysis, web software development to name a few examples. For this reason it is important for many people to have some working knowledge of NLP. The book “Natural Language Processing with Python” by Steven Bird, Ewan Klein and Edward Loper is a recent contribution to cover this demand. It introduces the freely available Natural Language Toolkit (NLTK) 1—a project by the same authors—that was designed with the following goals: simplicity, consistency, extensibility and modularity.
Pattern matching, or querying, over annotations is a general purpose paradigm for inspecting, navigating, mining, and transforming annotation repositories—the common representation basis for modern pipelined text processing architectures. The open-ended nature of these architectures and expressiveness of feature structure-based annotation schemes account for the natural tendency of such annotation repositories to become very dense, as multiple levels of analysis get encoded as layered annotations. This particular characteristic presents challenges for the design of a pattern matching framework capable of interpreting ‘flat’ patterns over arbitrarily dense annotation lattices. We present an approach where a finite state device applies (compiled) pattern grammars over what is, in effect, a linearized ‘projection’ of a particular route through the lattice. The route is derived by a mix of static grammar analysis and runtime interpretation of navigational directives within an extended grammar formalism; it selects just the annotations sequence appropriate for the patterns at hand. For expressive and efficient pattern matching in dense annotations stores, our implemented approach achieves a mix of lattice traversal and finite state scanning by exposing a language which, to its user, provides constructs for specifying sequential, structural, and configurational constraints among annotations.