Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
The chapter presents a meta-search tool developed in order to deliver search results structured according to the specific interests of users. Meta-search means that for a specific query, several search mechanisms could be simultaneously applied. Using the clustering process, thematically homogenous groups are built up from the initial list provided by the standard search mechanisms. The results are more user oriented, as a result of the ontological approach of the clustering process. After the initial search made on multiple search engines, the results are pre-processed and transformed into vectors of words. These vectors are mapped into vectors of concepts, by calling an educational ontology and using the WordNet lexical database. The vectors of concepts are refined through concept space graphs and projection mechanisms, before applying the clustering procedure. Implementation details and early experimentation results are also provided.
This thesis introduces an experimental and quantitative approach to language through the study of the concept of soft constraints and its application to two phenomena of order in French: the position of the attributive adjective and the ordering of verbal complements occurring in postverbal position. Soft constraints are defined as affecting the acceptability rather than the grammaticality of the sentences. Our main hypothesis is that these constraints are properties of the language and thus must studied in syntax. These constraints raise a methodological issue: since they do not affect the grammaticality of the sentences, they cannot be investigated using the traditional tools of syntax (introspection and grammaticality judgment). It is therefore necessary to define tools for their description and analysis. The proposed methods are statistical analysis of corpus date, inspired by the work of Bresnan et al. (2007) and Bresnan & Ford (2010) and, to a lesser extent, psycholinguistic experiment. Regarding the position of the adjective, we test most of the constraints encountered in the literature and we propose a statistical analysis of the data extracted from the French Treebank corpus. We show the importance of the adjectival item and the nominal item with which it combines. Other constraints linked to the internal syntax of the adjectival phrase and the noun phrase also play a significant role in the choice of position. The work on the relative order of the verbal complements is conducted on a sample of sentences extracted from two newspapers corpora (French Treebank and Est-Républicain) and two corpora of spoken French (ESTER and C-ORAL-ROM). We show the significant influence of the constituent weight over the ordering: short before long order which is a feature of SVO languages like French, is observed in over 86% of cases. We also identify the important role of the verbal lemma associated with its semantic class (annotated with the dictionary of Dubois & Dubois-Charlier, 1997). Finally, building on the analysis of corpus data as well a two questionnaires eliciting acceptability judgments, it seems that nor animacy neither information structure (given/new, Prince, 1981) have a significant effect on the postverbal complement ordering.
This paper introduces the error corpus of Korean learner English and mal rules to detect the errors. Based on the corpus, we classified 42 error types. Our criteria for error classification are more general in order to enhance agreement rate and decrease errors. For generating mal rule, we testified two different grammars. One is Context Free Grammar (CFG) from Penn Treebank. The other is the typed feature structure grammars based on the Head-Driven Phrase Structure Grammar (HPSG), using Natural Language ToolKit (Bird et al. 2009). We advanced grammatical formalism from CFG to HPSG since CFG needs abundant phrasal markers causing over-generation, structural ambiguity and complexity.
We investigate aspects of interoperability between a broad range of common annotation schemes for syntacto-semantic dependencies. With the practical goal of making the LinGO Redwoods Treebank accessible to broader usage, we contrast seven distinct annotation schemes of functor‐argument structure, both in terms of syntactic and semantic relations. Drawing examples from a multi-annotated gold standard, we show how abstractly similar information can take quite different forms across frameworks. We further seek to shed light on the representational ‘distance’ between pure bilexical dependencies, on the one hand, and full-blown logical-form propositional semantics, on the other hand. Furthermore, we propose a fully automated conversion procedure from (logical-form) meaning representation to bilexical semantic dependencies. †
Based on Kachru’s Three Circle Model on the spread of English in different parts of the world, I question how foreign university students migrating from Expanding Circle countries to Singapore deal with the “clash” of linguistic norms set by different circles. This research hence explores English speaking behavioural intentions of these foreign tertiary students in Singapore and how they account for their language behavioural plans. The collected data reveal that most students held a notion on the native English-Singlish dichotomy, seeing native English as standard and superior while regarding Singlish as improper and non-standard. Considering language behavioural intentions, most respondents claimed to adopt three main strategies: speech maintenance, adapting to the formality level of the communicative situation, and speech convergence. Looking into respondents’ accounts for these intended strategies, I argue that speakers orient their language use not only towards language perceptions but also the communicative situations they are in. However, the relative influences of these two factors on language use vary among different cases.
The neological issue in the OSLL is a component of the wider framework of the research plan Creation of a Lexical Database of the Czech Language of the Beginning of the 21st Century (2005–2011, head – K. Oliva). The construction of the lexical collections (archives of lexical dynamics) is dealt with by the excerption section; the theoretical understanding and lexicographic treatment are assured by an independent working group of lexicographers. Thanks to the research plan being resolved, it was possible to ensure the continual complementation of the neological excerption, modify the method of the accumulation of the material in connection with the new tasks of the department and modernise the software equipment (in connection with that to make part of the neological material accessible to the wider public). These results are built on by the theoretical and practical activities of the neological working group, focusing on the treatment of new material (2002–2010).
The neological issue in the OSLL is a component of the wider framework of the research plan Creation of a Lexical Database of the Czech Language of the Beginning of the 21st Century (2005–2011, head – K. Oliva). The construction of the lexical collections (archives of lexical dynamics) is dealt with by the excerption section; the theoretical understanding and lexicographic treatment are assured by an independent working group of lexicographers. Thanks to the research plan being resolved, it was possible to ensure the continual complementation of the neological excerption, modify the method of the accumulation of the material in connection with the new tasks of the department and modernise the software equipment (in connection with that to make part of the neological material accessible to the wider public). These results are built on by the theoretical and practical activities of the neological working group, focusing on the treatment of new material (2002–2010).
Here we describe work on learning the subcategories of verbs in a morphologically rich language using only minimal linguistic resources. Our goal is to learn verb subcategorizations for Quechua, an under-resourced morphologically rich language, from an unannotated corpus. We compare results from applying this approach to an unannotated Arabic corpus with those achieved by processing the same text in treebank form. The original plan was to use only a morphological analyzer and an unannotated corpus, but experiments suggest that this approach by itself will not be effective for learning the combinatorial potential of Arabic verbs in general. The lower bound on resources for acquiring this information is somewhat higher, apparently requiring a a part-of-speech tagger and chunker for most languages, and a morphological disambiguater for Arabic.
The lexical database of the humanistic and baroque Czech MADLA covers vocabulary from 1500–1780. It contains about 750 000 hand-excerpted documents stemming from dictionaries, herbaria, chronicles and other literary documents from this historical period. Using this database, it is possible to monitor the development of meaning, word formation, paradigm and other grammatical categories, and it serves as a basis for further research on the vocabulary from the 16th–18th centuries. The text describes a representative sample of change in the meaning of the substantive kredenc („cupboard“), the meaning of which has shifted from the original „tasting“ to the contemporary „dresser“. Emphasis is placed on the proof and comparison of different meanings of this word.
This article deals with the use or relative que with a circumstantial complement of time as an antecedent, and introducing clauses in which the relative, not preceded by a preposition, performs also in the subordinate clause the grammatical function of time adjunct. The adjunct that plays the role of an antecedent can be formed by a prepositional phrase, an adverb or a subordinate clause of time. In these cases, the relative has an adverbial function within the clause that introduces. These uses seem to continue that of undeclinable QUOD in Late Latin, which was used following a noun expressing time, to introduce a clause that indicates simultaneity. Although this type of clauses introduced by que is documented since medieval times, only when the antecedent is a time adverb or certain prepositional phrases its use is accepted in modern normative Spanish, while in other cases, specially if the antecedent is a time clause, there is a trend to reject it in the written linguistic norm, though it is still documented in colloquial speech.
This article presents an interpretive study of âSiapa Menyuruh?â, a poem by Indonesiaâs contemporary poet Mustofa Bisri. The study is carried out within the framework of Sperber and Wilsonâs relevance theory (RT), which is based on the principles that human cognition tends to the maximization of relevance (I), and that every act of communication presumes its own optimal relevance (II). This suggests that in the relevance-theoretic perspective, Bisri wrote the poem âSiapa Menyuruhâ not because he wanted to violate certain linguistic norms or any communicative maxims, but because this was the most relevant utterance he could produce. He intentionally raised the effort to process his poem because he promised greater cognitive effects to the reader. The reader who is willing to process his utterances in the poem further is granted not with one, single strong implicature, but with a number of weak implicatures.
The goal of this paper is to expose the character of ‘Grammar’ in curriculum standards of Korean Language Education revision in 2011. To achieve this goal, we examined the contents of grammar education in this curriculum, comparing with previous curricula and viewing different points of view on grammar education in the Korean Language Education. The grammar curriculum revision in 2011 tried to raise the value of grammar education through connecting the goal of grammar education to language skills. For example, the contents items chosen from the field of ‘Linguistic Norms’ and ‘Vocabulary’ was increased comparing with previous curricula. But in terms of contents, this new curriculum is unsatisfactory, because it selected its contents from linguistic content system as ever and it was unsuccessful in pursuing some important value of grammar education. Therefore, it is necessary to try constructing contents system of grammar education with the items based on the intrinsic value of language.
In this paper we introduce a new approach to transition-based depen- dency parsing. We propose that the parser construct an undirected graph during the parsing process, instead of a standard directed dependency structure. A pos- teriori, the output undirected structure is converted into a dependency tree. This alleviates error propagation, a characteristic problem of these systems. We apply this approach to obtain undirected variants of the Planar and 2-Planar parsers and of Covington's non-projective parser. We perform experiments on several treebanks from the CoNLL-X shared task, showing that these variants outperform the original directed algorithms in most of the cases.
In the article, the currently existing codified orthographic norms are presented with reference to four selected problem topics and paralleled with certain other norms significantly influencing the synchronous usage of written language. The expression “norm” or “norms” is to be understood in its broadest sense: at one end of the continuum as a formalized set of rules and regulations which, if ignored, may result in sanctions, and, at the other end, as a set of principles and guidelines organized according to unified rules and affecting, or regulating, any one sphere of human activity and behaviour. The discussion includes other possible reasons for the discrepancies between orthographically prescribed use and use deviating from it. Despite the weight of these reasons, the non-linguistic norms presented in this article appear to be an influential factor that can by no means be overlooked by contemporary normative linguistics.
This paper discusses hybridization in contemporary Russian language. In particular, the work focuses on Dina Rubina�s novel Here comes the Messiah! (?????????????!.....).Her prose is characterized by a wide range of linguistic and communicative means; a key role is played by linguistic hybridization which alienates the text from any contexts, notions of time, as well as the ordinary and accepted linguistic norms. In her work, Rubina - a migrant writer living in Israel - describes a carnival atmosphere (based on Bachtin�s idea of carnival) as one of the main leitmotivs of the Russian community in Jerusalem. Words - which are are the primary means of expression in a world considered as a stage - often undergo hybridization processes. Therefore, new meanings, neologisms, original lexemes decorate Rubina�s multicultural text.
According to the questionnaire held in England, in 2000s, 333 from 547 of the informers admit the usage of split infinitive constructions in casual speech only, while 108 — in standard literary language and 106 consider these constructions to be non-standard. Such discrepancy in appraisal from the author’s point of view evidently proves the fact, that linguistic norm of this phenomenon is in the process of taking up its position. The debate history on the subject has deep roots and goes back to the Old English. Today this process is still not completed. In the author’s opinion the question is revealed and actual in modern linguistics. The work is devoted to a certain research of the problem, where the author applies to the sources of appearance and development of disputes, analyzies deferent positions and puts forward her own view point.
Learning vocabulary and understanding texts present difficulty for language learners due to, among other things, the high degree of lexical ambiguity. By developing an intelligent tutoring system, this dissertation examines whether automatically providing enriched sense-specific information is effective for vocabulary learning and reading comprehension of second language learners. The system developed in this study contributes to an extended understanding of how NLP techniques can be applied more effectively in an educational environment. The system allows learners to upload texts and click on any content word in order to obtain sense-appropriate lexical information for unfamiliar or unknown words during reading. The system consists of three components: (1) the system manager controls the interaction among each learner, the NLP server, and the lexical database; (2) the NLP server converts a raw input text to a linguistically-analyzed text; (3) the lexical database is used to provide a sense-appropriate definition and example sentences of a word to the learner. To obtain the sense-appropriate information, the system first performs word sense disambiguation (WSD) on the input text. Pointing to appropriate examples tuned for language learners, however, is complicated by the fact that the database of examples is from one repository (COBUILD), while automatic WSD systems generally rely on senses from another (WordNet). The lexical database, then, is indexed by WordNet senses, each of which points to an appropriate corresponding COBUILD sense. The fact that every sense inventory has its own standards of sense distinction poses a serious problem in integrating these inventories into one. To redirect an input WordNet sense to a corresponding COBUILD sense, thus, a word sense alignment algorithm was developed, following a heuristic of favoring flatter alignment structures. With this system, an empirical study was conducted with 60 intermediate learners of English as a second language to examine whether this system can lead learners to improve their vocabulary acquisition and reading comprehension. The findings show that learners demonstrated higher performance when receiving sense-specific information. Furthermore, the qualitative examination of the effect of automatic system errors show that, although learners showed learning regardless of the appropriateness of lexical information, they still showed relatively greater learning when given appropriate lexical information.
The objective of textology (text linguistics) is to analyse and describe interpersonal communication in all its aspects, as it happens now and did in the past, in various discourse communities. The basic categories of so defined textology are: text, utterance and discourse, regarded as different perspectives on the phenomenon of interpersonal communication. This phenomenon manifests itself in established works, with their specific semantic and syntactic structure, but also in genre affinity, in interactive events fixed in a widely understood situational context, and in social, cultural and linguistic norms which regulate communication activities in individual human communities and make interpersonal communication possible. These three categories can be said to reify in definite communication phenomena: in written and recorded text, and in spoken utterances, i.e. in actualized discourses.
Semantic verbal fluency (SVF) often shows early and disproportionate decline in AD relative to other language, attention, and executive abilities. Successful performance on SVF depends on the ability to organize conceptual information into related clusters and efficiently access these clusters. Current methods for clustering and switching assessment are labor-intensive and subjective. We developed an automated computational linguistic approach to quantify the semantic content of SVF responses. Neuropsychological and resting state fMRI data were obtained from the work-up of 52 patients presenting to the Minneapolis VAMC GRECC Memory Loss Clinic. Participants included had a clinical diagnosis of MCI or AD, a completed MRI protocol with good quality data, and neuropsychological evaluation including the SVF task (animals). Imaging data were collected on a Philips 1.5T system at the Minneapolis VAMC. Semantic indices based on pairs of words on the SVF task were quantified in two ways: based on the length of their hierarchical relations in WordNet, an electronic lexical database of English (similarity), or calculated using a computerized algorithm based on a variant of principal components analysis (relatedness). Mean cumulative and sequential indices were produced for each method: cumulative similarity and relatedness were computed between all possible pairs of words produced regardless of order; and sequential similarity and relatedness were computed only between pairs of adjacent words. Higher scores reflect larger clusters and reduced switching. Several resting state fMRI network measures were related to the four automated semantic fluency indices. Nodal diversity, local efficiency, and the mean clustering coefficient, were all significantly correlated with cumulative and sequential measures of semantic similarity and relatedness. Pearson r values ranged from.323-.407 with corresponding p-values of.012-.003. All correlations survived multiple comparison correction. The traditional SVF score was not significantly related to imaging indices. We found that computational linguistic measurements of similarity and relatedness were significantly related to network measures obtained from resting state fMRI. These results suggest automated assessment of SVF has correlates with brain function in MCI and AD, and may outperform the traditional SVF score. This approach provides an easy way to standardize clustering and switching assessment without adding burden.
Development of interpersonal relationships is a fundamental human motivation, and behaviors facilitating social bonding are prized. Some individuals experience enhanced reward from alcohol in social contexts and may be at heightened risk for developing and maintaining problematic drinking. There has been little systematic research conducted in group settings, though, and no prior studies have tried to link genetic variation to alcohol’s socially reinforcing effects. This research investigated whether the rewarding effects of alcohol in a group setting are associated with genetic variation implicated in the development of alcohol use disorders. Specifically, this study tested the moderating influence of genes encoding the dopamine D2 and D4 receptors, the serotonin transporter, and the alpha receptor for gamma-aminobutryic acid (GABAA) on the effects of alcohol on social bonding. Social drinkers (N=427; males=50.12%) were assembled into three-person unacquainted groups, and given a moderate dose of alcohol, placebo, or a non-alcohol (control) beverage, which they consumed over 36-min. To assess social bonding, participants completed the Perceived Group Reinforcement Scale immediately after the group drinking period. In addition, their social interaction was video-recorded, and the duration of facial behaviors was systematically coded using the Facial Action Coding System. After applying the Bonferroni correction to control for false positives in multiple genotype comparisons, there was one significant gene x environment interaction. Results showed that carriers of at least one copy of the 7-repeat allele of the DRD4 VNTR reported higher perceived social bonding in the alcohol, relative to placebo or control conditions, whereas alcohol did not affect ratings of 7-absent allele carriers. Findings indicate that carriers of the 7-repeat allele were especially sensitive to alcohol’s effects on social bonding. These data converge with other recent gene-environment interaction findings implicating the DRD4 polymorphism in the development of alcohol use disorders, and results suggest a specific pathway by which social factors may increase risk for problematic drinking among 7-repeat carriers.
Medical discharge documents are summaries written by a physician about the patient’s condition and aim at transferring information to other health care personnel but also to the patient. According to the legislation, the patient should be able to understand the document. In practice, however, this has been shown to be problematic. This paper studies discharge documents from the patients’ perspective and examines how they fulfil the legislation’s demands on understandability. Concentrating on the vocabulary of the texts, we analyse the frequency of domainadapted terms, abbreviations and foreign words. The material consists of 23 528 heart patients’ discharge documents (5 747 126 words). The analysis is performed with the morphological analyser FinTWOL (http://www2.lingsoft.fi/cgi-bin/fintwol). Altogether, FinTWOL analyses 24% of the corpus as unknown or foreign words, abbreviations or medical terms. The most common category, unknown words, includes misspellings and medical terms, such as l.dex. Of these, 100 most common cover for 43% of the total. These terms thus seem to be relatively fixed. Of the words analysed as abbreviations, some are common also in standard language, but others are still very domain-specific, such as I.V. (intravenous). Also the used abbreviations are very fixed: the 100 most common ones cover for 94% of the total. This, however, does not help the patient who probably reads only one document. Similarly, even though misspellings are globally infrequent, they still occur more than once per document. In order to place the obtained results in a context, we performed a similar analysis on general Finnish university newspaper text from Turku Dependency Treebank. In comparison with the 24% obtained with the discharge documents, from the total of 10 687 words, 8,6% were given a special tag. The results show that that terms and abbreviations are considerably more used in discharge documents than in general newspaper text. It is clear that a text with such a vocabulary is domain-specific and distinct from the language that the patient is used to. Also e.g. the varying use of upper and lower case letters (dg and DG for diagnosis) emphasize the particularity of the language. In standard language texts such writing would not be acceptable. Standard writing would, however, help the patients to better understand the texts.
Early-latency theories of emotional processing state that at least coarse monitoring of the emotional valence (a pleasure-displeasure continuum) of facial expressions should be both rapid and highly automated (LeDoux, 1995; Russell, 1980). Research has largely substantiated early-latency differential processing of emotional versus non-emotional facial expressions; however, the effect of valence on early-latency processing of emotional facial expression remains unclear. In an effort to delineate the effects of valence on early-latency emotional facial expression processing, the current investigation compared ERP responses to positive (happy and surprise), neutral, and negative (afraid and sad) basic facial expression photographs as well as to positive (happy-surprise), neutral (afraid-surprise, happy-afraid, happy-sad, sad-surprise), and negative (sad-afraid) morph facial expression photographs during a valence-rating task. Morphing manipulations have been shown to decrease the familiarity of facial patterns and thus preclude any overlearned responses to specific facial codes. Accordingly, it was proposed that morph stimuli would disrupt more detailed emotional identification to reveal a valence response independent of a specific identifiable emotion (Balconi & Lucchiari, 2005; Schweinberger, Burton & Kelly, 1999). ERP results revealed early-latency differentiation between positive, neutral, and negative morph facial expressions approximately 108 milliseconds post-stimulus (P1) within the right electrode cluster; negative morph facial expressions continued to elicit significantly smaller ERP amplitudes than other valence categories approximately 164 milliseconds post-stimulus (N170). Consistent with previous imaging research on emotional facial expression processing, source localization revealed substantial dipole activation within regions of the mesolimbic dopamine system. Thus, these findings confirm rapid valence processing of facial expressions and suggest that negative valence processing may continue to modulate subsequent structural facial processing.
在句法分析中,已有研究工作表明,词汇依存信息对短语结构句法分析是有帮助的,但是已有的研究工作都仅局限于使用一阶的词汇依存信息.提出了一种使用高阶词汇依存信息对短语结构树进行重排序的模型,该模型首先为输入句子生成有约束的搜索空间(例如,N-best 句法分析树列表或者句法分析森林),然后在约束空间内获取高阶词汇依存特征,并利用这些特征对短语结构候选树进行重排序,最终选择出最优短语结构分析树.在宾州中文树库上的实验结果表明,该模型的最高 F1 值达到了 85.74%,超过了目前在宾州中文树库上的最好结果.另外,在短语结构分析树的基础上生成的依存结构树的准确率也有了大幅提升.;The existing works on parsing show that lexical dependencies are helpful for phrase tree parsing.However, only first-order lexical dependencies have been employed and investigated in previous research. Thispaper proposes a novel method for employing higher-order lexical dependencies for phrase tree evaluation. Themethod is based on a parse reranking framework, which provides a constrained search space (via N-best lists orparse forests) and enables the parser to employ relatively complicated lexical dependency features. The models areevaluated on the UPenn Chinese Treebank. The highest F1 score reaches 85.74% and has outperformed allpreviously reported state-of-the-art systems. The dependency accuracy of phrase trees generated by the parser hasbeen significantly improved as well.
Evidence on the impact of nature images has been found in research with hospital patients (Ulrich, 2008, Nanda, Hathorn & Neumann, 2007). The use of art in healthcare environments has become increasingly common (Nanda, Eisen & Baladandayuthapani, 2008). Art is viewed as a positive distraction from stress of the hospital among patients and possibly staff (Ulrich et. al. 1991; Ulrich, Zimring, Quan, & Joseph, 2006). In a previous study art preference study (Nanda, Eisen & Baladandayuthapani, 2008) showed significant difference in the ratings of design students and patients. Findings showed that there was a significant difference in the ratings of the two groups. Furthermore, the emotional rating scale (how does the art picture make you feel) was highly correlated to the selection scale (would you put this art picture in your room) for hospital patients- while this was not the case with the design students. What is the role of culture in the above questions and in how does it impact healthcare design?A total of more than 600 design and non-design students from National University of Mexico, National University of Singapore and University of Texas San Antonio rated images of visual art included abstract, representational and nature images from Mexico, Singapore and Texas representative of the unique cultural contexts, in addition to images that strictly adhere to the evidence-based guidelines for healthcare art laid down by Ulrich & Gilpin (2003) and examples of classic high art.At the end of the survey students re-rated the images again as if they were hospitalized and lying in a patient room. An analysis of preferences across cultures, design disciplines and emotion and selection was undertaken.Results show a surprising amount of agreement across cultures on image rating for hospital rooms. Level of agreement for art selection for personal rooms is significantly lower. This is true in both design and non-design students. Landscapes with a high depth of field, bright colors and verdant foliage were rated consistently high across all cultures, regardless of indigenous elements, with few exceptions, that suggests that there is a certain universal appeal for restorative images of nature that go beyond cultural and educational boundaries. The study showed that empathy (how this art would make you feel) is a stronger determinant of selection than culture, or education, when it comes to art selection for hospitals.
Cet article décrit les étapes qui composent notre analyse du discours, en partant du texte brut, et pour en produire une représentation sémantique dans le cadre de la Discourse Representation Theory, désormais DRT (Kamp and Reyle, 1993). Une chaîne complète de traitement est proposée et testée sur le corpus Itipy, "Itinéraires Pyrénéens", lequel a été proposé par la médiathèque de Pau. Le premier but applicatif consiste à attacher un lieu aux portions de texte narrant une action dans ce lieu. Nous exploitons alors ce corpus de récits de voyage du XIXème siècle dans l'objectif d'extraire automatiquement les itinéraires décrits et afin d'indexer les portions de texte prenant effectivement pour décors les lieux géographiques en question. Notre outil, Grail est un parser pour grammaire logique de types avec un ensemble restreint de règles fixes et utilisant un lexique riche. Tout d'abord, la première phase a consisté en l'acquisition de la grammaire sur un corpus annoté (Paris 7 Treebank). Ce corpus nous a permis d'obtenir les informations grammaticales propres aux unités du lexique de la langue française présentes dans le corpus, le lexique produit ne contient donc pas la totalité des mots du français et contient plusieurs catégories pour les entrées les plus fréquentes. Dans la chaine de traitement, la méthode d'attribution de la catégorie intègre une approche statistique: lors- qu'un mot est absent du lexique, l'analyse propose une catégorie ou lorsqu'il présente plusieurs catégories possibles, elle sélectionne la plus appropriée. Chaque mot du texte est taggé, puis supertaggé en fonction des autres unités se trouvant dans son contexte proche (la phrase). Le supertagger propose plusieurs formules qui correspondent à une analyse syntaxique partielle pour chaque phrase du texte dans le cadre des grammaires catégorielles, et plus précisément du calcul de Lambek. S'ensuit une étape de combinaison de toutes les analyses partielles pour donner l'analyse globale. La structure obtenant la meilleure probabilité étant sélectionnée, on garde cette structure comme organisation du calcul de la représentation sémantique en fonction des unités qui la composent. On associe alors à chaque mot son λ-terme à partir du lexique sémantique cette fois et dont la formule correspond à celle présente dans le lexique grammatical pour cette même entrée (Moot, 2010). Le λ -terme pour chaque unité sémantique est saisi à la main dans le style de la λ -DRT. La représentation sémantique étant produite automatiquement à partir de l'analyse syntaxique, nous obtenons une représentation logique sémantique bien formée. La dimension pragmatique quant à elle ne peut être reléguée à un plan inférieur dans l'interprétation du discours. En effet, une analyse du discours impose de fait une interaction entre la sémantique des unités de langue dont on doit interpréter le sens en discours et la prise en compte de la dimension pragmatique de ce qui est dit. Notre approche s'inspire de l'approche de Busquets et al. (2001), "une théorie de l'interprétation des discours doit être aussi en fait une théorie de la sémantique, de la pragmatique, et de leur interaction, c'est-à-dire une théorie de l'interface pragmatique-sémantique". Certains phénomènes sémantiques restent cependant difficiles à traiter, certains cas de glissement de sens montrent qu'une flexibilité dans le typage doit être permise, alors que dans les cas les plus courants le typage doit être rigide pour éviter une repré- sentation inappropriée. Nous donnerons quelques exemples à propos et proposons donc afin d'améliorer les résultats de notre chaîne traitement de traiter ces phénomènes par l'affinement des λ -termes du lexique dans le cadre du système F, λ -calcul d'ordre supérieur. Nous détaillerons ici notre corpus et nos objectifs applicatifs quant à celui-ci, nous présenterons les étapes de traitement du discours, commençant par l'acquisition de la grammaire du français sur corpus annoté, puis l'analyse syntaxique dans le cadre des grammaires catégorielles. Nous expliquerons plus amplement l'interface syntaxe-sémantique dans la théorie des types logiques permettant la construction de nos repré- sentations sémantiques en λ-DRT. Nous présenterons le système F et notre traitement des phénomènes discursifs mettant en jeu l'interaction sémantique-pragmatique puis nous présenterons les perspectives de ce travail.
New Irish speakers in Belfast play a crucial, complex part in the revitalization and change of both the city and Irish within Northern Ireland. This paper examines the role of new Irish speakers in transforming Belfast, whose emergence from a post-conflict period involves a reassessment of communal cultural expressions. Markers of ethno-national identity are bitterly contentious locally, and yet increasingly celebrated, in line with international trends, as high status cultural forms and potentially profitable tourist attractions. Irish in Belfast currently occupies an ambiguous position: divisive enough for a sign reading ‘Happy Christmas’ in Irish to be experienced as an insult by some city councillors, yet a secure enough part of the establishment for a neighbourhood to be officially rebranded as the Gaeltacht Quarter. <br/>When, how and where new Irish speakers use the language in Belfast has implications for the relationship of Irishness to the Northern Irish state and for the place of Belfast within regional frameworks across the UK, Ireland and Europe. Adult learners and young people exiting Irish medium education have an impact on life in Belfast beyond its small population of Irish speakers. Urbanisation fuelled by new speakers, which shifts the balance of Irish language resources and speakers away from traditional rural Gaeltacht areas and towards cities, also has implications for the language itself. Recent increase in new Irish speakers in Belfast is due to expansion in the Irish-medium sector as well as to adult learners, whose decisions contribute to the school expansion. <br/>Urbanisation, multilingualism and intergenerational shift combine in Belfast to produce new linguistic norms. Moreover, in a minority language community where hierarchies of ‘authenticity’ are weighted towards the rural and the native speaker, where the rural and the native have traditionally been conflated, and where indigeneity is a central concept to contested nationalisms, the emergence of a self-confident, youthful Irish speaking community in Northern Ireland’s biggest city involves a recalibration of the qualities signifying ‘gaelicness’. As students, professionals, hobbyists and activists, new Irish speakers in Belfast occupy a vital position at the crux of changing ideas about place, language and identity.<br/>
The task of automatic machine translation (MT) is the focus of a huge variety of active research efforts, both because of the intrinsic utility of this difficult task, and the theoretical and linguistic insights that arise from modeling relationships between natural languages. However, MT systems that leverage syntactic information are only recently becoming practical, and in a typical system of this sort, syntactic information is generated by monolingual parsers; the task of explicitly modeling syntactic relationships between target and source languages is yet to be fully explored. This thesis investigates the problem of finding syntactic parse trees of target and/or source sentences that are more appropriate for use in a syntactic MT system. Two basic methodologies are explored. First, we present a sequence of two statistical models that leverage bilingual information to improve the linguistic quality of syntactic parses, as measured by their ability to replicate human-generated gold-standard annotations. The first model uses word to word alignments as an external source of information, while the second models the alignments jointly. These models are both quite effective at improving the intrinsic quality of the parse trees, and the second model additionally improves word alignment performance. However, while the two models achieve similar parsing improvements, we find that improving parses in conjunction with word alignments is much more helpful for the downstream machine translation task. In the next part of the thesis, we explore this finding further by investigating the effects on MT performance of agreement between parse trees and word alignments. We present a simple method for transforming input trees in a way that ignores gold-standard annotations, concentrating instead on improving syntactic agreement directly. In experiments, we find that though we obviously lose fidelity to more linguistically informed treebank annotation guidelines, this transformation-based approach yields the strongest improvements in syntactic machine translation.
It is my pleasure to introduce this thematic issue dedicated to the lexicography of Japanese as a second or foreign language, the first thematic issue in Acta Linguistica Asiatica since its inception.Japanese has an outstandingly long and rich lexicographical tradition, but there have been relatively few dictionaries of Japanese targeted at learners of Japanese as a foreign or second language until the end of the twentieth century. With the growth of Japanese language teaching and learning around the world, the rapid development of very large scale linguistic resources and language processing technologies for Japanese, a new generation of aggregated, collectively developed or crowd-sourced resources evolving in the context of the social web, a shift from static paper to constantly developing electronic resources, the spread of internet access on hand-held devices, and new approaches to the use of language reference resources stemming from these developments, dictionaries and other reference resources for learners, teachers and users of Japanese as a foreign/second language are being developed and used in new ways in different user communities. However, information about such developments often does not reach researchers, lexicographers, dictionary users and language teachers in other user communities or research spheres. This special issues wishes to contribute to the spread of such information by presenting some recent developments in this growing field.Having received a very lively response to our call for papers, not all papers selected for publishing could fit into this issue, and part of them will be included in the December issue of ALA, which is also going to be dedicated to Japanese lexicography.The first round of papers included in this issue presents a varied cross-section of current JFL lexicographical work and research. All papers in this issue point out the relative scarcity of appropriate reference works for learners of Japanese as a foreign language, especially when compared to lexicographical resources for Japanese native speakers, and each of the endeavours presented here confronts this lack with its own original approach. Reflecting the paradigm shift in Japanese language research, where corpus research is again playing a central role, most papers presented here take advantage of the bounty of newly available corpora and web data, most prominent among which is the Balanced Corpus of Contemporary Written Japanese developed by the National Institute for Japanese Language and Linguistics in Tokyo, and which is used by Mogi, Pardeshi et al. and Sunakawa et al. in their lexicographical research and projects, while Blin taps data for his research from the web, another increasingly important linguistic resource.The first two papers offer two perspectives on existing Japanese dictionaries. Tom Gally in his paper Kokugo Dictionaries as Tools for Learners: Problems and Potential points out the drawbacks of currently available Japanese dictionaries from the perspective of learners of Japanese as a foreign language, but at the same time offers a very detailed and convincing explanation of the merits of monolingual Japanese dictionaries for native speakers (kokugo dictionaries), such as their comprehensiveness, detailedness and quantity of contextual information, when compared to bilingual dictionaries, which make them a potentially useful resource even for an audience they are not targeting - foreign language learners. His detailed explanation of possible uses and potential hurdles and pitfalls learners may encounter in using them, is not only accurate and informative, but also of immediate practical value for language teachers and lexicographers.Toshinobu Mogi, in his paper Towards the Lexicographic Description of the Grammatical Behaviour of Japanese Loanwords: A Case Study, investigates the lexicographic description of loanwords in Japanese reference works and notes how information offered by currently available dictionaries, especially regarding the grammatical aspects of loanword use, is not sufficient for learners of Japanese as a foreign language. After pointing our the deficiencies of current dictionary descriptions and noting how dictionaries sense divisions do not reflect the frequency of different senses in actual use, as reflected in a large-scale representative general corpus of Japanese, he uses a fascinatingly detailed analysis of the behaviour of a Japanese loanword verb to describe a corpus-based method of lexical description, based on the correspondence between usage forms and senses, which could be used for the compilation of Japanese learners' dictionaries meant for the reception and production of Japanese.The second part of this special issue is composed of four reports on particular aspects of ongoing lexicographical work targeted at learners of Japanese as a foreign language.Prashant Pardeshi, Shingo Imai, Kazuyuki Kiryu, Sangmok Lee, Shiro Akasegawa and Yasunari Imamura in their paper Compilation of Japanese Basic Verb Usage Handbook for JFL Learners: A Project Report, after pointing out - as other authors in this issue - the lack of a detailed and pedagogically sound lexicographical description of Japanese basic vocabulary for foreign learners, propose a corpus-based on-line system which incorporates insights from cognitive grammar, contrastive studies and second language acquisition research to solve this problem. They present their current implementation of such a system, which includes audio-visual material and translations into Chinese, Korean and Marathi. The system also uses natural language processing techniques to support lexicographers who need to process daunting amounts of corpus data in order to produce detailed lexical descriptions based on actual use.The next article by Marcella Maria Mariotti and Alessandro Mantelli, ITADICT Project and Japanese Language Learning, focus on the learner's perspective. They present a collaborative project in which Italian learners of Japanese compiled an on-line Japanese-Italian dictionary using a purposely developed on-line dictionary editing system, under the supervision of a small group of teachers. One practical and obvious outcome of the project is a Japanese-Italian freely accessible lexical database, but the authors also highlight the pedagogical value of such an approach, which stimulates students' motivation for learning, hones their ICT skills, makes them more aware of the structure and usability of existing lexicographic and language learning resources, and helps them learn to cooperate on a shared task and exchange peer support.The third project report by Raoul Blin, Automatic Addition of Genre Information in a Japanese Dictionary, focuses on the labelling of lexical genre, an aspect of word usage which is not satisfactorily presented in current Japanese dictionaries, despite its importance for foreign language learners when using dictionaries for production tasks. The article describes a procedure for automatic labelling of genre by means of a statistical analysis of internet-derived genre-specific corpora. The automatisation of the process simplifies its later reiteration, thus making it possible to observe lexical genre development over time.The final paper in this issue is a report on The Construction of a Database to Support the Compilation of Japanese Learners’ Dictionaries, by Yuriko Sunakawa, Jae-ho Lee and Mari Takahara. Motivated by the lack of Japanese bilingual learners' dictionaries for speakers of most languages in the world, the authors engaged in the development of a database of detailed corpus-based descriptions of the vocabulary needed by learners of Japanese from beginning to advanced level. By freely offering online the basic data needed for bilingual dictionary compilation, they are building the basis from which editors in under-resourced language areas will be able to compile richer and more up-to-date contents even with limited human and financial resources. This project is certainly going to greatly contribute to the solution of existing problems in Japanese learners' lexicography.
Individuals who effectively regulate or mildly increase their systolic blood pressure (SBP) in response to an orthostatic challenge exhibit healthier affective status, cognitive functioning, and better quality of life. Thus, increased SBP in response to an orthostatic challenge serves as a proxy for several underlying changes. This study examined the relationship between SBP regulation and self-esteem in children. Data were collected from 92 boys and girls, aged 8–11 years. Systolic, diastolic, and pulse measurements were obtained after 5 minutes of remaining supine and again after 1 minute of standing. Children also provided affective ratings on the Children's Depression Inventory. The Negative Self-Esteem subscale was examined for this study. A multiple regression analysis revealed that poorer orthostatic regulation was associated with higher levels of negative self-esteem among children aged 8–11 years. Thus, orthostatic BP regulation may serve as a biological marker for poor self-esteem in children. This may have further implications for children's emotional functioning as low self-esteem may serve as a risk factor for future negative affective states.
Prior semantic processing can enhance subsequent picture naming performance, yet the neurocognitive mechanisms underlying this effect and its longevity are unknown. This functional magnetic resonance imaging study examined whether different neurological mechanisms underlie short-term (within minutes) and long-term (within days) facilitation effects from a semantic task in healthy older adults. Both short- and long-term facilitated items were named significantly faster than unfacilitated items, with short-term items significantly faster than long-term items. Region of interest results identified decreased activity for long-term facilitated items compared to unfacilitated and short-term facilitated items in the midportion of the middle temporal gyrus, indicating lexical-semantic priming. Additionally, in the whole brain results, increased activity for short-term facilitated items was identified in regions previously linked to episodic memory and object recognition, including the right l)
Annotating linguistic data has become a major field of interest, both for supplying the necessary data for machine learning approaches to NLP applications, and as a research issue in its own right. This comprises issues of technical formats, tools, and methodologies of annotation. We provide a brief overview of these notions and then introduce the papers assembled in this special issue.
Watson, C. 2012. Tradition and Translation: Maciej Stryjkowski's Polish Chronicle in Seventeenth-Century Russian Manuscripts. Acta Universitatis Upsaliensis. Studia Slavica Upsaliensia 46. 358 pp. Uppsala. ISBN 978-91-554-8308-1. The object of this study is a translation from Polish to Russian of the Polish historian Maciej Stryjkowski’s Kronika Polska, Litewska, Żmodzka i wszystkiej Rusi, made at the Diplomatic Chancellery in Moscow in 1673–79. The original of the chronicle, which relates the origin and early history of the Slavs, was published in 1582. This Russian translation, as well as the other East Slavic translations that are also discussed here, is preserved only in manuscripts, and only small excerpts have previously been published. In the thesis, the twelve extant manuscripts of the 1673–79 translation are described and divided into three groups based on variant readings. It also includes an edition of three chapters of the translation, based on a manuscript kept in Uppsala University Library. There was no standardized written language in 17 th -century Russia. Instead, there were several co-existing norms, and the choice depended on the text genre. This study shows that the language of the edited chapters contains both originally Church Slavonic and East Slavic linguistic features, distributed in a way that is typical of the so-called hybrid register. Furthermore, some features vary greatly between Slovo. Journal of Slavic Languages and Literatures No. 53, 2012 130 manuscripts and between scribes within the manuscripts, which shows that the hybrid register allowed a certain degree of variation. The translation was probably the joint work of several translators. Some minor changes were made in the text during the translation work, syntactic structures not found in the Polish original were occasionally used to emphasize the bookish character of the text, and measurements, names etc. were adapted to Russian norms. Nevertheless, influence from the Polish original can sometimes be noticed on the lexical and syntactic levels. All in all, this thesis is a comprehensive study of the language of the translated chronicle, which is a representative 17 th -century text.
Summary: This paper is devoted to the inter- and intra-linguistic comparison of comics on the basis of the volume Le trésor de Rackham le Rouge (engl. Red Rackham’s Treasure) of the Franco-Belgian Tintin series and its two Catalan versions published in 1964 and 2002, respectively. These versions are analyzed with respect to morpho-syntactic (forms of address, clitic pronoun combinations, clause-initial que and future-oriented temporal adverbial clauses introduced by quan) and lexical criteria (technical vocabulary). The study reveals that the models for the reference variety applied by the translators diverge and that the translations reflect, to a certain extent, the sociolinguistic situation of Catalan at the moment of their creation. [Keywords: Tintin; inter- and intra-linguistic translation; linguistic norm; sociolinguistics of normativization; Joaquim Ventalló]
This paper explores interoperability for data represented using the Graph Annotation Framework (GrAF) (Ide and Suderman, 2007) and the data formats utilized by two general-purpose annotation systems: the General Architecture for Text Engineering (GATE) (Cunningham et al., 2002) and the Unstructured Information Management Architecture (UIMA) (Ferrucci and Lally in Nat Lang Eng 10(3–4):327–348, 2004). GrAF is intended to serve as a “pivot” to enable interoperability among different formats, and both GATE and UIMA are at least implicitly designed with an eye toward interoperability with other formats and tools. We describe the steps required to perform a round-trip rendering from GrAF to GATE and GrAF to UIMA CAS and back again, and outline the commonalities as well as the differences and gaps that came to light in the process.
The study was carried out in the mainstream of linguistic phenomena in philology. The article is devoted to the study of the speech of the tatars, enduring in the cities of Urumqi and Kuldja of the People´s Republic of China. Studied the general characteristic of the tatar speech of China diaspora and considered the features of the use of the tatar language in the region Investigated some of the lexical phenomenon, an old vocabulary, drawing, and synonyms in the language of the China tatars. Discusses some of the phonetic phenomena in the field of substitution of vowels and changes consonants in a speech, which are directly connected with lexical norms of the tatar literary language and its dialects. The same examples as from the oral speech of the inhabitants of Kuldja and Urumqi, as well as from folklore material of the Tatar Diaspora in China. Identified preconditions of use of the tatar language, similarities and peculiarities of use of native Turkic tokens and borrowed words in the speech of the tatars, living in the PRC.
Amazon’s Mechanical Turk is an online labor market where requesters post jobs and workers choose which jobs to do for pay. The central purpose of this article is to demonstrate how to use this Web site for conducting behavioral research and to lower the barrier to entry for researchers who could benefit from this platform. We describe general techniques that apply to a variety of types of research and experiments across disciplines. We begin by discussing some of the advantages of doing experiments on Mechanical Turk, such as easy access to a large, stable, and diverse subject pool, the low cost of doing experiments, and faster iteration between developing theory and executing experiments. While other methods of conducting behavioral research may be comparable to or even better than Mechanical Turk on one or more of the axes outlined above, we will show that when taken as a whole Mechanical Turk can be a useful tool for many researchers. We will discuss how the behavior of workers compares with that of experts and laboratory subjects. Then we will illustrate the mechanics of putting a task on Mechanical Turk, including recruiting subjects, executing the task, and reviewing the work that was submitted. We also provide solutions to common problems that a researcher might face when executing their research on this platform, including techniques for conducting synchronous experiments, methods for ensuring high-quality work, how to keep data private, and how to maintain code security.
In this talk, I will outline some of the myriad of challenges and opportunities that social media offer for natural language processing. I will present analysis of how pre-processing can be used to make social media data more amenable to natural language processing, and review a selection of tasks which attempt to harness the considerable potential of different social media services. There is no question that social media are fantastically popular and varied in form — ranging from user forums, to microblogs such as Twitter, to social networking sites such as Facebook — and that much of the content they host is in the form of natural language. This would suggest a myriad of opportunities for natural language processing (NLP), and yet much of the applied research on social media which uses language data is based on superficial analysis, often in the form of simple keyword search. This begs the question: Are NLP methods not suited to social media analysis? Conversely, is social media data too challenging for modern-day NLP? Alternatively, are simple term search-based methods sufficient for social media analysis, i.e. is NLP overkill for social media? In exploring these questions, I attempt to answer the overarching question of whether social media data is the friend or foe of NLP. I approach the question first from the perspective of what challenges social media language poses for NLP. The most immediate answer is the infamously free-form nature of language in social media, encompassing spelling inconsistencies, the free-form adoption of new terms, and regular violations of English grammar norms. Unsurprisingly, when NLP tools are applied directly to social media data, the results tend to be miserable when compared to data sets such as the Wall Street Journal component of the Penn Treebank. However, there have been recent successes in adapting parsers and POS taggers to social media data (Foster et al., 2011; Gimpel et al., 2011). Additionally, lexical normalisation and other preprocessing strategies have been shown to enhance the performance of NLP tools over social media data (Lui and Baldwin, 2012; Han et al., to appear). Furthermore, social media posts tend to be short and the content highly varied, meaning it is difficult to adapt a tool to the domain, or harness textual context to disambiguate the content. There is also the engineering challenge of real-time processing of the text stream, as much of NLP research is carried out offline with only secondary concern for throughput. As such, we might conclude that social media data is a foe of NLP, in that it challenges traditional assumptions made in NLP research on the nature of the target text and the requirements for real-time responsiveness. However, if we look beyond the immediate text content of social media, we quickly realise that there are various non-textual data sources that can be used to enhance the robustness and accuracy of NLP models, in a way which is not possible with static text corpora. For example, simple information on the author of a post can be used to develop authoradapted models based on the previous posts of the same individual (at least for users who post sufficiently large volumes of data). Links in the post can be used to disambiguate the textual content of the post, whether in the form of URLs and the content contained in the target document(s), hashtags and the content of other similarly-tagged posts, thread-
Principal component analysis identifies uncorrelated components from correlated variables, and a few of these uncorrelated components usually account for most of the information in the input variables. Researchers interpret each component as a separate entity representing a latent trait or profile in a population. However, the components are guaranteed to be independent and uncorrelated only when the multivariate normality of the variables is assumed. If the normality assumption does not hold, components are guaranteed to be uncorrelated, but not independent. If the independence assumption is violated, each component cannot be uniquely interpreted because of contamination by other components. Therefore, in the present study, we introduced independent component analysis, whose components are uncorrelated and independent even when the multivariate normality assumption is violated, and each component carries unique information.
International audience
The study presented in this article is dedicated to a syntactic parser for Romanian. The central goal of the presented technique is to learn a model which is able to discriminate between probability for a word to be head of another word in a dependency structure corresponding to a sentence in the considered language. The model described in this paper was trained on a dependency treebank linguistic resource and is intended to be used in order to develop a dependency syntactic parser.
Background:Some previous studies have revealed that while congenitally blind people have a tendency to refer to visual attributes ('verbalism'), references to auditory and tactile attributes are scarcer. However, this statement may be challenged by current theories claiming that cognition is linked to the perceptions and actions from which it derives. Verbal productions by the blind could therefore differ from those of the sighted because of their specific perceptual experience. The relative weight of each sense in oral descriptions was compared in three groups with different visual experience Congenitally blind (CB), late blind (LB) and blindfolded sighted (BS) adults. Methodology/Principal Findings:Participants were asked to give an oral description of their mother and their father, and of four familiar manually- explored objects. The number of visual references obtained when describing people was relatively high, and was the same in the CB and BS groups ("verbalism" in )
The goal of the presented parallel phrase extraction algorithm is to provide rich and robust set of translation syntactic patterns. To make this approach feasible, we consider the phrase-to-phrase alignments of a bilingual treebank annotated with syntactic constituents. For the intended purpose, the extracted phrasal nodes are encoded by the syntactical information of their components, highlighting some special constructs such as the functional words.
LOLspeak is a complex and systematic reimagining of the English language. It is most often associated with the popular, productive and long-lasting Internet meme ‘LOLcats’. This style of English is characterised by the simultaneous playful manipulation of multiple levels of language. Using community-generated web content as a corpus, we analyse some of the common language play strategies (Sherzer 2002) used in LOLspeak, which include morphological reanalysis, atypical sentence structure and lexical playfulness. The linguistic variety that emerges from these manipulations displays collaboratively constructed norms and tendencies providing a standard which may be meaningfully adhered to or subverted by users. We conclude with a discussion of why people may choose to participate in such language play, and suggest that the language play strategies used by participants allow for the construction of complex identity.
This paper explores the lexical semantic properties of five near-synonymous Chinese words expressing the emotion of SHAME. The concept of self-construal is vital in understanding emotions such as shame as it relies on the reflections of oneself. The interdependent self-construal is a view of the self through relationship with others and it is related to the characteristics of SHAME in Chinese context. The current study carried out an in-depth examination of how interdependent self-construal shapes the shame concept in Chinese, and how features concerning "self versus others" are encoded in Chinese shame words. The "self versus others" features that we look at include cause attribution (to self vs. others), probable relevant outcome (affect self vs. others), social relations (between self and others that cause shame), social norm (personal values vs. social norms), and presence (or absence) of audience. We examine whether and how these features play a crucial role in the Chinese SHAME concept and how they contribute to the differences between the shame words in Mandarin Chinese. The features can be described as a dimension that is implicit in the denotative meaning of these words.
Corpora with high-quality linguistic annotations are an essential component in many NLP applications and a valuable resource for linguistic research. For obtaining these annotations, a large amount of manual effort is needed, making the creation of these resources time-consuming and costly. One attempt to speed up the annotation process is to use supervised machine-learning systems to automatically assign (possibly erroneous) labels to the data and ask human annotators to correct them where necessary. However, it is not clear to what extent these automatic pre-annotations are successful in reducing human annotation effort, and what impact they have on the quality of the resulting resource. In this article, we present the results of an experiment in which we assess the usefulness of partial semi-automatic annotation for frame labeling. We investigate the impact of automatic pre-annotation of differing quality on annotation time, consistency and accuracy. While we found no conclusive evidence that it can speed up human annotation, we found that automatic pre-annotation does increase its overall quality.
French Sign Language has a far smaller specialised lexicon than French, which poses regular problems to interpreters working between the two. Four management control sessions attended by a deaf student and interpreted for him by four professional interpreters were recorded, and the interpreters' tactics when encountering the problems of missing signs in French Sign Language ('lexical gaps') were identified, counted and analyzed. Lexical gaps were found to be numerous in the corpus. The tactics often used elements of French spoken language, in contradiction with a strong sociolinguistic norm in the French deaf community. This can be explained by the interpreters' wish to cater to the needs of the deaf student, who needed to know the French terms when taking exams, and is in line with skopos theory.
Treebank is a basic language resource for training and testing syntactic parser which forms a key module in various NLP systems like machine translation system. This paper reports an ongoing research of building dependency treebank for Kashmiri (KashTreeBank) and discusses some main annotation issues. The paper is based on the pilot annotation of 500 sentences.