Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
We have implemented a rule-based prototype of a Spanish-to-Cuzco Quechua MT system enhanced through the addition of statistical components. The greatest difficulty during the translation process is to generate the correct Quechua verb form in subordinated clauses. The prototype has several rules that decide which verb form should be used in a given context. However, matching the context in order to apply the correct rule depends crucially on the parsing quality of the Spanish input. As the form of the subordinated verb depends heavily on the conjunction in the subordinated Spanish clause and the semantics of the main verb, we extracted this information from two treebanks and trained different classifiers on this data. We tested the best classifier on a set of 4 texts, increasing the correct subordinated verb forms from 80% to 89%.
Reliably coded treebanks are a goldmine for linguistics research. Answering a typical research question involves: (a) querying a treebank to extract sentences containing the feature to be investigated, (b) recognizing and keeping track of characteristics that determine the way in which the linguistic feature is encoded, and (c) using statistics to find out which (combination) of these characteristics determines the outcome of the linguistic feature. While sufficient tools are available for steps (a) and (c) in this process, step (b) has not received much attention yet. This paper describes how the programs “Cesax” and “CorpusStudio” can be used jointly to construct a “corpus research database”, a database that contains the sentences of interest selected in step (a), as well as user-definable pre-calculated characteristics for step (b).
In contemporary South Korean society, there is a strong emphasis on cultural homogeneity and, simultaneously, the development of English proficiency as a human resource. Since language is inextricably linked to identity, bilingual learners from English speaking countries may feel pressure to conform to Korean cultural and linguistic norms, leading to negative identity practices that discourage the use of English. Like the “model minority” stereotype which has been assigned to Asian learners in the United States, the pervasive belief that learners from English speaking countries are highly proficient in English may have adverse effects on students who do not meet the conceptualized standard. To explore educational problems associated with the English-Korean bilingual learner, a case study was conducted on an American-Korean elementary school student. Results revealed that the learner avoided speaking English in public, learning English in formal contexts, and talking about American ethnic traditions, which has resulted in significant deficiencies in English pronunciation and literacy. The avoidance of explicit instruction appears to have precluded the development of cognitive and metacognitive strategies useful in overcoming language deficiencies in an English as a Foreign Language (EFL) context. Recommendations for educational reform have been suggested.
Lexical databases such as AOO or EWN are built in order to enable machine translation of written word. The author’s intention is to test the applicability of the above listed databases in machine translation and to choose the one that proves to be more successful. The presentation of final results consists of two parts. Part I is the analysis of factors influencing the process of machine translation with the use of AOO and EWN that is to say: steps in database design, their theoretical aspects and the categorization of lexical items. In author’s belief, the above mentioned elements exert profound influence on the result of machine translation produced with the use of one of the herein described lexical databases. The second part of the presentation touches on the matters of hierarchy, semantic inheritance and word-sense disambiguation.
Tree-to-tree Statistical Machine Translation models require the use of syntactic tree structures of both the source and target side in learning rules to guide the translation process. In order to accomplish the task, available treebanks for different languages are used as the main resources to collect necessary information to handle the translation task. However, since each treebank has its own defined tags, a barrier is inherently created in highlighting alignment relationships at different syntactic levels for different tag-sets. Moreover, these models are typically over constrained. This paper presents a unified tagset for all languages at Part-of-Speech and Phrasal Category level in tree-to-tree models. Different experiments are conducted to study for its feasibility, efficiency, and translation quality.
Czech models for MorphoDiTa, providing morphological analysis, morphological generation and part-of-speech tagging. The morphological dictionary is created from MorfFlex CZ and the PoS tagger is trained on PDT (Prague Dependency Treebank).
This paper explores the possibility of automatically measuring and comparing affectiveness features in texts from different domains and sub-domains. More specifically our main research question concerns the distribution of affective words within the various subgenres of a newspaper corpus along three affective parameters: Valence (V), pleasantness, Arousal (A), the intensity of the emotion, and Dominance (D), the degree of control exerted by the perceiver over the stimulus. The study is largely based on work by Warriner et al. (in press), who recently divulged a study reporting affective ratings for approximately 14,000 English lemmas in terms of V, A and D. 100,000 token samples of newspaper language from 10 subsections of the Guardian newspaper were analyzed for the presence of these lemmas and the average V, A and D values for each subsection were calculated. Crime and Travel were seen to be those with the most atypical values.
Web services adoption is a major advance in the development of interoperable information systems. In particular, the composition of services can meet the needs increasingly complex of user, by a combination of web services within a single business process. However, despite this widespread adoption of Web services, many obstacles prevent their reconciliation in the composition, or may occur within a BPEL process in a state change, the context for example. ASWSCC Method (Adaptation of Semantic Web Service Composition to Context) is an implementation of a theoretical model made in our an earlier work. It focuses on composition process adaptation to use context (preferences, user type and its environment as the device used, location, access mode and many others). This context and request service matching should be taken into account while composing new services. Our goal is to develop a model which ensures, on the one hand, web services matching during composition process by using domain ontology as lexical database WordNet, its purpose is to identify, classify and relate in different ways semantic content and lexical language. On the other hand, this model allows management and taking into account the context that makes composition process adaptable to different instances of use context, which may change during the same session. For this reason, we are interested to capture and manage the context and its impact on basic services and composition process at once. Changes can affect the context of web services during their executions and the need to adapt their dynamically becomes increasingly crucial. From here comes the need for a coherent solution to adapt web services context. We exploit the benefits of aspect weaving tool in this approach to inject aspects of web services to adapt them to change of context. Keywords-Context definition and management; adaptation; web services composition.
The International Institute for the Portuguese Language, under the auspices of the Community of Portuguese Speaking Countries, is coordinating the development of the Common Orthographic Vocabulary (VOC). VOC will be a large on-line lexical database which will take into account the varieties of the eight Portuguese speaking countries (Angola, Brazil, Cape Verde, Guinea-Bissau, Mozambique, Portugal, Sao Tome and Principe, and East Timor). The project comprises two inter-related phases: 1) the merging of existing Portuguese and Brazilian vocabularies as an evidence of the lexicographic tradition; 2) the compilation of corpora for the development of national vocabularies for the remaining countries. This paper describes in detail the methodology of the tasked being pursued in the context of that project.
We present a technique to improve out-of-domain statistical parsing by reducing lexical data sparseness in a PCFG-LA architecture. We replace terminal symbols with unsupervised word clusters acquired from a large newspaper corpus augmented with target domain data. We also investigate the impact of guiding out-of-domain parsing with predicted part-of-speech tags. We provide an evaluation for French, and obtain improvements in performance for both non-technical and technical target domains. Though the improvements over a strong baseline are slight, an interesting result is that the proposed techniques also improve parsing performance on the source domain, contrary to techniques such as self-training, thus leading to a more robust parser overall. We also describe new target domain evaluation treebanks, freely available, that comprise a total of about 3,000 annotated sentences from the medical domain, regional newspaper articles, French Europarl and French Wikipedia.
The overall goal of our work is to build a dependency grammar-based human sentence processor for Hindi. As a first step towards this end, in this paper we present a dependency grammar that is motivated by psycholinguistic concerns. We describe the components of the grammar that have been automatically induced using a Hindi dependency treebank. We relate some aspects of the grammar to relevant ideas in the psycholinguistics literature. In the process, we also extract statistics and patterns for phenomena that are interesting from a processing perspective. We finally present an outline of a dependency grammar-based human sentence processor for Hindi. Mary, ‘John ’ and ‘Mary ’ are connected via arcs to ‘kissed’; the former arc bears the label ‘subject’ and the latter arc the label ‘object’. Taken together, these nodes with their connections form a tree. Figure 1 shows the dependency and the phrase structure trees for the above sentence. 1
This report reviews and analyzes HSU, CSU, State, Federal, Tribal and professional policies, regulations or protocols that concern the extensive body of recordings, linguistic databases, publications and manuscripts that have been created and distributed by the Center for Indian Community Development in the course of nearly four decades.
The increasing significance of political communication in the functioning of the modern socium actualizes the problem of the communicative essence of power, forms and methods of its manifestation, the correlation of communicative and power components. The author researches the communicative specificity of political interaction, its moral-ethical principles, and gives the interpretation of the linguo-communicative code as a system of linguistic and extra-linguistic norms and rules that determine the success of modern political interaction.
Recent research has found it useful to distinguish between the form and meaning of sounds. To investigate the relevance of meaning, naïve students and professional drivers listened to four levels of meaning neutralisation and four levels of spectral slope of recorded truck sound. Self-assessment of emotional reactions showed that professional drivers did not vary much in activation and rated over all lower activation than naïve participants whose affect ratings moved more or less along the annoyance correlation line in the upper left quadrant of the affect map. This gives some information about the importance of the source being recognisable and of previous user experience for product sound quality. It is further supported by that the overall difference between naïve participants’ and professional drivers’ ratings decreased with increasing meaning neutralisation. The methodology applied in the current study may be adopted to form homogenous panels of experts for sound evaluation.
This paper builds six dependence syntactic networks based on six treebanks of different styles and gives a comparative analysis of overall characteristics of the networks, including the number of edges, the number of the nodes, the average degree, the clustering coefficient, the average path length, the centralization, the diameter, and the index of power-law, coefficient of determination. After that, the paper uses the Euclideanthe shortest distancemethod, with characteristics as variables, to do clustering analysis of these networks. The results show that using some main parameters of networks, namely the number of the nodes, the clustering coefficient, the average path length, the centralization and the index of power-law, can do cluster analysis on texts. Compared with the traditional text clustering, the results are easier to explain in linguistic angle.
Conventional statistics-based methods for joint Chinese word segmentation and part-of-speech tagging (S&T) have generalization ability to recognize new words that do not appear in the training data. An undesirable side effect is that a number of meaningless words will be incorrectly created. We propose an effective and efficient framework for S&T that introduces features to significantly reduce meaningless words generation. A general lexicon, Wikepedia and a large-scale raw corpus of 200 billion characters are used to generate word-based features for the wordhood. The word-lattice based framework consists of a character-based model and a word-based model in order to employ our word-based features. Experiments on Penn Chinese treebank 5 show that this method has a 62.9% reduction of meaningless word generation in comparison with the baseline. As a result, the F1 measure for segmentation is increased to 0.984.
This study undertook a critical appraisal of the correlation between the intractable social conflicts like the Boko-Haram and the Niger Delta crises, where youths are the key players, on the international image of Nigeria and tourism development in the country. It is motivated by the avalanche of media reports that the country’s image is being seriously battered abroad by these internal social problems. The specific objectives sought were to: ascertain the correlation between the Boko-Haram crisis and the nation’s image ratings abroad; the Niger Delta crisis and the nation’s image ratings abroad and their impacts on tourism development in the country. Survey design was adopted in the study, where electronic questionnaires (E-questionnaire) via the Internet were used to gather the primary data. The data so sourced were statistically presented/analyzed with Likert’s 5-points scale, Spearman’s correlation coefficient and Friedman chi-square. Results obtained show that both the Boko Haram crisis and the Niger Delta crisis have adverse impact on the country’s international image and tourism development, consequently on youths’ unemployment rate. It was then recommended that proactive public relations crisis management strategies should be used in nipping such crisis in their buds in future. Keywords: Boko Haram crisis, Niger Delta crisis, National Image, Tourism Development.
The paper presents a research tool for studying semantic change and polysemy patterns in Russian adjectives and adverbs. It is based on a corpus analysis of high-freqency polysemous units. For each of them we describe the meanings it can have, assign to each meaning a corresponding taxonomic class, identify types of semantic shifts between individual meanings (metaphor, metonymy; besides, a full-scale approach reveals non-canonical cases of semantic shifts), describe context conditions of these shifts (semantic and grammatical restrictions on co-occurring words). The results gained from this analysis are implemented in a database, which allows for various generalizations on the regularities of change in adjective and adverb meaning. Several examples are given to illustrate what kinds of queries can be performed on the database. Keywords: polysemy; semantic shift; metaphor; metonymy; semantics of adjectives; Russian language; lexical database
The increasing significance of political communication in the functioning of the modern socium actualizes the problem of the communicative essence of power, forms and methods of its manifestation, the correlation of communicative and power components. The author researches the communicative specificity of political interaction, its moral-ethical principles, and gives the interpretation of the linguo-communicative code as a system of linguistic and extra-linguistic norms and rules that determine the success of modern political interaction.
Recognizing and classifying implicit discourse relations is a challenging task since hardly any strong indicators exist, and a variety of weak indicators has to be harnessed to yield evidence for a particular discourse relation or another. Most current approaches rely on a combination of shallow, surface-based features and rather specialized hand-crafted features, with a considerable gap in between which is partly due to the sheer complexity of combining evidence from different levels of linguistic description. As a way to avoid both the shallowness of word-based representations and the lack of coverage of specialized linguistic features, we use a graph-based representation of discourse segments, which allows for a more abstract (and hence generalizable) notion of syntactic (and partially of semantic) structure, we propose an approach to use a graph-structured representation of discourse units in order to improve the classification of implicit discourse relations. We validate this approach using implicit discourse relation data from the TüBa-D/Z treebank, providing an extended discussion and error analysis that looks at the impact of the graph-based representation on the different kinds of discourse relations. The empirical evaluation shows that our graph-based approach not only provides a suitable representation for the linguistic factors that are needed in disambiguating discourse relations, but also improves results over a strong state-of-the-art baseline by more accurately identifying Temporal, Comparison and (for the German data) Reporting discourse relations. 1.
An MT-oriented system using Conditional Random Fields (CRFs) is presented to identify English Prepositional Phrases (PPs) within business domain. For the purpose of English-Chinese Machine Translation (MT), we, under the guidance of the theory of Syntactic Functional Grammar (SFG), refine PP function chunks into four types instead of the binary attachment. In order to improve the identification of these chunk types, we revise the Penn Treebank tagset with four major changes being made. A small size of 998k English annotated corpus in business domain is semi-automatically built based on our new tagset employing the Maximum Entropy model. Experiments show that our system achieves an accuracy of 88.45%, higher than other reported approaches. The adjustments made in the PP chunk types and POS tagset give rise to 4.11%, 4.25% and 4.15% increase in the precision, recall and F-score respectively.
This paper addresses semantic search of Web services using natural language processing. First we survey various existing approaches, focusing on the fact that the expensive costs of current semantic annotation frameworks result in limited use of semantic search for large scale applications. We then propose a service search framework based on the vector space model to combine the traditional frequency weighted term-document matrix, the syntactical information extracted from a lexical database and a dependency grammar parser. In particular, instead of using terms as the rows in a term-document matrix, we propose using synsets from WordNet to distinguish different meanings of a word under different contexts as well as clustering different words with similar meanings. Also based on the characteristics of Web services descriptions, we propose an approach to identifying semantically important terms to adjust weightings. Our experiments show that our approach achieves its goal well.
Information management is an important requirement in today"s world. The anaphoric references hide the important information. The identification of anaphorically referred information is called as anaphora resolution which has significant impact on increasing the efficiency of information management, techniques including text summarization, information extraction etc. In this paper we have proposed a method for anaphora resolution to engineer information management. The proposed method acceptably determines the potential referents of the anaphora specially of the verb phrase form, distinguishes between pleonastic "it" and the anaphoric "it" and resolve the anaphora which is referred to after an interval of multiple sentences. The referents are stored in a list in the order of their occurrence in the discourse and eliminated from the list if they are not referred for to long. "Recency" is used as a salience factor to select the correct referent if other information like gender, number and type are not suffices to estimate the correct referent for an anaphor. To achieve a more precise resolution system WordNet lexical database is exploited to compare the synonyms of the anaphor with its possible referents.
Inspired by robust generalization and adversarial learning we describe a novel approach to learning structured perceptrons for part-ofspeech (POS) tagging that is less sensitive to domain shifts. The objective of our method is to minimize average loss under random distribution shifts. We restrict the possible target distributions to mixtures of the source distribution and random Zipfian distributions. Our algorithm is used for POS tagging and evaluated on the English Web Treebank and the Danish Dependency Treebank with an average 4.4 % error reduction in tagging accuracy. 1
SYNTAGMA is a rule-based parsing system, structured on two levels: a general parsing engine and a language specific grammar. The parsing engine is a language independent program, while grammar and language specific rules and resources are given as text files, consisting in a list of constituent structuresand a lexical database with word sense related features and constraints. Since its theoretical background is principally Tesniere's Elements de syntaxe, SYNTAGMA's grammar emphasizes the role of argument structure (valency) in constraint satisfaction, and allows also horizontal bounds, for instance treating coordination. Notions such as Pro, traces, empty categories are derived from Generative Grammar and some solutions are close to Government&Binding Theory, although they are the result of an autonomous research. These properties allow SYNTAGMA to manage complex syntactic configurations and well known weak points in parsing engineering. An important resource is the semantic network, which is used in disambiguation tasks. Parsing process follows a bottom-up, rule driven strategy. Its behavior can be controlled and fine-tuned.
The paper describes a broadly applicable method of designing multilingual semantics-syntactic analyzers of recommender systems. The user inputs may include the questions of many kinds formed with the help of interrogative words (or without interrogative words), verbs, nouns, attributes, prepositions, the designations of the digital values of various parameters. For the queries in English and German, the developed algorithm of semantic-syntactic analysis processes the questions of many kinds, the commands, and the statements from a restricted sublanguage of NL. For the queries in Russian, the algorithm is additionally able to process the requests with participle constructions and attributive clauses. As a semantic intermediary language, the algorithm uses the SK-language determined by the considered linguistic database. The class of SK-languages is introduced by the theory of K-representations (knowledge representations), its current version is mainly stated in a monograph of the author published by Springer in 2010. The developed algorithm is implemented by means of the programming language PYTHON.
This experiment investigated the role of acoustic correlates of the singing voice in the perception of broad affect dimensions using the two-dimensional model of affect. The dataset consisted of vocal and glottal recordings of a sung vowel interpreted in different singing expressions. Listeners were asked to rate the sounds according to four perceived affect dimensions. A cross-tabulation was done between the singing expressions and affect judgments. A one-way ANOVA was performed for 11 acoustic cues with the affect ratings. It was found that the singing power ratio (SPR), mean intensity, brightness, mean pitch, jitter, shimmer, mean harmonic-to-noise ratio (HNR), and mean autocorrelation discriminate broad affect dimensions. Principal component analysis (PCA) was performed on the acoustic correlates. Two components were retained that explained 78.1% of the total variance of vocal cues and 73.5% of that of the glottal cues.
Automatic syntactic analysis of natural language is one of the fundamental problems in natural language processing. Dependency parses (directed trees in which edges represent the syntactic relationships between the words in a sentence) have been found to be particularly useful for machine translation, question answering, and other practical applications. For English dependency parsing, we show that models and features compatible with how conjunctions are represented in treebanks yield a parser with state-of-the-art overall accuracy and substantial improvements in the accuracy of conjunctions. For languages other than English, dependency parsing has often been formulated as either searching over trees without any crossing dependencies (projective trees) or searching over all directed spanning trees. The former sacrifices the ability to produce many natural language structures; the latter is NP-hard in the presence of features with scopes over siblings or grandparents in the tree. This thesis explores alternative ways to simultaneously produce crossing dependencies in the output and use models that parametrize over multiple edges. Gap inheritance is introduced in this thesis and quantifies the nesting of subtrees over intervals. The thesis provides O( n6) and O(n 5) edge-factored parsing algorithms for two new classes of trees based on this property, and extends the latter to include grandparent factors. This thesis then defines 1-Endpoint-Crossing trees, in which for any edge that is crossed, all other edges that cross that edge share an endpoint. This property covers 95.8% or more of dependency parses across a variety of languages. A crossing-sensitive factorization introduced in this thesis generalizes a commonly used third-order factorization (capable of scoring triples of edges simultaneously). This thesis provides exact dynamic programming algorithms that find the optimal 1-Endpoint-Crossing tree under either an edge-factored model or this crossing-sensitive third-order model in O(n 4) time, orders of magnitude faster than other mildly non-projective parsing algorithms and identical to the parsing time for projective trees under the third-order model. The implemented parser is significantly more accurate than the third-order projective parser under many experimental settings and significantly less accurate on none.
1 Kateřina Rysová Annotation The presented thesis is focused on the Czech word order of contextually non-bound verbal modifications. It monitors whether there is a basic order in the contextually non-bound part of the sentence (significantly predominant in frequency) in the surface word order (cf. narodit se v Brně v roce 1950 vs. narodit se v roce 1950 v Brně; literally to be born in Brno in 1950 vs. to be born in 1950 in Brno). At the same time, we try to find out the factors influencing the word order (such as the form of modifications, their lexical expression or the effect of verbal valency). Finally, we briefly compare the word order tendencies in Czech and German. For the verification of the objectives, mainly the data from the Prague Dependency Treebank are used. The work is based on the theoretical principles of Functional Generative Description. Research results demonstrate that, at least in some cases, it is possible to detect certain general tendencies to use preferably one of two possible surface word order sequences in Czech. Abstract The aim of the doctoral thesis is to describe particular aspects of the Czech (and partly also German) word order in the sentences coming mainly from journalistic texts. The first part examines the role of different types of verbal modifications in sentence...
Historical linguistics, among other things, aims at understanding the principles and factors that cause changes in languages. The Dravidian comparative linguistics in the last few decades has arrived at excellent results at different levels of language change: phonology, morphology and etymology. However, the field of historical syntax remains to be explored in detail. The linguistic analysis of Tamil inscriptions and classical and ancient literary texts will shed light on the historical linguistics of Tamil and will try to fill a gap in the historical linguistics of the Dravidian family of languages. An in-depth linguistic analysis of Tamil epigraphic texts will show how the Tamil language used in Tamil inscriptions constitutes an important diachronic evidence of both sociolinguistic and linguistic evolution. I will concentrate here on the following three aspects: 1) Historical sociolinguistics: Maṇipravāḷa style and the development of Tamil as Inscriptional Language, 2) Historical linguistics, Syntax and Information structure, and 3) Construction of a fine-grained linguistic database and demonstrate how ‘corpus analysis’ can help us mapping the process of language change and language use.
The dissertation is a three-paper investigation of potential methods for improving exposure therapy outcomes for anxiety disorders. If proven effective, these methods can be utilized to modify exposure therapy for anxiety disorders in order to enhance treatment effects and reduce relapse rates. Study 1 tested a prediction of the Rescorla-Wagner model: that presenting two fear-provoking stimuli simultaneously (compound extinction) would maximize learning during extinction. Participants were presented with single extinction trials only or single extinction trials followed by compound extinction trials. Additionally, participants were randomized to caffeine or placebo ingestion prior to extinction. Results indicated participants presented with compound trials demonstrated significantly less fear responding at spontaneous recovery whereas ingestion of caffeine provided limited protection (only on valence ratings). At reinstatement, only compound extinction trials predicted attenuated fear responding. Study 2 investigated whether occasional reinforced trials during extinction enhances learning, as has been demonstrated in recent literature in animal models. Participants were randomly assigned to typical exposure procedures (i.e., no CS-US pairings) or occasional CS-US pairings during extinction. Results indicated participants presented with occasional reinforced trials maintained elevated fear responding throughout exposure but demonstrated attenuated spontaneous recovery and rapid reacquisition effects at follow-up testing. Study 3 was based on recent evidence in animal models suggesting that sustained arousal and enhanced fear responding throughout extinction predicts better long-term outcomes, which is contrary to traditional exposure therapy in which reduction of fear responding is used as an index of learning. Participants completed exposure with or without the presence of additional excitatory stimuli which were intended to enhance arousal and fear responding throughout exposure. A set of regression analyses investigating whether any exposure process measures predicted outcome indicated that sustained arousal throughout exposure and variability in subjective fear responding throughout exposure predicted lower levels of fear at follow-up testing. In sum, these studies indicate that exposure therapy will be most effective when exposure sessions are unpredictable, variable, include multiple fear-provoking stimuli, and include some aversive events (e.g., social rejection, panic attack). Although fear responding may remain elevated throughout exposure, such procedures may predict better treatment outcomes and reduce the likelihood of relapse.
The topicality of material presented in the article is conditioned with the lack of research aimed at the formation of a culture of dialogue speech of students of nonphilological specialties. The author of the article has considered the dialog, voice, speech, cultural, communicative culture in forming the dialogues, the functioningof which depends on the ability to listen, hear and understand the interlocutor, which is aimed at the development ofskills to produce dialogic speech. Keywords: culture of dialogic speech, linguistic features of dialogs, vocal and linguistic norms, vocal etiquette, communication.
Lucas (3;6) is playing with Lego animals on the carpet. He takes the lions and puts them into the compound he prepared for them. He comments: “Kuck, e Léiw!” (Look, a lion!) Then he whirls them all around and says: “Tout mélanger!” (Mixing everything!) Lucas here seems not only to mix the animals but the languages, too… Linguistic diversity is not only an integral part of Luxembourgian society in general but also of the everyday practice in early childcare settings. This is nothing exceptional. Rather, the increasing diversification of languages, cultures, and identities – or the increasing acknowledgement that these concepts have never been simple and fixed – is a central characteristic of contemporary societies worldwide. Nevertheless, education political and media discourses keep on positing multilingualism as a special challenge that pedagogical practice has to cope with. They call for early language promotion, school preparation, and for the advancement of social integration and equality. While these discourses instantly turn to the programmatic question how multilingualism should be dealt with, thus presupposing a normative understanding of language in education, the present paper asks how these complex demands are actually met in everyday pedagogical practice and how, along the way, linguistic norms are practically accomplished as well as caught into question. The paper draws on ethnographic material from three Luxembourgian daycare centers that were investigated during 18 months as part of my doctoral research. The choice and interpretation of field notes is guided by the central question how linguistic diversity is dealt with in the centers’ everyday routines. The empirical exploration reveals how pedagogical practice is itself constituted within a field of tension between monolingualist agendas and the actors’ translingual practices. The three centers manage this tension differently, thus demonstrating the multiplicity of possible pathways when dealing with multilingualism in early education.
Chinese divides into simplified Chinese and traditional Chinese. Some treebank resources like Penn Chinese Treebank: CTB had been built for training simplified Chinese parser (Yu, et al. 2010) while Sinica Treebank was developed for parsing traditional Chinese (Chen et al., 1999). Limit to our knowledge, there are still not grammatical resources that analyze both simplified Chinese and traditional Chinese. A rule-based Chinese grammatical resource --Chinese Sentence Structure Grammar: CSSG had been developed based on the idea of Sentence Structure Grammar: SSG (Wang et al., 2012). We assume that a rule-based grammatical resource should analyze both simplified Chinese and traditional Chinese if there are no obvious differences between their grammatical constructions. Aiming at verifying this assumptions, we parse the test sentences from the simplified Chinese parsing task (task 3) and the traditional Chinese parsing task (task 4) of CLP 2012 with the same rule-based parser that was implemented the grammatical resource CSSG. CSSG includes two parts of resources: the grammatical rules and a simplified Chinese morphological dictionary. We transfer the simplified Chinese characters of the dictionary to traditional Chinese characters for obtaining a traditional Chinese morphological dictionary. We parse the test sentences of task 3 and task 4 with the same CSSG rules but different morphological dictionaries (simplified or traditional Chinese characters). We convert CSSG parsing trees to TCT-style trees and Sinica-style trees to participate in the evaluations of the two tasks. The experiments show that the CSSG rules can parse both simplified Chinese and traditional Chinese, but the performance of the latter is lower than the former. We noticed that a few traditional Chinese constructions are different from simplified Chinese.
Objective – To assess how the age, gender, and race characteristics of library users affect their perceptions of the approachability of reference librarians with similar or different demographic characteristics.
 
 Design – Image rating survey.
 
 Setting – Large, three-campus university system in the United States.
 
 Subjects – There were 449 students, staff, and faculty of different ages, gender, and race.
 
 Methods – In an online survey respondents were presented with images of hypothetical librarians and asked to evaluate their approachability, using a scale from 1 to 10. The images showed librarians with neutral emotional expressions against a standardized, neutral background. The librarians’ age, gender, and race were systematically varied. Only White, African American, and Asian American librarians were shown. Afterwards respondents were asked to identify their own age, gender, race, and status.
 
 Main Results – Respondents perceived female librarians as more approachable than male librarians, maybe due to expectations caused by the female librarian stereotype. They found librarians of their own age group more approachable. African American respondents scored African American librarians as more approachable, whereas Whites expressed no significant variation when rating the approachability of librarians of different races. Thus, African Americans demonstrated strong in-group bias but Whites manifested colour blindness – possibly a strategy to avoid the appearance of racial bias. Asian Americans rated African American librarians lower than White librarians.
 
 Conclusion – This study demonstrates that visible demographic characteristics matter in people’s first impressions of librarians. Findings confirm that diversity initiatives are needed in academic libraries to ensure that all users feel welcome and are encouraged to approach librarians. Regarding gender, programs that deflate the female librarian stereotype may help improve the approachability image of male librarians. Academic libraries should staff the reference desk with individuals covering a wide range of ages, including college-aged interns, whom traditional age students find most approachable. Libraries should also build a racially diverse staff to meet the needs of a racially diverse user population. Since first impressions have lasting effects on the development of social relationships, structural diversity should be a priority for libraries’ diversity programs.
Introduction The processing of nouns and verbs and their differences have been an area of intense interest among psycholinguists (see review by Vigliocco et al., 2010), for the obvious reasons that they are two major word classes across languages and convey the most basic information in communication. Word class effects are often reflected in response latency and/or accuracy in naming tasks. However, single word production does not resemble daily communication in which linguistic contexts may facilitate word finding. Previous studies directly comparing lexical retrieval between naming and narrative tasks have obtained mixed results (e.g. Berndt et al., 2002; Pashek & Tompkins, 2002), despite the fact that nouns and verbs were rarely matched for relevant psycholinguistic variables. This study minimized the influence of confounding factors and employed neuropsychological data to examine retrieval of nouns and verbs in confrontation naming and connected speech. Method The participants were 19 Cantonese-speaking adults with anomic aphasia and 19 age-, gender- and education-matched controls. Production of nouns and verbs was obtained from confrontation picture naming and narrative tasks from the Cantonese AphasiaBank database (Kong et al., 2009). At least 20 items in each condition were chosen with comparable age of acquisition and familiarity estimates; however, imageability ratings were higher in naming than narrative tasks and higher for nouns than verbs, and verbs in naming were longer than those in narrative task. Results and Discussion Significant main effects of speaker group, word class, and task, as well as a two-way interaction between task and word class were found (p < 0.01). Better performance in nouns than verbs and naming in picture than narrative tasks? was observed in normal speakers. The difference in accuracy between word classes was greater in naming than narrative tasks. A hierarchical multiple regression was also carried out to assess the effects of word class and task after the influence of imageability had been taken into consideration. Only “task” remained a significant predictor (p < 0.05). Our results have shown that when the influence of confounding factors is reduced, there is no evidence for word class specific deficits among fluent aphasic speakers (but see Matzig et al., 2009) or contextual support for word production.
The Electronic Health Record (EHR) contains information useful for clinical, epidemiological and genetic studies. This information of patient symptoms, history, medication and treatment is not completely captured in the structured part of the EHR but is often found in the form of freetext narrative. A major obstacle for clinical studies is finding patients that fit the eligibility criteria of the study. Using EHR in order to automatically identify relevant cohorts can help speed up both clinical trials and retrospective studies (Restificar, Korkontzelos et al. 2013). While the clinical criteria for inclusion and exclusion from the study are explicitly stated in most studies, automating the process using the EHR database of the hospital is often impossible as the structured part of the database (age, gender, ICD9/10 medical codes, etc.’) rarely covers all of the criteria. Many resources such as UMLS (Bodenreider 2004), cTakes (Savova, Masanz et al. 2010), MetaMap (Aronson and Lang 2010) and recently richly annotated corpora and treebanks (Albright, Lanfranchi et al. 2013) are available for processing and representing medical texts in English. Resource poor languages, however, suffer from lack in NLP tools and medical resources. Dictionaries exhaustively mapping medical terms to the UMLS medical meta-thesaurus are only available in a limited number of languages besides English. NLP annotation tools, when they exist for resource poor languages, suffer from heavy loss of accuracy when used outside the domain on which they were trained, as is well documented for English (Tsuruoka, Tateishi et al. 2005; Tateisi, Tsuruoka et al. 2006). In this work we focus on the problem of classifying patient eligibility for inclusion in retrospective study of the epidemiology of epilepsy in Southern Israel. Israel has a centralized structure of medical services which include advanced EHR systems. However, the free text sections of these EHR are written in Hebrew, a resource poor language in both NLP tools and handcrafted medical vocabularies. Epilepsy is a common chronic neurologic disorder characterized by seizures. These seizures are transient signs and/or symptoms of abnormal, excessive, or hyper synchronous neuronal activity in the brain. Epilepsy is one of the most common of the serious neurological disorders (Hirtz, Thurman et al. 2007).
Objective To explore the parental perception of school-age children's body images and its influence factors among parents of school children in Changsha city,Hunan province,and to provide reference for childhood obesity prevention. Methods A survey with anthropometric measurement(height,w eight) was conducted among 2 224 elementary school students of grade 4-6 in Changsha city in April,2012. Body image was categorized based on the WHO 2007 body mass index(BMI) reference standard. The parental perception of the body image was examined with a questionnaire among the parents of the school children and the influence factors of the perception were analyzed with logistic regression. Results The body image rating scale results show ed that 56. 1% of the parents underestimated their children' s body image size. Multiple-factors logistic regression show ed that the risk factors for underestimating the children's body image was child's big body image(odds ratio [OR]= 35. 763,95% confidence interval [95% CI]= 23. 745-53. 863), w hile the protective factor was female parent(OR = 0. 623,95% CI = 0. 400-0. 969). Residing in the countryside(OR =1. 464,95%CI =1. 090-1. 966),with the child in senior grade(OR =1. 272,95%CI =1. 067-1. 517),and with the child having big body images(OR = 32. 089,95% CI = 16. 810-61. 255) were risk factors for parental underestimation of children's body image,w hile high parental education(OR = 0. 870,95% CI = 0. 773-0. 979) and overw eight or obese of the parents(OR = 0. 578,95% CI = 0. 403-0. 830) were protective factors. Conclusion Parental perception of school children's body image prensents an underestimation trend among the parents in Changsha. We should reinforce the communication and education in the parents to promote the correct evaluation on their children's body images for the control of overw eight and obesity in the school children.
We describe in this paper how different learning strategies can be applied on the same NLP task, namely chunking. The reference corpus is extracted from the French Treebank, the symbolic learning strategy used is grammatical inference and the statistical one is CRFs (Conditional Random Fields). As expected, the symbolic approach allows readability but is less effective than the statistical one. We then propose two distinct ways to combine both approaches and show that in both cases they benefit from one another.
Ziele: To evaluate the possibilities in applying a novel algorithm to correct for beam hardening artifacts caused by metal implants in computed tomography performed on a flat panel equipped c-arm angiography system (FP-CT). Methode: 35 datasets of cerebral FP-CT acquisitions (XperCT, Allura Xper FD 20/20, Philips) with present implants (32 patients; 27 FP-CT after coil-embolization; 8 after aneurysm clipping, these with intravenous application of contrast media) have been reconstructed applying smooth (soft tissue) and sharp (implant) kernels with and without a novel reconstruction filter for metal artifact correction, resulting in high-resolution isotropic datasets (edge length: 0.15 mm for sharp kernel and 0.3 mm for smooth kernel, respectively). Image viewing was performed in multiplanar reformations (MPR) in both average and maximum intensity projection (MIP) mode, sharp kernel images additionally in 3D MIP and 3D volume rendering (VR) on a dedicated radiological workplace (PACS IW, GE). Two independent radiologists performed image rating following a defined scale in direct comparison of the image data with and without metal artifact correction, weighted kappa statistics were calculated. Ergebnis: Inter-rater agreement was at least substantial in all regards. Soft tissue image quality at the level of the implants was substantially improved; the additional metal artifact correction algorithm did not induce relevant falsification of the data. In addition, two factors were identified to be having an influence on the quality of the results: Firstly, volume/amount of the implants; secondly, orientation of a non-spherical implant relatively to the plane of acquisition. Schlussfolgerung: Adding metal artifact correction to FP-CT may in future help to spread the applicability of this technique, especially regarding non-invasive follow-up after clipping or coil embolization of intracranial aneurysms with FP-CT with intravenous application of contrast media.
The accuracy of Chinese parsers trained on Penn Chinese Treebank is evidently lower than that of the English parsers trained on Penn Treebank. It is plausible that the essential reason is the lack of surface syntactic constraints in Chinese. In this paper, we present evidences to show that strict deep syntactic constraints exist in Chinese sen-tences and such constraints cannot be effectively described with context-free phrase structure rules as in the Penn Chinese Treebank annotation; we show that such constraints may be described pre-cisely by the idea of Sentence Structure Grammar; we introduce how to develop a broad-coverage rule-based grammar for Chinese based on this idea; we evaluated the grammar and the evaluation results show that the coverage of the current grammar is 94.2%. 1
Analyzes the characteristics of the use of the normalized and non-normalized order of words in journalistic texts on the material of some periodicals of Chelyabinsk. It is shown that non-linguistic norms leads to difficulty perception of published materials.
Стаття представляє спробу науково-обґрунтованого введення терміну «квант лінгвістичної інформації». Такий підхід дає змогу впритул підійти до встановлення принципів формального структурування різних типів з метою їх комп’ютерної обробки. Запропоновано аналіз квантів лінгвістичної інформації у лексичній базі даних DANTE.The article presents scientific and practical reasons for introducing the new term “quantum of linguistic information”. The notion of the quantum of linguistic information gives way to approaching the problem of discriminating formal structures for decoding linguistic information. Such decoding enables computational processing of linguistic information in lexical databases. The quantum of linguistic information is formal and computationally explicit fragment of linguistic information, which serves as a field marker in lexical database. The field markers of the DANTE Lexical Database of Modern English have been analyzed as the basic quanta of linguistic information to be recognized and processed by computer.
The article deals with the problem of norm and normative approach to the diachronic language study. It identifies specificity of the normative approach to linguistic means within various linguistic traditions and determines the main features of the linguistic norm. The author points out that specific norms appear at each stage of language development as the result of correlation of the existing language means.
The semantic web is a synergetic movement led by International standards body, the WWW Consortium (W3C).It aims at converting the current web dominated by unstructured and semi structured documents into a "web of data". Here two techniques of semantic web crawling are reviewed, one is ontology based and other is based on Lexical database.For this, architecture has been proposed which is a combination of above two techniques. The future of WWW is semantic web where Ontology and Lexical database are used for effective and fast searching by the web crawler. It is used for Information retrieval and question answering system. Ontology is a formal designation of shared approach it is basically approach of entities and their attributes.
In this paper we describe the expansion of probabilistic context free grammar for Urdu language, we did some experiments in probabilistic context free grammar for Urdu language, extraction of CFG rules through The Penn Treebank, evaluation of PCFG from CFG. This PCFG is further useful for parsing the Urdu Sentences. The Tree-bank based grammar is the best technique for building up the PCFG for language through some easy and understandable steps as compared to theoretical. A PCFG can be used to estimate a number of useful probabilities concerning a sentence and its parse-tree(s). The resulting Penn Treebank based PCFG is widely used in natural language processing, speech recognition, and integrated spoken language systems as well as in theoretical linguistics.
The Penn Discourse Treebank (PDTB) was released to the public in 2008 and remains the largest corpus of manually annotated discourse relations — both relations that are signaled explicitly (e.g., by a coordinating or subordinating conjunction, or by a discourse adverbial or other construction) and ones that otherwise appear implicit. The Penn Discourse TreeBank also diverges from other discourse-annotated corpora in permitting more than one discourse relation to be annotated as holding concurrently. Annotators could indicate this by assigning multiple sense labels to an explicit connective. Or, in those cases where adjacent sentences had no explicit connective, annotators could indicate concurrent discourse relations by either annotating a single implicit connective that concurrently conveyed multiple senses or annotating multiple implicit connectives, each conveying one of the concurrent relation(s). Subsequent experiments carried out using Mechanical Turk showed that, when a discourse adverbial explicitly signalled a discourse relation, there was often a separate concurrent relation that could be associated with an implicit coordinating or subordinating conjunction. There are different circumstances in which different sets of concurrent discourse relations are taken to hold. I will go through these, and conclude with what I take the implications of this to be for various language technologies, including statistical machine translation.
We provide an overview of forty years of work with language corpora by the research group that started in 1972 as the Norwegian Computing Centre for the Humanities. A brief history highlights major corpora and tools that have been developed in numerous collaborations, including corpora of literature, dialect recordings, learner language, parallel texts, newspaper articles, blog posts and tweets. Current activities are also described, with a focus on corpus analysis tools, treebanks and social media analysis.
In this paper we present tools prepared for morphological and syntactic processing of Slovak: a model trained for tagging by the RFTagger and two syntactic analyzers Synt and SET for which we adapted their Czech grammars for Slovak. We describe the training process of RFTagger using the r-mak corpus and modifications of both parsers that have been performed partially in the lexical analysis and mainly in the formal grammars used in both systems. Finally we provide an evaluation of both tagging and parsing, the latter on two datasets – a phrasal and dependency treebank of Slovak.