Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
sponsorship: University of Leuven
The article deals with the functional treatment of the concept of number in the English noun system and looks into semantic features of substantive units, their lexical and grammatical collocation and ability to express the linguacultural peculiarities of communication. The authors focus on extralinguistic parametres aimed at better understanding the semantics of the unit and construing more effectively one’s own statement in accordance with the aim set and keeping within the varieties of the Modern English norm.
International audience
Large-scale linguistically annotated cor-pora have played a crucial role in advanc-ing the state of the art of key natural lan-guage technologies such as syntactic, se-mantic and discourse analyzers, and they serve as training data as well as evaluation benchmarks. Up till now, however, most of the evaluation has been done on mono-lithic corpora such as the Penn Treebank, the Proposition Bank. As a result, it is still unclear how the state-of-the-art analyzers perform in general on data from a vari-ety of genres or domains. The completion of the OntoNotes corpus, a large-scale, multi-genre, multilingual corpus manually annotated with syntactic, semantic and discourse information, makes it possible to perform such an evaluation. This paper presents an analysis of the performance of publicly available, state-of-the-art tools on all layers and languages in the OntoNotes v5.0 corpus. This should set the bench-mark for future development of various NLP components in syntax and semantics, and possibly encourage research towards an integrated system that makes use of the various layers jointly to improve overall performance. 1
ABSTRACT Online travel reviews are emerging as a powerful source of information affecting tourists' pre-purchase evaluation of a hotel organization. This trend has highlighted the need for a greater understanding of the impact of online reviews on consumer attitudes and behaviors. In view of this need, we investigate the influence of online hotel reviews on consumers' attributions of service quality and firms' ability to control service delivery. An experimental design was used to examine the effects of four independent variables: framing; valence; ratings; and target. The results suggest that in reviews evaluating a hotel, remarks related to core services are more likely to induce positive service quality attributions. Recent reviews affect customers' attributions of controllability for service delivery, with negative reviews exerting an unfavorable influence on consumers' perceptions. The findings highlight the importance of managing the core service and the need for managers to act promptly in addressing customer service problems.
This chapter examines the way in which Aristophanes introduces obscene words into his comedies both at the beginning of the plays and subsequently, following more heightened and/or more sober sequences. The Aristophanic norm is to introduce obscenity unsignalled, the 'obscenity out of nowhere' technique, often employed to signal abuse, crudeness, buffoonery and/or freedom from inhibitions. Alternatively, the poet sometimes employs the 'build-up' technique, in which double entendres and sexual allusions occur with increasing intensity before a climactic primary obscenity is finally introduced. Examples of both techniques are analysed, and some of the challenges that Aristophanic obscenity present and the relationship between obscenity and paratragedy are explored.
The Treebanks as the sets of syntactically annotated sentences, are the most widely used language resource in the application of Natural Language Processing. The occurrence of errors in the automatically created Treebanks is one of the main obstacles limiting the using of these resources in the real world applications. This paper aims to introduce an statistical method for diminishing the amount of errors occurred in a specific English LTAG-Treebank proposed in Basirat and Faili (2013). The problem has been formulated as a classification problem and has been tackled by using several classifiers. The experiments show that by using this approach, about 95% of the errors could be detected and more than 77% of them could successfully be corrected in the case of using Adaboost classifier. In addition, it has been shown that the new treebank could reach a high of 76% F-measure which is 8% higher than the original treebank.
We describe a novel approach to detecting empty categories (EC) as represented in de-pendency trees as well as a new metric for measuring EC detection accuracy. The new metric takes into account not only the position and type of an EC, but also the head it is a dependent of in a dependency tree. We also introduce a variety of new features that are more suited for this approach. Tested on a sub-set of the Chinese Treebank, our system im-proved significantly over the best previously reported results even when evaluated with this more stringent metric. 1
on, Tweets Abstract: The lexical richness and its ease of access to large volumes of information converts the Web 2.0 into an important resource for Natural Language Processing. Nevertheless, the frequent presence of non-normative linguistic phenomena that can make any automatic processing challenging. In this paper is described the partici- pation in the Text Normalisation Workshop at the SEPLN conference (Tweet-norm 2013). The Workshop includes one unique task focused on the normalisation of Spa- nish tweets. For this task we have used TENOR, a multilingual lexical normalisation tool for Web 2.0 texts.
This paper presents a reranking approach to combining constituent and dependency parsing, aimed at improving parsing performance on both sides. Most previous combination methods rely on complicated joint decoding to integrate graph- and transition-based dependency models. Instead, our approach makes use of a high-performance probabilistic context free grammar (PCFG) model to output k-best candidate constituent trees, and then a dependency parsing model to rerank the trees by their scores from both models, so as to get the most probable parse. Experimental results show that this reranking approach achieves the highest accuracy of constituent and dependency parsing on Chinese treebank (CTB5.1) and a comparable performance to the state of the art on English treebank (WSJ).
peer reviewed
Recent developments in Natural Language Processing (NLP) are heading towards knowledge rich resources and technology. Integration of linguistically sound grammars, sophisticated machine learning settings and world knowledge background is possible given the availability of the appropriate resources: deep multilingual treebanks, representing detailed syntactic and semantic information; and vast quantities of world knowledge information encoded within ontologies and Linked Open Data datasets (LOD). Thus, the addition of world knowledge facts provides a substantial extension of the traditional semantic resources like WordNet, FrameNet and others. This extension comprises numerous types of Named Entities (Persons, Locations, Events, etc.), their properties (Person has a birthDate; birthPlace, etc.), relations between them (Person works for an Organization), events in which they participated (Person participated in war, etc.), and many other facts. This huge amount of structured knowledge can be considered the missing ingredient of the knowledgebased NLP of 80’s and the beginning of 90’s. The integration of world knowledge within language technology is defined as an ontology-to-text relation comprising different language and world knowledge in a common model. We assume that the lexicon is based on the ontology, i.e. the word senses are represented by concepts, relations or instances. The problem of lexical gaps is solved by allowing the storage of not only lexica, but also free phrases. The gaps in the ontology (a missing concept for a word sense) are solved by appropriate extensions of the ontology. The mapping is partial in the sense that both elements (the lexicon and the ontology) are artefacts and thus — they are never complete. The integration of the interlinked ontology and lexicon with the grammar theory, on the other hand, requires some additional and non-trivial reasoning over the world knowledge. We will discuss phenomena like selectional constraints, metonymy, regular polysemy, bridging relations, which live in the intersective areas between world facts and their language reflection. Thus, the actual text annotation on the basis of ontology-to-text relation requires the explication of additional knowledge like co-occurrence of conceptual information, discourse structure, etc. Such knowledge is mainly present in deeply processed language resources like HPSG-based (LFG-based) treebanks (RedWoods treebank, DeepBank, and others). The inherent characteristics of these language resources is their dynamic nature. They are constructed simultaneously with the development of a deep grammar in the corresponding linguistic formalism. The grammar is used to produce all potential analyses of the sentences within the treebank. The correct analyses are selected manually on the base of linguistic discriminators which would determine the correct linguistic production. The annotation process of the sentences provides feedback for the grammar writer to update the grammar. The life cycle of a dynamic language resource can be naturally supported by the semantic technology behind the ontology and LOD modeling the grammatical knowledge as well as the annotation knowledge; supporting the annotation process; reclassification after changes within the grammar; querying the available resources; exploitation in real applications. The addition of a LOD component to the system would facilitate the exchange of language resources created in this way and would support the access to the existing resources on the web.
International audience
The present paper focuses on ways in which the pragmatic (functional) meaning that arises from various contextual features, known in corpus linguistics as semantic prosody, can become an integral part of lexicographical descriptions as they are represented in the Slovene Lexical Database (SLD). This is particularly important for the treatment of phraseology and idiomatics. First, the theoretical background is provided, with the focus on the prototype theory and its practical implications for monolingual lexicography. A parallel is drawn with the model of meaning analysis in the SLD. The second part begins with a brief introduction to semantic prosody and continues with an analysis of monolingual meaning descriptions in the SLD against a number of authentic corpus examples, investigating how their pragmatic components have been identified. The analysis of corpus data shows that pragmatics is an important contributor to the process of sense discrimination in works of lexical and lexicographic relevance.
If language comprehension requires a sensorimotor simulation, how can abstract language be comprehended? We show that preparation to respond in an upward or downward direction affects comprehension of the abstract quantifiers “more and more” and “less and less” as indexed by an N400-like component. Conversely, the semantic content of the sentence affects the motor potential measured immediately before the upward or downward action is initiated. We propose that this bidirectional link between motor system and language arises because the motor system implements forward models that predict the sensory consequences of actions. Because the same movement (e.g., raising the arm) can have multiple forward models for different contexts, the models can make different predictions depending on whether the arm is raised, for example, to place an object or raised as a threat. Thus, different linguistic contexts invoke different forward models, and the predictions constitute different understandings)
With the growing interest in statistical parsing, special attention has recently been devoted to the problem of comparing different treebanks to assess which languages or domains are more difficult to parse relative to a given model. A common methodology for comparing parsing difficulty across treebanks is based on the use of the standard labeled precision and recall measures. As an alternative, in this article we propose an information-theoretic measure, called the expected conditional cross-entropy (ECC). One important advantage with respect to standard performance measures is that ECC can be directly expressed as a function of the parameters of the model. We evaluate ECC across several treebanks for English, French, German, and Italian, and show that ECC is an effective measure of parsing difficulty, with an increase in ECC always accompanied by a degradation in parsing accuracy.
National audience
This paper discusses the extension of a sys-tem developed for automatic discovery of tree-bank annotation inconsistencies over an entire corpus to the particular case of evaluation of inter-annotator agreement. This system makes for a more informative IAA evaluation than other systems because it pinpoints the incon-sistencies and groups them by their structural types. We evaluate the system on two corpora- (1) a corpus of English web text, and (2) a corpus of Modern British English. 1
This paper reports our ongoing project for constructing an English multiword expression (MWE) dictionary and NLP tools based on the developed dictionary. We extracted functional MWEs from the English part of Wiktionary, annotated the Penn Treebank (PTB) with MWE information, and conducted POS tagging experiments. We report how the MWE annotation is done on PTB and the results of POS and MWE tagging experiments. 1
Various aspects of motherese also known as infant-directed speech (IDS) have been studied for many years. As it is a widespread phenomenon, it is suspected to play some important roles in infant development. Therefore, our purpose was to provide an update of the evidence accumulated by reviewing all of the empirical or experimental studies that have been published since 1966 on IDS driving factors and impacts. Two databases were screened and 144 relevant studies were retained. General linguistic and prosodic characteristics of IDS were found in a variety of languages, and IDS was not restricted to mothers. IDS varied with factors associated with the caregiver (e.g., cultural, psychological and physiological) and the infant (e.g., reactivity and interactive feedback). IDS promoted infants’ affect, attention and language learning. Cognitive aspects of IDS have been widely studied whereas affective ones still need to be developed. However, during interactions, the following two observati)
The meaning of person names is determined by their associated information. This study used event related potentials to investigate the time course of integrating the newly constructed meaning of person names into discourse context. The meaning of person names was built by two-sentence descriptions of the names. Then we manipulated the congruence of person names relative to discourse context in a way that the meaning of person names either matched or did not match the previous context. ERPs elicited by the names were compared between the congruent and the incongruent conditions. We found that the incongruent names elicited a larger N400 as well as a larger P600 compared to the congruent names. The results suggest that the meaning of unknown names can be effectively constructed from short linguistic descriptions and that the established meaning can be rapidly retrieved and integrated into contexts. [ABSTRACT FROM AUTHOR], Copyright of PLoS ONE is the property of Public Library of Scienc)
Speech processing inherently relies on the perception of specific, rapidly changing spectral and temporal acoustic features. Advanced acoustic perception is also integral to musical expertise, and accordingly several studies have demonstrated a significant relationship between musical training and superior processing of various aspects of speech. Speech and music appear to overlap in spectral and temporal features; however, it remains unclear which of these acoustic features, crucial for speech processing, are most closely associated with musical training. The present study examined the perceptual acuity of musicians to the acoustic components of speech necessary for intra-phonemic discrimination of synthetic syllables. We compared musicians and non-musicians on discrimination thresholds of three synthetic speech syllable continua that varied in their spectral and temporal discrimination demands, specifically voice onset time (VOT) and amplitude envelope cues in the temporal domain. M)
Background:The composite abuse scale (CAS) is a comprehensive tool used to measure intimate partner violence (IPV). The aim of the present study is to translate the CAS from English to Arabic. Methods:The translation of the CAS was conducted in four stages using a multi-method approach: 1) preliminary forward translation, 2) discussion with a panel of bilingual experts, 3) focus groups discussion, and 4) back-translation of the CAS. The discussion included a linguistic validation by a comparison of the Arabic translation with the original English by assessing conceptual and content equivalence. Findings:In all the stages of translation, there was an agreement to remove the question from the CAS that asked women about the use of objects in the vagina. Wording, format and order of the items were refined according to comments and suggestions made by the experts’ panel and focus groups’ members. The back-translated CAS showed similar wording and language of the original English version. C)
Spoken words carry linguistic and indexical information to listeners. Abstractionist models of spoken word recognition suggest that indexical information is stripped away in a process called normalization to allow processing of the linguistic message to proceed. In contrast, exemplar models of the lexicon suggest that indexical information is retained in memory, and influences the process of spoken word recognition. In the present study native Spanish listeners heard Spanish words that varied in grammatical gender (masculine, ending in -o, or feminine, ending in -a) produced by either a male or a female speaker. When asked to indicate the grammatical gender of the words, listeners were faster and more accurate when the sex of the speaker “matched” the grammatical gender than when the sex of the speaker and the grammatical gender “mismatched.” No such interference was observed when listeners heard the same stimuli, but identified whether the speaker was male or female. This finding sug)
This paper presents our preliminary conclusions as part of an ongoing effort to construct a new dependency representation framework for Turkish.We aim for this new framework to accommodate the highly agglutinative morphology of Turkish as well as to allow the annotation of unedited web data, and shape our decisions around these considerations.In this paper, we firstly describe a novel syntactic representation for morphosyntactic sub-word units (namely inflectional groups (IGs) in Turkish) which allows inter-IG relations to be discerned with perfect accuracy without having to hide lexical information.Secondly, we investigate alternative annotation schemes for coordination structures and present a better scheme (nearly 11% increase in recall scores) than the one in Turkish Treebank (Oflazer et al., 2003) for both parsing accuracies and compatibility for colloquial language.
Human communication relies on words—spoken or written labels for the concepts we intend to convey. These linguistic units map meanings onto forms that can be recognized and produced by others within a shared communication system. To make this possible, words are stored in long-term memory within what is often called the mental lexicon. This repository includes orthographic, phonological, morphological, and semantic information, and enables retrieval whenever comprehension or production demands it. The act of retrieving such information is what researchers describe as lexical access. In reading, the orthographic stimulus must be matched with its stored representation, just as the phonological form of the acoustic signal must be matched during speech comprehension. In production, by contrast, the intended meaning serves as the entry point, giving access to the phonological or orthographic form required for speech or writing. The term lexical access was first popularized in studies of visual word recognition, but its use has since expanded. In current literature, especially on word recognition, alternative terms such as lexical retrieval or lexical processing are often preferred, since access implies a discrete lexical entry that can be “looked up.” This assumption is at odds with many contemporary models, which favor distributed, parallel-activation accounts where sublexical units such as letters, phonemes, or morphemes contribute dynamically to recognition. In contrast, in word production the notion of access is less contentious, because selecting the correct lexical item from meaning necessarily requires pinpointing a specific representation. Research into lexical processing has focused on two main questions: the nature of the stored representations in the mental lexicon and the cognitive procedures through which they are retrieved. Much of this work has been conducted by cognitive psychologists, leading to an emphasis on mechanisms of retrieval rather than linguistic content. Moreover, explanations have focused on the cognitive rather than the neural level, though psycholinguistic theories are increasingly informed by neuroscience. Indeed, although the present article emphasizes cognitive perspectives—as suggested by its title—key findings on neural and electrophysiological correlates of lexical processing are also acknowledged, since they have substantially contributed to refining and constraining cognitive theories. The bulk of empirical research has concentrated on visual word recognition, not only because reading experiments offer precise control and measurement, but also because of their pedagogical and societal importance. Nevertheless, the same theoretical questions extend to spoken word recognition, speech production, and writing, each of which poses its own challenges for models of lexical access. In recent years, important advances have reshaped the field: large-scale megastudies and open-access lexical databases now allow researchers to examine the joint influence of multiple lexical and semantic variables, moving beyond traditional factorial designs. Computational modeling has also become more diverse, integrating Bayesian frameworks, hybrid connectionist approaches, and deep learning architectures, while empirical work has expanded to a wider range of languages and writing systems. The references selected throughout this article represent either foundational studies that shaped the field or recent contributions that capture the current state of debate, providing a framework for understanding how humans connect word forms to meanings in real time. Updated in September 2025 by Maria Fernández-López.
Theoretical linguists claim that the notorious reflexive ziji ‘self’ in Mandarin Chinese, if occurring more than once in a single sentence, can take distinct antecedents. This study tackles possibly the most interesting puzzle in the linguistic literature, investigating how two occurrences of ziji in a single sentence are interpreted and whether or not there are mixed readings, i.e., these zijis are interpretively bound by distinct antecedents. Using 15 Chinese sentences each having two zijis, we conducted two sentence reading experiments based on a modified self-paced reading paradigm. The general interpretation patterns observed showed that the majority of participants associated both zijis with the same local antecedent, which was consistent with Principle A of the Standard Binding Theory and previous experimental findings involving a single ziji. In addition, mixed readings also occurred, but did not pattern as claimed in the theoretical linguistic literature (i.e., one ziji is bo)
Although several syntactically annotated corpora (or treebanks) exist for Dutch, they are seldomly used for descriptive linguistic research because there are no easy-to-use exploitation tools available.This demonstration paper describes GrETEL, a linguistic search engine (http:// nederbooms.ccl.kuleuven.be/eng/gretel)that enables non-technical users to consult treebanks in a user-friendly way.Instead of a formal search expression, a natural language example is used as input to the system, allowing users to search for similar constructions as the example they provide.In the first version of GrETEL, only written Dutch (LASSY) was included.Based on user requests we have now included the Spoken Dutch Corpus (CGN) as well.
The article appraises gender representation in the 1999 Nigerian Constitution using insights from critical discourse analysis, feminism and systemic functional linguistics, with particular emphasis on grammatical cohesion. Specifically, it examines lexical and grammatical expressions that encode gender in the Constitution, the ideological positions evident in these expressions, and their impact on gender parity and socio-political equity. The focus is on the reference-antecedent cohesion of gender-marked pronouns and nouns used to refer to individuals and social/political positions. Our findings show a preponderance of generic masculine noun and pronoun references, tracking antecedents that refer to social and political positions open to eligible individuals in Nigeria, while the single feminine referent was a marked case. These findings buttress the ‘male-as-norm’ ideology and the relegation to anonymity of the female gender in this important national document. For equity and fairness, the article recommends revising the Constitution with epicene expressions to expunge gender biases.
Factors Related to Undergraduate Psychology Majors Learning Statistics Tamarah Faye Smith Doctor of Philosophy: Educational Psychology Major Advisor: Dr. Frank Farley The American Psychological Association (APA) has outlined goals for psychology undergraduates. These goals are aimed at several objectives including the need to build skills for interpreting and conducting psychological research (APA, 2007). These skills allow psychologists to conduct research that is covered in the media (Farley et al. 2009) and influences policy and law (Fischer, Stein & Heikkinen, 2009; Steinberg, Cauffman, Woolard, Graham & Banich, 2009a; Steinberg, Cauffman, Woolard, Graham & Banich, 2009b). One of the fundamental courses required for building these skills is statistics, a course that begins at the undergraduate level. Research has suggested that performance after completing statistics courses is weak for many students (Garfield, 2003; Hirsch & O'Donnell, 2001; Konold et al. 1993; Mulhern & Wylie, 2005; Schau & Mattern, 1997). The current study examined factors that may be related to performance on a statistical test. A sample of 231 students enrolled in or having already completed a statistics course for psychology majors completed a statistical skill questionnaire, built by the author, to measure performance with four APA outlined goals. To measure student attitudes the Survey of Attitudes Toward Statistics (SATS-36; Schau, 2003) was completed with adapted questions to measure perceived attitudes of peers and faculty toward statistics. Finally, questions pertaining to classroom techniques and content areas covered were assessed. Building off of social cognitive theory (SCT; Bandura, 1986) and expectancy-value theory (Eccles & Wigfield, 2002), it was expected that lower attitudes, such as low value and low interest, among the students and those perceived to be held by faculty and peers would be related to lower performance on the statistical test. A series of linear regressions were conducted and revealed no significant relationship between perceived faculty attitudes and performance. Students' own liking and positive affect ratings were positive predictors of performance indicating a gain of 3-4% on the statistical test. However, an interesting negative relationship emerged with respect to students' value of statistics and peer interest scores where performance on the statistical test decreased as value and peer interest increased. This may be demonstrating issues pertaining to the SATS-36 validity when measuring students' value as well as issues with the items created to measure perceived peer interest. The results of a factor analysis on perceived attitude measures for peers and faculty suggest that the need for more items is necessary, particularly for faculty attitudes. Finally, this study provides a first look at the performance of a sample of psychology students with APA goals for quantitative reasoning. Results showed that students performed best at reading basic descriptive statistics (M=74.5%), and worst when choosing statistical tests for a given research hypothesis (M=30%). Performance on questions pertaining to confidence intervals (M=38%) and discriminating between statistical and practical significance (M=39%) was also low. Future research can address limitations of this study by expanding the sample to include a broader range of psychology undergraduates and including additional items for measuring perceived attitudes. Other methodological approaches, such as experimental design and directly measuring faculty attitudes, should also be considered. Finally, further research and replication are necessary to determine if scores on the statistical test will continue to be low with other samples and varying question formats. These results can then be used to generate conversation about why and how students are, or are not, learning the appropriate quantitative skills.
We present, here, our analysis of systematic divergences in parallel English-Hindi dependency treebanks based on the Computational Paninian Grammar (CPG) framework. Study of structural divergences in parallel treebanks not only helps in developing larger treebanks automatically, but can also be useful for many NLP applications such as data-driven machine translation (MT) systems. Given that the two treebanks are based on the same grammatical model, a study of divergences in them could be of advantage to such tasks, along with making it more interesting to study how and where they diverge. We consider two parallel trees divergent based on differences in constructions, relations marked, frequency of annotation labels and tree depth. Some interesting instances of structural divergences in the treebanks have been discussed in the course of this paper. We also present our task of alignment of the two treebanks, wherein we talk about our extraction of divergent structures in the trees, and discuss the results of this exercise. 1
La medecine actuelle se base sur les principes de l'EBM dans l'optique destandardiser les pratiques. Les soins palliatifs sont eux aussi soumis a unereglementation et il apparait alors une norme dans les soins palliatifs. Or les soinspalliatifs sont une discipline a part qui a pour objet principal la prise en chargeindividualisee. A partir d'une revue de la litterature, les normativites theorique etpratique ont ete definies. L'objectif de ce travail est de voir quelle est la conceptiondes soignants du travail en soins palliatifs et de voir comment normativite et ideauxs'articulent au quotidien dans le travail en USP. Pour repondre a cette problematique, une etude qualitative par entretiens semi-diriges de soignants exercant en USP a ete realisee. Ils ont ete soumis a une double analyse: une analyse par le logiciel d'analyse lexical Alceste© et une analyse manuelle. Le travail en equipe, le dialogue, l'accompagnement ressortent comme des points importants du travail en USP. Il existe egalement une grande adaptabilite au patient meme s'il existe un certain rythme dans l'organisation de la journee. La norme est donc necessaire en soins palliatifs pour creer un cadre et assurer un acces egal a tous et garantir l'individualisation des prises en charge. Cette norme est en constante evolution pour s'adapter au patient
This paper introduces an advanced, efficient approach for rule based English to Bengali (E2B) machine translation (MT), where Penn-Treebank parts of speech (PoS) tags, HMM (Hidden Markov Model) Tagger is used.Fuzzy-If-Then-Rule approach is used to select the lemma from rule-based-knowledge. The proposed E2B-MT has been tested through F-Score measurement, and the accuracy is more than eighty percent.
This paper investigates the appropriateness of using lexical cohesion analysis to assess Chinese readability. In addition to term frequency features, we derive features from the result of lexical chaining to capture the lexical cohesive information, where E-HowNet lexical database is used to compute semantic similarity between nouns with high word frequency. Classification models for assessing readability of Chinese text are learned from the features using support vector machines. We select articles from textbooks of elementary schools to train and test the classification models. The experiments compare the prediction results of different sets of features.
The problem of Vietnamese syntactic parsing, especially constituency parsing, has recently been tackled by several research groups. A common effort of the Vietnamese language processing community has allowed the creation of VietTreebank, a reference parsed corpus containing about 10,000 sentences for the constituency parsing task. In this paper, we present our work to build a reference treebank, based on VietTreebank, for the dependency parsing task, which has not yet been very well studied for Vietnamese. First we define a dependency label set by adapting the dependency schema developed by the NLP group at Stanford university and taking into account the particularities of Vietnamese grammar. Then we propose an algorithm to convert a constituency treebank to a dependency one. The algorithm is tested on a set of 100 sentences of VietTreebank corpus and gives very good results. Finally, we carry out an experiment on Vietnamese dependency parsing using MaltParser tool and the dependency treebank converted from VietTreebank.
Psychophysiological evidence suggests that music and language are intimately coupled such that experience/training in one domain can influence processing required in the other domain. While the influence of music on language processing is now well-documented, evidence of language-to-music effects have yet to be firmly established. Here, using a cross-sectional design, we compared the performance of musicians to that of tone-language (Cantonese) speakers on tasks of auditory pitch acuity, music perception, and general cognitive ability (e.g., fluid intelligence, working memory). While musicians demonstrated superior performance on all auditory measures, comparable perceptual enhancements were observed for Cantonese participants, relative to English-speaking nonmusicians. These results provide evidence that tone-language background is associated with higher auditory perceptual performance for music listening. Musicians and Cantonese speakers also showed superior working memory capacity )
Question Time is a distinctive daily parliamentary routine. Its aim is to hold Ministers of the State accountable for the actions and decisions of the Government. However, in many Parliaments, including the New Zealand and Australian Federal Houses of Representatives, it is more of a theatrical performance where parties try their best to score political points. As any performance, Question Time is governed by certain rules and regulations outlined in an official document Standing Orders. As there is not much action, Standing Orders mainly describe language norms and specify „unparliamentary language‟. This research looks at and analyses the use of formulaic vocabulary used by MPs in the year preceding general elections in New Zealand and Australia. The formulaic language includes phrasal lexical items and formulae for asking / answering questions, for raising points of order and the Speakers‟ idiolectal phrasal vocabulary for quelling disorder in the Chambers and regulating the work of the House. The framework developed for this research consisted of the following steps: an ethnographic study of Question Time as a communicative performance which included the development of a database containing all the empirical material; a xii linguistic study of Question Time including genrelect study, parliamentary formulae study and disorder analysis before the elections. As a result this research has shown that Question Time is a communicative performance event in New Zealand and Australia with significant cultural, historic and linguistic differences in spite of the common origins of the two Parliaments. It has identified 60 Question Time genre-specific phrasal lexical items that MPs use in the two Parliaments, studied their structure and meaning (where necessary). It has also looked at the strategies the MPs employ for creating disorder in the House, and the ways of quelling disorder by the Speakers of the two Parliaments.
Size is an important visuo-spatial characteristic of the physical world. In language processing, previous research has demonstrated a processing advantage for words denoting semantically “big” (e.g., jungle) versus “small” (e.g., needle) concrete objects. We investigated whether semantic size plays a role in the recognition of words expressing abstract concepts (e.g., truth). Semantically “big” and “small” concrete and abstract words were presented in a lexical decision task. Responses to “big” words, regardless of their concreteness, were faster than those to “small” words. Critically, we explored the relationship between semantic size and affective characteristics of words as well as their influence on lexical access. Although a word’s semantic size was correlated with its emotional arousal, the temporal locus of arousal effects may depend on the level of concreteness. That is, arousal seemed to have an earlier (lexical) effect on abstract words, but a later (post-lexical) effect on)
В статье рассматривается проблема нормы и нормативного подхода к языку в диахроническом плане.Определяется специфика нормативного похода к языковым средствам в различных лингвистических традициях и выявляются основные характеристики лингвистической нормы.В статье указывается, что на каждом этапе развития языка складываются свои нормы как резуль
The paper presents the process of constructing a publicly available treebank of public messages written in Croatian. The messages were collected from various electronic sources – e-mail, blog, Facebook and SMS – and published on the Zagreb Museum of Contemporary Art LED facade within the Babel art project. The project aimed to use the facade as an open-space blog or social interface for enabling citizens to publicly express their views. Construction and current state of the treebank is presented along with future work plans. A comparison of Babel Treebank with Croatian Dependency Treebank and SETimes.HR treebank regarding differing domains and annotation schemes is briefly sketched. The treebank is used as a test platform for introducing a new standard for syntactic annotation of Croatian texts. An experiment with morphosyntactic tagging and dependency parsing of the treebank is conducted, providing first insight to computational processing of non-standard text in Croatian.
Background: Determining the semantic relatedness of two biomedical terms is an important task for many text-mining applications in the biomedical field. Previous studies, such as those using ontology-based and corpus-based approaches, measured semantic relatedness by using information from the structure of biomedical literature, but these methods are limited by the small size of training resources. To increase the size of training datasets, the outputs of search engines have been used extensively to analyze the lexical patterns of biomedical terms. Methodology/Principal Findings: In this work, we propose the Mutually Reinforcing Lexical Pattern Ranking (ReLPR) algorithm for learning and exploring the lexical patterns of synonym pairs in biomedical text. ReLPR employs lexical patterns and their pattern containers to assess the semantic relatedness of biomedical terms. By combining sentence structures and the linking activities between containers and lexical patterns, our algorithm can )
Code-mixing involves the deliberate mixing of two languages without an associated topic change. It is primarily used as a solidarity marker. It is not something brought about by laziness or ignorance as such, rather it requires the conversant to have a good knowledge of the grammar of the two languages and to be well aware of societal norms. It is a source of pride to bilinguals (Wardhaugh 1986). In this paper we examine code mixing from the interference dimension, looking at it from the phonological and inter-lingual angles. We discuss code-mixing because Efik people are not monolingual. To them, substantial command of English is a passport to the arena of globalization and competitive white-collar job market. Therefore mixing Efik and English is inevitable. Urbanization, education, government business and multilingualism have triggered the Efik people to learn English. A combination of research principles using unstructured forms of data collection research methods is used for this study which are (i) Participant Observation and (ii) In-depth Interview. This paper is rooted in the phonemic theory which models what happens to the languages when there is a mixing and interference. We used aspects of morphological and sociolinguistic models in the analysis. We have come up with the key findings which state that the grammatical items, rather than the lexical ones, are crucial to the identity of a language. Also a language may borrow lexical items freely assimilating or not assimilating them. We can again add that a language is on its way to losing its identity once it starts borrowing grammatical items from another language.
The ways in which literacy in English is taught in school generally subscribe to and perpetuate the notion of a homogenous, unvaried set of writing conventions associated with the language they represent, especially in relation to spelling and punctuation as well as grammar. Such teaching also perpetuates the myth that there is one correct way of language use which is fixed and invariant, and that any deviation is at best incorrect or illiterate and at worst, a threat to social stability. It is also very clear that the linguistic norms associated with standard English are predicated upon and replicate white, cultural hegemony. Yet, at the same time, there are plenty of literary and creative works written by authors from all kinds of different cultural, ethnic and linguistic backgrounds, including canonical ones, where spelling and punctuation are varied and championed as a sign of creativity. In the world beyond school, pupils are also surrounded by variational use of written language, especially in public displays such as shop signs, writing on mugs and t-shirts, posters, graffiti and so on, which link language to place. Equally, the voices we hear in entertainment and public broadcasting, far from being homogenous, celebrate diversity in Englishes. The homes and backgrounds of pupils in our schools, including their linguistic backgrounds, may also be very different either in terms of a different variation of English or languages spoken other than English. Since the emphasis is usually upon correct and fixed ways of teaching writing in English, it has often been difficult for teachers and pupils to reconcile the kind of English taught in school as the correct way and thus, by definition, all others as incorrect. However, narrow definitions of linguistic correctness are becoming increasingly difficult to uphold given that the public spaces with which we are surrounded are peppered by examples of variational use in writing. Recent sociolinguistic research into variation points to an increasing fluidity of linguistic use, especially when it comes to public displays of writing, particularly in media such as newspapers, websites, shop signs, TV channel logos and so on. Linguistic variability can thus be seen as a resource in creating unique voices and marking allegiance to, for example, a particular place and culture. Such research is indicative of the fact that variational use of English, far from being incorrect or illiterate, is increasingly being drawn upon creatively to mark a place identity. It also points to a shift in our conceptual thinking about language(s) and varieties from being perceived as static, fixed, totalised and immobile to being thought of as dynamic, fragmented and mobile, with the focus upon mobile resources rather than immobile languages. At the same time, the teaching of literacy centres upon the teaching of linguistic norms of spelling and grammar as fixed. There is a tension then, between creative expression of linguistic use often linked to place and those linked to standard English. This article explores those tensions and discusses the implications and possibilities for the teaching of English and literacy. © 2013.