Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
This paper deals with Spain’s recent linguistic policies devoted to the promotion of the Spanish language, particularly its spread in Brazil. A renewed interest in Spanish has been taking place in Brazil since the 1990’s. Such an interest is due to the creation of Mercosur and to the fact that many Spanish businesses have set up in the country. Based on a corpus of interviews of teachers at the Cervantes Institutes in Brazil, this paper focuses on the issue of the linguistic norm to be taught and to be promoted in this particular context.
In the context of processing Bengali words through a computer, there may arise several issues that are directly linked with surface structure of words. These issues may create problems in manual and computer-based counting of number of words in a corpus. They can also create problems in morphological processing of words. These issues come up because there is hardly any consistency in orthographic representation of words in written Bengali texts. The high irregularities in writing of inflected words, proper names, adjectival forms, adverbial forms, compound words, reduplicated words, onomatopoeic words, hyphenated words, etc. present a daunting task before an investigator in normalizing the surface forms of words for generating a lexical database as well as developing a word processing system for the works of language technology.
The French Lexical Network (fr-LN) is a global model of the French lexicon presently under construction. The fr-LN accounts for lexical knowledge as a lexical network structured by paradigmatic and syntagmatic relations holding between lexical units. This paper describes how morphological knowledge is presently being introduced into the fr-LN through the implementation and lexicographic exploitation of a dynamic morphological model. Section 1 presents theoretical and practical justifications for the approach which we believe allows for a cognitively sound description of morphological data within semantically-oriented lexical databases. Section 2 gives an overview of the structure of the dynamic morphological model, which is constructed through two complementary processes: a Morphological Process--section 3--and a Lexicographic Process--section 4.
Over the past 15 years, one of the basic functions of the Linguistic Landscape (LL) has been to measure and assess the visibility of multiple languages in public spaces. In research carried out around the world, English has recurrently appeared as one of the most important features of the LL, though its classification, and the extent to which its presence renders a space multilingual, remains controversial. In the light of Seargeant’s (2011) and Yano’s (2011) recent claims that English may be considered a feature of the local landscape rather than a language in its own right, this paper explores code classification in the LL of Toulouse. It challenges the concept of ‘global’ languages (Crystal, 1997), and illustrates the ways in which various codes become affected, consumed, and enveloped by the dominant language, French. Not only is this the case for English, but also the regional tongue, Occitan, whose visibility, despite appearing autonomous, may be more accurately considered a part of the French landscape. This illustrates the shortcomings in classifying languages according to fixed, subjective (and often ad hoc) criteria, because even the presentation of so-called ‘foreign’ languages in Toulouse is influenced by, and subject to, local linguistic norms. The implications of this are profound, as we see that French is the driving force behind every language in the city. Moreover, this calls into question the very definition of multilingualism, as LL authors exhibit a preference for the national code even when writing in others. As such, this research will have a meaningful impact on language coding in the LL, and contribute to the advancement of the methodologies by which we measure and analyse multilingualism, language beliefs, and language practices in public spaces.
Prior research suggested the possibility of establishing systematic linkages between some intrinsic features of a presupposition and textual and pragmatic functions that it can carry out with greater probability. This study aims, firstly, to provide an organic view of semantics of presupposition triggers, thanks to a lexical database comprising 19,500 entries. Secondly, the database was used to investigate a corpus of chat conversations including about 200,000 tokens. The results show that triggers occur mainly as non-informative, maintaining an information already known by all participants of the communication; but, depending on their different features, some of them are systematically associated to a function of anaphora and textual cohesion; others to strengthen social conventions and stereotypes. The informative function, although in a minority proportion, is quantitatively significant only in correspondence to a single class of presupposition triggers.
Incorporating knowledge for training a parser has been shown to remedy the weaknesses of probabilistic context-free grammar. Previous parsing systems have exploited content words semantic resource and word-formation knowledge. However, they are limited in that they do not take into account conjunction category refinement, which stands out to be helpful in predicting the syntactic structure and syntactic label in Chinese. We define a conjunction taxonomy representing intrinsic syntactic constraints, and show that refined categories in the taxonomy for conjunctions contribute to improved parsing performance. The taxonomy is used to supervise the splitting of these refined tags, and the automatic hierarchical state-split approach is employ to compensate the limitation in the scope and refinement degree of the taxonomy. The experiments are carried out on Penn Chinese Treebank, which show that our method can improve parsing performance significantly.
In this paper we present a system for experimenting with combinations of dependency parsers. The system supports initial training of different parsing models, creation of parsebank(s) with these models, and different strategies for the construction of ensemble models aimed at improving the output of the individual models by voting. The system employs two algorithms for construction of dependency trees from several parses of the same sentence and several ways for ranking of the arcs in the resulting trees. We have performed experiments with state-of-the-art dependency parsers including MaltParser (Nivre et al., 2006), MSTParser (McDonald, 2006), TurboParser (Martins et al., 2010), and MATEParser (Bohnet, 2010), on the data from the Bulgarian treebank – BulTreeBank. Our best result from these experiments is slightly better then the best result reported in the literature for this language (Martins et al.,
Discourse relations bind smaller linguistic elements into coherent texts. However, automatically identifying discourse relations is difficult, because it requires understanding the semantics of the linked sentences. A more subtle challenge is that it is not enough to represent the meaning of each sentence of a discourse relation, because the relation may depend on links between lower-level elements, such as entity mentions. Our solution computes distributional meaning representations by composition up the syntactic parse tree. A key difference from previous work on compositional distributional semantics is that we also compute representations for entity mentions, using a novel downward compositional pass. Discourse relations are predicted not only from the distributional representations of the sentences, but also of their coreferent entity mentions. The resulting system obtains substantial improvements over the previous state-of-the-art in predicting implicit discourse relations in the Penn Discourse Treebank.
Many languages, including Modern Stan-dard Arabic (MSA), insert resumptive pro-nouns in relative clauses, whereas many others, such as English, do not, using empty categories instead. This discrep-ancy is a source of difficulty when trans-lating between these languages because there are words in one language that cor-respond to empty categories in the other, and these words must either be inserted or deleted—depending on translation di-rection. In this paper, we first examine challenges presented by resumptive pro-nouns in MSA-English translations and re-view resumptive pronoun translations gen-erated by a popular online MSA-English MT engine. We then present what is, to the best of our knowledge, the first system for automatic identification of resumptive pronouns. The system achieves 91.9 F1 and 77.8 F1 on Arabic Treebank data when using gold standard parses and automatic parses, respectively. 1
In this study, a dictionary-based method is used to extract expressive concepts from documents. So far, there have been many studies concerning concept mining in English, but this area of study for Turkish, an agglutinative language, is still immature. We used dictionary instead of WordNet, a lexical database grouping words into synsets that is widely used for concept extraction. The dictionaries are rarely used in the domain of concept mining, but taking into account that dictionary entries have synonyms, hypernyms, hyponyms and other relationships in their meaning texts, the success rate has been high for determining concepts. This concept extraction method is implemented on documents, that are collected from different corpora.
Ratings of previously ignored visual stimuli reveal affective devaluation of such items when compared to ratings of novel items or the targets of attention. Growing evidence suggests this effect may reflect negative affective associations elicited by attentional inhibition of visual distractors. Here we investigate whether such 'inhibitory devaluation' is limited to situations involving visual-spatial selection of environmental stimuli (i.e., external attention) or extends to the selection of competing visual representations held solely in memory (i.e., internal attention). A two-item target-localization task in Experiment 1 utilized a delayed target-category cue ('circles' or 'squares') to ensure attentional selection occurred from the contents of working memory. An n-back task in Experiment 2 was used to examine the affective consequences of rejecting continually-updated visual representations when items held in memory did not match the corresponding visual display. And a Think/No-think paradigm employed in Experiment 3 was designed to explore the affective consequences of actively suppressing longer-term visual object memories. Across this relatively-wide range of memory-based selection tasks, the ignored/rejected/suppressed visual patterns consistently received more negative affective ratings than target items. Our results are consistent with prior suggestions that similar mechanisms are involved in the attentional selection of environmental stimuli and the selection of internally-maintained information that occurs even in the absence of external sensory stimulation. The similarity in these mechanisms appears to extend not only to processes of attentional selection, per se, but also to their affective consequences. Meeting abstract presented at VSS 2014
Opinion mining is becoming of high importance with the availability of opinionated data on the Internet and the different applications it can be used for. Intensive efforts have been made to develop opinion mining systems, and in particular for the English language. However, models for opinion mining in Arabic remain challenging due to the complexity and rich morphology of the language. Previous approaches can be categorized into supervised approaches that use linguistic features to train machine learning classifiers, and unsupervised approaches that make use of sentiment lexicons. Different features have been exploited such as surface-based, syntactic, morphological, and semantic features. However, the semantic extraction remains shallow. In this paper, we propose to go deeper into the semantics of the text when considered for opinion mining. We propose a model that is inspired by the cognitive process that humans follow to infer sentiment, where humans rely on a database of preconceived notions developed throughout their life experiences. A key aspect for the proposed approach is to develop a semantic representation of the notions. This model consists of a combination of a set of textual representations for the notion (Ti), and a corresponding sentiment indicator (Si). Thus <Ti, Si> denotes the representation of a notion. However, notions can be constructed at different levels of text granularity ranging from ideas covered by words to ideas covered in full documents. The range also includes clauses, phrases, sentences, and paragraphs. To demonstrate the use of this new semantic model of preconceived notions, we develop the full representation of one-word notions by including the following set of syntactic features for Ti: word surfaces, stems, and lemmas represented by binary presence and TFIDF. We also include morphological features such as part of speech tags, aspect, person, gender, mood, and number. As for the notion sentiment indicator Si, we create a new set of features that indicate the words' sentiment scores based on an internally-developed Arabic sentiment lexicon called ArSenL, and using a third-party lexicon called Sifaat. The aforementioned features are extracted at the word-level, and are considered as raw features. We also investigate the use of additional "engineered" features that reflect the aggregated semantics of a sentence. Such features are derived from word-level information, and include count of subjective words, average of sentiment scores per sentence. Experiments are conducted on a benchmark dataset collected from the Penn Arabic TreeBank (PATB) already annotated with sentiment labels. Results reveal that raw word-level features do not achieve satisfactory performance in sentiment classification. Feature reduction was also explored to evaluate the relative importance of the raw features, where the results showed low correlations between individual raw features and sentiment labels. On the other hand, the inclusion of engineered features had a significant impact on classification accuracy. The outcome of these experiments is a comprehensive set of features that reflect the one-word notion or idea representation in a human mind. The results from one-word also show promises towards higher level context with multi-word notions.
The intent of this study is to determine what sorts of images are considered more interesting by which demographic groups. Specifically, we attempt to identify images whose interestingness ratings are influenced by the demographic attribute of the viewer’s gender. To that end, we use the data from an experiment where 18 participants (9 women and 9 men) rated several hundred images based on “visual interest” or preferences in viewing images. The images were selected to represent the consumer “photo-space” - typical categories of subject matter found in consumer photo collections. They were annotated using perceptual and semantic descriptors. In analyzing the image interestingness ratings, we apply a multivariate procedure known as forced classification, a feature of dual scaling, a discrete analogue of principal components analysis (similar to correspondence analysis). This particular analysis of ratings (i.e., ordered-choice or Likert) data enables the investigator to emphasize the effect of a specific item or collection of items. We focus on the influence of the demographic item of gender on the analysis, so that the solutions are essentially confined to subspaces spanned by the emphasized item. Using this technique, we can know definitively which images’ ratings have been influenced by the demographic item of choice. Subsequently, images can be evaluated and linked, on one hand, to their perceptual and semantic descriptors, and, on the other hand, to the preferences associated with viewers’ demographic attributes.
Imagine saying to Hegel that his dialectic patterns occur in the behavior of magnets and rivers and trees, and that the self-understanding of spirit is an operation going on in brain tissue. That the descriptions in his Phenomenology of Spirit picture the behavior of social groups and changes in cultural norms and memes. Hegel's reaction would be complex but not hostile. You would find yourself in a discussion with him about different kinds and levels of categories, and the relation of physical science to his overall logic.Now imagine saying to Heidegger that his description of the care structure of Dasein is a sophisticated reworking of folk psychology, and that it is the result of what amounts to a software program. That his Fourfold is a description of the appearing of a world, based upon neurological and social processes. That his history of being is open to sociological analysis and historical and economic explanation. Heidegger would react to such claims more sharply. You would find yourself in a discussion with him about being caught in das Gestell, and the need to step back from attempts to absolutize one language and one revelation of beings.How would these different reactions play out in a discussion of the ontology of the self? This present essay approaches these major continental thinkers with a question from analytic philosophy, to see how they might respond.In standard mind/body discussions we find two rival descriptive languages with different sets of entities. One talks about ideas, purposes, awareness, meanings, concepts, norms, intentions, and so on. The other talks about the behavior of cells and electrical currents, brain activities, and entities described by physics.There are thinkers who argue that folk talk about thoughts and meanings could be in principle abandoned even in everyday life, replaced by descriptions that involved only physical scientific entities. Others argue that folk talk is not dispensable. Daniel Dennett and Wilfrid Sellars would say that we need to take up the intentional stance and avow social and linguistic norms in order to function in a meaningful social world. Kant had already argued that there is a practical necessity to view ourselves as free agents, no matter what our science may say. These considerations suggest that the tension between the modes of discussion cannot be easily wished away by eliminating one of them.DIFFERENT ONTOLOGIESOne could say that these two languages are using different ontologies. I am using ontology here in a sense derived from analytic philosophers who ask your ontology include... [relations, sets, second level properties, mereological wholes, abstract entities, Cartesian souls, etc.]? An ontology in this sense is the list of approved types of entities that are being affirmed as ultimately real. In most such discussions there is little talk about ontology in an older sense, namely, about the mode of being of those beings which analytic discussions often presuppose to be immediate factual presence. Much more active modes of being are affirmed in Whitehead or in Bergson or Deleuze, or in Aristotle's doctrine of potentiality. For none of these thinkers does being equal simple, positive presence. For them such presence is a result, and it is so independently of whatever processes bring it to presence for us in experience.But if we do ask about modes of being, beyond the factual lists, then other more traditional ontological questions arise: In Aristotle's terms, are the movements of Hegel's dialectic substantial changes or accidental, or relational, or what? Is Heidegger's Dasein or Hegel's spirit a substance? Or a set of emergent properties? How do we individuate Dasein(s)? Does Hegel's spirit have its own individuated self-consciousness? What kind of being do the components of the Fourfold have, and how do they relate to everyday objects and scientific entities?Ontology matters. When David Hume looks for his self, he can't find it. …
ABSTRACTIndividual differences in estimates of respiratory sinus arrhythmia (RSA) have been implicated in the ability to respond efficiently to challenges encountered in different environmental contexts. Studies have used resting RSA and the suppression of RSA to emotional or challenging cognitive tasks to assess parasympathetic contributions to emotional and attentional regulation. In the current study, estimates of resting RSA, and RSA suppression were used in conjunction with measures of attentional control and anxiety to predict self -rated measures of executive functions and task performance on a mental math exercise. The data suggest that resting RSA, in conjunction with good attentional control and low anxiety are better predictors of global measure of a participant's ability to employ executive processes compared to RSA suppression. Both resting RSA and RSA suppression predicted more efficient task performance on the mental math exercise. The suppression of RSA, however, was more predictive of better performance in participants with poor attentional control and high anxiety. Resting RSA may be a better predictor of the ability of an individual to respond to challenges encountered in the environment. The degree of RSA suppression may be task specific and related to the interpretation of the degree of challenge associated with the task requirements.KEYWORDS: executive function, attention, emotion, anxiety, heart rate variabilityThe ability to engage various executive processes has been shown to be directly related to the regulation of attention and negative affect (Healy, Treadwell, & Reagan, 2011). Initial work by Derryberry and Rothbart (1988) defined attentional control as the ability to sustain attention to aspects of the environment in addition to an ability to shift one's attentional focus as dictated by the demands of differing environmental events. This flexibility in attentional processing has been seen as representing the functioning of an anterior attentional system linked to frontal lobe function which would allow an individual to engage and disengage attention to aspects of the environment. At the other extreme, attentional rigidity has been linked to the prevalence of anxiety disorders whereby individuals have difficulty disengaging from aspects of the environment associated with threat (Derakshan & Eysenck, 2009; Wells & Matthews, 1994). Research suggests that the ability to allocate attention to different aspects of the environment is useful in situations that require a reduction in negative affect (Lonigan, Vasey, Phillips, & Hazen, 2004).From a physiological perspective, estimates of the parasympathetic contribution to the beat-to-beat variation in heart rate (i.e., respiratory sinus arrhythmia, or RSA) have been used as a marker of emotional self-regulation and attentional control (Applehans & Luekan, 2006). Initial research had placed an emphasis on resting measures of RSA as representing an individual's level of responsiveness with higher resting estimates predicting more positive outcomes and/or behaviors from infancy to adulthood. In this regard, there exits an extensive literature that points towards high resting levels of RSA as being related to approach behavior, and emotional self-regulation in infants and children (Calkins & Keene, 2004; Kagan, Resnick, & Snidman; 1987; Stifter & Fox, 1991), increased attentional control in infants and children (Richards & Casey, 1992), positive emotional expressiveness in children (Calkins, 1997), modulation of negative affect, ratings of lower anxiety in adults, and increases in positive mood (Greaves-Lord, Ferdinand, Sondeijker, Dietrich, Oldehinkel, Rosmalen, Ormel, & Verhulst, 2007; Gyurak & Ayduk, 2008; Thayer, Friedman, & Borkovec, 1996). In contrast, lower resting RSA has been linked to chronic anxiety and the prevalence of anxiety disorders such as generalized anxiety disorder and post-traumatic stress disorder in adults (Friedman, 2007). …
The author of the article suggests the taxonomy of the objectivation methods of pragmatic and linguistic norms of professional sublanguage: subjective method, semantic method (by enumerating referent's nominations), method of limiting the domain of the possible and the method of extensional definition. Professional sublanguage demonstrates the implementation of pragmatic and linguistic norm primarily as pointing out the domain of the possible (dysphemisation and the derisionof referent). The article shows that the relevance of norm is increased in the context of external contacts and the elimination of communication.
Context \nGrETEL and Treebanks \nExample Parse \nSearching in treebanks \nSearching with GrETEL \nSearching with GrETEL: Limitations \nComparison with Google \nConclusions and Invitation
HamleDT 2.0 is a collection of 30 existing treebanks harmonized into a common annotation style, the Prague Dependencies, and further transformed into Stanford Dependencies, a treebank annotation style that became popular recently. We use the newest basic Universal Stanford Dependencies, without added language-specific subtypes.
This article deals with the problem of deviation of linguistic norm in the artistic text by the example of the novel by A. Doblin «Berlin Alexanderplatz». The aims of this article were to define the means of the violation of the text structure, to consider expressive constructions used for the realization of authors certain intensions.
The purpose of our work is to explore the possibility of using sentence diagrams produced by schoolchildren as training data for automatic syntactic analysis. We have implemented a sentence diagram editor that schoolchildren can use to practice morphology and syntax. We collect their diagrams, combine them into a single diagram for each sentence and transform them into a form suitable for training a particular syntactic parser. In this study, the object language is Czech, where sentence diagrams are part of elementary school curriculum, and the target format is the annotation scheme of the Prague Dependency Treebank. We mainly focus on the evaluation of individual diagrams and on their combination into a merged better version.
This paper presents a Naxi dependency parsing method which combined rules and statistics, based on the characteristic of the Naxi language. Firstly, the annotation standard of Naxi Dependency Treebank was established and dependencies were defined which are based on the syntactic structure characteristics of Naxi language, then Naxi Dependency Treebank was constructed; Secondly, the constructed Naxi dependency Treebank was used as the basis, and the Naxi language phrases were analyzed independently by means of rules, then the boundaries and categories of phrases were determined. The features of core words postposition in the Naxi language were regarded as a condition of constraint. Then the dependencies between phrases were analyzed. Finally, the dependencies between phrases were combined, and interdependent probabilistic models were used to parse the syntax of Naxi language, and finally a completed Naxi language dependency parsing system was achieved. The results of dependency parsing comparative experiments show that the use of Naxi dependency rules improves the performance of Naxi language dependency parsing system. Moreover, the performance of the proposed method is better than the dependency parsing methods which are fully based on statistics.
Correspondence analysis (CA) was conceived by Benzécri (1977b; Benzécri et al., 1973, 1981) as an inductive method able to infer from their own texts of interest the linguistic norms that rule them. He wanted to offer a powerful tool to tackle all kinds of issues related to the form, meaning, and style of texts.
On the basis of studying the lexicographic sources a definition of cynicism the term “linguocinism” is given in the article. The meaning of this term is specified by its correlation with similar to its meaning, but not synonymous to the terms “vulgarisms” and “obscenism”. It became clear that vulgarisms and obscenisms being violation of ethic-linguistic norm are not always linguocinisms. Linguistic means of expressing linguocinism and the role of context in their realization are revealed.
Completely data-driven grammar training is prone to over-fitting. Human-defined word class knowledge is useful to address this issue. However, the manual word class taxonomy may be unreliable and irrational for statistical natural language processing, aside from its insufficient linguistic phenomena coverage and domain adaptivity. In this paper, a formalized representation of function word subcategorization is developed for parsing in an automatic manner. The function word classification representing intrinsic features of syntactic usages is used to supervise the grammar induction, and the structure of the taxonomy is learned simultaneously. The grammar learning process is no longer a unilaterally supervised training by hierarchical knowledge, but an interactive process between the knowledge structure learning and the grammar training. The established taxonomy implies the stochastic significance of the diversified syntactic features. The experiments on both Penn Chinese Treebank and Tsinghua Treebank show that the proposed method improves parsing performance by 1.6 % and 7.6 % respectively over the baseline. 1
The present contribution represents the first step in comparing the nature of syntactico-semantic relations present in the sentence structure to their equivalents in the discourse structure. The study is carried out on the basis of Czech manually annotated material collected in the Prague Dependency Treebank (PDT). According to the analysis of the underlying syntactic structure of a sentence (tectogrammatics) in the PDT, we distinguish various types of relations that can be expressed both within a single sentence (i.e. in a tree) and in a larger text, beyond the sentence boundary (between trees). We suggest that, on the one hand, semantic nature of each type of these relations corresponds both within a sentence and in a larger text (i.e. a causal relation remains a causal relation) but, on the other hand, according to the semantic properties of the relations, their distribution in a sentence or between sentences is very diverse. In this study, this observation is analyzed in detail for three cases (relations of condition, specification and opposition ) and further supported by similar behaviour of the English data from the Penn Discourse Treebank.
The domain of air traffic control is the perfect example for the analysis of a linguistic norm. In this field, phraseology is a specialised language created to cover the most common situations encountered in air navigation in order to secure and optimise radiotelephony communications. When phraseology proves inadequate, a more natural form of language, called plain language, is required: it was recently introduced in the field and is a difficult notion to implement. A comparative analysis between a reference corpus and a real-communication corpus allows the description and categorisation of different types of variations used on the radio frequency as well as a discussion on the notions of norm and usage in the field of air traffic control.
The article discusses lexical borrowings from German to Czech, especially to Middle Czech. It compares (using particularly statistical quantification) results of research in literature about the topic published so far with further possibilities of research based on the material of the Lexical database of humanistic and baroque Czech of OVJ ÚJČ AV ČR. The article also deals with thematic spheres, in which the lexical borrowings have been used.
We suggest a new annotation scheme for unlexicalized PCFGs that is inspired by formal language theory and only depends on the structure of the parse trees. We evaluate this scheme on the TüBa-D/Z treebank w.r.t. several metrics and show that it improves both parsing accuracy and parsing speed considerably. We also show that our strategy can be fruitfully com-bined with known ones like parent annota-tion to achieve accuracies of over 90 % la-beled F1 and leaf-ancestor score. Despite increasing the size of the grammar, our annotation allows for parsing more than twice as fast as the PCFG baseline. 1
This work presents an improvement on a novel Selection Method to develop applications in the context of the Service-Oriented Computing paradigm. We have defined an Interface Compatibility procedure to assess Web Services by exploring the available information from WSDL documents. Such information involves data types from return, parameters and exceptions, and identifiers from parameters and operation names. The lexical database WordNet was originally used as a semantic basis to assess terms from identifiers. In this paper we use the DISCO database as an alternative, to evaluate the independence of the approach w.r.t. the semantic basis. We made a comparative analysis through different experiments with a data-set of 465 real-life Web Services and measured the results using metrics from the Information Retrieval field.
Discourse relations bind smaller linguistic elements into coherent texts. However, automatically identifying discourse relations is difficult, because it requires understanding the semantics of the linked sentences. A more subtle challenge is that it is not enough to represent the meaning of each sentence of a discourse relation, because the relation may depend on links between lower-level elements, such as entity mentions. Our solution computes distributional meaning representations by composition up the syntactic parse tree. A key difference from previous work on compositional distributional semantics is that we also compute representations for entity mentions, using a novel downward compositional pass. Discourse relations are predicted not only from the distributional representations of the sentences, but also of their coreferent entity mentions. The resulting system obtains substantial improvements over the previous state-of-the-art in predicting implicit discourse relations in the Penn Discourse Treebank.
Bilingual Base Noun Phrase (BaseNP) extraction is one of the key tasks of Natural Language Processing (NLP). This task is more challenging for the pair of English-Vietnamese due to the lack of available Vietnamese language resources such as treebanks, part-of-speech taggers, and parsers. In this paper, we propose a combination model that uses language characteristics based on statistics and the projection method to extract BaseNP correspondences from a bilingual corpus. The language characteristics used in this model include the word segmentation, word order and word classification [1]. Our model overcomes not only the lack of resources of Vietnamese, but also improves the performance of miss-alignment, null-alignment, overlap and conflict projection of the existing methods. The proposed model can be easily applied to other language pairs. Experiment on 66,646 pairs of sentences in the English-Vietnamese bilingual corpus shows that our proposed model is very satisfactory.
The corpus for training a parser consists of sentences of heterogeneous grammar usages. Previous parser domain adaptation work has concentrated on adaptation to the shifts in vocabulary rather than grammar usage. In this paper, we focus on exploiting the diversity of training date separately and then accumulates their advantages. We propose an approach that grammar is biased toward relevant syntactic style, and the complementary grammar usage are combined for inference. Multiple grammars with partly complementary points of strength are induced individually. They capture complementary data representation, and we accumulates their advantages in a joint model to assemble the complementary depicting powers. Despite its compatibility with many other methods, out product model achieves 85.20% F <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sub> score on Penn Chinese Treebank, higher than previous systems.
The article analyzes main trends in the transformation of modern media that influence the subject and methods of media-linguistics. It examines integration of different media platforms and integration of the national mass media into the sphere of global mass media; transformation of structural components of mass media system, and institutionalization of mass media. According to the author, these processes change the traditional model of mass media and causes fundamental changes in the structure of media texts. Among these changes are a significant complication of the semantic load of each text element; responsiveness; and visuality.The article presents the analysis of communication matrix used to construct media texts. It is shown that we talk about such linguistic characteristics of modern media, as a simplification of the content based on the perception of editorial boards target audience; structural-compositional simplicity of the media texts, due to the fact that members of audience must decide on the desirability for themselves certain information in a single glance at it; the transformation of the language of communication, which entailed the removal of a large number of linguistic norms, especially for Internet communications. In particular, incorrect speech transforms from a communication error into the source of creating distinctive author’s individuality.
In a large scale study on 843 transcripts of Technology, Entertainment and Design (TED) talks, the authors address the relation between word usage and categorical affective ratings of lectures by a large group of internet users. Users rated the lectures by assigning one or more predefined tags which relate to the affective state evoked in the audience (e. g., ‘fascinating', ‘funny', ‘courageous', ‘unconvincing' or ‘long-winded'). By automatic classification experiments, they demonstrate the usefulness of linguistic features for predicting these subjective ratings. Extensive test runs are conducted to assess the influence of the classifier and feature selection, and individual linguistic features are evaluated with respect to their discriminative power. In the result, classification whether the frequency of a given tag is higher than on average can be performed most robustly for tags associated with positive valence, reaching up to 80.7% accuracy on unseen test data.
Bilingual Base Noun Phrase (BaseNP) extraction is one of the key tasks of Natural Language Processing (NLP). This task is more challenging for the pair of English-Vietnamese due to the lack of available Vietnamese language resources such as treebanks, part-of-speech taggers, and parsers. In this paper, we propose a combination model that uses language characteristics based on statistics and the projection method to extract BaseNP correspondences from a bilingual corpus. The language characteristics used in this model include the word segmentation, word order and word classification [1]. Our model overcomes not only the lack of resources of Vietnamese, but also improves the performance of miss-alignment, null-alignment, overlap and conflict projection of the existing methods. The proposed model can be easily applied to other language pairs. Experiment on 66,646 pairs of sentences in the English-Vietnamese bilingual corpus shows that our proposed model is very satisfactory.
This chapter explores the challenges of developing the field of Latin Computational Linguistics. Computational Linguistics aims at designing, implementing, and applying computational models for natural languages. A large part of Computational Linguistics research has been developed for English, or at least tested on this language. A crucial aspect of a fruitful exchange between the disciplines of Latin Linguistics and Computational Linguistics concerns the way Latin texts are collected, accessed, and investigated for linguistic analyses. So, any attempt into Latin Computational Linguistics is likely to start from corpora. Annotation provides each word form in a sentence with one or more labels that mark its attributes; for example, a morpho-syntactic annotation would add a 'genitive' tag to puellarum. The chapter advocates the use of corpora in Latin Linguistics by reporting on research based on Latin treebanks, to show their potential for Historical Linguistics research.Keywords: historical corpora; historical languages; historical Linguistics research; Latin Computational Linguistics; morpho-syntactic annotation
Abstract While some research has been carried out on linguistic innovations in some parts of Africa, studies are yet to attend to how linguistic innovations can be used in strengthening democratic institutions for good governance in Nigeria. This paper addresses this gap by investigating the challenges and strategies in approaching the use of English language in various democratic institutions through the use of loan words, coinages/neologisms, pidginization and code switching. Some sampled expressions that are routinely used in Nigerian social contexts were subjected to analyses, using insights from the theoretical coverage of linguistic determinism/relativity. This paper notes that deviation from the international linguistic norm helps citizens to express and understand themselves properly, and this promotes good governance as it unites people together, in and outside any institution. Although challenges such as public scepticism, apathy and persuasive distrust of politicians may be prevalent in politics, the paper contends that linguistic innovations in democratic institutions offer an outlet for creativity in language for beneficial effects. The paper concludes that in this age of globalisation, Nigerians cannot deny themselves of good governance by adopting the bookish imported language with unending rules. They should consider reconstructing English as a medium of communication for strengthening democratic institutions
Recent studies such as Manova and Aronoff (2010) and Zirkel (2010) provide evidence that suffix combinations are regulated by grammatical and processing factors. However, some languages like Athapaskan languages exhibit the affix ordering which cannot be constrained by linguistic principles. In order to account for the type of affix combinations, a template has been presented as a morphological device for word formation (Kari 1989, Rice 2000). In Old English, up to three derivational suffixes can occur to a single base in order, and I have found 23 suffix combinations through the Old English lexical database Nerthus. As a morphological means for ordering the suffixes of the combinations, I present a template and attempt to provide a templatic analysis of the combinations. I argue that the OE template is also useful to explain some morphological processes like inflectional endings with derivational function and the redundant occurrence of the adjectival suffix?lic. It will be shown that the morphological processes as well as the suffix combinations can be accounted for in a unified way in terms of the template.
Abstract—The article contains the concept of developing a motivation model aimed at supporting activity of both students and teachers in the process of implementing and using an open and distance learning system. Proposed motivation model is focused on the task of filling the knowledge repository with high quality didactic material. Open and distance learning system assures a computer space for the teaching/learning process in open environment. The structure of the motivation model and formal assumptions are described. Additionally, there is presented a structure of the linguistic database, helping the teacher to assess the student&apos;s motivation and the basic simulation model to analysis the teaching/learning process constrains. The proposed approach is based on the games theory and simulation approach. Keywords-motivation model; computer learning platform; knowledge repository; non-cooperative game. I.
This paper reports on a corpus-based quantitative study of the use of nominalizations across China English and British English in two comparable media corpora. In contrast to previous corpus-based studies of nominalizations, we start by using a syntactic approach and proceed with some methodological innovations incorporating large lexical databases and syntactically annotated corpora. The data show that there are significant differences in the use of nominalizations across these two English varieties. It is hoped that this research will offer useful insights on variations in nominalization across different English varieties and also on the understanding of the two English varieties in question. 1
The aim of this study was to evaluate imaging-based predictive validity to Alzheimer's disease in amnestic Mild Cognitive Impairment (MCI) patients using 18F-fluorodeoxyglucose Positron Emission Tomography (FDG PET) by an automatic computer-assisted system and visual evaluation rating. Baseline FDG PET images were compared between MCI converter to Alzheimer's disease group (MCI-C) and MCI non-converter group (MCI-NC) to get predictive validity by automatic computer-assisted system and two blinded human nuclear medicine professionals. To obtain the optimal cut-off point to discriminate MCI-C from MCI-NC, Minimal distance method from Receiver Operating Characteristic (ROC) curve was used. Fifty-nine patients with MCI received baseline FDG PET examinations when enrolled and were followed up for two years. After two years, 22 MCI patients (37.29%) were converted to Alzheimer's disease. When using automatic computer-assisted system, distance from ROC is the shortest in parietal lobe and at this point, sensitivity is 0.63 and specificity is 0.66. When using visual evaluation rating, distance from ROC is the shortest in parietal lobe and at this point, sensitivity is 0.64 and specificity is 0.70. Kappa value in inter-rater reliability is high (0.84). In the predictive validity from MCI to Alzheimer's disease, computer-assisted system showed similar accuracy to visual imaging rating and cannot replace visual imaging evaluation.
This dissertation attempts to link the linguistic information contained in REDES Diccionario combinatorio del espanol contemporaneo (Bosque, 2004), or REDES, to the ontological framework of Functional Grammar Knowledge Base, or FunGramKB. REDES is a dictionary that gathers systematic restrictions imposed by some 4,000 Spanish predicates to their selection of lexical arguments, while FunGramKB is a multilingual and multipurpose lexical conceptual knowledge base (KB) designed for Natural Language Processing (NLP). This work is thus part of the field known as electronic lexicography for the XXIst century or the third millennium (Fuertes and Tarp, 2011).This new lexicography comprises electronic lexicographical resources which are much more complex than the commonly used digitalized versions of traditional dictionaries. It is composed of lexical databases or lexical knowledge bases, of variable complexity and depth, built on electronic platforms. Even though electronic lexicography can also serve the typical dictionary queries posed by humans, the best-known NLP applications include machine translation (MT), question and answer systems, information extraction, and voice recognition programs. However, the main problem facing NLP continues to be Word Sense Disambiguation (WSD). Most words in a language are polysemic and a successful NLP application will depend on its ability to assign the correct sense to each word in a given context. A great deal of the work that is taking place in NLP is therefore focused on finding strategies to effectively disambiguate words in their context.In working with REDES and FunGramKB, it could seem at first glance that we are dealing with two distant fields –print lexicography and knowledge engineering, respectively–, but this thesis assumes that linking these resources would yield significant benefits for both. Furthermore, it would combine two sources of valuable linguistic and theoretical data: on the one hand, REDES contributes thousands of patterns of systematic predicate-argument word combinations, taken from real use in Spanish-language corpora. These combinations are not presented as “collocations”, that is, combinations between a single predicate and a single argument, or vice versa, but rather, between predicates and “lexical classes”, groups of arguments that share a common semantic basis. On the other hand, FunGramKB contributes a multilevel electronic platform designed for NLP, which includes a conceptual, a lexical and a grammatical model, and has been built around a hierarchical and taxonomical ontology of universal cognitive concepts.
While some cross-modal associations might have psychobiological basis, other patterns of association might be acquired or cultural (cf. [5], [11], [4]). Stimuli perceived through different sensory organs via parallel brain pathways may be associated at a higher level if they both happen to have the same effect on emotional state, mood, or affective state ([14]). If the perceived input under-specifies an event, more complex cognitive processing mechanisms kick in ([7]). This process is not primarily ecological and might be mediated by emotion ([12]). \nResearch on crossmodal matching has provided evidence that many non-arbitrary and universal correspondences exist. Audio-visual correspondences may be based on amodal correspondences, for example, the loudness of sound and the luminosity of light ([14]). [13] showed that most cultures display word clusters near ‘red’, ‘green’, ‘yellow’, and ‘blue’ (in addition to ‘white’ and ‘black’), and argued that “focal colours” really are universal. \nBresin ([2]) derived 24 colours from a scheme of selecting approximately equal distances in colour parameters in HSL (Hue, Saturation, Lightness) “space”. This produced a set where the colour patches are arguably more evenly distributed, from a perceptual point of view, than those in the two studies mentioned above. Bresin found correlations between colour parameters and the affective intent in music excerpts, i.e. listeners matched colours to music excerpt played with a certain ‘feeling’. As in [1] but more general, colour brightness was associated with positive emotion and darkness with negative emotion. \nPalmer ([12]) investigated colour association to classical music excerpt where tempo and tonal mode were manipulated. The authors found that colours of high saturation and brightness, and colours more towards yellow (‘warmth’) were selected for music stimuli in fast tempo, and that conversely, de-saturated (‘grayer’), ‘darker’, and blue colours were selected for music of slow tempo in minor mode. Furthermore, they claimed strong support for emotion as a mediating mechanism for the cross-modal associations. \nA review of the research provoked the idea that colour association to sound might be context-dependent. When associating colour to music, natural soundscapes, and ‘soundscape compositions’, do people use different strategies? Which musical features influence colour association? Can emotion mediate between musical features and colour association? \nWe designed an experiment to investigate a) correlations between visual colours defined by linear parameters and music stimuli with previously validated affect; b) correlations between the colour parameters and computational acoustic and musical features; and c) the multiple regressions onto colour parameters of affective ratings (emotions), psychoacoustic descriptors, and musical features.
It is impossible not to communicate.(Buda Bela)Communication implies transmission of information, ideas, feelings by means of symbols (words, images, graphics etc.).(Berelson-Steiner)As each language has several variants and is in a continuous process of change, I should like to compare the correct and incorrect text, which is used mainly on radio and television, by presenting authentic texts from mass media in Romania. Once television came into being, the question was whether this would not put the radio into shade. In time, it was proved that both media sources are needed. They function not in parallel, they complete each other, but each of them has its own well established role in the audio-visual mass-media. Nevertheless, it is also true that radio journalists should revise their journalistic genre - right under the influence of television. It has been necessary for them to renew their radio phonic text.Television can transmit information without a text, only by means of the language of images. While on the radio, the audio effect creates satisfaction. The nuance, the power, intonation, articulation and sound on the radio are part of the given information, having also a special role in the transmission of information. These have been the reasons for the radio phonic text and the text of the radio announcer to have developed as a distinct branch of the media science.Any person in good health can speak, but to be able to communicate in different situations, to express the nuance, rhythm, intonation of a text, you should have a certain training, you should be a specialist. Speaking in public has precise rules: how much of the text should be assimilated by the speakerwhich is the best variant, if he assimilates the text in its entirety or if he keeps a distance from it. Can he have control over his emotions or, because of these, he breathes faster and the rhythm of the speech is more rapid; does he raise the pitch of his voice and speaks louder? (e.g. during sports transmissions). What is the role of the logical intonation and of the emotive one? How does the melody of speech influence the intonation of the word, of the sentence, of the text? These and other are the questions I will try to answer. Mass-media has a decisive role in modem society, because people of our days get information about everything going on in the world by means of it. In this way, a person acquires knowledge about the world. It is very important for the language, for the speech, by means of which information is transmitted to the receptor, to observe the correct linguistic norms, to be beautiful. Speaking is the most complex activity, which is not inborn, it is learnt from parents and in school. Speaking is a system of communication by several channels. The general content of words is changing according to intonation, rhythm, volume, articulation etc. In the development of language and speaking, mass-media has a decisive role. On the radio and television, quite often we can hear texts where we can sense that the presenter concentrates a lot on the articulation of consonants and vowels, on the tone of the voice, on intonation and punctuation. We can hear all these but we cannot hear the idea in itself. Although, radio cannot transmit by images - as it is the case with televisionit has also its own informational language: the human voice, the speech of the presenter. By the harmony between the meaning of the text and the phonetic elements, a new quality can be achieved: the auditive influence leads to the visual effect. This inner visuality has a special effect. Today, it is evident that television does not put the radio into the shade.The acoustic language (and the read text), especially the spontaneous live one is related to the personality of the speaker, to the situation and the receptor. As a rule it is emphatic, it is a complex communication.Besides articulation, the live speech contributes to communication by the common and alternative use of phonetic instruments. …
哈金喜歡把玩中英文不同的語言特色,在創作中將兩者合而為一。他以本土化的英語論述重塑中文的諺語、譬喻與句構,翻譯、挪用並重組中文的語法,創造出饒富中文色彩的英語創作。閱讀哈金的詩集《殘骸》(Wreckage),受過中文教育的讀者很容易就可以發現詩中有許多中國古典詩詞的翻譯及改寫。作品中的詩人引述古典詩詞,以抒發個人情懷,而讀者不禁會思考:哈金的詩作源頭到底有多大的成分是來自中國文學?我們能夠依據哪些特點以及哪些標準來辨識所謂的創意?我們要如何區分「創新」與「改造」?而這樣的區分是否會改變我們對於哈金詩作的解讀與評斷?在哈金的詩作中,對中國古典詩詞的引述與對美國現今情境的感觸,二者或並置、或融合、或相互對照、或互相影響,這其中透露著何種訊息?被翻譯的是什麼,是中國詩詞還是美國情感?被移植的又是什麼,是文學還是文化?當中國的詩文被翻譯成英文、移植至美國之後,意義是否會有所轉化?是否會因應哈金複雜的美國移民經驗而呈現出某種迥異於以往的新義?而哈金的三重身份-詩者、評者、譯者-是彼此互補或彼此衝突?而這種種文學、文化再現,又訴說了何種華人離散的意涵,這些將是本文討論的重點。 Ha Jin's works bespeak an impressive linguistic creativity. Reconfigurating Chinese language through a nativized discourse of English, Ha Jin has translated, appropriated, and reconstructed Chinese linguistic norms and specifics into English-language literature in remarkable fashion. Studying Ha Jin's "Wreckage", a reader with Chinese education will be ready to identity fragments of renowned Chinese classical verse. Thus the reader is invited to ponder: To what extent does Ha Jin draw his poetic inspiration from the corpus of Chinese literature? How shall we measure accredited creativity? How do we distinguish innovation from renovation, and do those distinctions change our reading of the poems? Although Ha Jin has written exclusively of the reality of Chinese politics and society (with "A Free Life" the only exception so far), this material does not obviate the possibility of reading his works as belonging to a tradition of US immigrant literature. In Ha Jin's poetry, the juxtaposition, interaction and fusion of classical Chinese verse and contemporary American sensibility can be telling. Which has been translated, the Chinese verse or the American sensibility? What has been transplanted and translated? Have the Chinese poetry texts, after being transplanted into the English verse, undergone a transformation of meaning and resurfaced with new significance corresponding to the complexity of Ha Jin's immigrant experiences in America? This paper aims to explore how Ha Jin's Chinese poetry texts show significance corresponding to the complexity of his immigrant experiences in America.