Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
The corpus for training a parser consists of sentences of heterogeneous grammar usages. Previous parser domain adaptation work has concentrated on adaptation to the shifts in vocabulary rather than grammar usage. In this paper, we focus on exploiting the diversity of training date separately and then accumulates their advantages. We propose an approach that grammar is biased toward relevant syntactic style, and the complementary grammar usage are combined for inference. Multiple grammars with partly complementary points of strength are induced individually. They capture complementary data representation, and we accumulates their advantages in a joint model to assemble the complementary depicting powers. Despite its compatibility with many other methods, out product model achieves 85.20% F <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sub> score on Penn Chinese Treebank, higher than previous systems.
Bilingual Base Noun Phrase (BaseNP) extraction is one of the key tasks of Natural Language Processing (NLP). This task is more challenging for the pair of English-Vietnamese due to the lack of available Vietnamese language resources such as treebanks, part-of-speech taggers, and parsers. In this paper, we propose a combination model that uses language characteristics based on statistics and the projection method to extract BaseNP correspondences from a bilingual corpus. The language characteristics used in this model include the word segmentation, word order and word classification [1]. Our model overcomes not only the lack of resources of Vietnamese, but also improves the performance of miss-alignment, null-alignment, overlap and conflict projection of the existing methods. The proposed model can be easily applied to other language pairs. Experiment on 66,646 pairs of sentences in the English-Vietnamese bilingual corpus shows that our proposed model is very satisfactory.
Discourse relations bind smaller linguistic elements into coherent texts. However, automatically identifying discourse relations is difficult, because it requires understanding the semantics of the linked sentences. A more subtle challenge is that it is not enough to represent the meaning of each sentence of a discourse relation, because the relation may depend on links between lower-level elements, such as entity mentions. Our solution computes distributional meaning representations by composition up the syntactic parse tree. A key difference from previous work on compositional distributional semantics is that we also compute representations for entity mentions, using a novel downward compositional pass. Discourse relations are predicted not only from the distributional representations of the sentences, but also of their coreferent entity mentions. The resulting system obtains substantial improvements over the previous state-of-the-art in predicting implicit discourse relations in the Penn Discourse Treebank.
This work presents an improvement on a novel Selection Method to develop applications in the context of the Service-Oriented Computing paradigm. We have defined an Interface Compatibility procedure to assess Web Services by exploring the available information from WSDL documents. Such information involves data types from return, parameters and exceptions, and identifiers from parameters and operation names. The lexical database WordNet was originally used as a semantic basis to assess terms from identifiers. In this paper we use the DISCO database as an alternative, to evaluate the independence of the approach w.r.t. the semantic basis. We made a comparative analysis through different experiments with a data-set of 465 real-life Web Services and measured the results using metrics from the Information Retrieval field.
Part-of-speech (POS) taggers can be quite accurate, but for practical use, accuracy often has to be sacrificed for speed. For example, the maintainers of the Stanford tagger (Toutanova et al., 2003; Manning, 2011) recommend tagging with a model whose per tag error rate is 17% higher, relatively, than their most accurate model, to gain a factor of 10 or more in speed. In this paper, we treat POS tagging as a single-token independent multiclass classification task. We show that by using a rich feature set we can obtain high tagging accuracy within this framework, and by employing some novel feature-weight-combination and hypothesis-pruning techniques we can also get very fast tagging with this model. A prototype tagger implemented in Perl is tested and found to be at least 8 times faster than any publicly available tagger reported to have comparable accuracy on the standard Penn Treebank Wall Street Journal test set.
This study aims to demand the improvement of a broadcasting language and raise the necessity for a nation-wide campaign for improving linguistic cultures with the correct and classy use of language, by investigating the actual usage of lower class language in media and showing the serious damage of the standard language. The categories of usage consist of vulgar expressions and terms against Eomun-gyubeom(Korean linguistic norm). Vulgar expressions mean poor quality in their contents and are divided into personality-insulting, discriminative, violent, and suggestive expressions, slang and chatting language, coarse language, unnecessary foreign language as well as loanwords. Terms against Eomun-gyubeom involve a nonstandard dialect, ungrammatical expressions, and errors in subtitles. Media language has a tremendous ripple effect on everyday language not just of the ordinary masses but also of the youth. Since media language is a public language targeted at unspecified masses, the lower class language that might damage our national language should be sublated in broadcasting.
A grammar-driven dependency parsing has been attempted for Bangla (Bengali). The free-word order nature of the language makes the development of an accurate parser very difficult. The Paninian grammatical model has been used to tackle the free-word order problem. The approach is to simplify complex and compound sentences and then to parse simple sentences by satisfying the Karaka demands of the Demand Groups (Verb Groups). Finally, parsed structures are rejoined with appropriate links and Karaka labels. The parser has been trained with a Treebank of 1000 annotated sentences and then evaluated with un-annotated test data of 150 sentences. The evaluation shows that the proposed approach achieves 90.32% and 79.81% accuracies for unlabeled and labeled attachments, respectively.
This paper describes experiments for statistical dependency parsing using two different parsers trained on a recently extended dependency treebank for Greek, a language with a moderately rich morphology. We show how scores obtained by the two parsers are influenced by morphology and dependency types as well as sentence and arc length. The best LAS obtained in these experiments was 80.16 on a test set with manually validated POS tags and lemmas. 1
This study aims to demand the improvement of a broadcasting language and raise the necessity for a nation-wide campaign for improving linguistic cultures with the correct and classy use of language, by investigating the actual usage of lower class language in media and showing the serious damage of the standard language. The categories of usage consist of vulgar expressions and terms against Eomun-gyubeom(Korean linguistic norm). Vulgar expressions mean poor quality in their contents and are divided into personality-insulting, discriminative, violent, and suggestive expressions, slang and chatting language, coarse language, unnecessary foreign language as well as loanwords. Terms against Eomun-gyubeom involve a nonstandard dialect, ungrammatical expressions, and errors in subtitles. Media language has a tremendous ripple effect on everyday language not just of the ordinary masses but also of the youth. Since media language is a public language targeted at unspecified masses, the lower class language that might damage our national language should be sublated in broadcasting.
In this study, a dictionary-based method is used to extract expressive concepts from documents. So far, there have been many studies concerning concept mining in English, but this area of study for Turkish, an agglutinative language, is still immature. We used dictionary instead of WordNet, a lexical database grouping words into synsets that is widely used for concept extraction. The dictionaries are rarely used in the domain of concept mining, but taking into account that dictionary entries have synonyms, hypernyms, hyponyms and other relationships in their meaning texts, the success rate has been high for determining concepts. This concept extraction method is implemented on documents, that are collected from different corpora.
This article is devoted to the problems of linguistic and extra-linguistic peculiarities of texts in sport reporting as a kind of sport discourse. Texts of sports report, presenting an implementation of oral public speech, are prepared beforehand as compared to colloquial speech, so the content possesses logical-semantic structure, determined by a specific topic. Logical-semantic structure of the text is shown in the selection of language means. Compositional and stylistic features of a sports report are based on the rules of its construction. Violation of these rules leads to a variety of errors in the speech of sports journalists. The article provides examples of various violations of linguistic norms of the literary language and the requirements to the utterance. The speech of sports journalists is considered in two forms: monologue and dialogue. Results of described texts of sports report can be used in practice during the study of the discipline “Russian language and culture of speech”, as well as in teaching Russian as a foreign language (for teaching listening).
Lexical databases of semantic frames have been shown to be useful in problems related to natural language processing.However, creation of such databases is a task that is time consuming and involves many manual steps.One of these steps is selection and grouping of sentences to identify frames.However, we advocate that if sentences were previously annotated with ontological information, this grouping could be executed automatically.In this article we present tests performed with clustering sentences containing the lexeme Travel (noun and verb).Tests showed that the use of clustering algorithms on ontologically annotated sentences is a promising step towards automating construction of semantic frames databases.
For languages such as English, several constituent-to-dependency conversion schemes are pro-posed to construct corpora for dependency parsing. It is hard to determine which scheme is better because they reflect different views of dependency analysis. We usually obtain dependen-cy parsers of different schemes by training with the specific corpus separately. It neglects the correlations between these schemes, which can potentially benefit the parsers. In this paper, we study how these correlations influence final dependency parsing performances, by proposing a joint model which can make full use of the correlations between heterogeneous dependencies, and finally we can answer the following question: parsing heterogeneous dependencies jointly or separately, which is better? We conduct experiments with two different schemes on the Penn Treebank and the Chinese Penn Treebank respectively, arriving at the same conclusion that joint-ly parsing heterogeneous dependencies can give improved performances for both schemes over the individual models.
P.L. Grokhovskiy, V.P. Zakharov, A.M. Popov, M.O. Smirnova, M.V. Khokhlova The paper describes the process of developing a parallel Tibetan-Russian corpus. The corpus represents the collection of Tibetan grammar treatises from VII to XX cc. and their translation into Russian. The Tibetan linguistic tradition is mainly based on Buddhist grammars composed by Indian scholars. Thus it is closely connected with the Indian grammatical tradition. The project particularly aims at the creation of specific grammatical lexical database on the basis of corpus. The database will be useful both to tibetologists and general linguistics specialists.
Assisted by the Chinese-English parallel corpus with one source text and its three different English translations,the paper analyzes the Chinese construction of BU A BU B. The study shows that the construction consisting of opposite morphemes is more popular than that of similar morphemes,and the former is often translated asneither… nor…while the latter is often literally translated. The order of A and B in the English translation depends upon the usual linguistic norms of the English expressions. The English translations of the construction are usually equivalent to their Chinese source texts,and the interlingual explicitation or implicitation is not so obvious. The translators' native language backgrounds have little affect on their translations,but the choice of words in their translations do reflect their language accomplishments to some extent.
We examined the potential cost of practicing suppression of negative thoughts for subsequent performance in an unrelated task.Cues for previously suppressed and baseline responses in a think/no-think procedure were displayed as irrelevant flankers for neutral words to be judged for emotional valence.These critical flankers were homographs with one negative meaning denoted by their paired responses during learning.Suppression cues as flankers delayed responding to the targets, compared to baseline cues and new negative homographs, but only following directsuppression instructions and not when benign substitutes had been provided to aid suppression.On the final recall test, suppression-induced forgetting (SIF) following direct suppression and the flanker task was positively correlated with the flanker effect.Experiment 2 replicated these findings.Finally, valence ratings of neutral targets were influenced by the valence of the flankers but not by the prior role of the negative flankers.
We propose Diverse Embedding Neural Network (DENN), a novel architecture for language models (LMs). A DENNLM projects the input word history vector onto multiple diverse low-dimensional sub-spaces instead of a single higher-dimensional sub-space as in conventional feed-forward neural network LMs. We encourage these sub-spaces to be diverse during network training through an augmented loss function. Our language modeling experiments on the Penn Treebank data set show the performance benefit of using a DENNLM.
We examine the consequences, of integrating large minorities into productivity-relevant majority ethno-linguistic norms, for distribution, ethnic conflict and crime. We develop a two-community model where such assimilation generates social gains by: (a) facilitating economic interaction, and (b) dampening religious or racial conflict over symbolic and normative contents of the public sphere. However, integration shifts the distribution of both material and symbolic goods against the minority. It also expands income inequality within the minority community. This incentivizes decentralized attempts to expropriate producers which, through cumulative causation, both immiserize and criminalize the minority. An underclass thus results, with disproportionate minority presence.
The varieties of diachronic aspect occur in a language under the influence of external circumstances and they are the evidence of how the respective language users have lived and spoken in the past. Due to the fact that they do not correspond to the linguistic norms, they are often difficult to translate, so the translators use various methods in order to avoid possible wrong understanding or incorrect transfer into the target language. This paper researches some diachronic varieties in the German and the Macedonian vocabulary, compared to the translations of these phenomena into Macedonian and German language, respectively. The analysis is performed on selected literary texts, in parallel with their translations in the other respective language. The focus is on the Turkish words in the Macedonian language and the Latin and French words in the German language.
The authors focus on how to segment semantic units in Chinese discourse and how to label relations among semantic units automatically. During the parsing process, several sequence labelling methods are compared for discourse segmentation, while a maximum entropy-based training and decoding algorithm is specially proposed. Experiments are done based on Tsinghua Chinese Treebank, which is annotated with logical and semantic relations at complex-sentence level. Experimental results show that F-score of discourse segmentation reaches 89.1%. When parsing discourses with no more than 6 relations included, the labeling F-score can achieve 63%.
The article discusses lexical borrowings from German to Czech, especially to Middle Czech. It compares (using particularly statistical quantification) results of research in literature about the topic published so far with further possibilities of research based on the material of the Lexical database of humanistic and baroque Czech of OVJ ÚJČ AV ČR. The article also deals with thematic spheres, in which the lexical borrowings have been used.
The domain of air traffic control is the perfect example for the analysis of a linguistic norm. In this field, phraseology is a specialised language created to cover the most common situations encountered in air navigation in order to secure and optimise radiotelephony communications. When phraseology proves inadequate, a more natural form of language, called plain language, is required: it was recently introduced in the field and is a difficult notion to implement. A comparative analysis between a reference corpus and a real-communication corpus allows the description and categorisation of different types of variations used on the radio frequency as well as a discussion on the notions of norm and usage in the field of air traffic control.
The present contribution represents the first step in comparing the nature of syntactico-semantic relations present in the sentence structure to their equivalents in the discourse structure. The study is carried out on the basis of Czech manually annotated material collected in the Prague Dependency Treebank (PDT). According to the analysis of the underlying syntactic structure of a sentence (tectogrammatics) in the PDT, we distinguish various types of relations that can be expressed both within a single sentence (i.e. in a tree) and in a larger text, beyond the sentence boundary (between trees). We suggest that, on the one hand, semantic nature of each type of these relations corresponds both within a sentence and in a larger text (i.e. a causal relation remains a causal relation) but, on the other hand, according to the semantic properties of the relations, their distribution in a sentence or between sentences is very diverse. In this study, this observation is analyzed in detail for three cases (relations of condition, specification and opposition ) and further supported by similar behaviour of the English data from the Penn Discourse Treebank.
Completely data-driven grammar training is prone to over-fitting. Human-defined word class knowledge is useful to address this issue. However, the manual word class taxonomy may be unreliable and irrational for statistical natural language processing, aside from its insufficient linguistic phenomena coverage and domain adaptivity. In this paper, a formalized representation of function word subcategorization is developed for parsing in an automatic manner. The function word classification representing intrinsic features of syntactic usages is used to supervise the grammar induction, and the structure of the taxonomy is learned simultaneously. The grammar learning process is no longer a unilaterally supervised training by hierarchical knowledge, but an interactive process between the knowledge structure learning and the grammar training. The established taxonomy implies the stochastic significance of the diversified syntactic features. The experiments on both Penn Chinese Treebank and Tsinghua Treebank show that the proposed method improves parsing performance by 1.6 % and 7.6 % respectively over the baseline. 1
On the basis of studying the lexicographic sources a definition of cynicism the term “linguocinism” is given in the article. The meaning of this term is specified by its correlation with similar to its meaning, but not synonymous to the terms “vulgarisms” and “obscenism”. It became clear that vulgarisms and obscenisms being violation of ethic-linguistic norm are not always linguocinisms. Linguistic means of expressing linguocinism and the role of context in their realization are revealed.
Correspondence analysis (CA) was conceived by Benzécri (1977b; Benzécri et al., 1973, 1981) as an inductive method able to infer from their own texts of interest the linguistic norms that rule them. He wanted to offer a powerful tool to tackle all kinds of issues related to the form, meaning, and style of texts.
This paper presents a Naxi dependency parsing method which combined rules and statistics, based on the characteristic of the Naxi language. Firstly, the annotation standard of Naxi Dependency Treebank was established and dependencies were defined which are based on the syntactic structure characteristics of Naxi language, then Naxi Dependency Treebank was constructed; Secondly, the constructed Naxi dependency Treebank was used as the basis, and the Naxi language phrases were analyzed independently by means of rules, then the boundaries and categories of phrases were determined. The features of core words postposition in the Naxi language were regarded as a condition of constraint. Then the dependencies between phrases were analyzed. Finally, the dependencies between phrases were combined, and interdependent probabilistic models were used to parse the syntax of Naxi language, and finally a completed Naxi language dependency parsing system was achieved. The results of dependency parsing comparative experiments show that the use of Naxi dependency rules improves the performance of Naxi language dependency parsing system. Moreover, the performance of the proposed method is better than the dependency parsing methods which are fully based on statistics.
Resumen Este artículo analiza las actitudes lingüísticas de hablantes nativos de español de la Ciudad Autónoma de Buenos Aires, hacia al español de la Argentina y el español de los otros países hispanohablantes. El artículo es parte de los resultados del Proyecto LIAS (Linguistic Identity and Attitudes in Spanish-speaking Latin America), financiado por El Consejo Noruego de Investigación (RCN). La recolección de los datos se realizó en la capital del país, entrevistando a una muestra de 400 informantes previamente estratificada con las variables de edad, sexo y nivel socioeconómico. El procesamiento estadístico de los datos de campo recolectados arrojó resultados de interés en torno a la mayoría de los tópicos analizados y especialmente en lo referente a aspectos tales como la valoración positiva de la propia variedad lingüística; la resistencia a identificar a España como la única fuente de la norma lingüística de la lengua española; el rechazo a la unificación de la lengua y, por consiguiente, la defensa de la diversidad lingüística como portadora de riqueza cultural. Abstract This article analyzes the linguistic attitudes of native Spanish speakers from Buenos Aires City, towards Spanish spoken in Argentina and in the other Spanish-speaking countries. It is a result of the LIAS-Project (Linguistic Identity and Attitudes in Spanish-speaking Latin America), funded by The Research Council of Norway (RCN). The data were gathered in the capital of the country, interviewing a stratified sample of 400 respondents, based on the variables of age, sex and socioeconomic status. The analysis of the data rendered interesting results on most of the analyzed topics; especially important was the positive appraisal of Argentineans' own linguistic variety; the strong resistance against identifying Spain as the only source of the linguistic norm for the Spanish language; and the rejection of language unification, defending in this way linguistic diversity as an important conveyor of cultural richness.
This article deals with the problem of deviation of linguistic norm in the artistic text by the example of the novel by A. Doblin «Berlin Alexanderplatz». The aims of this article were to define the means of the violation of the text structure, to consider expressive constructions used for the realization of authors certain intensions.
HamleDT 2.0 is a collection of 30 existing treebanks harmonized into a common annotation style, the Prague Dependencies, and further transformed into Stanford Dependencies, a treebank annotation style that became popular recently. We use the newest basic Universal Stanford Dependencies, without added language-specific subtypes.
Context \nGrETEL and Treebanks \nExample Parse \nSearching in treebanks \nSearching with GrETEL \nSearching with GrETEL: Limitations \nComparison with Google \nConclusions and Invitation
The author of the article suggests the taxonomy of the objectivation methods of pragmatic and linguistic norms of professional sublanguage: subjective method, semantic method (by enumerating referent's nominations), method of limiting the domain of the possible and the method of extensional definition. Professional sublanguage demonstrates the implementation of pragmatic and linguistic norm primarily as pointing out the domain of the possible (dysphemisation and the derisionof referent). The article shows that the relevance of norm is increased in the context of external contacts and the elimination of communication.
ABSTRACTIndividual differences in estimates of respiratory sinus arrhythmia (RSA) have been implicated in the ability to respond efficiently to challenges encountered in different environmental contexts. Studies have used resting RSA and the suppression of RSA to emotional or challenging cognitive tasks to assess parasympathetic contributions to emotional and attentional regulation. In the current study, estimates of resting RSA, and RSA suppression were used in conjunction with measures of attentional control and anxiety to predict self -rated measures of executive functions and task performance on a mental math exercise. The data suggest that resting RSA, in conjunction with good attentional control and low anxiety are better predictors of global measure of a participant's ability to employ executive processes compared to RSA suppression. Both resting RSA and RSA suppression predicted more efficient task performance on the mental math exercise. The suppression of RSA, however, was more predictive of better performance in participants with poor attentional control and high anxiety. Resting RSA may be a better predictor of the ability of an individual to respond to challenges encountered in the environment. The degree of RSA suppression may be task specific and related to the interpretation of the degree of challenge associated with the task requirements.KEYWORDS: executive function, attention, emotion, anxiety, heart rate variabilityThe ability to engage various executive processes has been shown to be directly related to the regulation of attention and negative affect (Healy, Treadwell, & Reagan, 2011). Initial work by Derryberry and Rothbart (1988) defined attentional control as the ability to sustain attention to aspects of the environment in addition to an ability to shift one's attentional focus as dictated by the demands of differing environmental events. This flexibility in attentional processing has been seen as representing the functioning of an anterior attentional system linked to frontal lobe function which would allow an individual to engage and disengage attention to aspects of the environment. At the other extreme, attentional rigidity has been linked to the prevalence of anxiety disorders whereby individuals have difficulty disengaging from aspects of the environment associated with threat (Derakshan & Eysenck, 2009; Wells & Matthews, 1994). Research suggests that the ability to allocate attention to different aspects of the environment is useful in situations that require a reduction in negative affect (Lonigan, Vasey, Phillips, & Hazen, 2004).From a physiological perspective, estimates of the parasympathetic contribution to the beat-to-beat variation in heart rate (i.e., respiratory sinus arrhythmia, or RSA) have been used as a marker of emotional self-regulation and attentional control (Applehans & Luekan, 2006). Initial research had placed an emphasis on resting measures of RSA as representing an individual's level of responsiveness with higher resting estimates predicting more positive outcomes and/or behaviors from infancy to adulthood. In this regard, there exits an extensive literature that points towards high resting levels of RSA as being related to approach behavior, and emotional self-regulation in infants and children (Calkins & Keene, 2004; Kagan, Resnick, & Snidman; 1987; Stifter & Fox, 1991), increased attentional control in infants and children (Richards & Casey, 1992), positive emotional expressiveness in children (Calkins, 1997), modulation of negative affect, ratings of lower anxiety in adults, and increases in positive mood (Greaves-Lord, Ferdinand, Sondeijker, Dietrich, Oldehinkel, Rosmalen, Ormel, & Verhulst, 2007; Gyurak & Ayduk, 2008; Thayer, Friedman, & Borkovec, 1996). In contrast, lower resting RSA has been linked to chronic anxiety and the prevalence of anxiety disorders such as generalized anxiety disorder and post-traumatic stress disorder in adults (Friedman, 2007). …
Imagine saying to Hegel that his dialectic patterns occur in the behavior of magnets and rivers and trees, and that the self-understanding of spirit is an operation going on in brain tissue. That the descriptions in his Phenomenology of Spirit picture the behavior of social groups and changes in cultural norms and memes. Hegel's reaction would be complex but not hostile. You would find yourself in a discussion with him about different kinds and levels of categories, and the relation of physical science to his overall logic.Now imagine saying to Heidegger that his description of the care structure of Dasein is a sophisticated reworking of folk psychology, and that it is the result of what amounts to a software program. That his Fourfold is a description of the appearing of a world, based upon neurological and social processes. That his history of being is open to sociological analysis and historical and economic explanation. Heidegger would react to such claims more sharply. You would find yourself in a discussion with him about being caught in das Gestell, and the need to step back from attempts to absolutize one language and one revelation of beings.How would these different reactions play out in a discussion of the ontology of the self? This present essay approaches these major continental thinkers with a question from analytic philosophy, to see how they might respond.In standard mind/body discussions we find two rival descriptive languages with different sets of entities. One talks about ideas, purposes, awareness, meanings, concepts, norms, intentions, and so on. The other talks about the behavior of cells and electrical currents, brain activities, and entities described by physics.There are thinkers who argue that folk talk about thoughts and meanings could be in principle abandoned even in everyday life, replaced by descriptions that involved only physical scientific entities. Others argue that folk talk is not dispensable. Daniel Dennett and Wilfrid Sellars would say that we need to take up the intentional stance and avow social and linguistic norms in order to function in a meaningful social world. Kant had already argued that there is a practical necessity to view ourselves as free agents, no matter what our science may say. These considerations suggest that the tension between the modes of discussion cannot be easily wished away by eliminating one of them.DIFFERENT ONTOLOGIESOne could say that these two languages are using different ontologies. I am using ontology here in a sense derived from analytic philosophers who ask your ontology include... [relations, sets, second level properties, mereological wholes, abstract entities, Cartesian souls, etc.]? An ontology in this sense is the list of approved types of entities that are being affirmed as ultimately real. In most such discussions there is little talk about ontology in an older sense, namely, about the mode of being of those beings which analytic discussions often presuppose to be immediate factual presence. Much more active modes of being are affirmed in Whitehead or in Bergson or Deleuze, or in Aristotle's doctrine of potentiality. For none of these thinkers does being equal simple, positive presence. For them such presence is a result, and it is so independently of whatever processes bring it to presence for us in experience.But if we do ask about modes of being, beyond the factual lists, then other more traditional ontological questions arise: In Aristotle's terms, are the movements of Hegel's dialectic substantial changes or accidental, or relational, or what? Is Heidegger's Dasein or Hegel's spirit a substance? Or a set of emergent properties? How do we individuate Dasein(s)? Does Hegel's spirit have its own individuated self-consciousness? What kind of being do the components of the Fourfold have, and how do they relate to everyday objects and scientific entities?Ontology matters. When David Hume looks for his self, he can't find it. …
The intent of this study is to determine what sorts of images are considered more interesting by which demographic groups. Specifically, we attempt to identify images whose interestingness ratings are influenced by the demographic attribute of the viewer’s gender. To that end, we use the data from an experiment where 18 participants (9 women and 9 men) rated several hundred images based on “visual interest” or preferences in viewing images. The images were selected to represent the consumer “photo-space” - typical categories of subject matter found in consumer photo collections. They were annotated using perceptual and semantic descriptors. In analyzing the image interestingness ratings, we apply a multivariate procedure known as forced classification, a feature of dual scaling, a discrete analogue of principal components analysis (similar to correspondence analysis). This particular analysis of ratings (i.e., ordered-choice or Likert) data enables the investigator to emphasize the effect of a specific item or collection of items. We focus on the influence of the demographic item of gender on the analysis, so that the solutions are essentially confined to subspaces spanned by the emphasized item. Using this technique, we can know definitively which images’ ratings have been influenced by the demographic item of choice. Subsequently, images can be evaluated and linked, on one hand, to their perceptual and semantic descriptors, and, on the other hand, to the preferences associated with viewers’ demographic attributes.
Opinion mining is becoming of high importance with the availability of opinionated data on the Internet and the different applications it can be used for. Intensive efforts have been made to develop opinion mining systems, and in particular for the English language. However, models for opinion mining in Arabic remain challenging due to the complexity and rich morphology of the language. Previous approaches can be categorized into supervised approaches that use linguistic features to train machine learning classifiers, and unsupervised approaches that make use of sentiment lexicons. Different features have been exploited such as surface-based, syntactic, morphological, and semantic features. However, the semantic extraction remains shallow. In this paper, we propose to go deeper into the semantics of the text when considered for opinion mining. We propose a model that is inspired by the cognitive process that humans follow to infer sentiment, where humans rely on a database of preconceived notions developed throughout their life experiences. A key aspect for the proposed approach is to develop a semantic representation of the notions. This model consists of a combination of a set of textual representations for the notion (Ti), and a corresponding sentiment indicator (Si). Thus <Ti, Si> denotes the representation of a notion. However, notions can be constructed at different levels of text granularity ranging from ideas covered by words to ideas covered in full documents. The range also includes clauses, phrases, sentences, and paragraphs. To demonstrate the use of this new semantic model of preconceived notions, we develop the full representation of one-word notions by including the following set of syntactic features for Ti: word surfaces, stems, and lemmas represented by binary presence and TFIDF. We also include morphological features such as part of speech tags, aspect, person, gender, mood, and number. As for the notion sentiment indicator Si, we create a new set of features that indicate the words' sentiment scores based on an internally-developed Arabic sentiment lexicon called ArSenL, and using a third-party lexicon called Sifaat. The aforementioned features are extracted at the word-level, and are considered as raw features. We also investigate the use of additional "engineered" features that reflect the aggregated semantics of a sentence. Such features are derived from word-level information, and include count of subjective words, average of sentiment scores per sentence. Experiments are conducted on a benchmark dataset collected from the Penn Arabic TreeBank (PATB) already annotated with sentiment labels. Results reveal that raw word-level features do not achieve satisfactory performance in sentiment classification. Feature reduction was also explored to evaluate the relative importance of the raw features, where the results showed low correlations between individual raw features and sentiment labels. On the other hand, the inclusion of engineered features had a significant impact on classification accuracy. The outcome of these experiments is a comprehensive set of features that reflect the one-word notion or idea representation in a human mind. The results from one-word also show promises towards higher level context with multi-word notions.
Ratings of previously ignored visual stimuli reveal affective devaluation of such items when compared to ratings of novel items or the targets of attention. Growing evidence suggests this effect may reflect negative affective associations elicited by attentional inhibition of visual distractors. Here we investigate whether such 'inhibitory devaluation' is limited to situations involving visual-spatial selection of environmental stimuli (i.e., external attention) or extends to the selection of competing visual representations held solely in memory (i.e., internal attention). A two-item target-localization task in Experiment 1 utilized a delayed target-category cue ('circles' or 'squares') to ensure attentional selection occurred from the contents of working memory. An n-back task in Experiment 2 was used to examine the affective consequences of rejecting continually-updated visual representations when items held in memory did not match the corresponding visual display. And a Think/No-think paradigm employed in Experiment 3 was designed to explore the affective consequences of actively suppressing longer-term visual object memories. Across this relatively-wide range of memory-based selection tasks, the ignored/rejected/suppressed visual patterns consistently received more negative affective ratings than target items. Our results are consistent with prior suggestions that similar mechanisms are involved in the attentional selection of environmental stimuli and the selection of internally-maintained information that occurs even in the absence of external sensory stimulation. The similarity in these mechanisms appears to extend not only to processes of attentional selection, per se, but also to their affective consequences. Meeting abstract presented at VSS 2014
This web service performs dependency parsing in Spanish using a Malt Parser instance.It parses plain texts introduced by the user and generates linguistically annotated Treebank instances based on a data-driven parsing model.The parsing model is induced from de dependency-annotated IULA Treebank (Marimon et al, 2012) using the language-independent MaltParser 2 system as a dependency model trainer (Nivre et al, 2007).This Treebank contains 589,542 tokens in 42,099.In order to achieve optimal performance, the training corpus was previously analyzed with MaltOptimizer 3 (Ballesteros and Nivre, 2012), a tool developed to set the best parameters for MaltParser. Inputs, outputs and formats InputsThe input to be parsed is a plain text encoded in UTF-8.It can be introduced directly as a text instance in the dialogue box, as a text file or as a URL.The input language available at this moment in the web service is Spanish (es).
Using fMRI, we investigated the behavioral and neural consequences of Compassion meditation when employed as an emotion regulation strategy and compared this to Reappraisal based cognitive emotion regulation. 15 expert meditators were scanned while either passively viewing, or using Compassion meditation or Reappraisal to regulate their emotional reactions to short film clips depicting people in distress. Subjective affect ratings showed that Compassion meditation primarily increased positive affect while Reappraisal primarily decreased negative affect. Neuroimaging results showed that the Compassion vs. Passive Viewing contrast was associated with increased activation in regions involved in affiliation and positive affect (ventral striatum, mOFC), in addition to cognitive (left IFG, TPJ, pre-SMA) and affective (sgACC) control regions. Mirroring behavioral results, the Compassion vs Reappraisal contrast showed higher activation in regions involved in negative (amygdala, insula) and positive (ventral striatum, mOFC) affect and emotional control regions (sgACC). Relatively lower activation was observed in cognitive control regions (Frontoparietal network, IFG). Our findings demonstrate the efficacy of Compassion meditation as an emotion regulation strategy, suggesting that the active regulatory mechanism of Compassion is primarily the up-regulation of positive affect. Thus, Compassion is markedly different from other coping strategies in relying less on cognitive effort when faced with stressors, suggesting it could be powerful strategy for the fostering of resilience.
BACKGROUND: Careful observation of the longitudinal course of bipolar disorders is pivotal to finding optimal treatments and improving outcome. A useful tool is the daily prospective Life-Chart Method, developed by the National Institute of Mental Health. However, it remains unclear whether the patient version is as valid as the clinician version. METHODS: We compared the patient-rated version of the Lifechart (LC-self) with the Young-Mania-Rating Scale (YMRS), Inventory of Depressive Symptoms-Clinician version (IDS-C), and Clinical Global Impression-Bipolar version (CGI-BP) in 108 bipolar I and II patients who participated in the Naturalistic Follow-up Study (NFS) of the German centres of the Bipolar Collaborative Network (BCN; formerly Stanley Foundation Bipolar Network). For statistical evaluation, levels of severity of mood states on the Lifechart were transformed numerically and comparison with affective scales was performed using chi-square and t tests. For testing correlations Pearson´s coefficient was calculated. RESULTS: Ratings for depression of LC-self and total scores of IDS-C were found to be highly correlated (Pearson coefficient r = -.718; p <.001), whilst the correlation of ratings for mania with YMRS compared to LC-self were slightly less robust (Pearson coefficient r =.491; p =.001). These results were confirmed by good correlations between the CGI-BP IA (mania), IB (depression) and IC (overall mood state) and the LC-self ratings (Pearson coefficient r =.488, r =.721 and r =.65, respectively; all p <.001). CONCLUSIONS: The LC-self shows a significant correlation and good concordance with standard cross sectional affective rating scales, suggesting that the LC-self is a valid and time and money saving alternative to the clinician-rated version which should be incorporated in future clinical research in bipolar disorder. Generalizability of the results is limited by the selection of highly motivated patients in specialized bipolar centres and by the open design of the study.
A growing number of studies in adults document critical relationships between sleep and emotional processing based on responses to affective images from the International Affective Picture System (IAPS; Lang, Bradley, & Cuthbert, 2005). Our aim was to extend examination of the interrelationships between sleep and emotional processing to a sample of healthy girls, ages 10 to 16 years. A total of 86 girls (M = 12.88 years, SD = 1.92) without psychiatric disorders were recruited. In addition to structured diagnostic interviews, report of sleep quality was examined in relation to valence and arousal ratings of pleasant, neutral and unpleasant IAPS images. Overall, picture ratings were consistent with findings from previous research showing pleasant images to produce high arousal and valence ratings in childhood and that these relationships decrease with age. Regression models revealed poor sleep quality to be associated with decreased subjective arousal in response to negative/unpleasant images, but not pleasant or neutral images. Findings are discussed in terms of a need for more research aimed at better elucidating how sleep quality during the childhood years relates to the processing of emotional information.
This study examined the factors that advertising practitioners in Nigeria consider when selecting media for advertising campaigns. The survey research method was applied to collect data with the structured questionnaire as the research instrument. 120 copies of the questionnaire were administered on the respondents who were members of the Media Independents Association of Nigeria (MIPAN) sectoral group of the Nigerian advertising industry. The purposive sampling technique was employed for selecting respondents in order to eliminate waste. 110 copies of the questionnaire were properly completed and returned representing 91.6 percent response rate. Findings revealed that reach was the most important factor that influenced media selection among advertising practitioners in Nigeria as indicated by 40.5 percent of the respondents. This was followed by cost of media as well as available budget for advertising campaigns. Other factors include prestige, image, rating and share of media as well as circulation with respect to the print media. Radio emerged as the most preferred advertising medium as indicated by 37 percent of the respondents. It was followed by newspaper, television and online media respectively. The major conclusion drawn from this study is that traditional media of mass communication are still the preferred advertising media among advertising practitioners in Nigeria. It was recommended that private media owners should increase the reach of their media outlets in order to attract more advertising revenue. It was further recommended that future researchers should investigate how Nigerian advertising practitioners determine some of the variables utilized in media selection. These include reach, prestige, image, and share of advertising media. Furthermore, it was recommended that constant research be conducted to determine the actual reach of media outlets in Nigeria as opposed to unsubstantiated claims by media owners.
The objective of this study was to evaluate the response of the human frontal cortex when listening to cat meows and dog barks, accompanied with a non-verbal pictorial using an affective rating system that assesses the dimensions of valence and arousal. Each participant (24 students; 12 females and 12 males) sat individually in the middle of a room and listened to the sounds (cat meows, dog barks, and the sound of a train) through ceiling-mounted speakers with a near infrared spectroscopic (NIRS) device. Participants had significantly higher oxygenated hemoglobin (oxy-Hb) levels during the exposure to meows (p < 0.05) and barks (p < 0.01) compared with the train sound. A significant correlation was observed between the dimensions of valence and oxy-Hb activation when the participants listened to cat meows (r = 0.53, p < 0.017; Bonferroni correction). In conclusion, we found that the acoustic signals of companion animals lead to frontal cortex responses in humans, suggesting that their signals have an important function related to their long coexistence with people.
In order to examine whether Arabic has Heavy Noun Phrase Shifting (HNPS), I have extracted from the Prague Arabic Dependency Treebank a data set in which a verb governs either an object NP and an Adjunct Phrase (PP or AdvP) or a subject NP and an Adjunct Phrase. I have used binary logistic regression where the criterion variable is whether the subject/object NP shifts, and used as predictor variables heaviness (the number of tokens per NP, adjunct), part of speech tag, verb disposition (ie. whether the verb has a history of taking double objects or sentential objects), NP number, NP definiteness, and the presence of referring pronouns in either the NP or the adjunct. The results show that only object heaviness and adjunct heaviness are useful predictors of object HNPS, while subject heaviness, adjunct heaviness, subject part of speech tag, definiteness, and adjunct head POS tags are active predictors of subject HNPS. I also show that HNPS can in principle be predicted from sentence structure.
Grammatical Change Begins within the Word: Causal Modeling of the Co-evolution of Icelandic Morphology and Syntax Fermin Moscoso del Prado Martin (fmoscoso@linguistics.ucsb.edu) Department of Linguistics, UC – Santa Barbara Santa Barbara, CA 93106 USA Abstract I introduce a combination of information-theoretical and causal modeling to study the cascading of changes between the morphology and the syntax of a language on a diachronic scale. Through the analysis of a historical treebank of Icelandic language ranging from the XII to the XXI century, I show that it is changes in the inflectional morphology of the language that triggered changes in its syntax. This offers a novel and powerful approach to draw conclusions in historical linguistics from a macroscopic perspective. In addition, these findings have implications for the dynamical properties of the linguistic system. Keywords: Granger-causality; Historical Linguistics; Icelandic; Information Theory; Morphological paradigms; Syntactic complexity. Introduction A common observation in the field of Historical Linguistics is that grammatical changes in language tend to be cascaded (e.g., Biberauer & Roberts, 2008; Lightfoot, 2002). Changes at a given level of language (e.g., phonology, morphology, syntax, …) disturb the unstable equilibrium at which human languages reside, triggering a cascade of further linguistic changes as a result. These cascaded changes continue until the system reaches a new meta-stable state (e.g, Croft, 1995; Smith, 1996). For instance, it is widely documented that the loss of the grammatical case markers (morphology) in Late Middle English led to a more rigid word order (syntax) in Early Modern English (cf., Fisiak, 1984). Traditionally, historical linguists have tracked these cascaded changes by looking for the earliest document at which a particular grammatical innovation can be found or – conversely– at the latest time in which a later extinct construction was documented. Historical linguistics often relies on hard dichotomies on the presence or absence of individual words, affixes, or constructions. This approach is clearly useful for documenting the approximate time at which particular constructions appeared or disappeared. The approach is however limited in its power to detect more subtle forms of grammatical change. For instance, as I will argue in this study, Icelandic morphological paradigms are remarkably resilient, having survived with little change for a such a long period, that most of its current system can be directly traced –virtually unchanged– all the way up to the Old West Norse of the XII century. However, despite the striking conservativeness of the Icelandic paradigms, their patterns of usage have changed along this period. One can obtain a higher degree of sensitivity by studying the frequencies (and implicitly the probabilities) of usage of different constructions in diachronic scale. These often reveal gradual changes in the grammar of a language. For instance, Ellegard (1953) documented how the usage of the English periphrastic do construction (e.g., I do not speak vs. I speak not, or Do you speak? vs. Speak you?) arose gradually –rather than abruptly– during a period of two hundred years, from the late XV century to early XVIII century. With the wide availability of diachronic corpora in electronic form, these frequency-based methods have gained much prominence in recent years (see Hilpert & Gries, in press, for a recent survey of quantitative methods in historical linguistics). Admittedly, frequency methods offer an improved sensitivity to the gradualness of grammatical changes along time. Most often, these methods are applied to relate the evolution of the usage of a particular construction in different contexts. The natural tools to achieve this type of inferences are different types of regression analysis. These analyses study the correlations between different factors that might affect the emergence or demise of a construction. For instance, Hilpert (2013) uses such tools to analyze the evolution of (among many others) the patterns of usage of the English future constructions will vs. be going to. Using a logistic regression analysis, he identifies several factors that significantly co-occur with the uses of either construction. One must –however– keep in mind the old adagio: “correlation does not imply causation”. Even in the cases when one finds significant correlation between the frequencies of use and co-occurrence of different linguistic patterns and constructions, it still remains problematic to argue that one pattern causes another. Back to example above, the causal connection between the loss of grammatical case and the fixation of word order is in fact not so trivial on the basis of the historical data alone. Evidently, I cannot infer that I have just had dinner because it stopped raining, even if I observed both events in a sequence. To make such arguments, I would require some form of statistical evidence on how reliably do I start eating whenever it stops raining. Similarly, just the temporal sequentiality between the loss of case marking and the emergence of rigid word order does not –by itself alone– necessarily warrant causality. For instance, Kiparsky (1996) discusses how fixed word order arose also in Icelandic while the case markers were preserved.
‘Micro- and Macrolinguistics’ has the strategic purpose of introducing students to the various linguistic databases held in or under development in the department. ‘Micro- and Macrolinguistics’ is designed to draw on and draw in all structural levels of language — to integrate the various elements of the discipline as a kind of finishing course for students in the third year of a Linguistics major. The ‘learners’ for whom this course in Micro- and Macrolinguistics is designed are investigating linguistics rather than a particular language. Although morphology is the interface between items of the linguistic code and their meaning, the relationship isn’t necessarily one-to-one. Polysemy is the normal condition of words, as a glance at the dictionary would confirm. The E. Brill technique can also be usefully applied to differentiating homographic inflections and clitics. The processing of large amounts of written text requires a slightly different approach than the one taken in conventional linguistics.
With this research and design paper, we are proposing that Open Educational Resources (OERs) and Open Access (OA) publications give increasing access to high quality online educational and research content for the development of powerful domain-specific language collections that can be further enhanced linguistically with the Flexible Language Acquisition System (FLAX, http://flax.nzdl.org). FLAX uses the Greenstone digital library system, which is a widely used open-source software that enables end users to build collections of documents and metadata directly onto the Web (Witten, Bainbridge, & Nichols, 2010). FLAX offers a powerful suite of interactive text-mining tools, using Natural Language Processing and Artificial Intelligence designs, to enable novice collections builders to link selected language content to large pre-processed linguistic databases. An open methodology trialed at Queen Mary University of London in collaboration with the OER Research Hub at the UK Open University demonstrates how applying open corpus-based designs and technologies can enhance open educational practices among language teachers and subject academics for the preparation and delivery of courses in English for Specific Academic Purposes (ESAP).