Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
This paper focuses on improving a specific opinion spam detection task, deceptive spam. In addition to traditional word form and other shallow syntactic features, we introduce two types of deep level linguistic features. The first type of features are derived from a shallow discourse parser trained on Penn Discourse Treebank (PDTB), which can capture inter-sentence information. The second type is based on the relationship between sentiment analysis and spam detection. The experimental results over the benchmark dataset demonstrate that both of the proposed deep features achieve improved performance over the baseline.
The task of recognizing events from video has attracted a lot of attention in recent years. However, due to the complex nature of user-defined events, the use of purely audio- visual content analysis without domain knowledge has been found to be grossly inadequate. In this paper, we propose to construct a semantic-visual knowledge base to encode the rich event-centric concepts and their relationships from the well- established lexical databases, including FrameNet, as well as the concept-specific visual knowledge from ImageNet. Based on this semantic-visual knowledge bases, we design an effective system for video event recognition. Specifically, in order to narrow the semantic gap between the high-level complex events and low-level visual representations, we utilize the event-centric semantic concepts encoded in the knowledge base as the intermediate-level event representation, which offers both human-perceivable and machine-interpretable semantic clues for event recognition. In addition, in order to leverage the abundant ImageNet images, we propose a robust transfer learning model to learn the noise- resistant concept classifiers for videos. Extensive experiments on various real-world video datasets demonstrate the superiority of our proposed system as compared to the state-of-the-art approaches.
We present a new treebank of English and French technical forum content which has been annotated for grammatical errors and phrase structure. This double annotation allows us to empirically measure the effect of errors on parsing performance. While it is slightly easier to parse the corrected versions of the forum sentences, the errors are not the main factor in making this kind of text hard to parse.
Congenital amusia is a neurodevelopmental disorder characterized by impaired pitch processing. Although pitch simultaneities are among the fundamental building blocks of Western tonal music, affective responses to simultaneities such as isolated dyads varying in consonance/dissonance or chords varying in major/minor quality have rarely been studied in amusic individuals. Thirteen amusics and thirteen matched controls enculturated to Western tonal music provided pleasantness ratings of sine-tone dyads and complex-tone dyads in piano timbre as well as perceived happiness/sadness ratings of sine-tone triads and complex-tone triads in piano timbre. Acoustical analyses of roughness and harmonicity were conducted to determine whether similar acoustic information contributed to these evaluations in amusics and controls. Amusic individuals' pleasantness ratings indicated sensitivity to consonance and dissonance for complex-tone (piano timbre) dyads and, to a lesser degree, sine-tone dyads, whereas controls showed sensitivity when listening to both tone types. Furthermore, amusic individuals showed some sensitivity to the happiness-major association in the complex-tone condition, but not in the sine-tone condition. Controls rated major chords as happier than minor chords in both tone types. Linear regression analyses revealed that affective ratings of dyads and triads by amusic individuals were predicted by roughness but not harmonicity, whereas affective ratings by controls were predicted by both roughness and harmonicity. We discuss affective sensitivity in congenital amusia in view of theories of affective responses to isolated chords in Western listeners.
Neuroimaging while participants listen to audiobooks provides a rich data source for theories of incremental parsing. We compare nested regression models of these data. These mixed-effects models incorporate linguistic predictors at various grain sizes ranging from part-of-speech bigrams, through surprisal on context-free treebank grammars, to incremental node counts in trees that are derived by Minimalist Grammars. The fine-grained structures make an independent contribution over and above coarser predictors. However, this result only obtains with time courses from anterior temporal lobe (aTL). In analogous time courses from inferior frontal gyrus, only n-grams improve upon a non-syntactic baseline. These results support the idea that aTL does combinatoric processing during naturalistic story comprehension, processing that bears a systematic relationship to linguistic structure.
The enrichment of Arabic treebank with syntactic properties provides the increase of its use in different applications, the acquisition of new linguistic resources and the alleviation of the probabilistic parsing process by using statistics to limit the properties to satisfied ones. This method of enrichment requires two steps to follow starting by inducting a Property Grammar from a source treebank and generating finally the new syntactic property-based representation.
The scent of blood is potentially one of the most fundamental and survival-relevant olfactory cues in humans. This experiment tests the first human parameters of perceptual threshold and emotional ratings in men and women of an artificially simulated smell of fresh blood in contact with the skin. We hypothesize that this scent of blood, with its association with injury, danger, death, and nutrition will be a critical cue activating fundamental motivational systems relating to either predatory approach behavior or prey-like withdrawal behavior, or both. The results show that perceptual thresholds are unimodally distributed for both sexes, with women being more sensitive. Furthermore, both women and men's emotional responses to simulated blood scent divide strongly into positive and negative valence ratings, with negative ratings in women having a strong arousal component. For women, this split is related to the phase of their menstrual cycle and oral contraception (OC). Future research will investigate whether this split in both genders is context-dependent or trait-like.
This research explores the cultural and linguistic strategies of immigrant youth to negotiate inclusion/exclusion, including language discrimination in Vancouver, Canada. My theoretical framework draws upon the Arendtian notions of ‘public space’, and ‘action and speech’ as well as Bourdieu’s concepts of ‘symbolic violence’ and ‘habitus’. My methodology is a critical qualitative approach. Fourteen immigrant youth, aged 15–25, were involved in this research. The findings of this study indicate that unlike second-generation immigrants, first-generation immigrant youth face cultural and linguistic challenges. Non-recognition of youths’ distinct linguistic and social capitals, the imposition of official languages and the regulation of the education and language market according to the dominant linguistic norms include forms of discrimination against Turkish minority youth in Canada. Taken together, the findings suggest that immigrant youths’ cultural and linguistic experiences of inclusion and exclusion cannot be dissociated from the wider politics of the nation-state, popular hegemony and social inequalities in the host society.
We present the first dynamic programming (DP) algorithm for shift-reduce constituency parsing, which extends the DP idea of Huang and Sagae (2010) to context-free grammars. To alleviate the propagation of errors from part-of-speech tagging, we also extend the parser to take a tag lattice instead of a fixed tag sequence. Experiments on both English and Chinese treebanks show that our DP parser significantly improves parsing quality over non-DP baselines, and achieves the best accuracies among empirical linear-time parsers.
Our aim in the present study was to investigate the psychological mechanisms that underlie the disinhibiting effects of alcohol cues in social drinkers by contrasting motor and oculomotor inhibition after exposure to alcohol-related, emotional, and neutral pictures. We conducted 2 studies in which social drinkers completed modified stop-signal (laboratory) and antisaccade (online) tasks in which positive, negative, alcohol-related, and neutral pictures were embedded. We measured cue-specific disinhibition in each task, and investigated whether sex and drinking status moderated the effects of pictures on disinhibition. Across both studies, comparable increases in disinhibition were observed in response to both alcohol and negatively valenced pictures, relative to both positive and neutral pictures. These differences in disinhibition could not be explained by differences between picture sets in arousal or valence ratings. There was no clear evidence of moderation by sex or drinking status. Secondary analyses demonstrated that alcohol-specific disinhibition was not reliably associated with individual differences in alcohol consumption or craving. These results suggest that the disinhibiting properties of alcohol-related cues cannot be attributed solely to their valence or arousing properties, and that alcohol cues may have unique disinhibiting properties.
Vietnamese Treebank is a syntactically annotated corpus newly published in 2009. In this paper, we applied automated methods to detect errors in Vietnammese Treebank based on the concept of equivalence classes proposed by Dickinson. On this basis, we propose an improved method of error detection by transforming syntax trees based on vertical markovization. Our experimental results on Vietnamese Treebank showed that the scope of error detection was extended more than 2 times and the precision was improved more than 18.07% in comparison with the base line methods.
<p>We define a dynamic oracle for the Covington non-projective dependency parser. This is not only the first dynamic oracle that supports arbitrary non-projectivity, but also considerably more efficient (O(n)) than the only existing oracle with restricted non-projectivity support. Experiments show that training with the dynamic oracle significantly improves parsing accuracy over the static oracle baseline on a wide range of treebanks.</p>
Direct content analysis reveals important details about movies including those of gender representations and potential biases. We investigate the differences between male and female character depictions in movies, based on patterns of language used. Specifically, we use an automatically generated lexicon of linguistic norms characterizing gender ladenness. We use multivariate analysis to investigate gender depictions and correlate them with elements of movie production. The proposed metric differentiates between male and female utterances and exhibits some interesting interactions with movie genres and the screenplay writer gender.
We study non-deterministic oracles for training non-projective beam search parsers with swap transitions. We map out the spurious ambiguities of the transition system and present two non-deterministic oracles as well as a static oracle that minimizes the number of swaps. An evaluation on 10 treebanks reveals that the difference between static and non-deterministic oracles is generally insignificant for beam search parsers but that non-deterministic oracles can improve the accuracy of greedy parsers that use swap transitions.
Previous studies have shown that emotional facial expressions capture visual attention. However, it has been unclear whether attentional modulation is attributable to their emotional significance or to their visual features. We investigated this issue using a spatial cueing paradigm in which non-predictive cues were peripherally presented before the target was presented in either the same (valid trial) or the opposite (invalid trial) location. The target was an open dot and the cues were photographs of normal emotional facial expressions of anger and happiness, their anti-expressions and neutral expressions. Anti-expressions contained the amount of visual changes equivalent to normal emotional expressions compared with neutral expressions, but they were usually perceived as emotionally neutral. The participants were asked to localize the target as soon as possible. After the cueing task, they evaluated their subjective emotional experiences to the cue stimuli. Compared with anti-expressions, the normal emotional expressions decreased and increased the reaction times (RTs) in the valid and invalid trials, respectively. Shorter RTs in the valid trials and longer RTs in the invalid trials were related to higher subjective arousal ratings. These results suggest that emotional facial expressions accelerate attentional engagement and prolong attentional disengagement due to their emotional significance.
INTRODUCTION: Emotional behavioral disturbances are hallmarks of many dementias but their pathophysiology is poorly understood. Here we addressed this issue using the paradigm of emotionally salient sounds. METHODS: Pupil responses and affective valence ratings for nonverbal sounds of varying emotional salience were assessed in patients with behavioral variant frontotemporal dementia (bvFTD) (n = 14), semantic dementia (SD) (n = 10), progressive nonfluent aphasia (PNFA) (n = 12), and AD (n = 10) versus healthy age-matched individuals (n = 26). RESULTS: Referenced to healthy individuals, overall autonomic reactivity to sound was normal in Alzheimer's disease (AD) but reduced in other syndromes. Patients with bvFTD, SD, and AD showed altered coupling between pupillary and affective behavioral responses to emotionally salient sounds. DISCUSSION: Emotional sounds are a useful model system for analyzing how dementias affect the processing of salient environmental signals, with implications for defining pathophysiological mechanisms and novel biomarker development.
Both animal studies and studies using deep brain stimulation in humans have demonstrated the involvement of the subthalamic nucleus (STN) in motivational and emotional processes; however, participation of this nucleus in processing human emotion has not been investigated directly at the single-neuron level. We analyzed the relationship between the neuronal firing from intraoperative microrecordings from the STN during affective picture presentation in patients with Parkinson's disease (PD) and the affective ratings of emotional valence and arousal performed subsequently. We observed that 17% of neurons responded to emotional valence and arousal of visual stimuli according to individual ratings. The activity of some neurons was related to emotional valence, whereas different neurons responded to arousal. In addition, 14% of neurons responded to visual stimuli. Our results suggest the existence of neurons involved in processing or transmission of visual and emotional information in the human STN, and provide evidence of separate processing of the affective dimensions of valence and arousal at the level of single neurons as well.
This release adds the <em>Peregrinatio Aetheriae</em> to the treebank, adds further material from the <em>Gallic War</em> and <em>Letters to Atticus</em> and includes numerous corrections across the treebank.
In conversation, speakers tend to echo the linguistic style of the person they are interacting with. This paper contributes to a body of work that addresses how this linguistic style coordination is affected by the social context in which the interaction occurs. In particular, we investigate the effect that an agent's social network centrality has on the coordination exhibited in replies to their utterances. We find that linguistic coordination is positively correlated with social network centrality and that this effect is greater than previous results showing a similar connection between statusbased power and linguistic coordination. We conjecture that the social value of coordination may reside in the wish to conform to the linguistic norms of a community.
This release contains numerous corrections, in particular to the morphology. It also corrects a mistake in the previous release, which failed to alter the placement of PRO tokens in the <em>Peregrinatio Aetheriae</em> (see release 20150615).
Official releases of the PROIEL treebank of ancient Indo-European languages
Il saggio si propone di analizzare la narrativa di Nelida Milani, scrittrice e linguista contemporanea istriana (Croazia), che nei suoi lavori alterna volutamente codici diversi: italiano standard, croato, dialetto istro- veneto, dando origine ad un vero e proprio miscuglio linguistico e mettendo in discussione la norma linguistica. Le frequenti commutazioni di codice, l'uso di registri e codici diversi nei suoi Racconti di guerra sono una prova importante e tangibile dell'identita plurilinguistica e pluriculturale della penisola Istriana. E proprio grazie alla lingua e allo sconvolgimeto della norma che la Milani riesce a trasmettere al meglio la tensione e la drammaticita dei sui racconti e a dipingere la realta del territorio e della storia che racconta. Oggetto della nostra analisi, di approccio pragmalinguistico, saranno proprio i momenti in cui la scrittrice alterna i diversi codici.
Abstract. The purpose of the experiment was to test the relationship between attributes of color, self-rated arousal, and autonomic reactions to color stimuli. Sixteen colored backgrounds of different hue, saturation, and brightness were each viewed by 64 subjects (females, M age = 22.48) while skin conductance responses (SCRs) were recorded. Subjective judgments relating to pleasantness (valence) and arousal were also measured. Results show that among color attributes only saturation had an effect on SCR magnitude, F(1, 63) = 6.31, p <.05, η G 2 =.01. There was also significant correlation, r(14) =.64, p <.01, between aggregated SCR magnitude and arousal ratings. It confirms that SCR could be used as a marker of phasic arousal even in response to the abstract, devoid of content stimuli. Saturation seems to be the main property connected with color’s ability to elicit orienting response. More saturated stimuli are better in capturing attention regardless of hue, thus suggesting that at the first stage of color perception, color intensity is more important than qualitative properties. Such results clarify some incoherent findings known from previous studies on psychophysiological responses to color stimuli.
Building Parallel Treebanks for Underresourced Languages: a Georgian-Ukrainian Treebank ProposalWe present outcomes of an undertaking on building a parallel Treebank for the Georgian and the Ukrainian languages, which is a “side product” of the GRUG multilingual Treebank project. The GRUG acronym stands for the German- Russian-Ukrainian-Georgian Treebank intended for contrastive studies and translation memory systems. The monolingual Ukrainian and Georgian parallel sentences were syntactically annotated manually using the Synpathy tool. Tagsets for both languages follow an adapted version of the German TIGER guidelines with necessary changes relevant for the Georgian and the Ukrainian grammar formal description. An output of the monolingual syntactic annotation is in the TIGER-XML format. Alignment of monolingual resources into a bilingual Georgian-Ukrainian Treebank was done by the Stockholm TreeAligner software.
Speakers respond more slowly when naming pictures presented with taboo (i.e., offensive/embarrassing) than with neutral distractor words in the picture-word interference paradigm. Over four experiments, we attempted to localize the processing stage at which this effect occurs during word production and determine whether it reflects the socially offensive/embarrassing nature of the stimuli. Experiment 1 demonstrated taboo interference at early stimulus onset asynchronies of -150 ms and 0 ms although not at 150 ms. In Experiment 2, taboo distractors sharing initial phonemes with target picture names eliminated the interference effect. Using additive factors logic, Experiment 3 demonstrated that taboo interference and phonological facilitation effects do not interact, indicating that the two effects originate at different processing levels within the speech production system. In Experiment 4, interference was observed for masked taboo distractors, including those sharing initial phonemes with the target picture names, indicating that the effect cannot be attributed to a processing level involving responses in an output buffer. In two of the four experiments, the magnitude of the interference effect correlated significantly with arousal ratings of the taboo words. However, no significant correlations were found for either offensiveness or valence ratings. These findings are consistent with a locus for the taboo interference effect prior to the processing stage responsible for word form encoding. We propose a pre-lexical account in which taboo distractors capture attention at the expense of target picture processing due to their high arousal levels.
Space-delimited words in Turkish and Hebrew text can be further segmented into meaningful units, but syntactic and semantic context is necessary to predict segmentation. At the same time, predicting correct syntactic structures relies on correct segmentation. We present a graph-based lattice dependency parser that operates on morphological lattices to represent different segmentations and morphological analyses for a given input sentence. The lattice parser predicts a dependency tree over a path in the lattice and thus solves the joint task of segmentation, morphological analysis, and syntactic parsing. We conduct experiments on the Turkish and the Hebrew treebank and show that the joint model outperforms three state-of-the-art pipeline systems on both data sets. Our work corroborates findings from constituency lattice parsing for Hebrew and presents the first results for full lattice parsing on Turkish.
The article deals with the problem of language policy and planning (LPP) in terms of corpus. The research is focused on linguistic norm and standard language in Russia and the USA. The author gives historical background of developing standard language in the RF and the USA and describes its stages based on the role of LPP actors forming linguistic norm. The work provides description of the two LPP types characteristic of Russia and the United States.
The most renowned strategy utilized for perusing mind movement is electroencephalography (EEG). Electroencephalography is the neurophysiologic estimation of the electrical action of the cerebrum by recording from anodes put on the scalp, or in the exceptional cases on the cortex. The ensuing follows are known as an electroencephalogram (EEG) and speak to alleged brainwaves. This system is picking up prevalence as it is a non-intrusive interface and is giving a methodology to controlling machines through contemplations. The proposed linking and familiarity rating method classifies the music, video assessment responses of EEG-Signal. The metrics namely true positive, true negative, false positive, false negative, sensitivity, specificity and classification accuracy are chosen for evaluating the performance of the proposed classifier. The simulation result shows that the proposed classifier achieves 95.4 % accuracy which is better than other methods.
Abstract Using the framework of Language Management Theory (LMT), this article seeks to analyze the ways in which non-native speakers negotiate their position in English-language online discussion forums. Based on the material collected from four discussion forums, competing opinions have been identified regarding the acceptability of “bad English” and the need for language management, i.e. acting upon a perceived lack of compliance to linguistic norms. Some users propose that compliance to communication norms should be enforced in a top-down manner or based on an explicit set of rules, whereas others hold that the community of users can deal with potential communication problems individually in an emergent manner. While the applicability of native speaker norms to the discussion forums is being questioned, non-native speakers, especially in practically oriented forums, tend to perform pre-interaction language management, using disclaimers in their posts, such as “Excuse my poor English”, to avoid potential misunderstandings and to prevent native speaker norms from being applied to them. The article argues for the use of LMT in computer mediated communication research, as it offers a dynamic view of the process in which rules, conventions and norms of online communication are being continuously discussed, negotiated and applied.
We investigate the relation between general affective meaning and the use of particular phonological segments in poems, presenting a novel quantitative measure to assess the basic affective tone of a text based on foregrounded phonological units and their iconic affective properties. The novel method is applied to the volume of German poems “verteidigung der wölfe” (defense of the wolves) by Hans Magnus Enzensberger, who categorized these 57 poems as friendly, sad, or spiteful. Our approach examines the relation between the phonological inventory of the texts to both the author’s affective categorization and readers’ perception of the poems—assessed by a survey study. Categorical comparisons of basic affective tone reveal significant differences between the 3 groups of poems in accordance with the labels given by the author as well as with the affective rating scores given by readers. Using multiple regression, we show our sublexical measures of basic affective tone to account for a considerable part of variance (9.5%−20%) of ratings on different emotion scales. We interpret this finding as evidence that the iconic properties of foregrounded phonological units contribute significantly to the poems’ emotional perception—potentially reflecting an intentional use of phonology by the author. Our approach represents a first independent statistical quantification of the basic affective tone of texts. (PsycINFO Database Record (c) 2015 APA, all rights reserved)
BACKGROUND: Parsing, which generates a syntactic structure of a sentence (a parse tree), is a critical component of natural language processing (NLP) research in any domain including medicine. Although parsers developed in the general English domain, such as the Stanford parser, have been applied to clinical text, there are no formal evaluations and comparisons of their performance in the medical domain. METHODS: In this study, we investigated the performance of three state-of-the-art parsers: the Stanford parser, the Bikel parser, and the Charniak parser, using following two datasets: (1) A Treebank containing 1,100 sentences that were randomly selected from progress notes used in the 2010 i2b2 NLP challenge and manually annotated according to a Penn Treebank based guideline; and (2) the MiPACQ Treebank, which is developed based on pathology notes and clinical notes, containing 13,091 sentences. We conducted three experiments on both datasets. First, we measured the performance of the three state-of-the-art parsers on the clinical Treebanks with their default settings. Then we re-trained the parsers using the clinical Treebanks and evaluated their performance using the 10-fold cross validation method. Finally we re-trained the parsers by combining the clinical Treebanks with the Penn Treebank. RESULTS: Our results showed that the original parsers achieved lower performance in clinical text (Bracketing F-measure in the range of 66.6%-70.3%) compared to general English text. After retraining on the clinical Treebank, all parsers achieved better performance, with the best performance from the Stanford parser that reached the highest Bracketing F-measure of 73.68% on progress notes and 83.72% on the MiPACQ corpus using 10-fold cross validation. When the combined clinical Treebanks and Penn Treebank was used, of the three parsers, the Charniak parser achieved the highest Bracketing F-measure of 73.53% on progress notes and the Stanford parser reached the highest F-measure of 84.15% on the MiPACQ corpus. CONCLUSIONS: Our study demonstrates that re-training using clinical Treebanks is critical for improving general English parsers' performance on clinical text, and combining clinical and open domain corpora might achieve optimal performance for parsing clinical text.