Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Abstract The language of a speech community can only act as an identity marker for all of its speakers if linguistic norms are widely shared and if a minimal number of language varieties are spoken. This article examines briefly how a linguistic norm came to serve the whole of Iceland and how a situation of relative linguistic homogeneity was maintained for centuries. Sociolinguistic theory tells us that the speech community that we can reconstruct for early Iceland should lead to the establishment and maintenance of local norms. However, Iceland, arguably monodialectal, was certainly characterized by long-term linguistic homogeneity and remained a society where nucleated settlements barely formed over a thousand-year period. Scholars have argued that a mixture of dialects leveled shortly after the settlement of Iceland in the ninth century (Settlement). Studies show that dialect leveling requires dialect mixing, the convergence of people on one place, and sustained linguistic contact between the speakers. The settlement pattern of Iceland is indicative of population divergence (not convergence) and there is limited evidence of sustained contact. It is therefore proposed that the dialect leveling might be linked instead with significant population movements and social upheaval in mainland Scandinavia in the immediate pre-Viking period. The variety of Norse that was taken westward across the Atlantic might itself already have been the result of several earlier stages of mixing and koineization. It is only by combining linguistic, historical, and archaeological knowledge that this problem of how one linguistic norm came to serve the whole of Iceland can be understood.
Anaphora resolution is one of the most difficult tasks in NLP. The ability to identify non-referential pronouns before attempting an anaphora resolution task would be significant, since the system would not have to attempt resolving such pronouns and hence end up with fewer errors. In addition, the number of non-referential pronouns has been found to be non-trivial in many domains. The task of detecting non-referential pronouns could also be incorporated into a part-of-speech tagger or a parser, or treated as an initial step in semantic interpretation. In this article, I describe a machine learning method for identifying non-referential pronouns in an annotated subsegment of the Penn Arabic Treebank using three different feature settings. I achieve an accuracy of 97.22% with 52 different features extracted from a small window size of -5/+5 tokens surrounding each potentially non-referential pronoun.
Eighteenth-century language usage is markedly under-represented in the first two editions of the OED, whose quotations for this period were gathered almost entirely during the late nineteenth and early twentieth centuries. This article reviews some of the possible causes, characteristics and consequences of OED’s gap in eighteenth-century documentation and shows that female authors were particularly scanted. The role of quotations in the OED, as the evidential basis for the dictionary, is briefly considered, along with eighteenth-century (and Victorian/Edwardian) views on women and language, and the availability of female-authored texts for quotation by the lexicographers. The article reports sample reading in eighteenth-century female writers (especially Jean Adam, Penelope Aubin and Anna Seward), which shows that OED could easily have supplied its eighteenth-century deficiency from such authors, and that it often favoured distinctive usages in female-authored texts—innovative, eccentric or domestic vocabulary—rather than usage which exemplified linguistic norms (especially in poetry, where Seward’s case is examined). It also discusses revisions to the OED so far conducted in the third (ongoing) edition, and their implications for readers and editors of eighteenth-century texts.
Abstract The aim of computational semantics is to capture the meaning of natural language expressions in representations suitable for performing inferences, in the service of understanding human language in written or spoken form. First‐order logic is a good starting point, both from the representation and inference point of view. But even if one makes the choice of first‐order logic as representation language, this is not enough: the computational semanticist needs to make further decisions on how to model events, tense, modal contexts, anaphora and plural entities. Semantic representations are usually built on top of a syntactic analysis, using unification, techniques from the lambda‐calculus or linear logic, to do the book‐keeping of variable naming. Inference has many potential applications in computational semantics. One way to implement inference is using algorithms from automated deduction dedicated to first‐order logic, such as theorem proving and model building. Theorem proving can help in finding contradictions or checking for new information. Finite model building can be seen as a complementary inference task to theorem proving, and it often makes sense to use both procedures in parallel. The models produced by model generators for texts not only show that the text is contradiction‐free; they also can be used for disambiguation tasks and linking interpretation with the real world. To make interesting inferences, often additional background knowledge is required (not expressed in the analysed text or speech parts). This can be derived (and turned into first‐order logic) from raw text, semi‐structured databases or large‐scale lexical databases such as WordNet. Promising future research directions of computational semantics are investigating alternative representation and inference methods (using weaker variants of first‐order logic, reasoning with defaults), and developing evaluation methods measuring the semantic adequacy of systems and formalisms.
This chapter presents a meta-search approach, meant to deliver bibliography from the internet, according to trainees’ results obtained at an e-assessment task. The bibliography consists of web pages related to the knowledge gaps of the trainees. The meta-search engine is part of an education recommender system, attached to an e-assessment application for project management knowledge. Meta-search means that, for a specific query (or mistake made by the trainee), several search mechanisms for suitable bibliography (further reading) could be applied. The lists of results delivered by the standard search mechanisms are used to build thematically homogenous groups using an ontology-based clustering algorithm. The clustering process uses an educational ontology and WordNet lexical database to create its categories. The research is presented in the context of recommender systems and their various applications to the education domain.
A method for deriving an approximately labeled dependency treebank from the Thai Categorial Grammar Treebank has been implemented. The method involves a lexical dictionary for assigning dependency directions to the CG types associated with the grammatical entities in the CG bank, falling back on a generic mapping of CG types in case of unknown words. Currently, all but a handful of the trees in the Thai CG bank can unambiguously be transformed into directed dependency trees. Dependency labels can optionally be assigned with a learned classifier, which in a preliminary evaluation with a very small training set achieves 76.5% label accuracy. In the process, a number of annotation errors in the CG bank were identified and corrected. Although rather limited in its coverage, excluding e.g. long-distance dependencies, topicalisations and longer sentences, the resulting treebank is believed to be sound in terms of structural annotational consistency and a valuable complement to the scarce Thai language resources in existence.
A method for deriving an approximately labeled dependency treebank from the Thai Categorial Grammar Treebank has been implemented. The method involves a lexical dictionary for assigning dependency directions to the CG types associated with the grammatical entities in the CG bank, falling back on a generic mapping of CG types in case of unknown words. Currently, all but a handful of the trees in the Thai CG bank can unambiguously be transformed into directed dependency trees. Dependency labels can optionally be assigned with a learned classifier, which in a preliminary evaluation with a very small training set achieves 76.5% label accuracy. In the process, a number of annotation errors in the CG bank were identified and corrected. Although rather limited in its coverage, excluding e.g. long-distance dependencies, topicalisations and longer sentences, the resulting treebank is believed to be sound in terms of structural annotational consistency and a valuable complement to the scarce Thai language resources in existence.
Named Entity Recognition (NER) is an important first step for BioNLP tasks, e.g., gene normalization and event extraction. Employing supervised machine learning techniques for achieving high performance recent NER systems require a manually annotated corpus in which every mention of the desired semantic types in a text is annotated. However, great amounts of human effort is necessary to build and maintain an annotated corpus. This study explores a method to build a high-performance NER without a manually annotated corpus, but using a comprehensible lexical database that stores numerous expressions of semantic types and with huge amount of unannotated texts. We underscore the effectiveness of our approach by comparing the performance of NERs trained on an automatically acquired training data and on a manually annotated corpus. 1
Deep semantic parsing is the key to understand sentence meaning. This paper integrates some Chinese semantic relation systems given by different scholars, and presents a more comprehensive system for semantic dependency parsing. The new semantic relation system includes the definition for the situation that a verb acts as a modifier and a verbal noun acts as the center of the noun phrase. According to the relation system, a large scale Chinese semantic dependency relation tree bank is constructed by the combination of automatic and manual means. This semantic dependency tree bank will become a basis of studying deep semantic parsing.
We present a novel approach to Data-Oriented Parsing (DOP). Like other DOP models, our parser utilizes syntactic fragments of arbitrary size from a treebank to analyze new sentences, but, crucially, it uses only those which are encountered at least twice. This criterion al-lows us to work with a relatively small but representative set of fragments, which can be employed as the symbolic backbone of sev-eral probabilistic generative models. For pars-ing we define a transform-backtransform ap-proach that allows us to use standard PCFG technology, making our results easily replica-ble. According to standard Parseval metrics, our best model is on par with many state-of-the-art parsers, while offering some comple-mentary benefits: a simple generative proba-bility model, and an explicit representation of the larger units of grammar. 1
BACKGROUND: Traumatized individuals and particularly post-traumatic stress disorder (PTSD) patients are characterized by memory disturbances that suggest altered memory control. The present study investigated the issue using an item method, directed forgetting (DF) paradigm in 51 civil war victims in Uganda. All participants had been exposed to severe traumatic stress and 26 additionally suffered from PTSD. METHOD: In an item cued, DF paradigm photographs were presented, each followed by an instruction to either remember or forget it. A recognition test for all initially presented photographs and thematically similar distracters followed. DF patterns were compared between the non-PTSD and the PTSD groups. Post-experimental ratings of picture valence and arousal were collected and correlated with DF. RESULTS: Results revealed DF, that is, reduced recognition for 'to-be-forgotten' items in the non-PTSD but not in the PTSD group. Moreover, in the non-PTSD, but not in the PTSD group, false alarms were reduced for 'to-be-remembered' items. Finally, DF was reduced in those participants who rated the pictures as more arousing, the PTSD group giving, on average, higher arousal ratings. CONCLUSIONS: Data indicate that DF is reduced in PTSD and that the reduction is related to stimulus arousal. Furthermore, individuals with PTSD are characterized by a more global encoding style than individuals without PTSD, reflected in a higher false alarm rate. In sum, traumatized individuals with (but not without) PTSD are impaired in their ability to selectively control episodic memory encoding. This impairment may contribute to clinical features of the disorder such as intrusions and flashbacks.
Two experiments examined the effects of repetition on listeners' emotional response to music. Listeners heard recordings of orchestral music that contained a large section repeated twice. The music had a symmetric phrase structure (same-length phrases) in Experiment 1 and an asymmetric phrase structure (different-length phrases) in Experiment 2, hypothesized to alter the predictability of sensitivity to musical repetition. Continuous measures of arousal and valence were compared across music that contained identical repetition, variation (related), or contrasting (unrelated) structure. Listeners' emotional arousal ratings differed most for contrasting music, moderately for variations, and least for repeating musical segments. A computational model for the detection of repeated musical segments was applied to the listeners' emotional responses. The model detected the locations of phrase boundaries from the emotional responses better than from performed tempo or physical intensity in both experiments. These findings indicate the importance of repetition in listeners' emotional response to music and in the perceptual segmentation of musical structure.
We propose a relaxed correspondence assumption for cross-lingual projection of constituent syntax, which allows a supposed constituent of the target sentence to correspond to an unrestricted treelet in the source parse. Such a relaxed assumption fundamentally tolerates the syntactic non-isomorphism between languages, and enables us to learn the target-language-specific syntactic idiosyncrasy rather than a strained grammar directly projected from the source language syntax. Based on this assumption, a novel constituency projection method is also proposed in order to induce a projected constituent treebank from the source-parsed bilingual corpus. Experiments show that, the parser trained on the projected treebank dramatically outperforms previous projected and unsupervised parsers. 1
Historical linguistics aims at inferring the most likely language phylogenetic tree starting from information concerning the evolutionary relatedness of languages. The available information are typically lists of homologous (lexical, phonological, syntactic) features or characters for many different languages: a set of parallel corpora whose compilation represents a paramount achievement in linguistics. From this perspective the reconstruction of language trees is an example of inverse problems: starting from present, incomplete and often noisy, information, one aims at inferring the most likely past evolutionary history. A fundamental issue in inverse problems is the evaluation of the inference made. A standard way of dealing with this question is to generate data with artificial models in order to have full access to the evolutionary process one is going to infer. This procedure presents an intrinsic limitation: when dealing with real data sets, one typically does not know which model of evolution is the most suitable for them. A possible way out is to compare algorithmic inference with expert classifications. This is the point of view we take here by conducting a thorough survey of the accuracy of reconstruction methods as compared with the Ethnologue expert classifications. We focus in particular on state-of-the-art distance-based methods for phylogeny reconstruction using worldwide linguistic databases. In order to assess the accuracy of the inferred trees we introduce and characterize two generalizations of standard definitions of distances between trees. Based on these scores we quantify the relative performances of the distance-based algorithms considered. Further we quantify how the completeness and the coverage of the available databases affect the accuracy of the reconstruction. Finally we draw some conclusions about where the accuracy of the reconstructions in historical linguistics stands and about the leading directions to improve it.
Parallel treebanks with annotation of syntax, discourse, coreference, morphology, and semantics. Version 3 also includes the Danish Dependency Treebank (version 1) and the Danish-English Parallel Dependency Treebank (version 2).
We investigate how morphological features in the form of part-of-speech tags impact parsing performance, using Arabic as our test case. The large, fine-grained tagset of the Penn Arabic Treebank (498 tags) is difficult to handle by parsers, ultimately due to data sparsity. However, ad-hoc conflations of treebank tags runs the risk of discarding potentially useful parsing information. The main contribution of this paper is to describe several automated, language-independent methods that search for the optimal feature combination to help parsing. We first identify 15 individual features from the Penn Arabic Treebank tagset. Either including or excluding these features results in 32,768 combinations, so we then apply heuristic techniques to identify the combination achieving the highest parsing performance. Our results show a statistically significant improvement of 2.86 % for vocalized text and 1.88 % for unvocalized text, compared with the baseline provided by the Bikel-Bies Arabic POS mapping (and an improvement of 2.14 % using product models for vocalized text, 1.65 % for unvocalized text), giving state-of-the-art results for Arabic constituency parsing. 1
Based on research linking depressive symptoms and intimate partner aggression perpetration with negatively biased perception of social stimuli, the present authors examined biased perception of emotional expressions as a mechanism in the frequently observed relationship between depression and psychological aggression perpetration. In all, 30 university students made valence ratings (negative to positive) of emotional facial expressions and completed measures of depressive symptoms and psychological aggression perpetration. As expected, depressive symptoms were positively associated with psychological aggression perpetration in an individual's current relationship, and this relationship was mediated by ratings of negative emotional expressions. These findings suggest that negatively biased perception of emotional expressions within the context of elevated depressive symptoms may represent an early stage of information processing that leads to aggressive relationship behaviors.
BACKGROUND: Identification of discourse relations, such as causal and contrastive relations, between situations mentioned in text is an important task for biomedical text-mining. A biomedical text corpus annotated with discourse relations would be very useful for developing and evaluating methods for biomedical discourse processing. However, little effort has been made to develop such an annotated resource. RESULTS: We have developed the Biomedical Discourse Relation Bank (BioDRB), in which we have annotated explicit and implicit discourse relations in 24 open-access full-text biomedical articles from the GENIA corpus. Guidelines for the annotation were adapted from the Penn Discourse TreeBank (PDTB), which has discourse relations annotated over open-domain news articles. We introduced new conventions and modifications to the sense classification. We report reliable inter-annotator agreement of over 80% for all sub-tasks. Experiments for identifying the sense of explicit discourse connectives show the connective itself as a highly reliable indicator for coarse sense classification (accuracy 90.9% and F1 score 0.89). These results are comparable to results obtained with the same classifier on the PDTB data. With more refined sense classification, there is degradation in performance (accuracy 69.2% and F1 score 0.28), mainly due to sparsity in the data. The size of the corpus was found to be sufficient for identifying the sense of explicit connectives, with classifier performance stabilizing at about 1900 training instances. Finally, the classifier performs poorly when trained on PDTB and tested on BioDRB (accuracy 54.5% and F1 score 0.57). CONCLUSION: Our work shows that discourse relations can be reliably annotated in biomedical text. Coarse sense disambiguation of explicit connectives can be done with high reliability by using just the connective as a feature, but more refined sense classification requires either richer features or more annotated data. The poor performance of a classifier trained in the open domain and tested in the biomedical domain suggests significant differences in the semantic usage of connectives across these domains, and provides robust evidence for a biomedical sublanguage for discourse and the need to develop a specialized biomedical discourse annotated corpus. The results of our cross-domain experiments are consistent with related work on identifying connectives in BioDRB.
As the basis of the syntax analysis,BaseNP recognition is an important step in English machine translation.A method based on the maximum entropy model for BaseNP recognition is proposed in this paper.Firstly,this method uses English phrase structure characteristic and the context of the position to establish feature set,then uses frequency and average mutual information to extract effective features,which is expressed as the maximum entropy model,and finally recognition is carried out based on he maximum entropy principle.Simulation experiment is carried out based on Penn Treebank data,the accuracy and recall rate of this method are more than 90%,far higher than the traditional method,so the method is a simple,quick and efficient recognition method in English BaseNP.
We introduce dependency parsing schemata, a formal framework based on Sikkel's parsing schemata for constituency parsers, which can be used to describe, analyze, and compare dependency parsing algorithms. We use this framework to describe several well-known projective and non-projective dependency parsers, build correctness proofs, and establish formal relationships between them. We then use the framework to define new polynomial-time parsing algorithms for various mildly non-projective dependency formalisms, including well-nested structures with their gap degree bounded by a constant k in time O(n 5+2k ), and a new class that includes all gap degree k structures present in several natural language treebanks (which we call mildly ill-nested structures for gap degree k) in time O(n 4+3k ). Finally, we illustrate how the parsing schema framework can be applied to Link Grammar, a dependency-related formalism.
No annotation guidelines concerning substandard Latin are presently available. This paper describes an annotation style of substandard Latin that supplements the method designed for standard Latin by the Perseus Latin Dependency Treebank and the Index Thomisticus Treebank. Each word of the corpus can be assigned only one morphological analysis. In our system, the analysis can be either functional or formal. Functional analysis is applied when a form is language-evolutionarily deducible from the corresponding standard Latin form used in the same (semantico-)syntactic function (e.g. solidus pro solidos 'gold coins' as a direct object: analysis "accusative"). Formal analysis applies when no connection to the functionally required classical form exists (e.g. heredibus pro heredes 'heirs' as a subject: analysis "ablative" or "dative"). When running queries on the corpus, the formally analysed forms can be isolated, and percentages of standard and substandard forms can be counted. In addition, further principles concerning syntax and specific morphological issues are introduced.
This paper gives two contributions to depen-dency parsing in Korean. First, we build a Ko-rean dependency Treebank from an existing constituent Treebank. For a morphologically rich language like Korean, dependency pars-ing shows some advantages over constituent parsing. Since there is not much training data available, we automatically generate depen-dency trees by applying head-percolation rules and heuristics to the constituent trees. Second, we show how to extract useful features for dependency parsing from rich morphology in Korean. Once we build the dependency Tree-bank, any statistical parsing approach can be applied. The challenging part is how to ex-tract features from tokens consisting of multi-ple morphemes. We suggest a way of select-ing important morphemes and use only these as features to avoid sparsity. Our parsing ap-proach is evaluated on three different genres using both gold-standard and automatic mor-phological analysis. We also test the impact of fine vs. coarse-grained morphologies on de-pendency parsing. With automatic morpho-logical analysis, we achieve labeled attach-ment scores of 80%+. To the best of our knowledge, this is the first time that Korean dependency parsing has been evaluated on la-beled edges with such a large variety of data. 1
Pain catastrophizing is associated with enhanced temporal summation of pain (TS-Pain). However, because prior studies have found that pain catastrophizing is not associated with a measure of spinal nociception (nociceptive flexion reflex [NFR] threshold), this association may not result from changes in spinal nociceptive processes. The goal of the present study in healthy participants was to examine the relationship between trait (traditional) and state (situation-specific) pain catastrophizing and temporal summation of NFR (TS-NFR) and TS-Pain. A secondary goal was to replicate prior findings concerning relationships between catastrophizing and NFR threshold, electrocutaneous pain threshold, and sensory and affective ratings of electrocutaneous stimuli. All analyses controlled for depression symptoms, pain-related anxiety, and participant sex. As expected, multiple regression analyses indicated that neither trait nor situation-specific catastrophizing was associated with NFR threshold, but that situation-specific catastrophizing was associated with pain ratings. Multilevel linear growth models of TS data indicated that situation-specific catastrophizing was associated with TS-Pain but not TS-NFR. Trait catastrophizing was not related to TS-Pain or TS-NFR. Together, these results confirm prior studies that indicate that catastrophizing enhances pain via supraspinal processes rather than spinal processes. Moreover, because catastrophizing was associated with TS-Pain but not TS-NFR, caution is warranted when using pain ratings to infer temporal summation of spinal nociceptive processes.
We develop an open-source large-scale finitestate morphological processing toolkit (AraComLex) for Modern Standard Arabic (MSA) distributed under the GPLv3 license. The morphological transducer is based on a lexical database specifically constructed for this purpose. In contrast to previous resources, the database is tuned to MSA, eliminating lexical entries no longer attested in contemporary use. The database is built using a corpus of 1,089,111,204 words, a pre-annotation tool, machine learning techniques, and knowledge-based pattern matching to automatically acquire lexical knowledge. Our morphological transducer is evaluated and compared to LDC's SAMA (Standard Arabic Morphological Analyser).
m, 3abid_khan1961@y ahoo.com Abstract-- This paper is about the development of Pashto Treebank in the form of Extensible Markup Language (XML) code. A Chart Parser has been developed that uses Chart Parsing Algorithm (1) for building parse trees for Pashto sentences. The output of the parser is the parsed text which can be obtained in one of its three forms such as reduced graph, parse tree and XML code. For parsing, the parser needs Context Free Grammar (CFG) of Pashto language and Tagged Input Text as input. The system has been tested on real world text taken from Pashto novels and web sites and tagged manually. Eighty seven (87) sentences were parsed by the parser in which fifty four (54) were correctly parsed with a single parse tree and the rest 33 were parsed with multiple trees and thus the accuracy obtained is 62.06%.
We examined how individual differences in mood and anxiety in the early postpartum period are related to brain response to infant stimuli during fMRI, with particular focus on regions implicated in both maternal behavior and mood/anxiety, that is, the subgenual anterior cingulate cortex (sgACC) and the amygdala. At approximately 3 months postpartum, 22 mothers completed an affect-rating task (ART) during fMRI, where their affective response to infant stimuli was explicitly probed. Mothers viewed/rated four infant face conditions: own positive (OP), own negative (ON), unfamiliar positive (UP), and unfamiliar negative (UN). Mood and anxiety were measured by the Edinburgh Postnatal Depression Scale (EDPS) and the State-Trait Anxiety Inventory-Trait Version (STAI-T); maternal factors related to parental stress and attachment were also assessed. Brain-imaging data underwent a random-effects analysis, and cluster-based statistical thresholding was applied to the following contrasts: OP-UP, ON-UN, OP-ON, and UP-UN. Our main finding was that poorer quality of maternal experience was significantly related to reduced amygdala response to OP compared to UP infant faces. Our results suggest that, in human mothers, infant-related amygdala function may be an important factor in maternal anxiety/mood, in quality of mothering, and in individual differences in the motivation to mother. We are very grateful to the staff at the Imaging Research Center of the Brain-Body Institute for their contributions to this project. This work was supported by an Ontario Mental Health Foundation operating grant awarded to Alison Fleming and a postdoctoral fellowship awarded to Jennifer Barrett.
Abstract This article presents an approach to automatic language classification by means of linguistic networks. Networks of 11 languages were constructed from dependency treebanks, and the topology of these networks serves as input to the classification algorithm. The results match the genealogical similarities of these languages. In addition, we test two alternative approaches to automatic language classification – one based on n-grams and the other on quantitative typological indices. All three methods show good results in identifying genealogical groups. Beyond genetic similarities, network features (and feature combinations) offer a new source of typological information about languages. This information can contribute to a better understanding of the interplay of single linguistic phenomena observed in language.
This paper proposes a direct parsing of non-local dependencies in English. To this end, we use probabilistic linear context-free rewriting systems for data-driven parsing, following recent work on parsing German. In order to do so, we first perform a transformation of the Penn Treebank annotation of non-local dependencies into an annotation using crossing branches. The resulting treebank can be used for PLCFRS-based parsing. Our evaluation shows that, compared to PCFG parsing with the same techniques, PLCFRS parsing yields slightly better results. In particular when evaluating only the parsing results concerning long-distance dependencies, the PLCFRS approach with discontinuous constituents is able to recognize about 88% of the dependencies of type *T* and *T*-PRN encoded in the Penn Treebank. Even the evaluation results concerning local dependencies, which can in principle be captured by a PCFG-based model, are better with our PLCFRS model. This demonstrates that by discarding information on non-local dependencies the PCFG model loses important information on syntactic dependencies in general.
People show autonomic responses when they empathize with the suffering of another person. However, little is known about how these autonomic changes are related to prosocial behavior. We measured skin conductance responses (SCRs) and affect ratings in participants while either receiving painful stimulation themselves, or observing pain being inflicted on another person. In a later session, they could prevent the infliction of pain in the other by choosing to endure pain themselves. Our results show that the strength of empathy-related vicarious skin conductance responses predicts later costly helping. Moreover, the higher the match between SCR magnitudes during the observation of pain in others and SCR magnitude during self pain, the more likely a person is to engage in costly helping. We conclude that prosocial motivation is fostered by the strength of the vicarious autonomic response as well as its match with first-hand autonomic experience.
Dependency parsers are critical components within many NLP systems. However, currently available dependency parsers each exhibit at least one of several weaknesses, including high running time, limited accuracy, vague dependency labels, and lack of nonprojectivity support. Furthermore, no commonly used parser provides additional shallow semantic interpretation, such as preposition sense disambiguation and noun compound interpretation. In this paper, we present a new dependency-tree conversion of the Penn Treebank along with its associated fine-grain dependency labels and a fast, accurate parser trained on it. We explain how a non-projective extension to shift-reduce parsing can be incorporated into non-directional easy-first parsing. The parser performs well when evaluated on the standard test section of the Penn Treebank, outperforming several popular open source dependency parsers; it is, to the best of our knowledge, the first dependency parser capable of parsing more than 75 sentences per second at over 93 % accuracy. 1
This paper presents a simple yet effective semi-supervised method to improve Chinese word segmentation and POS tagging. We introduce novel features derived from large auto-analyzed data to enhance a simple pipelined system. The auto-analyzed data are generated from unlabeled data by using a baseline system. We evaluate the usefulness of our approach in a series of experiments on Penn Chinese Treebanks and show that the new features provide substantial performance gains in all experiments. Furthermore, the results of our proposed method are superior to the best reported results in the literature. 1
BACKGROUND: Major depressive disorder (MDD) is associated with deficits in recalling specific autobiographical memories (AMs). Extensive research has examined the functional anatomical correlates of AM in healthy humans, but no studies have examined the neurophysiological underpinnings of AM deficits in MDD. The goal of the present study was to examine the differences in the hemodynamic response between patients with MDD and controls while they engage in AM recall. METHOD: Participants (12 unmedicated MDD patients; 14 controls) underwent functional magnetic resonance imaging (fMRI) scanning while recalling AMs in response to positive, negative and neutral cue words. The hemodynamic response during memory recall versus performing subtraction problems was compared between MDD patients and controls. Additionally, a parametric linear analysis examined which regions correlated with increasing arousal ratings. RESULTS: Behavioral results showed that relative to controls, the patients with MDD had fewer specific (p=0.013), positive (p=0.030), highly arousing (p=0.036) and recent (p=0.020) AMs, and more categorical (p<0.001) AMs. The blood oxygen level-dependent (BOLD) response in the parahippocampus and hippocampus was higher for memory recall versus subtraction in controls and lower in those with MDD. Activity in the anterior insula was lower for specific AM recall versus subtraction, with the magnitude of the decrement greater in MDD patients. Activity in the anterior cingulate cortex was positively correlated with arousal ratings in controls but not in patients with MDD. CONCLUSIONS: We replicated previous findings of fewer specific and more categorical AMs in patients with MDD versus controls. We found differential activity in medial temporal and prefrontal lobe structures involved in AM retrieval between MDD patients and controls as they engaged in AM recall. These neurophysiological deficits may underlie AM recall impairments seen in MDD.
Facial expressions frequently involve multiple individual facial actions. How do facial actions combine to create emotionally meaningful expressions? Infants produce positive and negative facial expressions at a range of intensities. It may be that a given facial action can index the intensity of both positive (smiles) and negative (cry-face) expressions. Objective, automated measurements of facial action intensity were paired with continuous ratings of emotional valence to investigate this possibility. Degree of eye constriction (the Duchenne marker) and mouth opening were each uniquely associated with smile intensity and, independently, with cry-face intensity. In addition, degree of eye constriction and mouth opening were each unique predictors of emotion valence ratings. Eye constriction and mouth opening index the intensity of both positive and negative infant facial expressions, suggesting parsimony in the early communication of emotion.
How does expertise influence the perception of representational and abstract paintings? We asked 20 experts on art history and 20 laypersons to explore and evaluate a series of paintings ranging in style from representational to abstract in five categories. We compared subjective esthetic judgments and emotional evaluations, gaze patterns, and electrodermal reactivity between the two groups of participants. The level of abstraction affected esthetic judgments and emotional valence ratings of the laypersons but had no effect on the opinions of the experts: the laypersons' esthetic and emotional ratings were highest for representational paintings and lowest for abstract paintings, whereas the opinions of the experts were independent of the abstraction level. The gaze patterns of both groups changed as the level of abstraction increased: the number of fixations and the length of the scanpaths increased while the duration of the fixations decreased. The viewing strategies - reflected in the target, location, and path of the fixations - however indicated that experts and laypersons paid attention to different aspects of the paintings. The electrodermal reactivity did not vary according to the level of abstraction in either group but expertise was reflected in weaker responses, compared with laypersons, to information received about the paintings.
The possibility to analyse vast amounts of linguistic data has brought about changes both in methodology as well as in the ways we perceive certain language phenomena. A key insight gained by computational methods in language analysis is undoubtedly the importance of lexical co-occurrence and usage patterns for the description of lexical meaning. Corpus analysis and new methods in the analysis of pragmatic components of meaning have also yielded significant results in areas such as the treatment of semantic prosody. The present paper does not focus on what is traditionally subsumed under connotation or the speaker’s attitude (e.g., swear words, pejorative and offensive language, praise, excuses, requests, demands, etc.), but on ways in which the pragmatic (functional) meaning that arises from various contextual features can become an integral part of lexicographic descriptions. This is important for the treatment of all of those lexical items whose meanings reside in their function rather than in their bare lexical-semantic meaning, as this is particularly the case with phraseology and idiomatics. From another perspective, pragmatics turns out to be an effective means of sense discrimination in works of lexical and lexicographic relevance, as will be shown in the continuation.
In this paper we present a user-centered approach for defining the dependency syntactic specification for a treebank. We show that by collecting information on syntactic interpretations from the future users of the treebank, we can model so far dependency-syntactically undefined syntactic structures in a way that corresponds to the users ’ intuition. By consulting the users at the grammar definition phase we aim at better usability of the treebank in the future. We focus on two complex syntactic phenomena: elliptical comparative clauses and participial NPs or NPs with a verbderived noun as their head. We show how the phenomena can be interpreted in several ways and ask for the users ’ intuitive way of modeling them. The results aid in constructing the syntactic specification for the treebank. 1
We describe a method for training a semantic role labeler for CCG in the absence of gold-standard syntax derivations. Traditionally, semantic role labeling is performed by placing human-annotated semantic roles on gold-standard syntactic parses, identifying patterns in the syntaxsemantics relationship, and then predicting roles on novel syntactic analyses. The gold standard syntactic training data can be eliminated from the process by extracting training instances from semantic roles projected onto a packed parse chart. This process can be used to rapidly develop NLP tools for resource-poor languages of interest. 1
In this paper, we first analyze and classify the empty categories in a Hindi dependency tree-bank and then identify various discovery procedures to automatically detect the existence of these categories in a sentence. For this we make use of lexical knowledge along with the parsed output from a constraint based parser. Through this work we show that it is possible to successfully discover certain types of empty categories while some other types are more difficult to identify. This work leads to the state-of-the-art system for automatic insertion of empty categories in the Hindi sentence.
We outline a method of detecting ad hoc, or anomalous, rules in treebank grammars, by exploiting the fact that such rules do not fit with the rest of the grammar. Ad hoc rules are rules used for specific constructions in one data set and unlikely to be used again. These include ungeneralizable rules, erroneous rules, rules for ungrammatical text, and rules which are not consistent with the rest of the annotation scheme. Based on the idea that valid rules should receive support from other rules in the grammar, we develop two methods for detecting ad hoc rules in flat treebanks and show they are successful in detecting such rules. Although one can put some linguistic knowledge into determining rule similarity and dissimilarity, the methods work best by using a simple, modified Levenshtein distance. We illustrate this on the English Wall Street Journal treebank and the German TIGER treebank. For the latter, we extend the method to formalisms incorporating discontinuous constituents, employing CFG-like rules for the comparisons.