Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
ABSTRACT Online travel reviews are emerging as a powerful source of information affecting tourists' pre-purchase evaluation of a hotel organization. This trend has highlighted the need for a greater understanding of the impact of online reviews on consumer attitudes and behaviors. In view of this need, we investigate the influence of online hotel reviews on consumers' attributions of service quality and firms' ability to control service delivery. An experimental design was used to examine the effects of four independent variables: framing; valence; ratings; and target. The results suggest that in reviews evaluating a hotel, remarks related to core services are more likely to induce positive service quality attributions. Recent reviews affect customers' attributions of controllability for service delivery, with negative reviews exerting an unfavorable influence on consumers' perceptions. The findings highlight the importance of managing the core service and the need for managers to act promptly in addressing customer service problems.
We present a reformulation of the word pair features typically used for the task of disambiguating implicit relations in the Penn Discourse Treebank. Our word pair features achieve significantly higher performance than the previous formulation when evaluated without additional features. In addition, we present results for a full system using additional features which achieves close to state of the art performance without resorting to gold syntactic parses or to context outside the relation.
People believe that women are more emotionally intense than men, but the scientific evidence is equivocal. In this study, we tested the novel hypothesis that men and women differ in the neural correlates of affective experience, rather than in the intensity of neural activity, with women being more internally (interoceptively) focused and men being more externally (visually) focused. Adult men (n = 17) and women (n = 17) completed a functional magnetic resonance imaging study while viewing affectively potent images and rating their moment-to-moment feelings of subjective arousal. We found that men and women do not differ overall in their intensity of moment-to-moment affective experiences when viewing evocative images, but instead, as predicted, women showed a greater association between the momentary arousal ratings and neural responses in the anterior insula cortex, which represents bodily sensations, whereas men showed stronger correlations between their momentary arousal ratings and neural responses in the visual cortex. Men also showed enhanced functional connectivity between the dorsal anterior insula cortex and the dorsal anterior cingulate cortex, which constitutes the circuitry involved with regulating shifts of attention to the world. These results demonstrate that the same affective experience is realized differently in different people, such that women's feelings are relatively more self-focused, whereas men's feelings are relatively more world-focused.
We present an affective text analysis model that can directly estimate and combine affective ratings of multi-word terms, with application to the problem of sentence polarity/semantic orientation detection. Starting from a hierarchical compositional method for generating sentence ratings, we expand the model by adding multi-word terms that can capture non-compositional semantics. The method operates similarly to a bigram language model, using bigram terms or backing off to unigrams based on a (degree of) compositionality criterion. The affective ratings for n-gram terms of different orders are estimated via a corpus-based method using distributional semantic similarity metrics between unseen words and a set of seed words. N-gram ratings are then combined into sentence ratings via simple algebraic formulas. The proposed framework produces state-of-the-art results for word-level tasks in English and German and the sentence-level news headlines classification SemEval'07-Task14 task. The inclusion of bigram terms to the model provides significant performance improvement, even if no term selection is applied.
With the growing interest in statistical parsing, special attention has recently been devoted to the problem of comparing different treebanks to assess which languages or domains are more difficult to parse relative to a given model. A common methodology for comparing parsing difficulty across treebanks is based on the use of the standard labeled precision and recall measures. As an alternative, in this article we propose an information-theoretic measure, called the expected conditional cross-entropy (ECC). One important advantage with respect to standard performance measures is that ECC can be directly expressed as a function of the parameters of the model. We evaluate ECC across several treebanks for English, French, German, and Italian, and show that ECC is an effective measure of parsing difficulty, with an increase in ECC always accompanied by a degradation in parsing accuracy.
The present paper focuses on ways in which the pragmatic (functional) meaning that arises from various contextual features, known in corpus linguistics as semantic prosody, can become an integral part of lexicographical descriptions as they are represented in the Slovene Lexical Database (SLD). This is particularly important for the treatment of phraseology and idiomatics. First, the theoretical background is provided, with the focus on the prototype theory and its practical implications for monolingual lexicography. A parallel is drawn with the model of meaning analysis in the SLD. The second part begins with a brief introduction to semantic prosody and continues with an analysis of monolingual meaning descriptions in the SLD against a number of authentic corpus examples, investigating how their pragmatic components have been identified. The analysis of corpus data shows that pragmatics is an important contributor to the process of sense discrimination in works of lexical and lexicographic relevance.
We present a novel method ("waste") for the segmentation of text into tokens and sentences. Our approach makes use of a Hidden Markov Model for the detection of segment boundaries. Model parameters can be estimated from pre-segmented text which is widely available in the form of treebanks or aligned multi-lingual corpora. We formally define the waste boundary detection model and evaluate the system's performance on corpora from various languages as well as a small corpus of computer-mediated communication.
peer reviewed
This paper investigates the effect of the label bias problem of maximum entropy Markov models for part-of-speech tagging, a typical sequence prediction task in natural language processing. This problem has been underexploited and underappreciated. The investigation reveals useful information about the entropy of local transition probability distributions of the tagging model which enables us to exploit and quantify the label bias effect of part-of-speech tagging. Experiments on a Vietnamese treebank and on a French treebank show a significant effect of the label bias problem in both of the languages.
We describe a novel approach to detecting empty categories (EC) as represented in de-pendency trees as well as a new metric for measuring EC detection accuracy. The new metric takes into account not only the position and type of an EC, but also the head it is a dependent of in a dependency tree. We also introduce a variety of new features that are more suited for this approach. Tested on a sub-set of the Chinese Treebank, our system im-proved significantly over the best previously reported results even when evaluated with this more stringent metric. 1
This paper sets out to study the letters of Gaston B., a French prisoner of war held in captivity in the camp of Münster (Germany) from the beginning of the First World War until its end. These letters make possible a relativisation of linguistic macrohistory through microhistory, by focussing on the grassroots level and by using as sources the traces of people with no significant name or identity. They shed important light on how a member of a lower class acquired the prescriptive linguistic norm through his schooling at the end of the nineteenth century and how this affected his subsequent linguistic behaviour. An individual is exposed to the political and social dimension of language planning, and his language reflects its level of success, but also reveals what grammatical tools and rules have been focused on during his schooling.
Study of the situation in Lithuanian machine translation revealed, it is not possible to create the statistical machine translation system jet because of the lack of required amount of parallel texts. Thus the status of Lithuanian language is similar to that of the English language in the beginning of the computer – it is no sufficient computer resources. So we are now the best way to go which was the English language 50 years ago – to create rule-based machine translation system. The planed Treebank will serve creation of well working automatic syntactic analysis, which is needed for the rule-based machine translation.
Abstract This chapter presents the methods on which this book is based. Large quantities of data were generated from electronic sources. The most important sources used were the Corpus of Contemporary American English, the British National Corpus, various types of dictionary, websites and lexical databases. The chapter discusses the methodological problems with gathering and analyzing the data. Furthermore the conventions for citing data and for deciding which data to include are detailed. Problems of interpretation are also considered.
This package contains a partition of the Iula Spanish LSP Treebank into train and test sets to perform Machine Learning experiments. In that way the same partitions can be used by different researchers and their results can be directly compared. In this package we also deliver the Tibidabo Treebank (Marimon 2010) which contains a set of sentences extracted from Ancora corpus annotated in the same way than the Iula Treebank. Tibidabo Treebank is a very good test set for models trained with Iula Spanish LSP Treebank since the sentences that form it from a very different domain than those of the Iula Spanish LSP Treebank.
초록·키워드 목차 오류제보하기 本?文??在????的?程中,只有???准答案的束?,?生才能??自己的思???,展示?言表?能力和??。?文中?到的“故?”是在特定的?合下,有意出?的?言技巧,????非??,而是制造矛盾?到引??生思考的一?方式。在?????程中,故??含深意,能?迪?生的思?,?展?生的智力及?言表?能力,强化???容,?而提高?堂???量。 所?“?筋急??”是指?思?遇到特殊的阻碍?,不能按照原定思路思考,需要????路?,??的方面?思考??,?在泛指一些不能用通常的思路?回答的智力?答?。?筋急??巧妙地?用了??中的?音、??、?法以及文字、修?等?言?象。??益智??言游?,不?能?生喜?效果,??言表?含蓄幽默?趣,?能吸引?生激??生的???趣,使???得?松、?生?得愉快。 #汉语教学 #标准答案 #冒犯-故错 #反规范化 #脑筋急转 Ⅰ. 서론Ⅱ. 模範答案과 冒犯答案Ⅲ. 중국어의 ‘標準答案’과 탈규범화Ⅳ. 난센스 문답 속의 언어학Ⅴ. 결론참고문헌中文提要
Understanding, prioritizing and responding to infant affective cues is a key component of motherhood, with long-term implications for infant socio-emotional development. This important task includes identifying unique characteristics of one's own infant, as they relate to differences in affect valence-happy or sad-while monitoring one's own level of arousal. The amygdala has traditionally been understood to respond to affective valence; in the present study, we examined the potential effect of personal relevance on amygdala response, by testing whether mothers' amygdala response to happy and sad infant face cues would be modulated by infant identity. We used functional MRI to measure amygdala activation in 39 first-time mothers, while they viewed happy, neutral and sad infant faces of both their own and a matched unknown infant. Emotional arousal to each face was rated using the Self-Assessment Manikin Scales. Mixed-effects linear regression models were used to examine significant predictors of amygdala response. Overall, both arousal ratings and amygdala activation were greater when mothers viewed their own infant's face compared with unknown infant faces. Sad faces were rated as more arousing than happy faces, regardless of infant identity. However, within the amygdala, a highly significant interaction effect was noted between infant identity and valence. For own-infant faces, amygdala activation was greater for happy than sad faces, whereas the opposite trend was seen for unknown-infant faces. Our findings suggest that the amygdala response to positive or negative valenced cues is modulated by personal relevance. Positive facial expressions from one's own infant may play a particularly important role in eliciting maternal responses and strengthening the mother-infant bond.
Empty categories (EC) are artificial ele-ments in Penn Treebanks motivated by the government-binding (GB) theory to ex-plain certain language phenomena such as pro-drop. ECs are ubiquitous in languages like Chinese, but they are tacitly ignored in most machine translation (MT) work because of their elusive nature. In this paper we present a comprehensive treat-ment of ECs by first recovering them with a structured MaxEnt model with a rich set of syntactic and lexical features, and then incorporating the predicted ECs into a Chinese-to-English machine translation task through multiple approaches, includ-ing the extraction of EC-specific sparse features. We show that the recovered empty categories not only improve the word alignment quality, but also lead to significant improvements in a large-scale state-of-the-art syntactic MT system. 1
Are men from Mars and women from Venus (Gray, 1992) when it comes to marketing evaluations scales? The thought came as one of the authors was compiling a study that compared ratings by two student populations from different cultures. Could the differences in rating questions be due largely to the way these cultures perceive what a 1 to 10 scaling means? If so, what does this mean for other marketing segments like gender? The authors (a male and a female) wanted to explore if gender had any bias for scaling questions. After conducting a review of the literature, they examined several existing research studies to see if ratings are consistently higher for one gender or the other.
The problem of Vietnamese syntactic parsing, especially constituency parsing, has recently been tackled by several research groups. A common effort of the Vietnamese language processing community has allowed the creation of VietTreebank, a reference parsed corpus containing about 10,000 sentences for the constituency parsing task. In this paper, we present our work to build a reference treebank, based on VietTreebank, for the dependency parsing task, which has not yet been very well studied for Vietnamese. First we define a dependency label set by adapting the dependency schema developed by the NLP group at Stanford university and taking into account the particularities of Vietnamese grammar. Then we propose an algorithm to convert a constituency treebank to a dependency one. The algorithm is tested on a set of 100 sentences of VietTreebank corpus and gives very good results. Finally, we carry out an experiment on Vietnamese dependency parsing using MaltParser tool and the dependency treebank converted from VietTreebank.
Previous functional imaging studies have shown key roles of the dorsal anterior insula (dAI) and anterior midcingulate cortex (aMCC) in empathy for the suffering of others. The current study mapped structural covariance networks of these regions and assessed the relationship between networks and individual differences in empathic responding in 94 females. Individual differences in empathy were assessed through average state measures in response to a video task showing others' suffering, and through questionnaire-based trait measures of empathic concern. Overall, covariance patterns indicated that dAI and aMCC are principal hubs within prefrontal, temporolimbic, and midline structural covariance networks. Importantly, participants with high empathy state ratings showed increased covariance of dAI, but not aMCC, to prefrontal and limbic brain regions. This relationship was specific for empathy and could not be explained by individual differences in negative affect ratings. Regarding questionnaire-based empathic trait measures, we observed a similar, albeit weaker modulation of dAI covariance, confirming the robustness of our findings. Our analysis, thus, provides novel evidence for a specific contribution of frontolimbic structural covariance networks to individual differences in social emotions beyond negative affect.
The accuracy of similarity measurement between sentences is critical to the performance of several applications such as text mining, question answering, and text summarization. This paper focuses on calculating semantic similarities between sentences and performing a comparative analysis among identified similarity measurement techniques. Comparison between three popular similarity measurements which are Jaccard, Cosine and Dice similarity measures has been conducted. The performance of each identified measurement was evaluated and recorded. In this paper, we use a large lexical database of English known as WordNet to calculate the word-to-word semantic similarity. The result of this research concludes that the Jaccard and Dice performs better in measuring the semantic similarity between sentences.
Behavioral habituation during repeated exposure to aversive stimuli is an adaptive process. However, the way in which changes in self-reported emotional experience are related to the neural mechanisms supporting habituation remains unclear. We probed these mechanisms by repeatedly presenting negative images to healthy adult participants and recording behavioral and neural responses using functional magnetic resonance imaging. We were particularly interested in investigating patterns of activity in insula, given its significant role in affective integration, and in amygdala, given its association with appraisal of aversive stimuli and its frequent coactivation with insula. We found significant habituation behaviorally along with decreases in amygdala, occipital cortex and ventral prefrontal cortex (PFC) activity with repeated presentation, whereas bilateral posterior insula, dorsolateral PFC and precuneus showed increased activation. Posterior insula activation during image presentation was correlated with greater negative affect ratings for novel presentations of negative images. Further, repeated negative image presentation was associated with increased functional connectivity between left posterior insula and amygdala, and increasing insula-amygdala functional connectivity was correlated with increasing behavioral habituation. These results suggest that habituation is subserved in part by insula-amygdala connectivity and involves a change in the activity of bottom-up affective networks.
We investigated the reappraisal and the time course of negative emotion regulation by performing event-related potential (ERP) recordings. We found that negative pictures elicited more positive P2 and late positivity potential (LPP) deflections than neutral pictures. This effect occurred between 150–2000 ms post-stimulus. Compared to the emotion maintaining condition, the emotion enhancing condition was associated with higher arousal ratings and displayed increased P2 and LPP amplitudes. The decrease condition was also associated with reduced picture-induced arousal; however, it led to increased P2 and LPP amplitudes. Furthermore, when compared with the maintain condition, both the enhancing and decrease conditions significantly enhanced LPP in the early stage (350–750 ms). Compared to previous studies using western subjects, the negative emotion LPP effects of the present study were shorter in duration and the decrease-emotion condition elicited larger LPPs.
This paper tries to give answers for successful receptive multilingualism (RM) but also for its failure. It is mainly based on the results of two projects, one on inter-dialectal communication in the Baltic area during the era of the Hanseatic League and the other analyses inter-Scandinavian communication today. The main purpose of this survey is to outline the essential preconditions for successful RM, from a linguistic, social and environmental perspective. The historical project about communication in the Baltic focuses on long-term language contact based on common mutual trading interests whilst the contemporary project highlights the cultural factors (among others Pan-Scandinavism) as a common basis for using one's own mother tongue in transnational communication. Moreover, other relevant issues belonging to successful RM are touched upon, such as diglossia (i.e. the functional distribution of different languages/varieties in various settings), oral face-to-face communication, the absence of written norms and the non-existence of standardised forms, which result not only in a greater flexibility in communication but also support openness for divergent varieties. Disfavouring factors for RM are, however, taken into consideration as well, such as nationalism and the suppression of minorities (and thus indirectly multilingualism), the enforcement of strict linguistic norms by the society and finally the use of a lingua franca such as Latin in the Middle Ages or English today.
SMS messaging and communicating on social networks are increasingly widespread forms of informal communication. Mobile phones have almost all, and in addition they open profiles on the Internet social network, corresponding in this way with their peers. In writing messages is being recorded a large number of spelling errors, most of errors are those whose adoption is foreseen in the the lower grades of elementary school. In order to determine the level of mastery of linguistic norms, the message will be analysed as well as comments from the social networks of fourth-grade students.
ABSTRACT The aim of this article is to carry out a structural-functional analysis of the formation of Old English adjectives by means of affixation. By analysing the rules and operations that produce the 3,356 adjectives which the lexical database of Old English Nerthus (www.nerthusproject.com) turns out as affixal derivatives, a total of fourteen derivational functions have been identified. Additionally, the analysis yields conclusions concerning the relationship between affixes and derivational functions, the patterns of recategorization present in adjective formation and recursive word-formation.
This paper has two main objectives. The first is to provide an overview of the CDT annotation design with special emphasis on the modeling of the interface between syntactic and morphological structure. Against this background, the second objective is to explain the basic fundamentals of how CDT is marked-up with semantic relations in accordance with the dependency principles governing the annotation on the other levels of CDT. Specifically, focus will be on how Generative Lexicon theory has been incorporated into the unitary theoretical dependency framework of CDT by developing an annotation scheme for lexical semantics which is able to account for the lexico-semantic structure of complex NPs.
There is a substantial body of recent evidence showing ergogenic effects of carbohydrate (CHO) mouth rinsing on endurance performance. However, there is a lack of research on the dose-effect and the aim of this study was to investigate the effect of two different concentrations (6% and 12% weight/volume, w/v) on 90 minute treadmill running performance. Seven active males took part in one familiarization trial and three experimental trials (90-minute self-paced performance trials). Solutions (placebo, 6% or 12% CHO-electrolyte solution, CHO-E) were rinsed in the mouth at the beginning, and at 15, 30 and 45 minutes during the run. The total distance covered was greater during the CHO-E trials (6%, 14.6 ± 1.7 km; 12%, 14.9 ± 1.6 km) compared to the placebo trial (13.9 ± 1.7 km, P < 0.05). There was no significant difference between the 6% and 12% trials (P > 0.05). There were no between trial differences (P > 0.05) in ratings of perceived exertion (RPE) and feeling or arousal ratings suggesting that the same subjective ratings were associated with higher speeds in the CHO-E trials. Enhanced performance in the CHO-E trials was due to higher speeds in the last 30 minutes even though rinses were not provided during the final 45 minutes, suggesting the effects persist for at least 20-45 minutes after rinsing. In conclusion, mouth rinsing with a CHO-E solution enhanced endurance running performance but there does not appear to be a dose-response effect with the higher concentration (12%) compared to a standard 6% solution.
This article presents a new approach of us-ing dependency treebanks in theoretical syn-tactic research: the view of dependency treebanks as combined networks. This al-lows the usage of advanced tools for net-work analysis that quite easily provide novel insight into the syntactic structure of lan-guage. As an example of this approach, we will show how the network approach can provide clear structural distinctions among the Chinese function words, which are very difficult to obtain directly from the original treebank. We hope to illustrate the enor-mous potential of the language network ap-
The perception of emotional cues from voice and face is essential for social interaction. However, this process is altered in various psychiatric conditions along with impaired social functioning. Emotion communication trainings have been demonstrated to improve social interaction in healthy individuals and to reduce emotional communication deficits in psychiatric patients. Here, we investigated the impact of a non-verbal emotion communication training (NECT) on cerebral activation and brain structure in a controlled and combined functional magnetic resonance imaging (fMRI) and voxel-based morphometry study. NECT-specific reductions in brain activity occurred in a distributed set of brain regions including face and voice processing regions as well as emotion processing- and motor-related regions presumably reflecting training-induced familiarization with the evaluation of face/voice stimuli. Training-induced changes in non-verbal emotion sensitivity at the behavioral level and the respective cerebral activation patterns were correlated in the face-selective cortical areas in the posterior superior temporal sulcus and fusiform gyrus for valence ratings and in the temporal pole, lateral prefrontal cortex and midbrain/thalamus for the response times. A NECT-induced increase in gray matter (GM) volume was observed in the fusiform face area. Thus, NECT induces both functional and structural plasticity in the face processing system as well as functional plasticity in the emotion perception and evaluation system. We propose that functional alterations are presumably related to changes in sensory tuning in the decoding of emotional expressions. Taken together, these findings highlight that the present experimental design may serve as a valuable tool to investigate the altered behavioral and neuronal processing of emotional cues in psychiatric disorders as well as the impact of therapeutic interventions on brain function and structure.
Moral decision-making is a key asset for humans' integration in social contexts, and the way we decide about moral issues seems to be strongly influenced by emotions. For example, individuals with deficits in emotional processing tend to deliver more utilitarian choices (accepting an emotionally aversive action in favor of communitarian well-being). However, little is known about the association between emotional experience and moral-related patterns of choice. We investigated whether subjective reactivity to emotional stimuli, in terms of valence, arousal, and dominance, is associated with moral decision-making in 95 healthy adults. They answered to a set of moral and non-moral dilemmas and assessed emotional experience in valence, arousal and dominance dimensions in response to neutral, pleasant, unpleasant non-moral, and unpleasant moral pictures. Results showed significant correlations between less unpleasantness to negative stimuli, more pleasantness to positive stimuli and higher proportion of utilitarian choices. We also found a positive association between higher arousal ratings to negative moral laden pictures and more utilitarian choices. Low dominance was associated with greater perceived difficulty over moral judgment. These behavioral results are in fitting with the proposed role of emotional experience in moral choice.
The linguistic annotation of noun-verb complex predicates (also termed as light verb constructions) is challenging as these predicates are highly productive in Hindi. For semantic role labelling, each argument of the noun-verb complex predicate must be given a role label. For complex predicates, frame files need to be created specifying the role labels for each noun-verb complex predicate. The creation of frame files is usually done manually, but we propose an automatic method to expedite this process. We use two resources for this method: Hindi PropBank frame files for simple verbs and the annotated Hindi Treebank. Our method perfectly predicts 65 % of the roles in 3015 unique noun-verb combinations, with an additional 22 % partial predictions, giving us 87 % useful predictions to build our annotation resource. 1
We present an idiographic approach to modeling dyadic interactions using differential equations. Using data representing daily affect ratings from romantic relationships, we examined several models conceptualizing different types of dyadic interactions. We fitted each model to each of the dyads and the resulting AICc values were used to classify the most likely configuration of interaction for each dyad. Additionally, the AICc from the different models were used in parameter averaging across models. Averaged parameters were used in models involving predictors of relationship dynamics, as indexed by these parameters, as well as models wherein the parameters predicted distal outcomes of the dyads such as relationship satisfaction and status. Results indicated that, within our sample, the most likely interaction style was that of independence, without evidence of emotional interrelations between the two individuals in the couple. Attachment-related avoidance and anxiety showed significant relations with model parameters, such that ideal levels of affect for males were negatively influenced by higher levels of avoidance from their partner while their own levels of anxiety had positive effects on their levels of dyadic coregulation. For females coregulation was negatively influenced by both time in the relationship and their partner's level of avoidance. Analysis involving distal outcomes showed modest influences from the individual's level of ideal affect.
We explored the influence of implicit motives and activity inhibition (AI) on subjectively experienced affect in response to the presentation of six different facial expressions of emotion (FEEs; anger, disgust, fear, happiness, sadness, and surprise) and neutral faces from the NimStim set of facial expressions (Tottenham et al., 2009). Implicit motives and AI were assessed using a Picture Story Exercise (PSE) (Schultheiss et al., 2009b). Ratings of subjectively experienced affect (arousal and valence) were assessed using Self-Assessment Manikins (SAM) (Bradley and Lang, 1994) in a sample of 84 participants. We found that people with either a strong implicit power or achievement motive experienced stronger arousal, while people with a strong affiliation motive experienced less arousal and less pleasurable affect across emotions. Additionally, we obtained significant power motive × AI interactions for arousal ratings in response to FEEs and neutral faces. Participants with a strong power motive and weak AI experienced stronger arousal after the presentation of neutral faces but no additional increase in arousal after the presentation of FEEs. Participants with a strong power motive and strong AI (inhibited power motive) did not feel aroused by neutral faces. However, their arousal increased in response to all FEEs with the exception of happy faces, for which their subjective arousal decreased. These differentiated reaction patterns of individuals with an inhibited power motive suggest that they engage in a more socially adaptive manner of responding to different FEEs. Our findings extend established links between implicit motives and affective processes found at the procedural level to declarative reactions to FEEs. Implications are discussed with respect to dual-process models of motivation and research in motive congruence.
A major computational burden, while performing document clustering, is the calculation of similarity measure between a pair of documents. Similarity measure is a function that assigns a real number between 0 and 1 to a pair of documents, depending upon the degree of similarity between them. A value of zero means that the documents are completely dissimilar whereas a value of one indicates that the documents are practically identical. Traditionally, vector-based models have been used for computing the document similarity. The vector-based models represent several features present in documents. These approaches to similarity measures, in general, cannot account for the semantics of the document. Documents written in human languages contain contexts and the words used to describe these contexts are generally semantically related. Motivated by this fact, many researchers have proposed seman-tic-based similarity measures by utilizing text annotation through external thesauruses like WordNet (a lexical database). In this paper, we define a semantic similarity measure based on documents represented in topic maps. Topic maps are rapidly becoming an industrial standard for knowledge representation with a focus for later search and extraction. The documents are transformed into a topic map based coded knowledge and the similarity between a pair of documents is represented as a correlation between the common patterns (sub-trees). The experimental studies on the text mining datasets reveal that this new similarity measure is more effective as compared to commonly used similarity measures in text clustering.
Tree substitution grammar (TSG) is a generalization of context-free grammar (CFG) that permits non-terminals to rewrite as fragments of arbitrary size, instead of just depth-one productions. We discuss connections between the TSG framework and the larger family of usage-based approaches to language, showing how TSG allows us to make some of the claims of these approaches sufficiently concrete for computational modeling. A fundamental difficulty in defining a TSG is to determine the set of fragments for the grammar, because the set of possible fragments is exponential in the size of the parse trees from which TSGs are typically learned. We describe a model-based approach that learns a TSG using Gibbs sampling with a non-parametric prior to control fragment size, yielding grammars that contain mostly small fragments but that include larger ones. as the data permits. We evaluate these grammars on two tasks (parsing accuracy and grammaticality classification), and find that these Bayesian TSGs achieve excellent performance on two tasks relative to a set of heuristically extracted TSGs spanning the spectrum of representations, from a standard depth-one context-free Treebank grammar to explicit approximations of the Data-Oriented Parsing model.
In a large scale study on 843 transcripts of Technology, Entertainment and Design (TED) talks, the authors address the relation between word usage and categorical affective ratings of lectures by a large group of internet users. Users rated the lectures by assigning one or more predefined tags which relate to the affective state evoked in the audience (e. g., ‘fascinating’, ‘funny’, ‘courageous’, ‘unconvincing’ or ‘long-winded’). By automatic classification experiments, they demonstrate the usefulness of linguistic features for predicting these subjective ratings. Extensive test runs are conducted to assess the influence of the classifier and feature selection, and individual linguistic features are evaluated with respect to their discriminative power. In the result, classification whether the frequency of a given tag is higher than on average can be performed most robustly for tags associated with positive valence, reaching up to 80.7% accuracy on unseen test data.
Today, more than ten years after the resolution of the language controversy on a state level, it is far from resolved in the scholarly or the public sphere, where representatives of the two countries continue to debate the question of how distinct Macedonian is from Bulgarian. This chapter explains why this question is so hotly debated and so politicized. It deals with the process of codification of the contemporary Macedonian linguistic norm and with the conflicts between Bulgarians and Macedonians about the definition both of the Slavic vernacular dialects in geographic Macedonia and of the Macedonian norm itself. The chapter argues that the codification of the contemporary Macedonian idiom cannot be understood without examining a larger international context. Finally, it shows how the codification of a separate Macedonian norm has shaped Bulgarian nationalist representations-especially in the field of linguistics. Keywords:Bulgarian; Macedonian linguistic norm; Slavic vernacular dialects
The aim of the article is to identify the exponent for the semantic prime TOUCH in Old English. Therefore, this research contributes to the frame of the Natural Semantic Metalanguage Research Programme (NSMRP) by applying it to the study of a historical language. Throughout such an application several descriptive and methodological questions arise. On the descriptive side, it is necessary to propose a cluster of semantic, morphological, textual and syntactic criteria that allow for the identification of the prime at stake, given that the nature of the object of study is not compatible with the translation into the native language generally adopted by the NSMRP. The analysis focuses on the category actions, events, movement and contact, and relies on data retrieved from the Historical Thesaurus of the Oxford English Dictionary, the Dictionary of Old English Corpus and the lexical database of Old English Nerthus. Although the cluster of criteria evinces a clear candidate for semantic prime it also raises the methodological issue of the distinction between the semantic prime and the hyperonym because some of the criteria used in the search for the former also play a role in the process of identification of the latter. The conclusion is reached that the verb hrīnan is the main exponent for the semantic prime TOUCH in Old English because it satisfies the criteria of meaning, word-formation, textual frequency and syntactic complementation.
Abstract. Recent research and development have created the necessary ingredients for a major push in web-scale language understanding: large repositories of structured knowledge (DBpedia, the Google knowledge graph, Freebase, YAGO) progress in language processing (parsing, information extraction, computational semantics), linguistic knowledge resources (Treebanks, WordNet, BabelNet, UWN) and new powerful techniques for machine learning. A major goal is the automatic aggregation of knowledge from textual data. A central component of this endeavor is relation extraction (RE). In this paper, we will outline a new approach to connecting repositories of world knowledge with linguistic knowledge (syntactic and lexical semantics) via web-scale relation extraction technologies.
Morphology is the study of internal structure of words and is an essential early step in many NLP applications such as parsing and machine translation. Researchers working in Hindi NLP have either used the widely popular paradigm based analyzer (PBA) or extensions of it. In this work, we undertook a comprehensive evaluation of PBA using the data from the Hindi Treebank (HTB) and presented a new morphological analyzer trained on the HTB. Our morphological analyzer has better coverage and accuracy when compared to the existing analyzers for Hindi. An oracle system that takes the best values from the PBA’s output achieves only 63.41% for lemma, gender, number, person and case. Our statistical analyzer has an accuracy of 84.16% for these morphological attributes when evaluated on the test section of the Hindi Treebank.