Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Sundanese language is the second biggest local language used in Indonesia. Currently, Sundanese language is rarely used since we have the Indonesian language in everyday conversation and as the national language. We built a Sundanese lexical database based on WordNet and Indonesian WordNet as an alternative way to preserve the language as one of local culture. WordNet was chosen because of Sundanese language has three levels of word delivery, called language code of conduct. Web user participant involved in this research for specifying Sundanese semantic relations, and an expert linguistic for validating the relations. The merge methodology was implemented in this experiment. Some words are equivalent with WordNet while another does not have its equivalence since some words are not exist in another culture.
In geographic information science, semantic relatedness is important for Geographic Information Retrieval (GIR), Linked Geospatial Data, geoparsing, and geo-semantics. But computing the semantic similarity/relatedness of geographic terminology is still an urgent issue to tackle. The thesaurus is a ubiquitous and sophisticated knowledge representation tool existing in various domains. In this article, we combined the generic lexical database (WordNet or HowNet) with the Thesaurus for Geographic Science and proposed a thesaurus–lexical relatedness measure (TLRM) to compute the semantic relatedness of geographic terminology. This measure quantified the relationship between terminologies, interlinked the discrete term trees by using the generic lexical database, and realized the semantic relatedness computation of any two terminologies in the thesaurus. The TLRM was evaluated on a new relatedness baseline, namely, the Geo-Terminology Relatedness Dataset (GTRD) which was built by us, and the TLRM obtained a relatively high cognitive plausibility. Finally, we applied the TLRM on a geospatial data sharing portal to support data retrieval. The application results of the 30 most frequently used queries of the portal demonstrated that using TLRM could improve the recall of geospatial data retrieval in most situations and rank the retrieval results by the matching scores between the query of users and the geospatial dataset.
In view of the differences between the annotations of micro and macro discourse rela-tionships, this paper describes the relevant experiments on the construction of the Macro Chinese Discourse Treebank (MCDTB), a higher-level Chinese discourse corpus. Fol-lowing RST (Rhetorical Structure Theory), we annotate the macro discourse information, including discourse structure, nuclearity and relationship, and the additional discourse information, including topic sentences, lead and abstract, to make the macro discourse annotation more objective and accurate. Finally, we annotated 720 articles with a Kappa value greater than 0.6. Preliminary experiments on this corpus verify the computability of MCDTB.
This paper presents the Coptic Universal Dependency Treebank, the first dependency treebank within the Egyptian subfamily of the Afro-Asiatic languages. We discuss the composition of the corpus, challenges in adapting the UD annotation scheme to existing conventions for annotating Coptic, and evaluate inter-annotator agreement on UD annotation for the language. Some specific constructions are taken as a starting point for discussing several more general UD annotation guidelines, in particular for appositions, ambiguous passivization, incorporation and object-doubling.
This article presents the LIA treebank of transcribed spoken Norwegian dialects. It consists of dialect recordings made in the period between 1950--1990, which have been digitised, transcribed, and subsequently annotated with morphological and dependency-style syntactic analysis as part of the LIA (Language Infrastructure made Accessible) project at the University of Oslo. In this article, we describe the LIA material of dialect recordings and its transcription, transliteration and further morphosyntactic annotation. We focus in particular on the extension of the native NDT annotation scheme to spoken language phenomena, such as pauses and various types of disfluencies, and present the subsequent conversion of the treebank to the Universal Dependencies scheme. The treebank currently consists of 13,608 tokens, distributed over 1396 segments taken from three different dialects of spoken Norwegian. The LIA treebank annotation is an on-going effort and future releases will extend on the current data set.
The article analyzes the actual concept of linguistic expertise of translation and determines its place in linguistic expertology. The substantiation of the difference between the expert opinion and the usual translation quality assessment is a fundamentally new approach to this problem. In connection with this, modern concepts of equivalence, types and methods of determining translation errors in correlation with the objectives of the expert opinion are touched upon. An attempt is made to systematize linguistic (interlingual correspondence, directed equivalence, linguistic norms, etc.), communicative-psychological (communicative intention, translation creativity, emotive evaluation system, recipient’s response), information (text depth, static and dynamic information, etc.) and logical (opinion, judgment, statement) tools of linguistic expertise of translation. Preliminary conclusions based on the conducted research can contribute to increasing the effectiveness of forensic linguistic expertise of translation, and also prove useful in the process of preparing translators for expert activity.
espanolEste articulo se centra en la lexicografia del ingles antiguo y el analisis de corpus. El objetivo es definir un procedimiento de lematizacion para un tipo de corpus del ingles antiguo anotado y parseado conocido como treebank. Este estudio se centra en dos cuestiones, concretamente en indicar donde se encuentran los datos con los que se puede lematizar el treebank del ingles antiguo; y que procedimiento debe adoptarse para enlazar la lematizacion disponible en las fuentes con el treebank. A partir de las bases de conocimiento del Proyecto Nerthus, se disena, pone en practica y evalua un procedimiento semiautomatico para dotar The York-Toronto-Helsinki Parsed Corpus of Old English Prose de etiquetas de lemas. EnglishThis article deals with Old English lexicography and corpus analysis. It aims at devising a lemmatisation procedure for a type of annotated and parsed corpus of Old English known as treebank. This study addresses two questions, namely where to find the data with which an Old English treebank can be lemmatised; and what procedure should be adopted to link the lemmatisation available from the sources to the treebank. On the grounds of the set of knowledge bases compiled by the Nerthus Project, a semi-automatic procedure for annotating The York-Toronto-Helsinki Parsed Corpus of Old English Prose with lemma tags is devised, illustrated and assessed.
This paper (1) presents the first partially manually verified treebank for Dutch CHILDES corpora, the AnnCor CHILDES Treebank; (2) argues explicitly that it is useful to assign adult grammar syntactic structures to utterances of children who are still in the process of acquiring the language; (3) argues that human annotation and automatic checks on this annotation must go hand in hand; (4) argues that explicit annotation guidelines and conventions must be developed and adhered to and emphasises consistency of the annotations as an important desirable property for annotations. It also describes the tools used for annotation and automated checks on edited syntactic structures, as well as extensions to an existing treebank query application (GrETEL) and the multiple formats in which the resources will be made available
Research on food experience is typically challenged by the way questions are worded. We therefore developed the EmojiGrid: a graphical (language-independent) intuitive self-report tool to measure food-related valence and arousal. In a first experiment participants rated the valence and the arousing quality of 60 food images, using either the EmojiGrid or two independent visual analog scales (VAS). The valence ratings obtained with both tools strongly agree. However, the arousal ratings only agree for pleasant food items, but not for unpleasant ones. Furthermore, the results obtained with the EmojiGrid show the typical universal U-shaped relation between the mean valence and arousal that is commonly observed for a wide range of (visual, auditory, tactile, olfactory) affective stimuli, while the VAS tool yields a positive linear association between valence and arousal. We hypothesized that this disagreement reflects a lack of proper understanding of the arousal concept in the VAS condition. In a second experiment we attempted to clarify the arousal concept by asking participants to rate the valence and intensity of the taste associated with the perceived food items. After this adjustment the VAS and EmojiGrid yielded similar valence and arousal ratings (both showing the universal U-shaped relation between the valence and arousal). A comparison with the results from the first experiment showed that VAS arousal ratings strongly depended on the actual wording used, while EmojiGrid ratings were not affected by the framing of the associated question. This suggests that the EmojiGrid is largely self-explaining and intuitive. To test this hypothesis, we performed a third experiment in which participants rated food images using the EmojiGrid without an associated question, and we compared the results to those of the first two experiments. The EmojiGrid ratings obtained in all three experiments closely agree. We conclude that the EmojiGrid appears to be a valid and intuitive affective self-report tool that does not rely on written instructions and that can efficiently be used to measure food-related emotions.
The emotional valence of target information has been a centerpiece of recent false memory research, but in most experiments, it has been confounded with emotional arousal. We sought to clarify the results of such research by identifying a shared mathematical relation between valence and arousal ratings in commonly administered normed materials. That relation was then used to (a) decide whether arousal as well as valence influences false memory when they are confounded and to (b) determine whether semantic properties that are known to affect false memory covary with valence and arousal ratings. In Study 1, we identified a quadratic relation between valence and arousal ratings of words and pictures that has 2 key properties: Arousal increases more rapidly as function of negative valence than positive valence, and hence, a given level of negative valence is more arousing than the same level of positive valence. This quadratic function predicts that if arousal as well as valence affects false memory when they are confounded, false memory data must have certain fine-grained properties. In Study 2, those properties were absent from norming data for the Cornell-Cortland Emotional Word Lists, indicating that valence but not arousal affects false memory in those norms. In Study 3, we tested fuzzy-trace theory's explanation of that pattern: that valence ratings are positively related to semantic properties that are known to increase false memory, but arousal ratings are not. (PsycINFO Database Record (c) 2019 APA, all rights reserved).
Participants manually explored 47 solid, fluid, and granular materials and rated them according to a list of sensory and affective attributes. In principal component analyses (PCA) of sensory ratings, we extracted six dimensions: Fluidity, Roughness, Deformability, Fibrousness, Heaviness, and Granularity. PCAs on affective ratings revealed Valence, Arousal, and Dominance. PCAs explained 87 percent of variance or more. We found sensory dimensions beyond the surface characteristics on which many previous studies had focused, and the affective dimension of Dominance which previously had not been reported-probably due to our wide range of materials. Experiment 1 investigated a single sample, Experiment 2 distinguished between participants with more versus less outdoor experience during childhood. High correlations between scores of the two groups suggested that group differences were small. Across different experiments and groups greater Arousal was associated with more Fluidity, greater Dominance with increasing Heaviness and decreasing Deformability, and greater Granularity with more positive Valence. Participants with more outdoor experience associated fluid materials with unpleasant feelings, whereas participants with less outdoor experience rated rough materials as being unpleasant. Overall, we demonstrate that the range of affective responses to touched material is broader than previously assumed, and suggests systematic associations between specific affective and sensory dimensions.
The recognition of emotional facial expressions is a central aspect for an effective interpersonal communication. This study aims to investigate whether changes occur in emotion recognition ability and in the affective reactions (self-assessed by participants through valence and arousal ratings) associated with the viewing of basic facial expressions during preadolescence (n = 396, 206 girls, aged 11–14 years, Mage = 12.73, DS =.91). Our results confirmed that happiness is the best recognised emotion during preadolescence. However, a significant decrease in recognition accuracy across age emerged for fear expressions. Moreover, participants’ affective reactions elicited by the vision of happy facial expressions resulted to be the most pleasant and arousing compared to the other emotional expressions. On the contrary, the viewing of sadness was associated with the most negative affective reactions. Our results also revealed a developmental change in participants’ affective reactions to the stimuli. Implications are discussed by taking into account the role of emotion recognition as one of the main factors involved in emotional development.
We report on a pilot study involving emotion elicitation in virtual reality (VR) and assessment of emotional responses with a consumer-grade EEG device. The stimulation used HTC Vive VR system showing pictures from NAPS database within a specifically designed virtual environment. The stimulation consisted of two distinct sequences with 10 pictures of happiness and 10 pictures of fear. Each picture was contained in a separate virtual room that the participants traveled through along a preset path. The estimation employed EMOTIV EPOC+ 14-channel EEG headset and a custom-developed application. The software wirelessly received EEG signals from alpha, beta low, beta high, gamma and theta bands, time-stamped them and dynamically stored in a relational database for subsequent analysis. Our preliminary results show that statistically significant correlations between valence and arousal ratings of pictures and EEG bands are present but highly personalized. Simultaneous correct placement of VR and EEG headsets is demanding and precise localization of electrodes is difficult. In fact, if emotion estimation is not strictly necessary we recommend using devices with fewer electrodes. Nevertheless, we found the EEG to be effective. By acknowledging its limitations, and using the headset in the correct context, experiments involving emotions may be significantly amended.
In this paper we introduce SzegedKoref, a Hungarian corpus in which coreference relations are manually annotated. For annotation, we selected some texts of Szeged Treebank, the biggest treebank of Hungarian with manual annotation at several linguistic layers. The corpus contains approximately 55,000 tokens and 4000 sentences. Due to its size, the corpus can be exploited in training and testing machine learning based coreference resolution systems, which we would like to implement in the near future. We present the annotated texts, we describe the annotated categories of anaphoric relations, we report on the annotation process and we offer several examples of each annotated category. Two linguistic phenomena – phonologically empty pronouns and pronouns referring to subordinate clauses – are important characteristics of Hungarian coreference relations. In our paper, we also discuss both of them.
In dimensional affect recognition, the machine learning methods, which are used to model and predict affect, are mostly classification and regression. However, the annotation in the dimensional affect space usually takes the form of a continuous real value which has an ordinal property. The aforementioned methods do not focus on taking advantage of this important information. Therefore, we propose an affective rating ranking framework for affect recognition based on face images in the valence and arousal dimensional space. Our approach can appropriately use the ordinal information among affective ratings which are generated by discretizing continuous annotations. Specifically, we first train a series of basic cost-sensitive binary classifiers, each of which uses all samples relabeled according to the comparison results between corresponding ratings and a given rank of a binary classifier. We obtain the final affective ratings by aggregating the outputs of binary classifiers. By comparing the experimental results with the baseline and deep learning based classification and regression methods on the benchmarking database of the AVEC 2015 Challenge and the selected subset of SEMAINE database, we find that our ordinal ranking method is effective in both arousal and valence dimensions.
International Journal of Exercise Science 11(5): 609-624, 2018. An aversion to the sensations of physical exertion can deter engagement in physical activity. This is due in part to an associative focus in which individuals are attending to uncomfortable interoceptive cues. The purpose of this study was to test the effect of mindfulness on affective valence, ratings of perceived exertion (RPE), and enjoyment during treadmill walking. Participants (N=23; Mage=19.26, SD = 1.14) were only included in the study if they engaged in no more than moderate levels of physical activity and reported low levels of intrinsic motivation. They completed three testing sessions including a habituation session to determine the grade needed to achieve 65% of heart rate reserve (HRR); a control condition in which they walked at 65% of HRR for 10 minutes and an experimental condition during which they listened to a mindfulness track that directed them to attend to the physical sensations of their body in a nonjudgmental manner during the 10-minute walk. ANOVA results showed that in the mindfulness condition, affective valence was significantly more positive (p =.02, np2 =.22), enjoyment and mindfulness of the body were higher (p <.001, np2 =.36 and.40, respectively), attentional focus was more associative (p <.001, np2 =.67) and RPE was minimally lower (p =.06, np2 =.15). Higher mindfulness of the body was moderately associated with higher enjoyment (p <.05, r =.44) in the mindfulness but not the control condition. Results suggest that mindfulness during exercise is associated with more positive affective responses.
We encounter metaphors every day, but only a few jump out on us and make us stumble. However, little effort has been devoted to investigating more novel metaphors in comparison to general metaphor detection efforts. We attribute this gap primarily to the lack of larger datasets that distinguish between conventionalized, i.e., very common, and novel metaphors. The goal of this paper is to alleviate this situation by introducing a crowdsourced novel metaphor annotation layer for an existing metaphor corpus. Further, we analyze our corpus and investigate correlations between novelty and features that are typically used in metaphor detection, such as concreteness ratings and more semantic features like the Potential for Metaphoricity. Finally, we present a baseline approach to assess novelty in metaphors based on our annotations.
We introduce TED-Multilingual Discourse Bank, a corpus of TED talks transcripts in 6 languages (English, German, Polish, EuropeanPortuguese, Russian and Turkish), where the ultimate aim is to provide a clearly described level of discourse structure and semanticsin multiple languages. The corpus is manually annotated following the goals and principles of PDTB, involving explicit and implicitdiscourse connectives, entity relations, alternative lexicalizations and no relations. In the corpus, we also aim to capture the character-istics of spoken language that exist in the transcripts and adapt the PDTB scheme according to our aims; for example, we introducehypophora. We spot other aspects of spoken discourse such as the discourse marker use of connectives to keep them distinct from theirdiscourse connective use. TED-MDB is, to the best of our knowledge, one of the few multilingual discourse treebanks and is hoped tobe a source of parallel data for contrastive linguistic analysis as well as language technology applications. We describe the corpus, theannotation procedure and provide preliminary corpus statistics.
Down-regulation of negative emotions has been shown to reliably inhibit the emotion-modulated startle reflex, but it remains unclear whether the timing of the startle probe influences the quantification of emotion regulation with this measure. Moreover, it is not known whether the degree of startle inhibition corresponds to the subjective attenuation of negative emotions. Therefore, the two main goals of the study were, first, to systematically analyze the effect of probe time on startle inhibition and, second, to explore the association between subjectively perceived down-regulation of arousal and valence and the degree of startle inhibition. We presented negative and neutral pictures to N = 47 participants. Pictures were paired with the instruction to reappraise or to maintain the emotions elicited by these pictures. Probes were delivered at three different times during a 12.5-s regulation phase, and the startle response was measured with electromyography. Valence and arousal ratings were assessed after each trial. Results revealed no significant impact of probe time on startle inhibition during reappraisal. Startle inhibition and perceived down-regulation of arousal were significantly and positively correlated, whereas perceived down-regulation of valence was not. The results provide important implications for future studies in terms of startle probe timing and shed light onto the interpretation of startle inhibition as an indicator of subjective attenuation of negative emotions. (PsycINFO Database Record (c) 2018 APA, all rights reserved).
Research in the field of text analysis will always be related to words, either the selection of words to be used or the position of the words in a sentence. Furthermore, a hypothesis that each language difference can cause different meanings, makes some researchers interested in doing research classifying words based on emotion or affective words. Research focuses on affective states as a continuous numerical value to the dimensions of valence and arousal. Sentiment analysis that is usually done with positive and negative category approaches, nowadays, the dimensional approach can provide more analysis of grained sentiments. On the other hand, the affective words dataset with valence and arousal rating are still very rare, especially for the Indonesian language. Therefore, this research does an affective lexicon dataset called Indonesian Valence and Arousal Words (IVAW) containing 1024 words by Self-Assessment Manikin (SAM) surveys. Furthermore, for the next study, we will also crawls status in twitter based on selected words from IVAW to get Indonesian Valence and Arousal Text (IVAT). To predict VA rating for obtaining the advance of annotation quality, experiment will be compared by brain signal using EEG tool.
Semantic classification and annotation of satellite images are of great importance and require knowledge resources. The complexity of satellite scenes makes its classification and annotation hard tasks and we are still far from totally resolving the semantic gap problem. There are several knowledge resources such as semantic networks, taxonomies and ontologies. In this paper, we propose to enrich the SatelliteScene-Ontology using real hyperspectral scenes, the USGS spectral library and the WordNet lexical database. The resulting ontology would be published online for further exploitation by researchers.
The way information spreads through society has changed significantly over the past decade with the advent of online social networking. \nTwitter, one of the most widely used social networking websites, is known as the real-time, public microblogging network where news \nbreaks first. Most users love it for its iconic 140-character limitation and unfiltered feed that show them news and opinions in the \nform of tweets. Tweets are usually multilingual in nature and of varying quality. However, machine translation (MT) of twitter data \nis a challenging task especially due to the following two reasons: (i) tweets are informal in nature (i.e., violates linguistic norms), and \n(ii) parallel resource for twitter data is scarcely available on the Internet. In this paper, we develop FooTweets, a first parallel corpus of \ntweets for English–German language pair. We extract 4, 000 English tweets from the FIFA 2014 world cup and manually translate them \ninto German with a special focus on the informal nature of the tweets. In addition to this, we also annotate sentiment scores between 0 \nand 1 to all the tweets depending upon the degree of sentiment associated with them. This data has recently been used to build sentiment \ntranslation engines and an extensive evaluation revealed that such a resource is very useful in machine translation of user generated \ncontent.
Language regulation, as Hynninen explains it in this volume, ‘is a concept that is intended to depict the kind of language use that may lead to formation of (new) norms’ (p. 20). This concept draws attention to the usage of language in micro-level communication and its impact on norms development. Hynninen’s study explores the phenomenon of language regulation in spoken academic settings using English as a Lingua Franca (henceforth ELF), looking at it from both interactional and ideological perspectives. An important distinction in this context was made by Bartsch (1987), who distinguishes between grammatically-conforming language use and acceptable language use. The former is measured against established linguistic norms, whereas the latter is driven by the interactive norm to accomplish mutual understanding. This means that interaction becomes a site for negotiating norms, in the sense that if deviation from what is grammatically correct occurs repeatedly over time and is accepted in talk, new norms may be formed. In this book Hynninen constructs language norms as acceptable linguistic conduct, focusing particularly on the negotiation of acceptable language in interaction. Nevertheless, she does not restrict the exploration of language norms only to interactional behaviours of regulatory negotiations, as research indicates that the conformity to a norm is influenced by expectations of language use. In addition, we also know that speakers’ beliefs about what is correct usage of language is not necessarily reflected in their actual linguistic behaviour. This calls for a binary approach in exploring language norms. In this light, Hynninen conceptualises language norms as a combination of communicative linguistic behaviours (the interactive dimension of the study) and speakers’ beliefs about and expectations of others’ behaviours (the ideological element studied in this research).
We introduce a method to reduce constituent parsing to sequence labeling. For each word w t, it generates a label that encodes: (1) the number of ancestors in the tree that the words w t and w t+1 have in common, and (2) the nonterminal symbol at the lowest common ancestor. We first prove that the proposed encoding function is injective for any tree without unary branches. In practice, the approach is made extensible to all constituency trees by collapsing unary branches. We then use the PTB and CTB treebanks as testbeds and propose a set of fast baselines. We achieve 90.7% F-score on the PTB test set, outperforming the Vinyals et al. ( In addition, sacrificing some accuracy, our approach achieves the fastest constituent parsing speeds reported to date on PTB by a wide margin. 1
This paper describes a collection of modules for Italian language processing based on CoreNLP and Universal Dependencies (UD). The software will be freely available for download under the GNU General Public License (GNU GPL). Given the flexibility of the framework, it is easily adaptable to new languages provided with an UD Treebank.
There are many mechanisms to sense arousal. Most of them are either intrusive, prone to bias, costly, require skills to set-up or do not provide additional context to the user's measure of arousal. We present arousal detection through the analysis of pupillary response from eye trackers. Using eye-trackers, the user's focal attention can be detected with high fidelity during user interaction in an unobtrusive manner. To evaluate this, we displayed twelve images of varying arousal levels rated by the International Affective Picture System (IAPS) to 41 participants while they reported their arousal levels. We found a moderate correlation between the self-reported arousal and the algorithm's arousal rating, r(47)=0.46, p<.01. The results show that eye trackers can serve as a multi-sensory device for measuring arousal, and relate the level of arousal to the user's focal attention. We anticipate that in the future, high fidelity web cameras can be used to detect arousal in relation to user attention, to improve usability, UX and understand visual behaviour.
Abstract The aim of the article is to discuss the legal language transformations from a diachronic perspective taking into account the following factors: (i) spatial and temporal, (ii) linguistic norm changes, (iii) political, (iv) social (customs), and (v) globalization as well as (vi) EU-induced. Spatial and temporal factors include legal relations influenced by climate and the cycles of nature. Linguistic factors include spelling reforms and grammatical changes each language undergoes, for example, as a result of usage. As far as the law is concerned, normative changes can be observed when laws are amended. Other factors such as customs, usage, etc. cannot be neglected when discussing the language of the law. Analogously political correctness and usage can be observed in gender sensitive language and the introduction of such terms as chairperson instead of chairman. Social factors should not be overlooked. As a result of social changes, numerous terms have been introduced to legal lexicons in many countries starting with same-sex unions or same-sex-marriages. The so-called political correctness enforces some language changes and leads to the introduction of new terms and at the same time the abandonment of others. Consequently, some terms cease to be used and consequently become archaic. The aim of the article is to focus on diachronic changes in legal languages and present the communication problems resulting from them from intra- and inter-lingual perspectives.
Conventional lexical-clustering algorithms treat text fragments as a mixed collection of words, with a semantic similarity between them calculated based on the term of how many the particular word occurs within the compared fragments. Whereas this technique is appropriate for clustering large-sized textual collections, it operates poorly when clustering small-sized texts such as sentences. This is due to compared sentences that may be linguistically similar despite having no words in common. This chapter presents a new version of the original k-means method for sentence-level text clustering that is relay on the idea of use of the related synonyms in order to construct the rich semantic vectors. These vectors represent a sentence using linguistic information resulting from a lexical database founded to determine the actual sense to a word, based on the context in which it occurs. Therefore, while traditional k-means method application is relay on calculating the distance between patterns, the new proposed version operates by calculating the semantic similarity between sentences. This allows it to capture a higher degree of semantic or linguistic information existing within the clustered sentences. Experimental results illustrate that the proposed version of clustering algorithm performs favorably against other well-known clustering algorithms on several standard datasets.
Enriched discourse annotation of a subset of the Prague Discourse Treebank, adding implicit relations, entity based relations, question-answer relations and other discourse structuring phenomena.
This paper addresses and targets morpheme segmentation of Kannada words using supervised classification. We have used manually annotated Kannada treebank corpus, which is recently developed by us. Kannada bears resemblance to other Dravidian languages in morphological structure. It is an agglutinative language, hence its words have complex morphological form with each word comprising of a root and an optional set of suffixes. These suffixes carry additional meaning, apart from the root word in a context. This paper discusses the extraction of morphemes of a word by using Support Vector Machines for Classification. Additional features representing the properties of the Kannada words were extracted and the different letters were classified into labels that result in the morphological segmentation of the word. Various methods for evaluation were considered and an accuracy of 85.97% was achieved.
Despite the well-established benefits of regular participation in physical activity, many Australians still fail to maintain sufficient levels. More self-determined types of motivation and more positive affect during activity have been found to be associated with the maintenance of physical activity behaviour over time. Need-supportive approaches to physical activity behaviour change have previously been shown to improve quality of motivation and psychological well-being. This paper outlines the development of a need-supportive, person-centred physical activity program for frontline aged-care workers. The program emphasises the use of self-determined methods of regulating activity intensity (affect, rating of perceived exertion and self-pacing) and is aimed at increasing physical activity behaviour and psychological well-being. The development process was undertaken in six steps using guidance from the Intervention Mapping framework: (i) an in-depth needs assessment (including qualitative interviews where information was gathered from members of the target population); (ii) formation of change objectives; (iii) selecting theory-informed and evidence-based intervention methods and planning their practical application; (iv) producing program components and materials; (v) planning program adoption and implementation, and (vi) planning for evaluation. The program is based in Self-Determination Theory (SDT) and provides tools and elements to support autonomy (the use of a collaboratively developed activity plan and participant choice in activity types), competence (action/coping planning, goal-setting and pedometers), and relatedness (the use of a motivational interviewing-inspired appointment and ongoing support in activity).
We perform a fine-grained large-scale analysis of coreference projection. By projecting gold coreference from Czech to English and vice versa on Prague Czech-English Dependency Treebank 2.0 Coref, we set an upper bound of a proposed projection approach for these two languages. We undertake a detailed thorough analysis that combines the analysis of projection's subtasks with analysis of performance on individual mention types. The findings are accompanied with examples from the corpus.
In Word Sense Disambiguation (WSD), the predominant approach generally involves a supervised system trained on sense annotated corpora. The limited quantity of such corpora however restricts the coverage and the performance of these systems. In this article, we propose a new method that solves these issues by taking advantage of the knowledge present in WordNet, and especially the hypernymy and hyponymy relationships between synsets, in order to reduce the number of different sense tags that are necessary to disambiguate all words of the lexical database. Our method leads to state of the art results on most WSD evaluation tasks, while improving the coverage of supervised systems, reducing the training time and the size of the models, without additional training data. In addition, we exhibit results that significantly outperform the state of the art when our method is combined with an ensembling technique and the addition of the WordNet Gloss Tagged as training corpus.
Dependency is dynamic manifestation of valency, and dependency distance (DD) is closely related to syntactic structure. By introducing the concept degree in graph theory, we analyze the relationship between dynamic valency (DV) and DD based on the Chinese and English treebanks. Our findings are: (1) the mean dependency distance (MDD) of the Chinese treebank is greater than that of English, while the variance of DV of Chinese is lower than that of English; (2) some values of the variance of DV exist in Chinese but not in English; (3) at a specific sentence length, there is a linear relationship between MDD and the variance of DV in the syntactic structures in English and Chinese. These findings suggest: (1) Chinese may have some unique syntactic dependency structures that are not found in English; (2) high DV may contribute to high MDD, but this effect on MDD may not be stronger than some grammatical factors.
BACKGROUND: Neuroplastic underpinnings of meditation-induced changes in affective processing are largely unclear. METHODS: We included healthy older participants in an active-controlled experiment. They were involved a meditation training or a control relaxation training of eight weeks. Associations between behavioral and neural morphometric changes induced by the training were examined. RESULTS: The meditation group demonstrated a change in valence perception indexed by more neutral valence ratings of positive and negative affective images. These behavioral changes were associated with synchronous structural enlargements in a prefrontal network involving the ventromedial prefrontal cortex and the inferior frontal sulcus. In addition, these neuroplastic effects were modulated by the enlargement in the inferior frontal junction. In contrast, these prefrontal enlargements were absent in the active control group, which completed a relaxation training. Supported by a path analysis, we propose a model that describes how meditation may induce a series of prefrontal neuroplastic changes related to valence perception. These brain areas showing meditation-induced structural enlargements are reduced in older people with affective dysregulations. CONCLUSION: We demonstrated that a prefrontal network was enlarged after eight weeks of meditation training. Our findings yield translational insights in the endeavor to promote healthy aging by means of meditation.
We present ongoing work on data-driven parsing of German and French with Lexicalized Tree Adjoining Grammars. We use a supertagging approach combined with deep learning. We show the challenges of extracting LTAG supertags from the French Treebank, introduce the use of leftand right-sister-adjunction, present a neural architecture for the supertagger, and report experiments of n-best supertagging for French and German.