Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Kamusi has been developing a system to analyze texts on the source side and present users with sense-specified dictionary options. Similarly to spellcheck, the user selects the intended meaning. We then use a multilingual lexical database to bridge to matching vocabulary in other languages. When paired with Freeling, additional pre-processing is possible for several languages. Integration with MT via Moses and Apertium is planned, but not yet undertaken. MWEs treatment is important. An MWE is lexicalized in the Kamusi database and marked for separability, with a definition and translation equivalents (one or more words) in other languages. When the initial term of an MWE appears in the source text, Pre:D queries the database and scans the sentence for all MWEs that could follow. The user can select the relevant MWE rather than the component words. A user can submit a missing sense or MWE for inclusion in the lexicon. Named entities can also be identified from data sources or by users and rendered appropriately across languages. When users agree, we will also use sense-tagged sentences for machine learning. A prototype of the core system is already functional.
Unsupervised dependency parsing becomes more and more popular in recent years because it does not need expensive annotations, such as treebanks, which are required for supervised and semi-supervised dependency parsing. However, its accuracy is still far below that of supervised dependency parsers, partly due to the fact that their parsing model is insufficient to capture linguistic phenomena underlying texts. The performance for unsupervised dependency parsing can be improved by mining knowledge from the texts and by incorporating it into the model. In this article, syntactic knowledge is acquired from query logs to help estimate better probabilities in dependency models with valence. The proposed method is language independent and obtains an improvement of 4.1% unlabeled accuracy on the Penn Chinese Treebank by utilizing additional dependency relations from the Sogou query logs and Baidu query logs. Morever, experiments show that the proposed model achieves improvements of 8.07% on CoNLL 2007 English using the AOL query logs. We believe query logs are useful sources of syntactic knowledge for many natural language processing (NLP) tasks.
We investigate mutual benefits between syntax and semantic roles using neural network models, by studying a parsingSRL pipeline, a SRLparsing pipeline, and a simple joint model by embedding sharing. The integration of syntactic and semantic features gives promising results in a Chinese Semantic Treebank, demonstrating large potentials of neural models for joint parsing and semantic role labeling.
The article presents methodological analyses of topical ideas of the famous modern linguist – E.Cosseriu. The authors argue that incorporation of theoretical ideas of E.Cosseriu could substentially extand the euristic potential of the conept of norm in the sphere of linguistics. The key to solvation of the problem lies in the necessity of changing of modern theoretical context of the question. The authors of the article consider the history of operationalization of the concept of norm in linguistics from the point of view of a specific hermeneutic approach in relation to other linguistic techniques. In the course of study of the role of linguistic norms in interaction of content and expression the article presents examples of extrapolation of the concept of norm from one discipline to another. In case of extrapolation of the concept of norm from the other disciplines into linguistic investigations the structural and functional dichotomy of language turns out that it is impossible to avoid considering spiritual as the world of objects, so that speech and language are considered as two different things. The rapid development of information technologies made possible to calculate many of the aspects of Humboldtian ideas. Computer statistics created conditions where the idea of "language picture of the world" and of the "inner form of the language" is gradually losing its original romantic charge and turns in a very trivial thing. Yet despite the fact that global standardization significantly enhances the processing and automatic addition of translation, which once gave beginning to hermeneutics the hermeneutic potential of linguistic norm still preseves many promissing prospects from the epistemological point of view. Key words: Linguistic norm and variability, E. Cosseriu, hermeneutics, euristic potential. В статье представлен методологический анализ актуальных идей известного современного лингвиста - E. Коссериу. Авторы утверждают, что введение теоретических идей E. Коссериу могло бы существенно расширить эвристический потенциал нормы в области лингвистики. Ключ к решению проблемы заключается в необходимости изменения современного теоретического контекста вопроса. Авторы статьи рассматривают историю ввода в действие понятия нормы в лингвистике с точки зрения конкретного герменевтического подхода по отношению к другим языковым методам. В ходе изучения роли языковых норм во взаимодействии содержания и выражения, в статье представлены примеры экстраполяции понятия нормы от одной дисциплины к другой. В случае экстраполяции понятия нормы с других дисциплин в лингвистические исследования структурно-функциональной дихотомии языка оказывается, что нельзя не рассматривать духовное как мир объектов, так как речь и язык рассматриваются как две разные вещи. Быстрое развитие информационных технологий сделало возможным для расчета многие аспекты гумбольдтовских идей. Статистика компьютера создала условия, в которых идея «языковой картины мира» и «внутренней формы языка» постепенно теряет свой первоначальный романтический заряд и превращается в очень тривиальную вещь. Тем не менее, несмотря на то, что глобальная стандартизация значительно улучшает обработку и автоматический перевод, который когда-то дал начало герменевтике, герменевтический потенциал языковой нормы хранит еще много перспектив с гносеологической точки зрения.Ключевые слова: лингвистическая норма и изменчивость, E.Коссериу, герменевтика, эвристический по- тенциал.
An alternative method for deriving typicality judgments, applicable in young children that are not familiar with numerical values yet, is introduced, allowing researchers to study gradedness at younger ages in concept development. Contrary to the long tradition of using rating-based procedures to derive typicality judgments, we propose a method that is based on typicality ranking rather than rating, in which items are gradually sorted according to their typicality, and that requires a minimum of linguistic knowledge. The validity of the method is investigated and the method is compared to the traditional typicality rating measurement in a large empirical study with eight different semantic concepts. The results show that the typicality ranking task can be used to assess children’s category knowledge and to evaluate how this knowledge evolves over time. Contrary to earlier held assumptions in studies on typicality in young children, our results also show that preference is not so much )
Parsing Arabic language is a difficult task given the specificities of the language and given the scarcity of linguistic resources. Linguistic resources such as grammars are very important to any natural language processing application. Unfortunately, the manual construction of these resources is laborious and time-consuming. The use of annotated corpora as a knowledge database might be a solution to a fast construction of a grammar for a given language. In this paper, we began by presenting an overview of our method to automatically induce a probabilistic context free grammar from an Arabic annotated corpus (The Penn Arabic TreeBank). Then we tested the obtained grammar in the parsing task and we expose the evaluation results. Finally we present our vision of a hybrid method for parsing Modern Standard Arabic (MSA) that we believe that it could enhance obtained results.
This article deals with the problem of development of the dialectological base of the Machine fund of the Bashkir language. Along with the written monuments, folklore material, dialect is one of the most important sources for the study of the historical development and formation of the literary language. The main task of dialectologists is not only the collection but also the storage of dialect materials of the Bashkir language. Taking into account the importance of translation of dialect materials in electronic format and opening of free access to a wide audience to the information provided, the staff of the Laboratory of Linguistics and Information Technologies of the Institute of History, Language and Literature of the Ufa Scientific Center established a dialectological base as a part of the Machine fund of the Bashkir language consisting of three separate databases — lexical database, the database of dialectological atlas and textological database. The lexical database includes information about the dialectal lexicon of the Bashkir language. The database was developed on the basis of dialectal dictionaries compiled and published by the staff of the Institute of History, Language and Literature. The volume of the database is more than 52 000 dialectal units. The database of dialectological atlas is developed on the basis of materials of Dialectological atlas of the Bashkir language. The database allows to select the types of linguistic phenomena (phonetic, morphological, syntactic, lexical), for each type specific isoglosses were identified. Isoglosses allocated 250 strong points of the republic and neighboring regions. The textological database represents illustrative materials collected during numerous expeditions by the staff of the Institute of History, Language and Literature. To date, the database contains more than 500 texts in all dialects of the Bashkir language. Introduced in dialectological base material forms the basis of synchronic and diachronic study of the language features of dialects and sub-dialects of the Bashkir language.
It is recognized that both sources of and solutions to the big data challenges are human collective intelligence. Data are an abstract representation of the quantity of realistic entities and perceived objects. Big data play an indispensable role not only in a wide range of engineering applications, but also in the cognitive mechanisms of both humans and cognitive robots such as sensation, quantification, qualification, estimation, memory, and reasoning. This keynote lecture presents online big data analytic theories by machine learning as well as knowledge extraction by cognitive robotics. Data, information, knowledge, and intelligence are the four hierarchical layers of cognitive objects in the brain and cognitive systems from the bottom up. It is discovered by the author that, although the cognitive unit of data is bit, that of knowledge is bir, i.e., a binary relation. A resent finding towards big data science is that big data systems in nature are a recursive n-dimensional typed hyperstructure (RNDTHS). This topological property of big data reveals that the mathematical foundation of big data science is underpinned by big data algebra (BDA), which is a denotational mathematical structure for efficiently dealing with the inherited complexities and unprecedented challenges in big data engineering. This leads to a coherent theory for big data modeling, analyses, mining, information elicitation, knowledge extraction, and intelligent generation for cognitive robots. Latest experiments demonstrate that cognitive robots may autonomously transform big data in vast linguistic databases (corpuses) into sophistic knowledge bases by cognitive machine learning.
This paper presents the Universal Dependencies tagset (UD v1) as a new annotation scheme for Russian treebanks. The universal list of dependency relations was adopted and extended to comply with certain language-specific syntactic constructions. The tagset was validated, converting two Russian treebanks into the UD format, UD-Russian-SynTagRus and UD-Russian-Google.
We present a new, sizeable dataset of nounnoun compounds with their syntactic analysis (bracketing) and semantic relations. Derived from several established linguistic resources, such as the Penn Treebank, our dataset enables experimenting with new approaches towards a holistic analysis of noun-noun compounds, such as jointlearning of noun-noun compounds bracketing and interpretation, as well as integrating compound analysis with other tasks such as syntactic parsing.
This article deals with the problem of development of the dialectological base of the Machine fund of the Bashkir language. Along with the written monuments, folklore material, dialect is one of the most important sources for the study of the historical development and formation of the literary language. The main task of dialectologists is not only the collection but also the storage of dialect materials of the Bashkir language. Taking into account the importance of translation of dialect materials in electronic format and opening of free access to a wide audience to the information provided, the staff of the Laboratory of Linguistics and Information Technologies of the Institute of History, Language and Literature of the Ufa Scientific Center established a dialectological base as a part of the Machine fund of the Bashkir language consisting of three separate databases — lexical database, the database of dialectological atlas and textological database. The lexical database includes information about the dialectal lexicon of the Bashkir language. The database was developed on the basis of dialectal dictionaries compiled and published by the staff of the Institute of History, Language and Literature. The volume of the database is more than 52 000 dialectal units. The database of dialectological atlas is developed on the basis of materials of Dialectological atlas of the Bashkir language. The database allows to select the types of linguistic phenomena (phonetic, morphological, syntactic, lexical), for each type specific isoglosses were identified. Isoglosses allocated 250 strong points of the republic and neighboring regions. The textological database represents illustrative materials collected during numerous expeditions by the staff of the Institute of History, Language and Literature. To date, the database contains more than 500 texts in all dialects of the Bashkir language. Introduced in dialectological base material forms the basis of synchronic and diachronic study of the language features of dialects and sub-dialects of the Bashkir language.
Recently, several sets of standardized food pictures have been created, supplying both food images and their subjective evaluations. However, to date only the OLAF (Open Library of Affective Foods), a set of food images and ratings we developed in adolescents, has the specific purpose of studying emotions toward food. Moreover, some researchers have argued that food evaluations are not valid across individuals and groups, unless feelings toward food cues are compared with feelings toward intense experiences unrelated to food, that serve as benchmarks. Therefore the OLAF presented here, comprising a set of original food images and a group of standardized highly emotional pictures, is intended to provide valid between-group judgments in adults. Emotional images (erotica, mutilations, and neutrals from the International Affective Picture System/IAPS) additionally ensure that the affective ratings are consistent with emotion research. The OLAF depicts high-calorie sweet and savory foods and low-calorie fruits and vegetables, portraying foods within natural scenes matching the IAPS features. An adult sample evaluated both food and affective pictures in terms of pleasure, arousal, dominance, and food craving, following standardized affective rating procedures. The affective ratings for the emotional pictures corroborated previous findings, thus confirming the reliability of evaluations for the food images. Among the OLAF images, high-calorie sweet and savory foods elicited the greatest pleasure, although they elicited, as expected, less arousal than erotica. The observed patterns were consistent with research on emotions and confirmed the reliability of OLAF evaluations. The OLAF and affective pictures constitute a sound methodology to investigate emotions toward food within a wider motivational framework.
The paper evaluates the differences between two currently leading annotation schemes for dependency treebanks. By relying on four treebanks, we demonstrate that the treatment of conjunctions and adpositions represents the core difference between the two schemes and that this impacts the topological properties of the linguistic networks induced from the treebanks. We also show that such properties are reflected in the performances of four probabilistic dependency parsers trained on the treebanks. L’articolo valuta le differenze tra i due principali schemi di annotazione a dipenden-ze in uso. Sulla base di quattro treebank, l’articolo dimostra che il trattamento delle congiunzioni e delle pre/postposizioni rappresenta la differenza principale tra i due schemi e che ciò comporta delle conseguenze sulle proprietà topologiche dei net-work indotti dalle treebank. Inoltre, si dimostra come tali proprietà siano riflesse nell’accuratezza di quattro parser probabilistici a dipendenze addestrati sulle treebank.
Statistical parsers are trained on treebanks that are composed of a few thousand sentences. In order to prevent data sparseness and computational complexity, such parsers make strong independence hypotheses on the decisions that are made to build a syntactic tree. These independence hypotheses yield a decomposition of the syntactic structures into small pieces, which in turn prevent the parser from adequately modeling many lexico-syntactic phenomena like selectional constraints and subcategorization frames. Additionally, treebanks are several orders of magnitude too small to observe many lexico-syntactic regularities, such as selectional constraints and subcategorization frames. In this article, we propose a solution to both problems: how to account for patterns that exceed the size of the pieces that are modeled in the parser and how to obtain subcategorization frames and selectional constraints from raw corpora and incorporate them in the parsing process. The method proposed was evaluated on French and on English. The experiments on French showed a decrease of 41.6% of selectional constraint violations and a decrease of 22% of erroneous subcategorization frame assignment. These figures are lower for English: 16.21% in the first case and 8.83% in the second.
While linguistic theory posits an arbitrary relation between signifiers and the signified (de Saussure, 1916), our analysis of a large-scale German database containing affective ratings of words revealed that certain phoneme clusters occur more often in words denoting concepts with negative and arousing meaning. Here, we investigate how such phoneme clusters that potentially serve as sublexical markers of affect can influence language processing. We registered the EEG signal during a lexical decision task with a novel manipulation of the words' putative sublexical affective potential: the means of valence and arousal values for single phoneme clusters, each computed as a function of respective values of words from the database these phoneme clusters occur in. Our experimental manipulations also investigate potential contributions of formal salience to the sublexical affective potential: Typically, negative high-arousing phonological segments-based on our calculations-tend to be less frequent and more structurally complex than neutral ones. We thus constructed two experimental sets, one involving this natural confound, while controlling for it in the other. A negative high-arousing sublexical affective potential in the strictly controlled stimulus set yielded an early posterior negativity (EPN), in similar ways as an independent manipulation of lexical affective content did. When other potentially salient formal features at the sublexical level were not controlled for, the effect of the sublexical affective potential was strengthened and prolonged (250-650 ms), presumably because formal salience helps making specific phoneme clusters efficient sublexical markers of negative high-arousing affective meaning. These neurophysiological data support the assumption that the organization of a language's vocabulary involves systematic sound-to-meaning correspondences at the phonemic level that influence the way we process language.
Deaf or hard-of-hearing individuals usually face a greater challenge to learn to write than their normal-hearing counterparts. Due to the limitations of traditional research methods focusing on microscopic linguistic features, a holistic characterization of the writing linguistic features of these language users is lacking. This study attempts to fill this gap by adopting the methodology of linguistic complex networks. Two syntactic dependency networks are built in order to compare the macroscopic linguistic features of deaf or hard-of-hearing students and those of their normal-hearing peers. One is transformed from a treebank of writing produced by Chinese deaf or hard-of-hearing students, and the other from a treebank of writing produced by their Chinese normal-hearing counterparts. Two major findings are obtained through comparison of the statistical features of the two networks. On the one hand, both linguistic networks display small-world and scale-free network structures, but the network of the normal-hearing students' exhibits a more power-law-like degree distribution. Relevant network measures show significant differences between the two linguistic networks. On the other hand, deaf or hard-of-hearing students tend to have a lower language proficiency level in both syntactic and lexical aspects. The rigid use of function words and a lower vocabulary richness of the deaf or hard-of-hearing students may partially account for the observed differences.
Abstract syntax is a semantic tree representation that lies between parse trees and logical forms. It abstracts away from word order and lexical items, but contains enough information to generate both surface strings and logical forms. Abstract syntax is commonly used in compilers as an intermediate between source and target languages. Grammatical Framework (GF) is a grammar formalism that generalizes the idea to natural languages, to capture cross-lingual generalizations and perform interlingual translation. As one of the main results, the GF Resource Grammar Library (GF-RGL) has implemented a shared abstract syntax for over 30 languages. Each language has its own set of concrete syntax rules (morphology and syntax), by which it can be generated from the abstract syntax and parsed into it. This paper presents a conversion method from abstract syntax trees to dependency trees. The method is applied for converting GF-RGL trees to Universal Dependencies (UD), which uses a common set of labels for different languages. The correspondence between GF-RGL and UD turns out to be good, and the relatively few discrepancies give rise to interesting questions about universality. The conversion also has potential for practical applications: (1) it makes the GF parser usable as a rule-based dependency parser; (2) it enables bootstrapping UD treebanks from GF treebanks; (3) it defines formal criteria to assess the informal annotation schemes of UD; (4) it gives a method to check the consistency of manually annotated UD trees with respect to the annotation schemes; (5) it makes information from UD treebanks available.
Abstract Three studies examined gender differences in the effect of storytelling ability on perceptions of a person's attractiveness as a short‐term and long‐term romantic partner. In Study 1, information about a potential partner's storytelling ability was provided. Study 2 participants read a good or poor story supposedly written by a potential partner. Results suggested that only women's attractiveness assessments of men as a long‐term date increased for good storytellers. Storytelling ability did not affect men's ratings of women nor did it affect ratings of short‐term partners. Study 3 suggested that the effect of storytelling ability on long‐term attractiveness for male targets may be mediated by perceived status. Storytelling ability appears to increase perceived status and thus helps men attract long‐term partners.
Olfactory identification abilities in adolescents have been reported inferior compared with adults. Though this seems to be the case when comparing identification abilities using tests validated on-and for-adults, odor familiarity has been hypothesized to affect identification abilities in younger participants. However, this has never been thoroughly tested. The aims of this study were to investigate patterns in odor familiarity differences between adolescents and adults, and to investigate if an adolescent familiarity-based modification of an identification test could lead to similar identification scores in adolescents and adults. In total, 411 adolescent participants and 320 adult participants were included in the study. Odor familiarity ratings were obtained for 125 odors. A modified version of the "Sniffin' Sticks" identification test was created and validated on 72 adolescents based on adolescent familiarity scores. This test was applied to 82 normosmic adults and 167 normosmic adolescents. Results show a lower familiarity for spices and environmental odors, and a higher familiarity for candy odors in adolescents. The identification abilities in adults and adolescents were equal after familiarity-based modification. We conclude that changes in odor familiarity from adolescence to adulthood do not develop evenly for all odors, but are dependent on odor-object category.
We present a study on two key characteristics of human syntactic annotations: anchoring and agreement. Anchoring is a well known cognitive bias in human decision making, where judgments are drawn towards pre-existing values. We study the influence of anchoring on a standard approach to creation of syntactic resources where syntactic annotations are obtained via human editing of tagger and parser output. Our experiments demonstrate a clear anchoring effect and reveal unwanted consequences, including overestimation of parsing performance and lower quality of annotations in comparison with human-based annotations. Using sentences from the Penn Treebank WSJ, we also report systematically obtained inter-annotator agreement estimates for English dependency parsing. Our agreement results control for parser bias, and are consequential in that they are on par with state of the art parsing performance for English newswire. We discuss the impact of our findings on strategies for future annotation efforts and parser evaluations.
We train one multilingual model for dependency parsing and use it to parse sentences in several languages. The parsing model uses (i) multilingual word clusters and embeddings; (ii) token-level language information; and (iii) language-specific features (fine-grained POS tags). This input representation enables the parser not only to parse effectively in multiple languages, but also to generalize across languages based on linguistic universals and typological similarities, making it more effective to learn from limited annotations. Our parser’s performance compares favorably to strong baselines in a range of data scenarios, including when the target language has a large treebank, a small treebank, or no treebank for training.
The PDTB Annotator is a tool for annotating and adjudicating discourse relations based on the annotation framework of the Penn Discourse TreeBank (PDTB). This demo describes the benefits of using the PDTB Annotator, gives an overview of the PDTB Framework and discusses the tool’s features, setup requirements and how it can also be used for adjudication.
The increasing need of automated analyzing web texts especially the short texts on Social Network Services (SNS) brings new demands of computerized text analysis instruments. The psychometric properties are the basis of the extensive use of these instruments such as the Linguistic Inquiry and Word Count (LIWC). For this study, Sina Weibo statuses were analyzed via rater coding and Simplified Chinese version of LIWC (SCLIWC), in order to evaluate the validity of SCLIWC in detecting psychological expressions in Weibo statuses (n = 60) and in identifying the psychological meaning of a single Weibo status (n = 11). Significant correlations between human ratings and SCLIWC scores and the high sensitivities of capturing single statuses with certain expressions identified by raters, proved the validity of SCLIWC in detecting psychological expressions. The results also suggested that, the efficiency of SCLIWC in detecting psychological expressions of SNS short texts could be higher if using s)
We investigated whether lines and shapes that present face-like features would be associated with emotions. In Experiment 1, participants associated concave, convex, or straight lines with the words happy or sad. Participants found it easiest to associate the concave line with happy and the convex line with sad. In Experiment 2, participants rated (valence, pleasantness, liking, and tension) and categorised (valence and emotion words) two convex and concave lines that were paired with six distinct pairs of eyes. The presence of eyes affected participants' valence ratings and response latencies; more congruent eye-mouth matches produced more consistent ratings and faster reaction times. In Experiment 3, we examined whether dots that resembled eyes would be associated with emotional words. Participants found it easier to match certain sets of dots with specific emotions. These results suggest that facial gestures that are associated with specific emotions can be captured using relatively simple shapes and lines.
Due to the constant increasing of electronic textual information, modern society needs for the automatic processing of natural language (NL). The main purpose of NL automatic text processing systems is to analyze and create texts and represent their content. The purpose of the paper is the development of linguistic and software bases of an automatic system for processing English publicistic texts. This article discusses the examples of different approaches to the creation of linguistic databases for processing systems. The author gives a detailed description of basic building blocks for a new linguistic processor: lexical-semantic, syntactical and semantic-syntactical. The main advantage of the processor is using special semantic codes in the alphabetical dictionary. The semantic codes have been developed in accordance with a lexical-semantic classification. It helps to precisely define semantic functions of the keywords that are situated in parsing groups and allows the automatic system to avoid typical mistakes. The author also represents the realization of a developed linguistic database in the form of a training computer program.
OBJECTIVE: To investigate the effect of transcranial direct current stimulation (tDCS) on food craving, intake, binge eating desire, and binge eating frequency in individuals with binge eating disorder (BED). METHOD: N = 30 adults with BED or subthreshold BED received a 20-min 2 milliampere (mA) session of tDCS targeting the dorsolateral prefrontal cortex (DLPFC; anode right/cathode left) and a sham session. Food image ratings assessed food craving, a laboratory eating test assessed food intake, and an electronic diary recorded binge variables. RESULTS: tDCS versus sham decreased craving for sweets, savory proteins, and an all-foods category, with strongest reductions in men (p < 0.05). tDCS also decreased total and preferred food intake by 11 and 17.5%, regardless of sex (p < 0.05), and reduced desire to binge eat in men on the day of real tDCS administration (p < 0.05). The reductions in craving and food intake were predicted by eating less frequently for reward motives, and greater intent to restrict calories, respectively. DISCUSSION: This proof of concept study is the first to find ameliorating effects of tDCS in BED. Stimulation of the right DLPFC suggests that enhanced cognitive control and/or decreased need for reward may be possible functional mechanisms. The results support investigation of repeated tDCS as a safe and noninvasive treatment adjunct for BED. © 2016 Wiley Periodicals, Inc.(Int J Eat Disord 2016; 49:930-936).
In this paper, we study novel neural network structures to better model long term dependency in sequential data. We propose to use more memory units to keep track of more preceding states in recurrent neural networks (RNNs), which are all recurrently fed to the hidden layers as feedback through different weighted paths. By extending the popular recurrent structure in RNNs, we provide the models with better short-term memory mechanism to learn long term dependency in sequences. Analogous to digital filters in signal processing, we call these structures as higher order RNNs (HORNNs). Similar to RNNs, HORNNs can also be learned using the back-propagation through time method. HORNNs are generally applicable to a variety of sequence modelling tasks. In this work, we have examined HORNNs for the language modeling task using two popular data sets, namely the Penn Treebank (PTB) and English text8 data sets. Experimental results have shown that the proposed HORNNs yield the state-of-the-art performance on both data sets, significantly outperforming the regular RNNs as well as the popular LSTMs.
Content-basis image retrieval is important option to prevail within the difficulties of previous works and contains attracted an excellent concentration in past decades. The models according to graph-based ranking were mostly analysed and extensively functional in file recovery area. Within our work we concentrate on the novel in addition to efficient graph-based model for content based image retrieval, designed for out-of-sample recovery on extensive databases. We advise a scalable graph-based ranking representation referred to as effective Manifold Ranking, which address weak points of Manifold Ranking from two most significant viewpoints for example scalable graph construction in addition to effective ranking computation. We concentrate on a famous graph-based model known Manifold Ranking that is a well-known graph-based ranking representation that ranks data samples relevant to intrinsic geometrical structure uncovered with a huge data. The suggested model includes two separate stages just like an offline stage for structuring of ranking model plus an online stage for controlling of recent query. Using the suggested system, we are able to handle database by a million images and perform online retrieval inside a short instance.
<span>This work reports on ongoing research aimed at modeling a metonymic relationship in the FrameNet <span>Brasil database. This paper is based on a case study with the <span>Teams <span>frame. Both the frame and the <span>corpus consulted are part of a frame-based trilingual (Portuguese – Spanish – English) electronic <span>dictionary covering the soccer, tourism and World Cup domains developed by FrameNet Brasil. The<br /><span>basic infrastructure, analytical categories and methodology used were those developed for FrameNet <span>(Fillmore et al. 2003, Baker et al. 2003, Ruppenhofer et al. 2010), which can be defined as an<br /><span>application of Frame Semantics to practical lexicography.</span></span></span></span></span></span><br /></span></span></span>
This paper presents neural probabilistic parsing models which explore up to thirdorder graph-based parsing with maximum likelihood training criteria. Two neural network extensions are exploited for performance improvement. Firstly, a convolutional layer that absorbs the influences of all words in a sentence is used so that sentence-level information can be effectively captured. Secondly, a linear layer is added to integrate different order neural models and trained with perceptron method. The proposed parsers are evaluated on English and Chinese Penn Treebanks and obtain competitive accuracies.
In this paper we study different types of Recurrent Neural Networks (RNN) for sequence labeling tasks. We propose two new variants of RNNs integrating improvements for sequence labeling, and we compare them to the more traditional Elman and Jordan RNNs. We compare all models, either traditional or new, on four distinct tasks of sequence labeling: two on Spoken Language Understanding (ATIS and MEDIA); and two of POS tagging for the French Treebank (FTB) and the Penn Treebank (PTB) corpora. The results show that our new variants of RNNs are always more effective than the others.
We use reinforcement learning to learn tree-structured neural networks for computing representations of natural language sentences. In contrast with prior work on tree-structured models in which the trees are either provided as input or predicted using supervision from explicit treebank annotations, the tree structures in this work are optimized to improve performance on a downstream task. Experiments demonstrate the benefit of learning task-specific composition orders, outperforming both sequential encoders and recursive encoders based on treebank annotations. We analyze the induced trees and show that while they discover some linguistically intuitive structures (e.g., noun phrases, simple verb phrases), they are different than conventional English syntactic structures.
We compare different word embeddings from a standard window based skipgram model, a skipgram model trained using dependency context features and a novel skipgram variant that utilizes additional information from dependency graphs. We explore the effectiveness of the different types of word embeddings for word similarity and sentence classification tasks. We consider three common sentence classification tasks: question type classification on the TREC dataset, binary sentiment classification on Stanford's Sentiment Treebank and semantic relation classification on the SemEval 2010 dataset. For each task we use three different classification methods: a Support Vector Machine, a Convolutional Neural Network and a Long Short Term Memory Network. Our experiments show that dependency based embeddings outperform standard window based embeddings in most of the settings, while using dependency context embeddings as additional features improves performance in all tasks regardless of the classification method. Our embeddings and code are available at
The Kashmiri population is an ethno-linguistic group that resides in the Kashmir Valley in northern India. A longstanding hypothesis is that this population derives ancestry from Jewish and/or Greek sources. There is historical and archaeological evidence of ancient Greek presence in India and Kashmir. Further, some historical accounts suggest ancient Hebrew ancestry as well. To date, it has not been determined whether signatures of Greek or Jewish admixture can be detected in the Kashmiri population. Using genome-wide genotyping and admixture detection methods, we determined there are no significant or substantial signs of Greek or Jewish admixture in modern-day Kashmiris. The ancestry of Kashmiri Tibetans was also determined, which showed signs of admixture with populations from northern India and west Eurasia. These results contribute to our understanding of the existing population structure in northern India and its surrounding geographical areas. [ABSTRACT FROM AUTHOR], Copyright)
Our current understanding of pre-Columbian history in the Americas rests in part on several trends identified in recent genetic studies. The goal of this study is to reexamine these trends in light of the impact of post-Columbian admixture and the methods used to study admixture. The previously-published data consist of 645 autosomal microsatellite genotypes from 1046 individuals in 63 populations. We used STRUCTURE to estimate ancestry proportions and tested the sensitivity of these estimates to the choice of the number of clusters, K. We used partial correlation analyses to examine the relationship between gene diversity and geographic distance from Beringia, controlling for non-Native American ancestry (from Africa, Europe and East Asia), and taking into account alternative paths of migration. Principal component analysis and multidimensional scaling were used to investigate the relationships between Andean and non-Andean populations and to explore gene-language correspondence. We )
This paper presents a novel latent variable recurrent neural network architecture for jointly modeling sequences of words and (possibly latent) discourse relations between adjacent sentences.A recurrent neural network generates individual words, thus reaping the benefits of discriminatively-trained vector representations.The discourse relations are represented with a latent variable, which can be predicted or marginalized, depending on the task.The resulting model can therefore employ a training objective that includes not only discourse relation classification, but also word prediction.As a result, it outperforms state-ofthe-art alternatives for two tasks: implicit discourse relation classification in the Penn Discourse Treebank, and dialog act classification in the Switchboard corpus.Furthermore, by marginalizing over latent discourse relations at test time, we obtain a discourse informed language model, which improves over a strong LSTM baseline.
Normal-hearing listeners use acoustic cues in speech to interpret a speaker's emotional state. This study investigates the effect of hearing aids on the perception of the emotion dimensions arousal (aroused/calm) and valence (positive/negative attitude) in older adults with hearing loss. More specifically, we investigate whether wearing a hearing aid improves the correlation between affect ratings and affect-related acoustic parameters. To that end, affect ratings by 23 hearing-aid users were compared for aided and unaided listening. Moreover, these ratings were compared to the ratings by an age-matched group of 22 participants with age-normal hearing.For arousal, hearing-aid users rated utterances as generally more aroused in the aided than in the unaided condition. Intensity differences were the strongest indictor of degree of arousal. Among the hearing-aid users, those with poorer hearing used additional prosodic cues (i.e., tempo and pitch) for their arousal ratings, compared to those with relatively good hearing. For valence, pitch was the only acoustic cue that was associated with valence. Neither listening condition nor hearing loss severity (differences among the hearing-aid users) influenced affect ratings or the use of affect-related acoustic parameters. Compared to the normal-hearing reference group, ratings of hearing-aid users in the aided condition did not generally differ in both emotion dimensions. However, hearing-aid users were more sensitive to intensity differences in their arousal ratings than the normal-hearing participants.We conclude that the use of hearing aids is important for the rehabilitation of affect perception and particularly influences the interpretation of arousal.
Recurrent Neural Network (RNN) is one of the most popular architectures used in Natural Language Processsing (NLP) tasks because its recurrent structure is very suitable to process variable-length text. RNN can utilize distributed representations of words by first converting the tokens comprising each text into vectors, which form a matrix. And this matrix includes two dimensions: the time-step dimension and the feature vector dimension. Then most existing models usually utilize one-dimensional (1D) max pooling operation or attention-based operation only on the time-step dimension to obtain a fixed-length vector. However, the features on the feature vector dimension are not mutually independent, and simply applying 1D pooling operation over the time-step dimension independently may destroy the structure of the feature representation. On the other hand, applying two-dimensional (2D) pooling operation over the two dimensions may sample more meaningful features for sequence modeling tasks. To integrate the features on both dimensions of the matrix, this paper explores applying 2D max pooling operation to obtain a fixed-length representation of the text. This paper also utilizes 2D convolution to sample more meaningful information of the matrix. Experiments are conducted on six text classification tasks, including sentiment analysis, question classification, subjectivity classification and newsgroup classification. Compared with the state-of-the-art models, the proposed models achieve excellent performance on 4 out of 6 tasks. Specifically, one of the proposed models achieves highest accuracy on Stanford Sentiment Treebank binary classification and fine-grained classification tasks.
The Teacher Forcing algorithm trains recurrent networks by supplying observed sequence values as inputs during training and using the network's own one-step-ahead predictions to do multi-step sampling. We introduce the Professor Forcing algorithm, which uses adversarial domain adaptation to encourage the dynamics of the recurrent network to be the same when training the network and when sampling from the network over multiple time steps. We apply Professor Forcing to language modeling, vocal synthesis on raw waveforms, handwriting generation, and image generation. Empirically we find that Professor Forcing acts as a regularizer, improving test likelihood on character level Penn Treebank and sequential MNIST. We also find that the model qualitatively improves samples, especially when sampling for a large number of time steps. This is supported by human evaluation of sample quality. Trade-offs between Professor Forcing and Scheduled Sampling are discussed. We produce T-SNEs showing that Professor Forcing successfully makes the dynamics of the network during training and sampling more similar.
A defining trait of linguistic competence is the ability to combine elements into increasingly complex structures to denote, and to comprehend, a potentially infinite number of meanings. Recent magnetoencephalography (MEG) work has investigated these processes by comparing the response to nouns in combinatorial (blue car) and non-combinatorial (rnsh car) contexts. In the current study we extended this paradigm using electroencephalography (EEG) to dissociate the role of semantic content from phonological well-formedness (yerl car). We used event-related potential (ERP) recordings in order to better relate the observed neurophysiological correlates of basic combinatorial operations to prior ERP work on comprehension. We found that nouns in combinatorial contexts (blue car) elicited a greater centro-parietal negativity between 180-400ms, independent of the phonological well-formedness of the context word. We discuss the potential relationship between this ‘combinatorial’ effect and clas)
Online Voting Advice Applications (VAAs) are survey-like instruments that help citizens to shape their political preferences and compare them with those of political parties. Especially in multi-party democracies, their increasing popularity indicates that VAAs play an important role in opinion formation for citizens, as well as in the public debate prior to elections. Hence, the objectivity and transparency of VAAs are crucial. In the design of VAAs, many choices have to be made. Extant research in survey methodology shows that the seemingly arbitrary choice to word questions positively (e.g., ‘The city council should allow cars into the city centre’) or negatively (‘The city council should ban cars from the city centre’) systematically affects the answers. This asymmetry in answers is in line with work on negativity bias in other areas of linguistics and psychology. Building on these findings, this study investigated whether question polarity also affects the answers to VAA statemen)
We present ASCERTAIN-a multimodal databaASe for impliCit pERsonaliTy and Affect recognitIoN using commercial physiological sensors. To our knowledge, ASCERTAIN is the first database to connect personality traits and emotional states via physiological responses. ASCERTAIN contains big-five personality scales and emotional self-ratings of 58 users along with their Electroencephalogram (EEG), Electrocardiogram (ECG), Galvanic Skin Response (GSR) and facial activity data, recorded using off-the-shelf sensors while viewing affective movie clips. We first examine relationships between users' affective ratings and personality scales in the context of prior observations, and then study linear and non-linear physiological correlates of emotion and personality. Our analysis suggests that the emotion-personality relationship is better captured by non-linear rather than linear statistics. We finally attempt binary emotion and personality trait recognition using physiological features. Experimental results cumulatively confirm that personality differences are better revealed while comparing user responses to emotionally homogeneous videos, and above-chance recognition is achieved for both affective and personality dimensions.
[Resumen] En este trabajo presentamos una nueva estrategia para crear treebanks de lenguas con pocos recursos para el análisis sintáctico. El método consiste en la adaptación y combinación de diferentes treebanks anotados con dependencias universales de variedades lingüísticas próximas, con el objetivo de entrenar un analizador sintáctico para la lengua elegida, en nuestro caso el gallego. Durante el proceso de selección y adaptación de los treebanks de origen, analizamos el impacto de propiedades de tres niveles diferentes: (i) la distancia entre las lenguas de origen y destino, (ii) la adaptación de características léxico-ortográficas, y (iii) las directrices de anotación entre los treebanks. Usando la estrategia propuesta, entrenamos un analizador sintáctico estadístico para etiquetar, con resultados prometedores y sin datos previos de gallego, un pequeño corpus de esta lengua. La corrección manual de este corpus, usado como gold-standard, nos permitió probar la eficacia del método propuesto.
This paper describes the automatic procedure we developed to convert an Italian dependency treebank into a different format. We defined about 4,250 formal rules for rewriting dependencies and token tags as well as an algorithm for treebank rewriting able to avoid rule interference. At the end of this process a large portion of the whole treebank was automatically converted, with very few errors, leaving only a small amount of work to be done manually.