Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
We investigate mutual benefits between syntax and semantic roles using neural network models, by studying a parsingSRL pipeline, a SRLparsing pipeline, and a simple joint model by embedding sharing. The integration of syntactic and semantic features gives promising results in a Chinese Semantic Treebank, demonstrating large potentials of neural models for joint parsing and semantic role labeling.
The article presents methodological analyses of topical ideas of the famous modern linguist – E.Cosseriu. The authors argue that incorporation of theoretical ideas of E.Cosseriu could substentially extand the euristic potential of the conept of norm in the sphere of linguistics. The key to solvation of the problem lies in the necessity of changing of modern theoretical context of the question. The authors of the article consider the history of operationalization of the concept of norm in linguistics from the point of view of a specific hermeneutic approach in relation to other linguistic techniques. In the course of study of the role of linguistic norms in interaction of content and expression the article presents examples of extrapolation of the concept of norm from one discipline to another. In case of extrapolation of the concept of norm from the other disciplines into linguistic investigations the structural and functional dichotomy of language turns out that it is impossible to avoid considering spiritual as the world of objects, so that speech and language are considered as two different things. The rapid development of information technologies made possible to calculate many of the aspects of Humboldtian ideas. Computer statistics created conditions where the idea of "language picture of the world" and of the "inner form of the language" is gradually losing its original romantic charge and turns in a very trivial thing. Yet despite the fact that global standardization significantly enhances the processing and automatic addition of translation, which once gave beginning to hermeneutics the hermeneutic potential of linguistic norm still preseves many promissing prospects from the epistemological point of view. Key words: Linguistic norm and variability, E. Cosseriu, hermeneutics, euristic potential. В статье представлен методологический анализ актуальных идей известного современного лингвиста - E. Коссериу. Авторы утверждают, что введение теоретических идей E. Коссериу могло бы существенно расширить эвристический потенциал нормы в области лингвистики. Ключ к решению проблемы заключается в необходимости изменения современного теоретического контекста вопроса. Авторы статьи рассматривают историю ввода в действие понятия нормы в лингвистике с точки зрения конкретного герменевтического подхода по отношению к другим языковым методам. В ходе изучения роли языковых норм во взаимодействии содержания и выражения, в статье представлены примеры экстраполяции понятия нормы от одной дисциплины к другой. В случае экстраполяции понятия нормы с других дисциплин в лингвистические исследования структурно-функциональной дихотомии языка оказывается, что нельзя не рассматривать духовное как мир объектов, так как речь и язык рассматриваются как две разные вещи. Быстрое развитие информационных технологий сделало возможным для расчета многие аспекты гумбольдтовских идей. Статистика компьютера создала условия, в которых идея «языковой картины мира» и «внутренней формы языка» постепенно теряет свой первоначальный романтический заряд и превращается в очень тривиальную вещь. Тем не менее, несмотря на то, что глобальная стандартизация значительно улучшает обработку и автоматический перевод, который когда-то дал начало герменевтике, герменевтический потенциал языковой нормы хранит еще много перспектив с гносеологической точки зрения.Ключевые слова: лингвистическая норма и изменчивость, E.Коссериу, герменевтика, эвристический по- тенциал.
This paper presents the Universal Dependencies tagset (UD v1) as a new annotation scheme for Russian treebanks. The universal list of dependency relations was adopted and extended to comply with certain language-specific syntactic constructions. The tagset was validated, converting two Russian treebanks into the UD format, UD-Russian-SynTagRus and UD-Russian-Google.
Drawing on the work of Lawrence Abu Hamdan, a British-Lebanese artist and researcher currently based in Beirut, this essay examines the juridical and conceptual field of critical forensis which is situated at the juncture of security studies, art, and architecture. Abu Hamdan extends forensics to the area of “new audibilities,” with a focus on the politics of juridical hearing in situations of legal-identity profiling and voice authentication (the “shibboleth test”). Abu Hamdan's projects investigate how accent monitoring and audio surveillance, voice recognition, translation technologies, sovereign acts of listening, and court determinations of linguistic norms emerge as so many technical constraints on “freedom of speech,” itself a malleable term ascribed to discrepant claims and principles, yet taking on performative force in site-specific situations.
We present a new, sizeable dataset of nounnoun compounds with their syntactic analysis (bracketing) and semantic relations. Derived from several established linguistic resources, such as the Penn Treebank, our dataset enables experimenting with new approaches towards a holistic analysis of noun-noun compounds, such as jointlearning of noun-noun compounds bracketing and interpretation, as well as integrating compound analysis with other tasks such as syntactic parsing.
Objectives: This study investigated the role of response style biases in the assessment of positive and negative affect in aging research; it addressed whether response styles (a) are associated with age-related changes in cognitive abilities, (b) lead to distorted conclusions about age differences in affect, and (c) reduce the convergent and predictive validity of affect measures in relation to health outcomes. Method: A multidimensional item response theory model was used to extract response styles from affect ratings provided by respondents to the psychosocial questionnaire (n = 6,295; aged 50-100 years) in the Health and Retirement Study (HRS). Results: The likelihood of extreme response styles (disproportionate use of "not at all" and "very much" response categories) increased significantly with age, and this effect was mediated by age-related decreases in HRS cognitive test scores. Removing response styles from affect measures did not alter age patterns in positive and negative affect; however, it consistently enhanced the convergent validity (relationships with concurrent depression and mental health problems) and predictive validity (prospective relationships with hospital visits, physical illness onset) of the affect measures. Discussion: The results support the importance of detecting and controlling response styles when studying self-reported affect in aging research.
The paper evaluates the differences between two currently leading annotation schemes for dependency treebanks. By relying on four treebanks, we demonstrate that the treatment of conjunctions and adpositions represents the core difference between the two schemes and that this impacts the topological properties of the linguistic networks induced from the treebanks. We also show that such properties are reflected in the performances of four probabilistic dependency parsers trained on the treebanks. L’articolo valuta le differenze tra i due principali schemi di annotazione a dipenden-ze in uso. Sulla base di quattro treebank, l’articolo dimostra che il trattamento delle congiunzioni e delle pre/postposizioni rappresenta la differenza principale tra i due schemi e che ciò comporta delle conseguenze sulle proprietà topologiche dei net-work indotti dalle treebank. Inoltre, si dimostra come tali proprietà siano riflesse nell’accuratezza di quattro parser probabilistici a dipendenze addestrati sulle treebank.
Abstract We are investigating methods by which data from dependency syntax treebanks of ancient Greek can be applied to questions of authorship in ancient Greek historiography. From the Ancient Greek Dependency Treebank were constructed syntax words (sWords) by tracing the shortest path from each leaf node to the root for each sentence tree. This paper presents the results of a preliminary test of the usefulness of the sWord as a stylometric discriminator. The sWord data was subjected to clustering analysis. The resultant groupings were in accord with traditional classifications. The use of sWords also allows a more fine-grained heuristic exploration of difficult questions of text reuse. A comparison of relative frequencies of sWords in the directly transmitted Polybius book 1 and the excerpted books 9–10 indicate that the measurements of the two texts are generally very close, but when frequencies do vary, the differences are surprisingly large. These differences reveal that a certain syntactic simplification is a salient characteristic of Polybius’ excerptor, who leaves conspicuous syntactic indicators of his modifications.
PURPOSE: The focus of this study was to examine the influence of fundamental frequency (F0) and vocal tract length (VTL) modifications on speaker gender recognition in cochlear implant (CI) recipients for different stimulus types. METHOD: Single words and sentences were manipulated using isolated or combined F0 and VTL cues. Using an 11-point rating scale, CI recipients and listeners with normal hearing rated the maleness/femaleness of the corresponding voice. RESULTS: Speaker gender ratings for combined F0 and VTL modifications were similar across all stimulus types in both CI recipients and listeners with normal hearing, although the CI recipients showed a somewhat larger ambiguity. In contrast to listeners with normal hearing, F0-VTL and F0-only modifications revealed similar ratings in the CI recipients when using words as stimuli. However, when sentences were used, a difference was found between F0-VTL-based and F0-based ratings. Modifying VTL cues alone did not affect ratings in the CI group. CONCLUSIONS: Whereas speaker gender ratings by listeners with normal hearing relied on combined VTL and F0 cues, CI recipients made only limited use of VTL cues, which might be one reason behind problems with identifying the speaker on the basis of voice. However, use of the voice cues depended on stimulus type, with the greater information in sentences allowing a more detailed analysis than single words in both listener groups.
Statistical parsers are trained on treebanks that are composed of a few thousand sentences. In order to prevent data sparseness and computational complexity, such parsers make strong independence hypotheses on the decisions that are made to build a syntactic tree. These independence hypotheses yield a decomposition of the syntactic structures into small pieces, which in turn prevent the parser from adequately modeling many lexico-syntactic phenomena like selectional constraints and subcategorization frames. Additionally, treebanks are several orders of magnitude too small to observe many lexico-syntactic regularities, such as selectional constraints and subcategorization frames. In this article, we propose a solution to both problems: how to account for patterns that exceed the size of the pieces that are modeled in the parser and how to obtain subcategorization frames and selectional constraints from raw corpora and incorporate them in the parsing process. The method proposed was evaluated on French and on English. The experiments on French showed a decrease of 41.6% of selectional constraint violations and a decrease of 22% of erroneous subcategorization frame assignment. These figures are lower for English: 16.21% in the first case and 8.83% in the second.
It is well recognised in psychology that music has affective connotations and that musical stimuli can modify affective states. The aim of this study was to assess the affective connotations of 120 fifteen-second musical excerpts, covering both modern musical genres such as pop, rock, jazz, rap/R&B and electronic music (5 x N = 20), and classical music ( N = 20). Expert judges used predetermined criteria to select excerpts with positive or negative valence that induced high arousal or low arousal. The excerpts were assessed by 50 undergraduate students (25 women) from different academic departments, aged between 18 and 28 years ( M = 21.46 years, SD = 1.85). They listened to all 120 fragments and rated them with respect to six dimensions: valence, arousal, dominance, origin, subjective significance and imageability. Analyses showed that ratings were reliable, with high split-half correlations and Cronbach’s alpha estimates. We did not identify any gender differences concerning affective reactions to the music. Some music genre specificity was found for all measures, and initial music preference appeared to shape affective ratings. The results presented here will be of interest to researchers working on musical perception and the influence of music on affective outcomes and emotional regulation.
Treebanks have recently been released for a number of languages with the harmonized annotation created by the Universal Dependencies project. The representation of certain constructions in UD are known to be suboptimal for parsing and may be worth transforming for the purpose of parsing. In this paper, we focus on the representation of verb groups. Several studies have shown that parsing works better when auxiliaries are the head of auxiliary dependency relations which is not the case in UD. We therefore transformed verb groups in UD treebanks, parsed the test set and transformed it back, and contrary to expectations, observed significant decreases in accuracy. We provide suggestive evidence that improvements in previous studies were obtained because the transformation helps disambiguating POS tags of main verbs and auxiliaries. The question of why parsing accuracy decreases with this approach in the case of UD is left open.
We train a language-universal dependency parser on a multilingual collection of treebanks. The parsing model uses multilingual word embeddings alongside learned and specified typological information, enabling generalization based on linguistic universals and based on typological similarities. We evaluate our parser's performance on languages in the training set as well as on the unsupervised scenario where the target language has no trees in the training data, and find that multilingual training outperforms standard supervised training on a single language, and that generalization to unseen languages is competitive with existing model-transfer approaches.
This study investigates whether individual differences in attachment status can be detected by electrophysiological responses to loss-themed pictures. The Adult Attachment Interview (AAI) was used to identify discourse/reasoning lapses during the discussion of loss experiences via death that place speakers in the Unresolved/disorganized AAI category. In parents, Unresolved AAI status has been associated with Disorganized infant Strange Situation response, a known risk factor for psychopathology (e.g., internalizing/externalizing/dissociation). This association has been related to anomalous frightening (FR) parental behavior in the infant's presence, behavior presumed to be instigated by vulnerability to trauma-related fright. Here, psychophysiological methods were utilized to examine whether Unresolved AAI status could be detected in brain responses to subtle/symbolic reminders of loss. One year after AAI administration, 31 undergraduate women who had experienced loss (16 Unresolved) underwent continuous electroencephalogram (EEG) recording during a picture-viewing, valence-rating task. Picture onset-locked event-related potentials (ERPs) revealed millisecond responses to 4 picture categories: pleasant people, pleasant nature, cemetery (symbolic death), and gruesome death (dead or dying people). Participants' valence ratings did not differ between groups across picture categories. However, the N2 ERP, implicated in detecting stimulus salience, was selectively greater in Unresolved participants viewing cemetery scenes; it was in fact as high as the N2 for gruesome death images observed throughout the sample. Additionally, Unresolved participants exhibited a right-hemispheric P3 asymmetry across picture categories, suggestive of continuously heightened vigilance/arousal. Together, these results suggest that Unresolved AAI status is associated with greater neurophysiological sensitivity to subtle reminders of loss that may disrupt ongoing mental function. (PsycINFO Database Record
While linguistic theory posits an arbitrary relation between signifiers and the signified (de Saussure, 1916), our analysis of a large-scale German database containing affective ratings of words revealed that certain phoneme clusters occur more often in words denoting concepts with negative and arousing meaning. Here, we investigate how such phoneme clusters that potentially serve as sublexical markers of affect can influence language processing. We registered the EEG signal during a lexical decision task with a novel manipulation of the words' putative sublexical affective potential: the means of valence and arousal values for single phoneme clusters, each computed as a function of respective values of words from the database these phoneme clusters occur in. Our experimental manipulations also investigate potential contributions of formal salience to the sublexical affective potential: Typically, negative high-arousing phonological segments-based on our calculations-tend to be less frequent and more structurally complex than neutral ones. We thus constructed two experimental sets, one involving this natural confound, while controlling for it in the other. A negative high-arousing sublexical affective potential in the strictly controlled stimulus set yielded an early posterior negativity (EPN), in similar ways as an independent manipulation of lexical affective content did. When other potentially salient formal features at the sublexical level were not controlled for, the effect of the sublexical affective potential was strengthened and prolonged (250-650 ms), presumably because formal salience helps making specific phoneme clusters efficient sublexical markers of negative high-arousing affective meaning. These neurophysiological data support the assumption that the organization of a language's vocabulary involves systematic sound-to-meaning correspondences at the phonemic level that influence the way we process language.
Deaf or hard-of-hearing individuals usually face a greater challenge to learn to write than their normal-hearing counterparts. Due to the limitations of traditional research methods focusing on microscopic linguistic features, a holistic characterization of the writing linguistic features of these language users is lacking. This study attempts to fill this gap by adopting the methodology of linguistic complex networks. Two syntactic dependency networks are built in order to compare the macroscopic linguistic features of deaf or hard-of-hearing students and those of their normal-hearing peers. One is transformed from a treebank of writing produced by Chinese deaf or hard-of-hearing students, and the other from a treebank of writing produced by their Chinese normal-hearing counterparts. Two major findings are obtained through comparison of the statistical features of the two networks. On the one hand, both linguistic networks display small-world and scale-free network structures, but the network of the normal-hearing students' exhibits a more power-law-like degree distribution. Relevant network measures show significant differences between the two linguistic networks. On the other hand, deaf or hard-of-hearing students tend to have a lower language proficiency level in both syntactic and lexical aspects. The rigid use of function words and a lower vocabulary richness of the deaf or hard-of-hearing students may partially account for the observed differences.
The present study evaluated the efficacy of adding a virtual reality (VR) component to the treatment of compulsive hoarding (CH), following inference-based therapy (IBT). Participants were randomly assigned to either an experimental or a control condition. Seven participants received the experimental and seven received the control condition. Five sessions of 1 h were administered weekly. A significant difference indicated that the level of clutter in the bedroom tended to diminish more in the experimental group as compared to the control group F(2,24) = 2.28, p = 0.10. In addition, the results demonstrated that both groups were immersed and present in the environment. The results on posttreatment measures of CH (Saving Inventory revised, Saving Cognition Inventory and Clutter Image Rating scale) demonstrate the efficacy of IBT in terms of symptom reduction. Overall, these results suggest that the creation of a virtual environment may be effective in the treatment of CH by helping the compulsive hoarders take action over their clutter.
Abstract syntax is a semantic tree representation that lies between parse trees and logical forms. It abstracts away from word order and lexical items, but contains enough information to generate both surface strings and logical forms. Abstract syntax is commonly used in compilers as an intermediate between source and target languages. Grammatical Framework (GF) is a grammar formalism that generalizes the idea to natural languages, to capture cross-lingual generalizations and perform interlingual translation. As one of the main results, the GF Resource Grammar Library (GF-RGL) has implemented a shared abstract syntax for over 30 languages. Each language has its own set of concrete syntax rules (morphology and syntax), by which it can be generated from the abstract syntax and parsed into it. This paper presents a conversion method from abstract syntax trees to dependency trees. The method is applied for converting GF-RGL trees to Universal Dependencies (UD), which uses a common set of labels for different languages. The correspondence between GF-RGL and UD turns out to be good, and the relatively few discrepancies give rise to interesting questions about universality. The conversion also has potential for practical applications: (1) it makes the GF parser usable as a rule-based dependency parser; (2) it enables bootstrapping UD treebanks from GF treebanks; (3) it defines formal criteria to assess the informal annotation schemes of UD; (4) it gives a method to check the consistency of manually annotated UD trees with respect to the annotation schemes; (5) it makes information from UD treebanks available.
Abstract Three studies examined gender differences in the effect of storytelling ability on perceptions of a person's attractiveness as a short‐term and long‐term romantic partner. In Study 1, information about a potential partner's storytelling ability was provided. Study 2 participants read a good or poor story supposedly written by a potential partner. Results suggested that only women's attractiveness assessments of men as a long‐term date increased for good storytellers. Storytelling ability did not affect men's ratings of women nor did it affect ratings of short‐term partners. Study 3 suggested that the effect of storytelling ability on long‐term attractiveness for male targets may be mediated by perceived status. Storytelling ability appears to increase perceived status and thus helps men attract long‐term partners.
We propose a framework to model human comprehension of discourse connectives. Following the Bayesian pragmatic paradigm, we advocate that discourse connectives are interpreted based on a simulation of the production process by the speaker, who, in turn, considers the ease of interpretation for the listener when choosing connectives. Evaluation against the sense annotation of the Penn Discourse Treebank confirms the superiority of the model over literal comprehension. A further experiment demonstrates that the proposed model also improves automatic discourse parsing.
The Universal Dependencies (UD) Project seeks to build a cross-lingual studies of treebanks, linguistic structures and parsing. Its goal is to create a set of multilingual harmonized treebanks that are designed according to a universal annotation scheme. In this paper, we report on the conversion of the Uyghur dependency treebank to a UD version of the treebank which we term the Uyghur Universal Dependency Treebank (UyDT). We present the mapping of the Uyghur dependency treebank’s labelling scheme to the UD scheme, along with a clear description of the structural changes required in this conversion.
In accordance with the compositionality criterion and hierarchy principle of Rhetorical Structure Theory (RST), this study reframes each tree in the RST Discourse Treebank into three new dependency trees with ultimate nodes being clauses, sentences, and paragraphs, respectively, which also draw on an analogy between syntactic and discourse trees. Detailed percentages of various RST relations at the three granularity levels are examined, illuminating the discourse processes of organizing units of one granularity level into those of the next upper level and suggesting certain homogeneity and interaction across levels in the Treebank, particularly at the two upper levels. The study demonstrates the applicability of RST analysis between same-level terminal units. With unique analytical advantages, the newly constructed discourse dependency trees provide new research prospects.
We propose a classification framework for semantic type identification of compounds in Sanskrit. We broadly classify the compounds into four different classes namely, Avyayībhāva, Tatpuruṣa, Bahuvrīhi and Dvandva. Our classification is based on the traditional classification system followed by the ancient grammar treatise Adṣṭādhyāyī, proposed by Pāṇini 25 centuries back. We construct an elaborate features space for our system by combining conditional rules from the grammar Adṣṭādhyāyī, semantic relations between the compound components from a lexical database Amarakoṣa and linguistic structures from the data using Adaptor Grammars. Our in-depth analysis of the feature space highlight inadequacy of Adṣṭādhyāyī, a generative grammar, in classifying the data samples. Our experimental results validate the effectiveness of using lexical databases as suggested by Amba Kulkarni and Anil Kumar, and put forward a new research direction by introducing linguistic patterns obtained from Adaptor grammars for effective identification of compound type. We utilise an ensemble based approach, specifically designed for handling skewed datasets and we %and Experimenting with various classification methods, we achieve an overall accuracy of 0.77 using random forest classifiers.
The Internet-scale open source software (OSS) production in various communities are generating abundant reusable resources for software developers. However, how to retrieve and reuse the desired and mature software from huge amounts of candidates is a great challenge: there are usually big gaps between the user application contexts (that often used as queries) and the OSS key words (that often used to match the queries). In this paper, we define the scenario-based query problem for OSS retrieval, and then we propose a novel approach to reformulate the raw query by leveraging the crowd wisdom from millions of developers to improve the retrieval results. We build a software-specific domain lexical database based on the knowledge in open source communities, by which we can expand and optimize the input queries. The experiment results show that, our approach can reformulate the initial query effectively and outperforms other existing search engines significantly at finding mature software.
Prague Czech-English Dependency Treebank - Russian translation (PCEDT-R) is a project of translating a subset of Prague Czech-English Dependency Treebank 2.0 (PCEDT 2.0) to Russian and linguistically annotating the Russian translations with emphasis on coreference and cross-lingual alignment of coreferential expressions. Cross-lingual comparison of coreference means is currently the purpose that drives development of this corpus. The current version 0.5 is a preliminary version, which contains (+ denotes new features): * complete PCEDT 2.0 documents "wsj_1900"-"wsj_1949" * Czech-English word alignment of coreferential expressions annotated manually mainly on the t-layer + Russian translations of the original English sentences + automatic tokenization, part-of-speech tagging and morphological analysis for Russian + automatic word alignment between all Czech and Russian words + manual alignment between Russian and the other two languages on possessive pronouns
Background: High intensity interval training (HIIT) is a robust and time-efficient approach to improve multiple health indices including maximal oxygen uptake (VO2max). Despite the intense nature of HIIT, data in untrained adults report greater enjoyment of HIIT versus continuous exercise (CEX). However, this has yet to be investigated in persons with spinal cord injury (SCI).Objective: To examine differences in enjoyment in response to CEX and HIIT in persons with SCI.Design: Repeated measures, within-subjects design.Setting: University laboratory in San Diego, CA.Participants: Nine habitually active men and women (age = 33.3 ± 10.5 years) with chronic SCI.Intervention: Participants performed progressive arm ergometry to volitional exhaustion to determine VO2peak. During subsequent sessions, they completed CEX, sprint interval training (SIT), or HIIT in randomized order.Outcome Measures: Physical activity enjoyment (PACES), affect, rating of perceived exertion (RPE), VO2, and blood lactate concentration (BLa) were measured.Results: Despite a higher VO2, RPE, and BLa consequent with HIIT and SIT (P < 0.05), PACES was significantly higher (P = 0.03) in response to HIIT (107.4 ± 13.4) and SIT (103.7 ± 12.5) compared to CEX (81.6 ± 25.4). Fifty-five percent of participants preferred HIIT and 45% preferred SIT, with none identifying CEX as their preferred exercise mode.Conclusion: Compared to CEX, brief sessions of submaximal or supramaximal interval training elicit higher enjoyment despite higher metabolic strain. The long-term efficacy and feasibility of HIIT in this population should be explored considering that it is not viewed as more aversive than CEX.
abstract Metaphors are comparisons that link dissimilar conceptual domains. We hypothesized that the aptness of a metaphor is linked to the reader’s experience of beauty, and that age and expertise influence these aesthetic judgments. We had young adults, literary experts, and elderly adults rate metaphors for beauty or aptness. Experimental materials consisted of single-sentence novel metaphors whose familiarity, figurativeness, imageability, interpretability, and overall valence ratings were known. Results suggest that beauty and aptness of metaphors are linked for elderly adults but are orthogonal for young adults and literary experts. Elderly participants seem to conflate emotional content with aptness. Young adults are most swayed by a perceived feeling of familiarity when rating for aptness, but not for beauty. Literary experts are relatively unaffected by the psycholinguistic variables, suggesting an emotionally distanced approach to these sentences. Individual differences in literary training and life experience have varying effects on the aesthetic experience of metaphor in regard to beauty and aptness.
Parsing texts into universal dependencies (UD) in realistic scenarios requires infrastructure for the morphological analysis and disambiguation (MA&D) of typologically different languages as a first tier. MA&D is particularly challenging in morphologically rich languages (MRLs), where the ambiguous space-delimited tokens ought to be disambiguated with respect to their constituent morphemes, each morpheme carrying its own tag and a rich set features. Here we present a novel, language-agnostic, framework for MA&D, based on a transition system with two variants — word-based and morpheme-based — and a dedicated transition to mitigate the biases of variable-length morpheme sequences. Our experiments on a Modern Hebrew case study show state of the art results, and we show that the morpheme-based MD consistently outperforms our word-based variant. We further illustrate the utility and multilingual coverage of our framework by morphologically analyzing and disambiguating the large set of languages in the UD treebanks.
Statistical parsers are e ective but are typically limited to producing projective dependencies or constituents. On the other hand, linguisti- cally rich parsers recognize non-local relations and analyze both form and function phenomena but rely on extensive manual grammar development. We combine advantages of the two by building a statistical parser that produces richer analyses. We investigate new techniques to implement treebank-based parsers that allow for discontinuous constituents. We present two systems. One system is based on a string-rewriting Linear Context-Free Rewriting System (LCFRS), while using a Probabilistic Discontinuous Tree Substitution Grammar (PDTSG) to improve disambiguation performance. Another system encodes the discontinuities in the labels of phrase structure trees, allowing for efficient context-free grammar parsing.The two systems demonstrate that tree fragments as used in tree-substitution grammar improve disambiguation performance while capturing non-local relations on an as-needed basis. Additionally, we present results of models that produce function tags, resulting in a more linguistically adequate model of the data. We report substantial accuracy improvements in discontinuous parsing for German, English, and Dutch, including results on spoken Dutch.
Text diacritization is a critical task which plays an important role for improving the performance of many NLP tasks for languages that include diacritics in their orthographies. In this paper, we handle the problem of Arabic text diacritization such that our system diacritize input Arabic sequence of words both morphologically and syntactically. The operation of the system is divided into three layers: the first layer uses HMM for the morphological diacritization of previously seen words, the second layer uses an external morphological analyzer for the morphological diacritization of OOV words, and the third layer uses CRF for the syntactic diacritization of all words. To evaluate the performance of the system, we used the benchmark LDC Arabic Treebank Part 3 datasets used by the state-of-the-art systems. The proposed system achieved a morphological WER of 4.3%, and a syntactic WER of 9.4%.
In recent years, dependency parsing is a fascinating research topic and has a lot of applications in natural language processing. In this paper, we present an effective approach to improve dependency parsing by utilizing supertag features. We performed experiments with the transition-based dependency parsing approach because it can take advantage of rich features. Empirical evaluation on Vietnamese Dependency Treebank showed that, we achieved an improvement of 18.92% in labeled attachment score with gold supertags and an improvement of 3.57% with automatic supertags.
The Open Linguistics Working Group (OWLG) brings together researchers from various fields of linguistics, natural language processing, and information technology to present and discuss principles, case studies, and best practices for representing, publishing and linking linguistic data collections. A major outcome of our work is the Linguistic Linked Open Data (LLOD) cloud, an LOD (sub-)cloud of linguistic resources, which covers various linguistic databases, lexicons, corpora, terminologies, and metadata repositories. We present and summarize five years of progress on the development of the cloud and of advancements in open data in linguistics, and we describe recent community activities. The paper aims to serve as a guideline to introduce and involve researchers with the community and more generally with Linguistic Linked Open Data.
Sequential LSTM has been extended to model tree structures, giving competitive results for a number of tasks. Existing methods model constituent trees by bottom-up combinations of constituent nodes, making direct use of input word information only for leaf nodes. This is different from sequential LSTMs, which contain reference to input words for each node. In this paper, we propose a method for automatic head-lexicalization for tree-structure LSTMs, propagating head words from leaf nodes to every constituent node. In addition, enabled by head lexicalization, we build a tree LSTM in the top-down direction, which corresponds to bidirectional sequential LSTM structurally. Experiments show that both extensions give better representations of tree structures. Our final model gives the best results on the Standford Sentiment Treebank and highly competitive results on the TREC question type classification task.
Olfactory identification abilities in adolescents have been reported inferior compared with adults. Though this seems to be the case when comparing identification abilities using tests validated on-and for-adults, odor familiarity has been hypothesized to affect identification abilities in younger participants. However, this has never been thoroughly tested. The aims of this study were to investigate patterns in odor familiarity differences between adolescents and adults, and to investigate if an adolescent familiarity-based modification of an identification test could lead to similar identification scores in adolescents and adults. In total, 411 adolescent participants and 320 adult participants were included in the study. Odor familiarity ratings were obtained for 125 odors. A modified version of the "Sniffin' Sticks" identification test was created and validated on 72 adolescents based on adolescent familiarity scores. This test was applied to 82 normosmic adults and 167 normosmic adolescents. Results show a lower familiarity for spices and environmental odors, and a higher familiarity for candy odors in adolescents. The identification abilities in adults and adolescents were equal after familiarity-based modification. We conclude that changes in odor familiarity from adolescence to adulthood do not develop evenly for all odors, but are dependent on odor-object category.
With the advent of word embeddings, lexicons are no longer fully utilized for sentiment analysis although they still provide important features in the traditional setting. This paper introduces a novel approach to sentiment analysis that integrates lexicon embeddings and an attention mechanism into Convolutional Neural Networks. Our approach performs separate convolutions for word and lexicon embeddings and provides a global view of the document using attention. Our models are experimented on both the SemEval'16 Task 4 dataset and the Stanford Sentiment Treebank, and show comparative or better results against the existing state-of-the-art systems. Our analysis shows that lexicon embeddings allow to build high-performing models with much smaller word embeddings, and the attention mechanism effectively dims out noisy words for sentiment analysis.
International audience
We present a study on two key characteristics of human syntactic annotations: anchoring and agreement. Anchoring is a well known cognitive bias in human decision making, where judgments are drawn towards pre-existing values. We study the influence of anchoring on a standard approach to creation of syntactic resources where syntactic annotations are obtained via human editing of tagger and parser output. Our experiments demonstrate a clear anchoring effect and reveal unwanted consequences, including overestimation of parsing performance and lower quality of annotations in comparison with human-based annotations. Using sentences from the Penn Treebank WSJ, we also report systematically obtained inter-annotator agreement estimates for English dependency parsing. Our agreement results control for parser bias, and are consequential in that they are on par with state of the art parsing performance for English newswire. We discuss the impact of our findings on strategies for future annotation efforts and parser evaluations.
We train one multilingual model for dependency parsing and use it to parse sentences in several languages. The parsing model uses (i) multilingual word clusters and embeddings; (ii) token-level language information; and (iii) language-specific features (fine-grained POS tags). This input representation enables the parser not only to parse effectively in multiple languages, but also to generalize across languages based on linguistic universals and typological similarities, making it more effective to learn from limited annotations. Our parser’s performance compares favorably to strong baselines in a range of data scenarios, including when the target language has a large treebank, a small treebank, or no treebank for training.
The PDTB Annotator is a tool for annotating and adjudicating discourse relations based on the annotation framework of the Penn Discourse TreeBank (PDTB). This demo describes the benefits of using the PDTB Annotator, gives an overview of the PDTB Framework and discusses the tool’s features, setup requirements and how it can also be used for adjudication.
In this paper, a German verb resource for verb-centered sentiment inference is introduced and evaluated. Our model specifies verb polarity frames that capture the polarity effects on the fillers of the verb's arguments given a sentence with that verb frame. Verb signatures and selectional restrictions are also part of the model. An algorithm to apply the verb resource to treebank sentences and the results of our first evaluation are discussed.