Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
In this paper, we propose a new annotation approach to Chinese word segmentation, part-of-speech (POS) tagging and dependency labelling that aims to overcome the two major issues in traditional morphology-based annotation: Inconsistency and data sparsity. We re-annotate the Penn Chinese Treebank 5.0 (CTB5) and demonstrate the advantages of this approach compared to the original CTB5 annotation through word segmentation, POS tagging and machine translation experiments.
Natural language processing (NLP) refers to the study of systems performing natural language related tasks in an automatic manner, that is, without human supervision or interference. This thesis work considers NLP problems related to morphology analysis, that is, the description of internal structure of words. Acquiring knowledge of morphology is necessary in order for applications, such as search engines, machine translators, and speech recognizers, to successfully address rare and previously unseen word forms. In particular, we focus on two widely applied morphological analysis tasks, namely, morphological tagging and segmentation. In morphological tagging, the aim is to assign words in sentential contexts with word class labels describing their morphological properties. Meanwhile, morphological segmentation considers describing the inner word structure by splitting word forms into their smallest meaning-bearing units, morphemes. \nIn the scope of this thesis, we approach the morphological tagging and segmentation problems using statistical, data-driven machine learning methodology. Using this approach, the processing systems are learned (estimated) based on training data prepared manually by a human expert. In particular, we focus on the highly influential conditional random field (CRF) model proposed for sequence tagging and segmentation in the early 2000s. \nAs the first main contribution, the thesis discusses data-driven morphological segmentation employing the CRF model. A particular emphasis is placed on the semi-supervised learning setting, in which the available data consists of a small number of annotated segmentation examples and a large amount of unannotated raw word forms. The provided empirical evaluation on six languages shows that the proposed semi-supervised CRF-based approach is highly successful in the considered morphological segmentation task compared to earlier methods. In particular, the performed error analysis shows that closed class phenomena, such as suffixation of English and Finnish, can be learned already from a small number of annotated examples in a supervised manner. Meanwhile, open morpheme class phenomena, such as compounding of Finnish, can be learned by additionally exploiting the large unannotated word list using the semi-supervised approach. \nAs the second main contribution, the thesis contains a presentation of FinnPos, the first open-source statistical morphological tagging and lemmatization toolkit designed specifically for Finnish. The CRF-based FinnPos system is readily applicable for tagging and lemmatization of running text with models learned from the recently published Finnish Turku Dependency Treebank and FinnTreeBank.
Accurate automatic processing of Web queries is important for high-quality information retrieval from the Web. While the syntactic structure of a large portion of these queries is trivial, the structure of queries with question intent is much richer. In this paper we therefore address the task of statistical syntactic parsing of such queries. We first show that the standard dependency grammar does not account for the full range of syntactic structures manifested by queries with question intent. To alleviate this issue we extend the dependency grammar to account for segments -independent syntactic units within a potentially larger syntactic structure. We then propose two distant supervision approaches for the task. Both algorithms do not require manually parsed queries for training. Instead, they are trained on millions of (query, page title) pairs from the Community Question Answering (CQA) domain, where the CQA page was clicked by the user who initiated the query in a search engine. Experiments on a new treebank 1 consisting of 5,000 Web queries from the CQA domain, manually parsed using the proposed grammar, show that our algorithms outperform alternative approaches trained on various sources: tens of thousands of manually parsed OntoNotes sentences, millions of unlabeled CQA queries and thousands of manually segmented CQA queries.
This volume takes as its central organizing principle the foundational understanding about community knowledge that challenges “narrow conceptions of language, literacy, personal stories, bounded or contained learning contexts (e.g., home, community, schools), hegemonic cultural and linguistic norms, quantitative and static views of ‘resources,’ and limited attention to the agency, identities, and strategic actions of diverse students and their families as they traverse contexts” (daSilva Iddings, this volume). The touchstone to this approach is “Funds of Knowledge” as it has been conceptualized for nearly twenty-five years. This chapter will briefly summarize the approach as it has evolved and will lay out the programmatic implementation as it unfolded within CREATE.
Due to the constant increasing of electronic textual information, modern society needs for the automatic processing of natural language (NL). The main purpose of NL automatic text processing systems is to analyze and create texts and represent their content. The purpose of the paper is the development of linguistic and software bases of an automatic system for processing English publicistic texts. This article discusses the examples of different approaches to the creation of linguistic databases for processing systems. The author gives a detailed description of basic building blocks for a new linguistic processor: lexicalsemantic, syntactical and semantic-syntactical. The main advantage of the processor is using special semantic codes in the alphabetical dictionary. The semantic codes have been developed in accordance with a lexical-semantic classification. It helps to precisely define semantic functions of the keywords that are situated in parsing groups and allows the automatic system to avoid typical mistakes. The author also represents the realization of a developed linguistic database in the form of a training computer program.
One consequence of ignoring visual objects in our environment is that we subsequently like those objects less; a stimulus-devaluation effect linked to attentional inhibition. While such 'inhibitory devaluation' has been well studied using perceptual tasks, little is known about the affective consequences of inhibition at later stages of representation. In the present experiment we combined behavioral and electrophysiological measures to examine inhibitory devaluation of items maintained in visual working memory (VWM). We specifically used the contralateral delay activity (CDA) event-related potential to investigate the immediate consequences of ignoring objects stored in VWM and whether changes in CDA amplitude are associated with the subsequent devaluation of ignored objects. Each trial of the experiment consisted of two tasks: a VWM test and a stimulus affective rating. Participants memorized three colored squares located on one side of a lateralized memory array. During the retention interval, a retro-cue was presented specifying with 100% validity that their memory of the cued item would be tested, and that the two un-cued items could be ignored. We found significant devaluation of un-cued stimuli relative to cued stimuli. Moreover, individual differences in the magnitude of devaluation were correlated with post retro-cue CDA amplitude (r=-.60), and with an earlier negative-going potential that resembled the latency and scalp distribution of an attention-related N2pc component time-locked to the retro-cue (r=-.59). However, differences in stimulus devaluation did not correlate with components prior to the retro-cue. Together, these electrophysiological results converge with our recent demonstration that the affective consequences of inhibition are the same for items represented solely in visual working memory as they are for sensory stimuli appearing in the external environment. More broadly, these results add to the growing literature comparing internal and external attentional mechanisms by demonstrating that attentional inhibition leads to the devaluation of distracting stimuli within both domains. Meeting abstract presented at VSS 2016
The study examined the relationship between idiom familiarity, knowledge of idiom meaning and idiom transparency judgments in L2. A group of 23 intermediate Japanese learners of English were asked to provide familiarity ratings, transparency judgments, and definitions for 30 English idioms, 27 of which had semantically equivalent but compositionally different idiomatic counterparts in Japanese and 3 phrases for which semantic equivalents in L1 also shared the same structural properties. Transparency ratings were repeated after the instructional treatment. A comparison of pre-treatment and post-treatment transparency scores showed that knowledge of conventional idiom meanings had a strong effect on the learners’ perceptions of idiom transparency. Transparency judgments, however, were not found to be a reliable predictor of the learners’ ability to infer figurative meanings of the idiomatic phrases. Idiom familiarity was not found to have a significant effect on idiom comprehension or on transparency judgments either. A limited positive effect of language transfer on L2 idiom comprehension and transparency ratings was observed.
In many natural language processing (NLP) tasks, a document is commonly modeled as a bag of words using the term frequency-inverse document frequency (TF-IDF) vector. One major shortcoming of the frequency-based TF-IDF feature vector is that it ignores word orders that carry syntactic and semantic relationships among the words in a document, and they can be important in some NLP tasks such as genre classification. This paper proposes a novel distributed vector representation of a document: a simple recurrent-neural-network language model (RNN-LM) or a long short-term memory RNN language model (LSTM-LM) is first created from all documents in a task; some of the LM parameters are then adapted by each document, and the adapted parameters are vectorized to represent the document. The new document vectors are labeled as DV-RNN and DV-LSTM respectively. We believe that our new document vectors can capture some high-level sequential information in the documents, which other current document representations fail to capture. The new document vectors were evaluated in the genre classification of documents in three corpora: the Brown Corpus, the BNC Baby Corpus and an artificially created Penn Treebank dataset. Their classification performances are compared with the performance of TF-IDF vector and the state-of-the-art distributed memory model of paragraph vector (PV-DM). The results show that DV-LSTM significantly outperforms TF-IDF and PV-DM in most cases, and combinations of the proposed document vectors with TF-IDF or PV-DM may further improve performance.
Facial expressions frequently involve multiple individual facial actions. How do facial actions combine to create emotionally meaningful expressions? Infants produce positive and negative facial expressions at a range of intensities. It may be that a given facial action can index the intensity of both positive (smiles) and negative (cry-face) expressions. Objective, automated measurements of facial action intensity were paired with continuous ratings of emotional valence to investigate this possibility. Degree of eye constriction (the Duchenne marker) and mouth opening were each uniquely associated with smile intensity and, independently, with cry-face intensity. In addition, degree of eye constriction and mouth opening were each unique predictors of emotion valence ratings. Eye constriction and mouth opening index the intensity of both positive and negative infant facial expressions, suggesting parsimony in the early communication of emotion.
Syntactic information, obtainable through syntactical analysis, plays an important role in many areas of NLP. Researches of Indonesian constituent parser have been very limited with the currently available yields very poor performance of 38.89% and 47.22% using Earley and CYK algorithm respectively. With the availability of the newly introduced Indonesian treebank corpus, we evaluate the performance of Indonesian constituent parser using Trance parser, a language independent constituent parser which employs deep learning. The parser achieved a respectable f-score of 74.91%, a very significant improvement to previous researches.
This article invetigates some of the most relevant issues concerning the relation between language and gender, understood as the social elaboration of the properties associated to the sex. It is in this perspective that we can interpret the notion of 'linguistic sexism', that is the whjole of the linguistic devices that the literature deals with as indicators of the different roles and the different power associated to the sexual differences. The analysis of the relation between language and gender naturally implies a more general reflection on the relation between language and thought and on the link between language and spciety, and between linguistic norm and linguistic use.
Krokodyl is an experimental hybrid deep depencency parser of Polish. Krokodyl has been developed at the Institute of Computer Science, Polish Academy of Sciences (IPI PAN) within the CLARIN-PL project. It was create to evaluate a hybrid approach to parsing: combining syntactic, lexical and semantic features for dependency parsing. It uses a number of tools as components of the feature generation chain, namely the Spejd Grammar, MALT parsing engine, MATE, the Polish Wordnet, the Skladnica treebank.
Latin VALLEX is a valency lexicon for Latin. It was built in close connection with the semantic/pragmatic annotation of the Index Thomisticus Treebank and the Latin Dependency Treebank. Data are stored in a single XML file, whose structure is the same of that for the valency lexicon for Czech PDT-VALLEX.
This article proposes an ontology design pattern for leading knowledge providers to represent knowledge in more normalized, precise and interrelated ways, hence in ways that help the matching and exploitation of knowledge from different sources. This pattern is a knowledge sharing best practice that is domain and language independent. It can be used as a criteria for measuring the quality of an ontology. This pattern is: using binary relation types directly derived from concept types, especially role types or types of process. The article explains and illustrates this pattern, and relates it to other patterns and general ontology quality criteria. It also provides an ontology for automatically deriving relation types from concept types (e.g., those from lexical ontologies such as those derived from the WordNet lexical database). This derivation helps normalizing knowledge, reduces having to introduce new relation types and helps keeping all the types organized.
One of the widely used approaches to Sentiment Analysis (SA) is lexicon-based approach that depends on sentiment-annotated lexical resources (such as SentiWordNet (SWN)). A broad variety of such resources are Synsetbased Lexical Databases (SLDs) (e.g. SWN is based on WordNet (WN)) and represent sentiment degrees of synonym groups of LDs, called "synsets." However, synsets themselves were open to criticism because although, in reality, not all the members of a synset represent its meaning with the same degree, in SLDs, they are, identically, considered as members of their synset. Therefore, the fuzzy version of synsets was proposed in a small number of previous studies. Fuzzy synsets can upgrade such lexicon-based SA by which the future SA systems can discriminate between word-senses of a same synset, how much each of them contains the sentiment load of that synset. But, to the best of our knowledge, none of the studies on fuzzy synsets has proposed any algorithm for providing fuzzy versions of "predefined synsets" of an SLD. In this study, we present the idea of an algorithm for constructing fuzzy version of any SLD of any language, given a corpus of that language and a word-sense-disambiguation system of that language/SLD.
The internship is a critical part of graduate training and often the only opportunity to receive on-site clinical supervision during school psychology practice. Nonetheless, the process of pairing interns with field supervisors is not standardized and sometimes relies on factors such as logistics and supervisor credentials rather than a consideration of interpersonal variables that could optimize the internship experience. Related fields have found mixed evidence for a relationship between personality similarity within a supervisory dyad and outcomes such as a strong supervisory relationship, satisfaction with supervision, and supervisee effectiveness. This study examined the influence of personality similarity on ratings of supervisory working alliance, supervision satisfaction, and intern work readiness. This study also evaluated the predictive power of personality, supervisory working alliance, and systemic factors on intern work readiness and supervision satisfaction. Lastly, this study assessed the development of the supervisory working alliance and intern work readiness over time. Twenty-six dyads were recruited for participation in this study, including 24 practicing school psychologists serving as field supervisors and 26 school psychology interns. Data collection occurred at the midpoint and end of the internship year. Participants completed a demographic questionnaire, personality inventory, and measures of supervisory working alliance, supervision satisfaction, supervisee work readiness, and systemic factors. Results indicated that personality similarity among supervisors and interns is not related to supervisory working alliance, supervision satisfaction, or supervisee work readiness. However, supervisor ratings of supervisory working alliance were predictive of intern work readiness, and intern ratings of supervisory working alliance were predictive of supervision satisfaction. Systemic factors were not predictive of intern work readiness or supervision satisfaction. For supervisors, the supervisory working alliance significantly decreased over time, while intern ratings remained consistent from midyear to the end of the year. Intern development from midyear to the end of year could not be determined due to low scale reliability. Future studies should further examine factors that contribute to the supervisory working alliance and validate measures specific to the school context. More research is needed to establish the conditions and interpersonal characteristics that enable an optimal internship experience for both supervisors and supervisees in school psychology.
International audience
The paper compares the grammar handbook <i>Gyakorlati Ilir Nyelvtan</i> (Baja, 1874, <sup>2</sup>1881) by Mihálovics with Mažuranić's <i>Slovnica Hèrvatska</i> (<sup>4</sup>1869), as Mažuranić is mentioned in the foreword as the normative model used in the handbook. This comparison will include their terminology, purpose, structure, and normative prescription, and will determine in which cases Mihálovics follows Mažuranić’s grammar and in which he distances himself from it. Since the grammar handbook was published outside the Croatian (ethnic and linguistic) area, the paper will show to what extent the characteristics of the Croatian linguistic norm were preserved in the Hungarian part of the Danube Region in the late 19<sup>th</sup> century.
Research finds we make spontaneous trait inferences from facial appearance, even after brief exposures to a face (i.e., less than or equal to 100 ms). We examined spontaneous impressions of criminality from facial appearance, testing whether these impressions persist after repeated presentation (i.e., one to three exposures) and increased exposure duration (100, 500, or 1,000 ms) to the face. Judgement confidence and response times were recorded. Other participants viewed the faces for an unlimited period of time, rating trustworthiness, dominance and criminal appearance. We found evidence that participants spontaneously make criminal appearance attributions. These inferences persisted with repeated presentation and increased exposure duration, were related to trustworthiness and dominance ratings, and were made with high confidence. Implications are discussed.
Integration of models requires linking of components, which may be developed by different teams, using different tools, methodologies, and assumptions. Participating models may operate at different temporal and spatial scales. We describe and discuss the design and prototype of the Distributed Model Integration Framework (DMIF) that links models, which can be deployed on different hardware and software platforms. Distributed computing and service-oriented software development approaches are utilized to address the different aspects of interoperability. Web services are used to enable technical interoperability between models. To illustrate its operation, we developed reusable web service wrappers for models developed in NetLogo and GAMS based modeling languages. We also demonstrated that some of semantic mediation tasks can be handled by using openly available ontologies, and that this technique helps to avoid significant amount of reinvention by different framework developers. We investigated automated semantic mapping of text-based input-output data and attribute names of components using direct semantic matching algorithms and using an openly available lexical database. We found that for short text-based input-output data and attribute names of components direct semantic matching algorithms work much better than applying a lexical database. This holds true for both standardized and non-standardized short text and this is mainly because short text does not include contextual information of data. Furthermore, direct semantic matching algorithms can be applied to search for components that can possibly provide data for a given component (1) if a model repository uses standard names for attributes of components and (2) if metadata of components are made available through an API. As a proof of concept we implemented our design to integrate climate-energy-economy models. Our design can be applied by different modeling groups to link a wide range of models, and it can improve the reusability of models by making them available on the web.
Background Conventionally, it is believed that high-frequency auditory information is important for speech understanding. This is only partly true, as recent studies have demonstrated the importance of low-frequency information. This research was taken up to develop, standardize, and validate auditory low-frequency word lists in Hindi, an Indian language. Material and Methods The first phase of the study involved collection of bisyllabic words followed by verification by a native linguist. Words were then short-listed based on familiarity ratings given by 10 adult native speakers; those words were recorded and the best recorded words selected through subjective and objective analysis. Then, using Fast Fourier Transform and k-means clustering, words with more energy below 1.5 kHz were isolated. Finally, equally difficult 10 word lists were generated by obtaining psychometric function curves. Finally, lists were administered on 40 adult normal hearing particip Results Results showed a similar trend of increase in speech identification scores with increase in SL across all lists except list 4. During the final phase, developed lists were validated on 10 simulated low-frequency cochlear hearing loss participants. Hearing loss was simulated using Matlab and National Institute for Occupational Safety and Health (NIOSH) software. Results of validation revealed that auditory low-frequency word lists were sensitive enough to tap the speech understanding difficulty in the simulated condition. Conclusions The developed word lists can be used clinically to assess communication ability in individuals with rising hearing loss. The word lists also have the potential to assess the performance after amplification provided to individuals with rising hearing loss.
In NLP data drives research, as evidenced by the frequency with which seminal works of database engineering such as The Penn Treebank have been employed as a basis for experimentation. Traditionally large-scale expertly annotated corpora are expensive and time consuming to produce. This paradigm drove researchers to adopt automated methods for generating labelled data with available tools such as Freebase, DBpedia, and the "infoboxes" found on Wikipedia pages. These knowledge bases have been, or are in the process of being, subsumed by Wikidata, an initiative to concentrate such disparate data repositories in an organized machine readable format. This resource is an important research tool. In this paper, we review our experience using Wikidata in constructing a large annotated corpus under distant supervision, moreover we make the materials, the code used to generate our annotations, freely available to all interested parties.
Abstract Drawing on an analogy between discourse and syntactic trees, this paper chooses 359 Wall Street Journal articles with multiple paragraphs from the Rhetorical Structure Theory (RST) Discourse Treebank, and converts each discourse tree into three additional dependency ones, at discourse, paragraph and sentence levels, with exclusively elementary discourse units of clauses, sentences and paragraphs, respectively. It empirically tests and visually presents the genre-specific “summary+details” or “inverted pyramid” structuring of news discourse. It further extends the idea of inverted pyramid structuring to the paragraph and sentence levels. It proves that the body of the report also has a similar schematic top-down installment organization with macro-propositions on top. It also visually and statistically presents the rhetorical structures at sentence level, which differ to some extent from grammatical structures. Operated in line with the compositionality criterion and hierarchy principle of RST, the converted trees provide unique analytical advantages and constitute new research prospects.
BACKGROUND: Curious parallels between the processes of species and language evolution have been observed by many researchers. Retracing the evolution of Indo-European (IE) languages remains one of the most intriguing intellectual challenges in historical linguistics. Most of the IE language studies use the traditional phylogenetic tree model to represent the evolution of natural languages, thus not taking into account reticulate evolutionary events, such as language hybridization and word borrowing which can be associated with species hybridization and horizontal gene transfer, respectively. More recently, implicit evolutionary networks, such as split graphs and minimal lateral networks, have been used to account for reticulate evolution in linguistics. RESULTS: Striking parallels existing between the evolution of species and natural languages allowed us to apply three computational biology methods for reconstruction of phylogenetic networks to model the evolution of IE languages. We show how the transfer of methods between the two disciplines can be achieved, making necessary methodological adaptations. Considering basic vocabulary data from the well-known Dyen's lexical database, which contains word forms in 84 IE languages for the meanings of a 200-meaning Swadesh list, we adapt a recently developed computational biology algorithm for building explicit hybridization networks to study the evolution of IE languages and compare our findings to the results provided by the split graph and galled network methods. CONCLUSION: We conclude that explicit phylogenetic networks can be successfully used to identify donors and recipients of lexical material as well as the degree of influence of each donor language on the corresponding recipient languages. We show that our algorithm is well suited to detect reticulate relationships among languages, and present some historical and linguistic justification for the results obtained. Our findings could be further refined if relevant syntactic, phonological and morphological data could be analyzed along with the available lexical data.
Inferring implicit discourse relations in natural language text is the most difficult subtask in discourse parsing. Surface features achieve good performance, but they are not readily applicable to other languages without semantic lexicons. Previous neural models require parses, surface features, or a small label set to work well. Here, we propose neural network models that are based on feedforward and long-short term memory architecture without any surface features. To our surprise, our best configured feedforward architecture outperforms LSTM-based model in most cases despite thorough tuning. Under various fine-grained label sets and a cross-linguistic setting, our feedforward models perform consistently better or at least just as well as systems that require hand-crafted surface features. Our models present the first neural Chinese discourse parser in the style of Chinese Discourse Treebank, showing that our results hold cross-linguistically.
We present a novel annotation framework for representing predicate-argument structures, which uses dependency trees to encode the syntactic and semantic roles of a sentence simultaneously. The main contribution is a semantic role transmission model, which eliminates the structural gap between syntax and shallow semantics, making them compatible. A Chinese semantic treebank was built under the proposed framework, and the first release containing about 14K sentences is made freely available. The proposed framework enables semantic role labeling to be solved as a sequence labeling task, and experiments show that standard sequence labelers can give competitive performance on the new treebank compared with state-of-the-art graph structure models.
Morphological segmentation has traditionally been modeled with non-hierarchical models, which yield flat segmentations as output. In many cases, however, proper morphological analysis requires hierarchical structureespecially in the case of derivational morphology. In this work, we introduce a discriminative, joint model of morphological segmentation along with the orthographic changes that occur during word formation. To the best of our knowledge, this is the first attempt to approach discriminative segmentation with a context-free model. Additionally, we release an annotated treebank of 7454 English words with constituency parses, encouraging future research in this area. 1
Lexical information, including surface word form and part-of-speech (POS) information, plays a crucial role when predicting ambiguous dependency relationships in dependency parsing. However, for resolving dependency ambiguities, surface word information may be too sparse, while POS information may be too coarse. Supertags, which are lexical templates that represent rich syntactic information, have been shown to provide effective features at an intermediate level on the coarse-to-fine scale. In this work, we present a supertag design framework that allows us to instantiate various supertag sets based on the dependency structures. Using this framework, we instantiate various supertag sets and utilize them as features in transition-based dependency parsing systems. Performing experiments on the Penn Treebank and Universal Dependencies data sets, we show that our supertags are effective for transition-based parsers in multilingual parsing as well as English parsing. The comparison of the results of the different supertag sets shows that it is crucial to incorporate the head directionality, head labels, and dependent possession information in supertags to improve the parser performance.
Historical treebanks tend to be manually annotated, which is not surprising, since state-of-the-art parsers are not accurate enough to ensure high-quality annotation for historical texts. We test whether automatic parsing can be an efficient pre-annotation tool for Old East Slavic texts. We use the TOROT treebank from the PROIEL treebank family. We convert the PROIEL format to the CONLL format and use MaltParser to create syntactic pre-annotation. Using the most conservative evaluation method, which takes into account PROIEL-specific features, MaltParser by itself yields 0.845 unlabelled attachment score, 0.779 labelled attachment score and 0.741 secondary dependency accuracy (note, though, that the test set comes from a relatively simple genre and contains rather short sentences). Experiments with human annotators show that preparsing, if limited to sentences where no changes to word or sentence boundaries are required, increases their annotation rate. For experienced annotators, the speed gain varies from 5.80% to 16.57%, for inexperienced annotators from 14.61% to 32.17% (using conservative estimates). There are no strong reliable differences in the annotation accuracy, which means that there is no reason to suspect that using preparsing might lower the final annotation quality.
We investigate mutual benefits between syntax and semantic roles using neural network models, by studying a parsingSRL pipeline, a SRLparsing pipeline, and a simple joint model by embedding sharing. The integration of syntactic and semantic features gives promising results in a Chinese Semantic Treebank, demonstrating large potentials of neural models for joint parsing and semantic role labeling.
The article presents methodological analyses of topical ideas of the famous modern linguist – E.Cosseriu. The authors argue that incorporation of theoretical ideas of E.Cosseriu could substentially extand the euristic potential of the conept of norm in the sphere of linguistics. The key to solvation of the problem lies in the necessity of changing of modern theoretical context of the question. The authors of the article consider the history of operationalization of the concept of norm in linguistics from the point of view of a specific hermeneutic approach in relation to other linguistic techniques. In the course of study of the role of linguistic norms in interaction of content and expression the article presents examples of extrapolation of the concept of norm from one discipline to another. In case of extrapolation of the concept of norm from the other disciplines into linguistic investigations the structural and functional dichotomy of language turns out that it is impossible to avoid considering spiritual as the world of objects, so that speech and language are considered as two different things. The rapid development of information technologies made possible to calculate many of the aspects of Humboldtian ideas. Computer statistics created conditions where the idea of "language picture of the world" and of the "inner form of the language" is gradually losing its original romantic charge and turns in a very trivial thing. Yet despite the fact that global standardization significantly enhances the processing and automatic addition of translation, which once gave beginning to hermeneutics the hermeneutic potential of linguistic norm still preseves many promissing prospects from the epistemological point of view. Key words: Linguistic norm and variability, E. Cosseriu, hermeneutics, euristic potential. В статье представлен методологический анализ актуальных идей известного современного лингвиста - E. Коссериу. Авторы утверждают, что введение теоретических идей E. Коссериу могло бы существенно расширить эвристический потенциал нормы в области лингвистики. Ключ к решению проблемы заключается в необходимости изменения современного теоретического контекста вопроса. Авторы статьи рассматривают историю ввода в действие понятия нормы в лингвистике с точки зрения конкретного герменевтического подхода по отношению к другим языковым методам. В ходе изучения роли языковых норм во взаимодействии содержания и выражения, в статье представлены примеры экстраполяции понятия нормы от одной дисциплины к другой. В случае экстраполяции понятия нормы с других дисциплин в лингвистические исследования структурно-функциональной дихотомии языка оказывается, что нельзя не рассматривать духовное как мир объектов, так как речь и язык рассматриваются как две разные вещи. Быстрое развитие информационных технологий сделало возможным для расчета многие аспекты гумбольдтовских идей. Статистика компьютера создала условия, в которых идея «языковой картины мира» и «внутренней формы языка» постепенно теряет свой первоначальный романтический заряд и превращается в очень тривиальную вещь. Тем не менее, несмотря на то, что глобальная стандартизация значительно улучшает обработку и автоматический перевод, который когда-то дал начало герменевтике, герменевтический потенциал языковой нормы хранит еще много перспектив с гносеологической точки зрения.Ключевые слова: лингвистическая норма и изменчивость, E.Коссериу, герменевтика, эвристический по- тенциал.
This paper presents the Universal Dependencies tagset (UD v1) as a new annotation scheme for Russian treebanks. The universal list of dependency relations was adopted and extended to comply with certain language-specific syntactic constructions. The tagset was validated, converting two Russian treebanks into the UD format, UD-Russian-SynTagRus and UD-Russian-Google.
We present a new, sizeable dataset of nounnoun compounds with their syntactic analysis (bracketing) and semantic relations. Derived from several established linguistic resources, such as the Penn Treebank, our dataset enables experimenting with new approaches towards a holistic analysis of noun-noun compounds, such as jointlearning of noun-noun compounds bracketing and interpretation, as well as integrating compound analysis with other tasks such as syntactic parsing.
The paper evaluates the differences between two currently leading annotation schemes for dependency treebanks. By relying on four treebanks, we demonstrate that the treatment of conjunctions and adpositions represents the core difference between the two schemes and that this impacts the topological properties of the linguistic networks induced from the treebanks. We also show that such properties are reflected in the performances of four probabilistic dependency parsers trained on the treebanks. L’articolo valuta le differenze tra i due principali schemi di annotazione a dipenden-ze in uso. Sulla base di quattro treebank, l’articolo dimostra che il trattamento delle congiunzioni e delle pre/postposizioni rappresenta la differenza principale tra i due schemi e che ciò comporta delle conseguenze sulle proprietà topologiche dei net-work indotti dalle treebank. Inoltre, si dimostra come tali proprietà siano riflesse nell’accuratezza di quattro parser probabilistici a dipendenze addestrati sulle treebank.
Statistical parsers are trained on treebanks that are composed of a few thousand sentences. In order to prevent data sparseness and computational complexity, such parsers make strong independence hypotheses on the decisions that are made to build a syntactic tree. These independence hypotheses yield a decomposition of the syntactic structures into small pieces, which in turn prevent the parser from adequately modeling many lexico-syntactic phenomena like selectional constraints and subcategorization frames. Additionally, treebanks are several orders of magnitude too small to observe many lexico-syntactic regularities, such as selectional constraints and subcategorization frames. In this article, we propose a solution to both problems: how to account for patterns that exceed the size of the pieces that are modeled in the parser and how to obtain subcategorization frames and selectional constraints from raw corpora and incorporate them in the parsing process. The method proposed was evaluated on French and on English. The experiments on French showed a decrease of 41.6% of selectional constraint violations and a decrease of 22% of erroneous subcategorization frame assignment. These figures are lower for English: 16.21% in the first case and 8.83% in the second.
Feedforward Neural Network (FNN)-based language models estimate the probability of the next word based on the history of the last N words, whereas Recurrent Neural Networks (RNN) perform the same task based only on the last word and some context information that cycles in the network. This paper presents a novel approach, which bridges the gap between these two categories of networks. In particular, we propose an architecture which takes advantage of the explicit, sequential enumeration of the word history in FNN structure while enhancing each word representation at the projection layer through recurrent context information that evolves in the network. The context integration is performed using an additional word-dependent weight matrix that is also learned during the training. Extensive experiments conducted on the Penn Treebank (PTB) and the Large Text Compression Benchmark (LTCB) corpus showed a significant reduction of the perplexity when compared to state-of-the-art feedforward as well as recurrent neural network architectures.
While linguistic theory posits an arbitrary relation between signifiers and the signified (de Saussure, 1916), our analysis of a large-scale German database containing affective ratings of words revealed that certain phoneme clusters occur more often in words denoting concepts with negative and arousing meaning. Here, we investigate how such phoneme clusters that potentially serve as sublexical markers of affect can influence language processing. We registered the EEG signal during a lexical decision task with a novel manipulation of the words' putative sublexical affective potential: the means of valence and arousal values for single phoneme clusters, each computed as a function of respective values of words from the database these phoneme clusters occur in. Our experimental manipulations also investigate potential contributions of formal salience to the sublexical affective potential: Typically, negative high-arousing phonological segments-based on our calculations-tend to be less frequent and more structurally complex than neutral ones. We thus constructed two experimental sets, one involving this natural confound, while controlling for it in the other. A negative high-arousing sublexical affective potential in the strictly controlled stimulus set yielded an early posterior negativity (EPN), in similar ways as an independent manipulation of lexical affective content did. When other potentially salient formal features at the sublexical level were not controlled for, the effect of the sublexical affective potential was strengthened and prolonged (250-650 ms), presumably because formal salience helps making specific phoneme clusters efficient sublexical markers of negative high-arousing affective meaning. These neurophysiological data support the assumption that the organization of a language's vocabulary involves systematic sound-to-meaning correspondences at the phonemic level that influence the way we process language.
Deaf or hard-of-hearing individuals usually face a greater challenge to learn to write than their normal-hearing counterparts. Due to the limitations of traditional research methods focusing on microscopic linguistic features, a holistic characterization of the writing linguistic features of these language users is lacking. This study attempts to fill this gap by adopting the methodology of linguistic complex networks. Two syntactic dependency networks are built in order to compare the macroscopic linguistic features of deaf or hard-of-hearing students and those of their normal-hearing peers. One is transformed from a treebank of writing produced by Chinese deaf or hard-of-hearing students, and the other from a treebank of writing produced by their Chinese normal-hearing counterparts. Two major findings are obtained through comparison of the statistical features of the two networks. On the one hand, both linguistic networks display small-world and scale-free network structures, but the network of the normal-hearing students' exhibits a more power-law-like degree distribution. Relevant network measures show significant differences between the two linguistic networks. On the other hand, deaf or hard-of-hearing students tend to have a lower language proficiency level in both syntactic and lexical aspects. The rigid use of function words and a lower vocabulary richness of the deaf or hard-of-hearing students may partially account for the observed differences.
Abstract syntax is a semantic tree representation that lies between parse trees and logical forms. It abstracts away from word order and lexical items, but contains enough information to generate both surface strings and logical forms. Abstract syntax is commonly used in compilers as an intermediate between source and target languages. Grammatical Framework (GF) is a grammar formalism that generalizes the idea to natural languages, to capture cross-lingual generalizations and perform interlingual translation. As one of the main results, the GF Resource Grammar Library (GF-RGL) has implemented a shared abstract syntax for over 30 languages. Each language has its own set of concrete syntax rules (morphology and syntax), by which it can be generated from the abstract syntax and parsed into it. This paper presents a conversion method from abstract syntax trees to dependency trees. The method is applied for converting GF-RGL trees to Universal Dependencies (UD), which uses a common set of labels for different languages. The correspondence between GF-RGL and UD turns out to be good, and the relatively few discrepancies give rise to interesting questions about universality. The conversion also has potential for practical applications: (1) it makes the GF parser usable as a rule-based dependency parser; (2) it enables bootstrapping UD treebanks from GF treebanks; (3) it defines formal criteria to assess the informal annotation schemes of UD; (4) it gives a method to check the consistency of manually annotated UD trees with respect to the annotation schemes; (5) it makes information from UD treebanks available.
Abstract Three studies examined gender differences in the effect of storytelling ability on perceptions of a person's attractiveness as a short‐term and long‐term romantic partner. In Study 1, information about a potential partner's storytelling ability was provided. Study 2 participants read a good or poor story supposedly written by a potential partner. Results suggested that only women's attractiveness assessments of men as a long‐term date increased for good storytellers. Storytelling ability did not affect men's ratings of women nor did it affect ratings of short‐term partners. Study 3 suggested that the effect of storytelling ability on long‐term attractiveness for male targets may be mediated by perceived status. Storytelling ability appears to increase perceived status and thus helps men attract long‐term partners.
We propose a framework to model human comprehension of discourse connectives. Following the Bayesian pragmatic paradigm, we advocate that discourse connectives are interpreted based on a simulation of the production process by the speaker, who, in turn, considers the ease of interpretation for the listener when choosing connectives. Evaluation against the sense annotation of the Penn Discourse Treebank confirms the superiority of the model over literal comprehension. A further experiment demonstrates that the proposed model also improves automatic discourse parsing.
The Universal Dependencies (UD) Project seeks to build a cross-lingual studies of treebanks, linguistic structures and parsing. Its goal is to create a set of multilingual harmonized treebanks that are designed according to a universal annotation scheme. In this paper, we report on the conversion of the Uyghur dependency treebank to a UD version of the treebank which we term the Uyghur Universal Dependency Treebank (UyDT). We present the mapping of the Uyghur dependency treebank’s labelling scheme to the UD scheme, along with a clear description of the structural changes required in this conversion.
In accordance with the compositionality criterion and hierarchy principle of Rhetorical Structure Theory (RST), this study reframes each tree in the RST Discourse Treebank into three new dependency trees with ultimate nodes being clauses, sentences, and paragraphs, respectively, which also draw on an analogy between syntactic and discourse trees. Detailed percentages of various RST relations at the three granularity levels are examined, illuminating the discourse processes of organizing units of one granularity level into those of the next upper level and suggesting certain homogeneity and interaction across levels in the Treebank, particularly at the two upper levels. The study demonstrates the applicability of RST analysis between same-level terminal units. With unique analytical advantages, the newly constructed discourse dependency trees provide new research prospects.
We propose a classification framework for semantic type identification of compounds in Sanskrit. We broadly classify the compounds into four different classes namely, Avyayībhāva, Tatpuruṣa, Bahuvrīhi and Dvandva. Our classification is based on the traditional classification system followed by the ancient grammar treatise Adṣṭādhyāyī, proposed by Pāṇini 25 centuries back. We construct an elaborate features space for our system by combining conditional rules from the grammar Adṣṭādhyāyī, semantic relations between the compound components from a lexical database Amarakoṣa and linguistic structures from the data using Adaptor Grammars. Our in-depth analysis of the feature space highlight inadequacy of Adṣṭādhyāyī, a generative grammar, in classifying the data samples. Our experimental results validate the effectiveness of using lexical databases as suggested by Amba Kulkarni and Anil Kumar, and put forward a new research direction by introducing linguistic patterns obtained from Adaptor grammars for effective identification of compound type. We utilise an ensemble based approach, specifically designed for handling skewed datasets and we %and Experimenting with various classification methods, we achieve an overall accuracy of 0.77 using random forest classifiers.