Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
While gender identities in the Western world are typically regarded as binary, our previous work (Hicks et al., 2015) shows that there is more lexical variety of gender identity and the way people identify their gender.There is also a growing need to lexically represent this variety of gender identities.In our previous work, we developed a set of tools and approaches for analyzing Twitter data as a basis for generating hypotheses on language used to identify gender and discuss genderrelated issues across geographic regions and population groups in the U.S.A.In this paper we analyze the coverage and relative frequency of the word forms in our Twitter analysis with respect to the National Transgender Discrimination Survey data set, one of the most comprehensive data sets on transgender, gender non-conforming, and gender variant people in the U.S.A.We then analyze the coverage of WordNet, a widely used lexical database, with respect to these identities and discuss some key considerations and next steps for adding gender identity words and their meanings to WordNet.
In this contribution we illustrate the methodology and the results of an experiment we conducted by applying Distributional Semantics Models to the analysis of the Holy Quran. Our aim was to gather information on the potential differences in meanings that the same words might take on when used in Modern Standard Arabic w.r.t. their usage in the Quran. To do so we used the Penn Arabic Treebank as a contrastive corpus.
<p>In the recent years, globalization prepared a ground for English to be the lingua franca of the academia. Thus, most highly prestigious international journals have defined their medium of publications as English. However, even advanced language learners have difficulties in writing their research articles due to the lack of appropriate lexical knowledge and discourse conventions of academia. Considering the fact that the underuse, overuse and misuse of formulaic sequences or lexical bundles are often characterized with non-native writers of English, lexical bundle studies have recently been on the top of the agenda of corpus studies. Although the related literature has represented specific genres or disciplines, no study has scrutinized lexical bundles in the research articles that are written in the educational sciences. Therefore, the current study compared the structural and functional characteristics of the lexical-bundle use in L1 and L2 research articles in English. The results revealed the deviation of the usages of lexical bundles by the non-native speakers of English from the native speaker norms. Furthermore, the results indicated the overuse of clausal or verb-phrase based lexical bundles in the research articles of Turkish scholars while their native counterparts used noun and prepositional phrase-based lexical bundles more than clausal bundles.</p>
In this paper, we present our new experimental system of merging dependency representations of two parallel sentences into one dependency tree. All the inner nodes in dependency tree represent source-target pairs of words, the extra words are in form of leaf nodes. We use Universal Dependencies annotation style, in which the function words, whose usage often differs between languages, are annotated as leaves. The parallel treebank is parsed in minimally supervised way. Unaligned words are there automatically pushed to leaves. We present a simple translation system trained on such merged trees and evaluate it in WMT 2016 English-to-Czech and Czechto-English translation task. Even though the model is so far very simple and no language model and word-reordering model were used, the Czech-to-English variant reached similar BLEU score as another established tree-based system.
Purpose: This study was intended to evaluate a series of algorithms developed to perform automatic classification of paraphasic errors (formal, semantic, mixed, neologistic, and unrelated errors). Method: We analyzed 7,111 paraphasias from the Moss Aphasia Psycholinguistics Project Database (Mirman et al., 2010) and evaluated the classification accuracy of 3 automated tools. First, we used frequency norms from the SUBTLEXus database (Brysbaert & New, 2009) to differentiate nonword errors and real-word productions. Then we implemented a phonological-similarity algorithm to identify phonologically related real-word errors. Last, we assessed the performance of a semantic-similarity criterion that was based on word2vec (Mikolov, Yih, & Zweig, 2013). Results: Overall, the algorithmic classification replicated human scoring for the major categories of paraphasias studied with high accuracy. The tool that was based on the SUBTLEXus frequency norms was more than 97% accurate in making lexicality judgments. The phonological-similarity criterion was approximately 91% accurate, and the overall classification accuracy of the semantic classifier ranged from 86% to 90%. Conclusion: Overall, the results highlight the potential of tools from the field of natural language processing for the development of highly reliable, cost-effective diagnostic tools suitable for collecting high-quality measurement data for research and clinical purposes.
BACKGROUND: Studies of amnestic mild cognitive impairment (aMCI) and late-life depression (LLD) have examined the similarities and differences between these syndromes, but few have investigated how the cognitive profile of comorbid aMCI and subclinical depressive symptoms (aMCI/D+) may compare to that of aMCI or LLD. Memory biases for certain types of emotional information may distinguish these groups. METHODS: A total of 35 aMCI, 23 aMCI/D+, 13 LLD, and 17 elderly controls (CONT) rated the valence (positive, negative, or neutral) of 30 pictures from the International Affective Picture System. Mean percent positive, negative, and neutral images recalled was compared within groups immediately and 30 minutes later. RESULTS: Overall memory performance was comparable in aMCI and aMCI/D+, and both recalled fewer items than CONT and LLD. Group differences emerged when valence ratings were considered: at immediate and delayed recall, positive and negative pictures were generally better-remembered than neutral pictures by CONT, aMCI, and LLD, but valence was not associated with recall in aMCI/D+. Follow-up analyses suggested that the perceived intensity of stimuli may explain the emotional enhancement effect in CONT, aMCI, and LLD. CONCLUSIONS: Results support previous research suggesting that the neuropsychological profile of aMCI/D+ is different from that of aMCI and LLD. Although depressed and non-depressed individuals with aMCI recall comparable quantities of information, the quality of the recalled information differs significantly. On theoretical grounds, this suggests the existence of distinct neurobiological or neurofunctional manifestations in both groups. Practically, these differences may guide the development of personalized emotion-focused encoding strategies in cognitive training programs.
The selection of standards and norms constitutes the first and most important step for language standardisation. In this paper, we examine the standard establishment for Huayu (or Singapore Mandarin), a new Chinese variety that has emerged in Singapore as a result of centralised planning and inter-linguistic contact. Huayu is the officially designated mother tongue of the Chinese community and a second language in school education in Singapore. The overall linguistic features of Huayu largely conform to the norms practiced in mainland China, though this localised variety has developed a number of distinctive phonological, lexical and grammatical features. Singapore’s government takes a Tacit Compliance Approach to the Mandarin norms, that is, exonormative standards are followed in an implicit manner. This pragmatic approach has engendered some confusion and dilemmas for Chinese language (CL) education in Singapore. Given the fact that Huayu is approaching a stage of nativisation, we propose the adoption of an explicit endonormative standard to cater to the pressing needs in CL teaching, learning and assessment.
This paper presents on-going work on creating NLP tools for under-resourced languages from very sparse training data coming from linguistic field work. In this work, we focus on Ingush, a Nakh-Daghestanian language spoken by about 300,000 people in the Russian republics Ingushetia and Chechnya. We present work on morphosyntactic taggers trained on transcribed and linguistically analyzed recordings and dependency parsers using English glosses to project annotation for creating synthetic treebanks. Our preliminary results are promising, supporting the goal of bootstrapping efficient NLP tools with limited or no task-specific annotated data resources available.
Research on heritage language (HL) development and education has characterized the unique linguistic, sociocultural, and affective profiles of heritage-language (HL) students, yet foreign-language (FL) education has only begun to understand HL students in relation to non-heritage students (Carreira & Kagan, 2011; Felix, 2008). To deepen our understanding of the unique characteristics of HL and FL populations, this investigation examines: (1) how HL learners’ language learning goals, perceptions of language proficiency and learning needs, and habits of language use distinguish them from those of their English-speaking FL counterparts, and (2) how FL and HL students’ perceptions of (and attitudes toward) Spanish implicitly or explicitly reflect attitudes expressed in the classroom and elsewhere. This mixed-methods study explores and compares the perceptions and experiences of 109 university-level HL students and 138 FL students as reported in survey responses and ethnographic interviews. Socio-affective variables distinguishing HL and FL groups include attitudes toward Spanish, judgments of linguistic norms, learners’ efforts to achieve legitimacy as users of Spanish, and obstacles to interaction, including anxiety and social inhibition. Quantitative and qualitative analyses depict participants’ self-images as fragile users of Spanish, with many characterizations paralleling instructor and peer perceptions.
Abstract Some dependency treebanks use special sequences of dependencies where main arguments are mixed with separators. Classical Categorial Dependency Grammars (CDG) do not allow this construction because iterative dependency types only introduce the iterations of the same dependency. An extension of CDG is defined here that introduces a new construction for repeatable sequences of one or several dependency names. The learnability properties of the extended CDG when grammars are infered from a dependency treebank is also studied. It leads to the definition of new classes of grammars that are learnable in the limit from dependency structures.
The purpose of this paper is to discuss the social legitimacy of the non-dominant variety of French that is used in Belgium (henceforth ‘Belgian French’). As will be detailed, Francophone Belgians’ attitudes have shifted from early 19th c. – late 20th c. purism and subsequent linguistic subjection to France to more recent acceptation of endogenous traits and increasing distance from the Hexagonal model. Nevertheless, these attitudes remain characterized by a “double distance” from both Hexagonal and Belgian French. The idea that French is viewed by Francophone Belgians as a polycentric/polynomic language will thus be questioned: do they really consider that there is a legitimate Belgian variety of French? What is the relevance of the national criterion in the way they define linguistic norms? What other criteria lie behind the definition and legitimization of their linguistic norms?
The measure of sentence similarity is useful in various research fields, such as artificial intelligence, knowledge management, and information retrieval. Several methods have been proposed to measure the sentence similarity based on syntactic and/or semantic knowledge. Most proposals are evaluated on English sentences where the accuracy can decrease when these proposals are applied to other languages. Moreover, the results of these methods are unsatisfactory, as much relevant semantic knowledge, such as semantic class, thematic role and syntactico-semantic knowledge like the semantic predicates, are not taken into account. We must acknowledge that this kind of knowledge is rare in most of the lexical resources. Recently, the International Organization for Standardization (ISO) has published the Lexical Markup Framework (LMF) ISO-24613 norm for the development of lexical resources. This norm provides, for each meaning of a lexical entry, all the semantic and syntactico-semantic knowledge in a fine structure. Profiting from the availability of LMF-standardized dictionaries, we propose, in this paper, a generic method that enhances the measure of sentence similarity by applying semantic and syntactico-semantic knowledge. An experiment was carried out on Arabic, as this language is processed within our research team and an LMF-standardized Arabic dictionary is at hand where the semantic and the syntactico-semantic knowledge are accessible and well structured. Moreover, the experiments yielded better results, showing a high correlation with human ratings.
From the perspective of structural linguistics, we explore paradigmatic and syntagmatic lexical relations for Chinese POS tagging, an important and challenging task for Chinese language processing. Paradigmatic lexical relations are explicitly captured by word clustering on large-scale unlabeled data and are used to design new features to enhance a discriminative tagger. Syntagmatic lexical relations are implicitly captured by syntactic parsing in the constituency formalism, and are utilized via system combination. Experiments on the Penn Chinese Treebank demonstrate the importance of both paradigmatic and syntagmatic relations. Our linguistically motivated, hybrid approaches yield a relative error reduction of 18% in total over state-of-the-art baselines. Despite the effectiveness to boost accuracy, computationally expensive parsers make hybrid systems inappropriate for many realistic NLP applications. In this article, we are also concerned with improving tagging efficiency at test time. In particular, we explore unlabeled data to transfer the predictive power of hybrid models to simple sequence models. Specifically, hybrid systems are utilized to create large-scale pseudo training data for cheap models. Experimental results illustrate that the re-compiled models not only achieve high accuracy with respect to per token classification, but also serve as a front-end to a parser well.
Penn Discourse Treebank (PDTB)-style annotation focuses on labeling local discourse relations between text spans and typically ignores larger discourse contexts. In this paper we propose two approaches to infer discourse relations in a paragraph-level context from annotated PDTB labels. We investigate the utility of inferring such discourse information using the task of revision classification. Experimental results demonstrate that the inferred information can significantly improve classification performance compared to baselines, not only when PDTB annotation comes from humans but also from automatic parsers.
On this paper I discuss 30 usual mistakes found in academic writing within texts written by students, master or PhD candidates and junior academics. From a theoretical standpoint, I tell self-consciously disobedient-to-norms writing (that can be considered a political statement) from simply improper writing (that signals ignorance of such norms). I argue that, given the lack of training in academic writing in the course of most college careers, such ignorance should not come as a surprise. Both as a political statement and as a symptom of ignorance of linguistic norms and from an epistemologically antirealist standpoint, I argue that since theories (and, therefore, the world as something scientifically conceivable and actually conceived) exist within and through academic writing, every effort towards proper writing is valuable. Finally, I list and discuss 30 usual mistakes taken from the casuistry that I have gathered over the years and from books on normative talk and writing.
Semantic analysis of sentences can only be carried out using Dependency Parsing. Dependency parser accepts words in a sentence and builds dependency relation among the words resulting in a unique tree for each sentence. An Indian Panini is the first to develop semantic analysis for Sanskrit using a dependency framework. Western researchers in the near past have also deliberated on dependency parsing so that automated dependency parser can be generated. Dependency Parser is useful in information extraction, question-answering, text summarization etc. For many indian languages namely Bengali, Kannada, Malayalam and Marathi a dependency based Treebank is in the development stage. Also for Hindi, Telugu and Tamil dependency Treebanks are already developed. The Treebank data can be used by a dependency parser generator like Maltparser to develop a Dependency parser. In this paper few dependency parsing algorithms are discussed
Neuroimaging studies have demonstrated that the medial prefrontal cortex is involved in attributions on enduring and abstract trait characteristics of persons, but not in causal attributions of temporary here-and-now events. Moreover, the neural representation of trait information is thought to be located in the ventral part of the medial prefrontal cortex (vmPFC). In order to verify this latter finding, this study compared the performance of 8 patients with hypoperfusion of the vmPFC, 10 with hypoperfusion excluding the vmPFC and 15 healthy controls on trait and causal attribution questionnaires consisting of several events presented in brief written scenarios. We also investigated whether vmPFC hypoperfusion influenced the experienced intensity of the negative or positive valence of the events. Our results showed that patients with ventral hypoperfusion performed significantly worse on trait attributions in comparison with the non-vmPFC group and healthy controls. All groups performed equally well on causal attributions. These findings support previous research suggesting that the vmPFC is critically involved in enduring trait attribution, but not in temporary causal attribution. Considering the emotional experience of valence, the findings showed more intense valence ratings for negative events and persons. This confirms the role of the vmPFC in the modulation and regulation of negative emotions.
One of the most pressing questions in cognitive science remains unanswered: what cognitive mechanisms enable children to learn any of the world’s 7000 or so languages? Much discovery has been made with regard to specific learning mechanisms in specific languages, however, given the remarkable diversity of language structures (Evans and Levinson, 2009; Bickel, 2014) the burning question remains: what are the underlying processes that make language acquisition possible, despite substantial cross-linguistic variation in phonology, morphology, syntax, etc.? To investigate these questions, a comprehensive cross-linguistic database of longitudinal child language acquisition corpora from maximally diverse languages has been built.
The ontological approach to the creation of the learning process support systems is proposed. We research the method for automated construction of learning ontologies based on computational linguistics algorithms, the method of analysis of lexical-semantic fields of text corpora in Russian and English, frequency dictionaries of terms. The prototype ontology is developed using the lexical database WordNet and terminological dictionaries and uploaded into the “OntoMASTER-Ontology” software tool for editing by experts. The developed method is used for constructing ontologies to support the learning process of students in the field of “Information systems and technologies”.
This journal article carries out a structural-functional analysis of the formation of Old English nouns by means of affixation. The data comprise a total of 4,370 nouns which result from either prefixation or suffixation, retrieved from the lexical database of Old English Nerthus. Twenty-five derivational functions, inspired by functional grammars and Pounder’s (2000) paradigmatic morphology are proposed to explain the relationship holding between affixes and their bases of derivation. These functions have been divided into split and unified, the former being realized by both prefixes and suffixes and the latter by either prefixation or suffixation. The conclusion is reached that the main target of prefixation is the modification of meaning, in such a way that the meaning of the derivative is less predictable from the input category whereas the main target of suffixation is the change of lexical category, given that the meaning of the derivative is more predictable from the the input category.
This article introduces the use of the Amazon Mechanical Turk (MTurk) crowdsourcing platform as a resource for R users to leverage crowdsourced human intelligence for preprocessing "messy" data into a form easily analyzed within R. The article first describes MTurk and the MTurkR package, then outlines how to use MTurkR to gather and manage crowdsourced data with MTurk using some of the package's core functionality. Potential applications of MTurkR include construction of manually coded training sets, human transcription and translation, manual data scraping from scanned documents, content analysis, image classification, and the completion of online survey questionnaires, among others. As an example of massive data preprocessing, the article describes an image rating task involving 225 crowdsourced workers and more than 5500 images using just three MTurkR function calls.
INTRODUCTIONThe Strategic Integrated Management Seminar (SIMS) course is mandatory for every senior student in the school of business at a mid-size private university in the northeastern United States. The course allows students to integrate their accumulated knowledge and apply this knowledge to issues from a strategic perspective. It examines a firm from the position of top level management, focusing on the role of the general manager in formulating and implementing corporate and business level strategy. Strategic issues of an entire athletic (hereon, footwear company or company) and industry are analyzed. Students are expected to draw their accumulated knowledge of the functional areas of their majors into a homogenous team effort. Each individual student works on developing his/her ability to analyze information, draw logical conclusions, and offer sound supporting evidence for their arguments in written form and classroom discussions. The course is highly interactive with students taking the lead and the professors sharing knowledge and offering supplementary support.The SIMS course uses the Business Strategy Game (BSG) simulation to enable students to experience a top management team perspective in running a and experiencing competitive conditions in the athletic industry. In the BSG, students compete in teams (each team constitutes a company, hereon team/company will be synonymous) within a global arena that encompasses four regions - Europe-Africa, North America, Asia-Pacific, and Latin America (The Business Strategy Game, 2016). They compete against teams in their individual classes and compare/contrast data with teams/companies worldwide. Each competes head-to-head against companies run by other teams in the course, hence competition plays an important role in the experience. Each sells its brand of to retailers worldwide and to individuals buying online at the company's website.Competing in the BSG requires a series of complex decisions by the students, taking into account the team's strategy for their and the competitive conditions in the industry and the strategies of their competitors. The simulation allows for numerous decisions for each round, requiring students to choose which decisions are most important to implement their strategy and which areas of the business must receive attention in order for their firm to be its most competitive. Beyond overall strategy (corporate, competitive) are several key functional areas for decision making. Decision areas in operations include capacity planning (either adding to existing plants or building new plants in new geographic locations), production quality decisions for the athletic footwear, plant operations efficiency, and labor decisions. Footwear must be shipped to distribution centers around the world and students must choose where it is best to manufacture the and where to ship taking into consideration demand, shipping costs, tariffs, and exchange rates. Marketing decisions include pricing the product in a wholesale and a retail environment, advertising and use or non-use of celebrity endorsements, rebates, and incentives to retailers. Financial decisions include funding the capital structure of the firm using debt, equity, and/or cash. Dividend payouts and stock repurchases may be used by the companies.The simulation has students take control of an athletic that has been in operation for ten years. Teams make in total eight years of decisions (years 11 - 18), approximately one per week. Each decision rollover represents one year and includes many decisions within the decision. The simulation evaluates team performance based on five investor expectation performance targets: Earnings Per Share (EPS), Return on Equity (ROE), credit rating, image rating (a combination of market share and shoe quality), and stock price. Each measure of performance is equally weighted at 20% of the total score (The Business Strategy Game, 2016). …
This study describes a change in which relative clause extraposition is in the process of being lost in English, Icelandic, French, and Portuguese. This current change in progress has never been observed before, probably because it is so slow that it is undetectable without the aid of multiple diachronic parsed corpora (treebanks) with time depths of over 500 years each. Building on insights from Kiparsky (1995), the study shows that the change may date as far back as the innovation of Proto-Germanic and Proto-Romance relative clauses, as these varieties differentiated from Proto-Indo-European. It also shows that the unusually slow speed of the change is due to partial specialization of the construction along the dimension of prosodic weight, following the argument made at greater length in Fruehwald & Wallenberg 2016. Finally, the change is shown to have important consequences for the syntax of extraposition, supporting the adjunction analysis of Culi- cover and Rochemont (1990). The article also discusses the implications of Sauerland's (2003) analysis of English relative clauses, and while modern English data supports his analysis, the diachronic extraposition data is not yet fine-grained enough to bear on the ‘raising’ analysis of relatives in general. This is identified as an important question for further research on this change.
We present an investigation of the perception of authenticity in audiovisual laughter, in which we contrast spontaneous and volitional samples and examine the contributions of unimodal affective information to multimodal percepts. In a pilot study, we demonstrate that listeners perceive spontaneous laughs as more authentic than volitional ones, both in unimodal (audio-only, visual-only) and multimodal contexts (audiovisual). In the main experiment, we show that the discriminability of volitional and spontaneous laughter is enhanced for multimodal laughter. Analyses of relationships between affective ratings and the perception of authenticity show that, while both unimodal percepts significantly predict evaluations of audiovisual laughter, it is auditory affective cues that have the greater influence on multimodal percepts. We discuss differences and potential mismatches in emotion signalling through voices and faces, in the context of spontaneous and volitional behaviour, and highlight issues that should be addressed in future studies of dynamic multimodal emotion processing.
Tiedemann K. Linguistic norms in mathematics lessons. In: Krainer K, Vondrová N, eds. <em>Proceedings of the Ninth Congress of the European Society for Research in Mathematics Education</em>. 2016: 1503-1509.
Speakers respond more slowly when naming pictures presented with taboo (i.e., offensive/embarrassing) than with neutral distractor words in the picture-word interference paradigm. Over four experiments, we attempted to localize the processing stage at which this effect occurs during word production and determine whether it reflects the socially offensive/embarrassing nature of the stimuli. Experiment 1 demonstrated taboo interference at early stimulus onset asynchronies of -150 ms and 0 ms although not at 150 ms. In Experiment 2, taboo distractors sharing initial phonemes with target picture names eliminated the interference effect. Using additive factors logic, Experiment 3 demonstrated that taboo interference and phonological facilitation effects do not interact, indicating that the two effects originate at different processing levels within the speech production system. In Experiment 4, interference was observed for masked taboo distractors, including those sharing initial phonemes with the target picture names, indicating that the effect cannot be attributed to a processing level involving responses in an output buffer. In two of the four experiments, the magnitude of the interference effect correlated significantly with arousal ratings of the taboo words. However, no significant correlations were found for either offensiveness or valence ratings. These findings are consistent with a locus for the taboo interference effect prior to the processing stage responsible for word form encoding. We propose a pre-lexical account in which taboo distractors capture attention at the expense of target picture processing due to their high arousal levels.
In this paper, we propose a new annotation approach to Chinese word segmentation, part-of-speech (POS) tagging and dependency labelling that aims to overcome the two major issues in traditional morphology-based annotation: Inconsistency and data sparsity. We re-annotate the Penn Chinese Treebank 5.0 (CTB5) and demonstrate the advantages of this approach compared to the original CTB5 annotation through word segmentation, POS tagging and machine translation experiments.
The author focuses on the methods of graphic (visual) coding of narrative polyphony in English postmodern fiction text. Resting on integrative interdisciplinary approach applied to the study of the issue under analysis, the author substantiates that the graphic surface of postmodern fiction text bears a particular layout, perspective and store of graphic (visual) signs originating from heterogeneous semiotic modes, both verbal and nonverbal. In this paper, the author defines and classifies visual signs as units of graphic coding of narrative polyphony in English postmodern multimodal fiction text. Graphic signs do not only change the graphic surface of the text, but also shape its new narrative structure, both of which construct polysemantic content of the text. Sinking into multimodal research, the author provides a study of the nature of graphic innovations as text units functioning on various text levels, both as attractors and distractors of its coded content that mirror linguistic norm democratization. By illustrating the applied methods in a text fragment from the English multimodal polyphonic fiction narrative, the author concludes that graphic coding units manifest postmodern fiction text as a synergetic whole in that the latter is a self-organized non-linear dissipative system that generates multiple ways of deconstructive text interpretation manipulating the reader. The author strongly believes that the findings may generate consequent research in the field of multimodal narratology.
Accurate automatic processing of Web queries is important for high-quality information retrieval from the Web. While the syntactic structure of a large portion of these queries is trivial, the structure of queries with question intent is much richer. In this paper we therefore address the task of statistical syntactic parsing of such queries. We first show that the standard dependency grammar does not account for the full range of syntactic structures manifested by queries with question intent. To alleviate this issue we extend the dependency grammar to account for segments -independent syntactic units within a potentially larger syntactic structure. We then propose two distant supervision approaches for the task. Both algorithms do not require manually parsed queries for training. Instead, they are trained on millions of (query, page title) pairs from the Community Question Answering (CQA) domain, where the CQA page was clicked by the user who initiated the query in a search engine. Experiments on a new treebank 1 consisting of 5,000 Web queries from the CQA domain, manually parsed using the proposed grammar, show that our algorithms outperform alternative approaches trained on various sources: tens of thousands of manually parsed OntoNotes sentences, millions of unlabeled CQA queries and thousands of manually segmented CQA queries.
Several studies exist in the literature that address the problem of emotion classification of visual stimuli but less effort has been devoted to emotion classification of audio stimuli. The most of these studies start from the analysis of physiological signals such as EEG data [1]. The aim of this work is to evaluate if it is possible to classify audio signals according to elicited emotions using only objective features. In our analysis we adopt the IADS (International Affective Digitized Sound) database [2], composed of 167 auditory stimuli. The database provides pleasure, arousal and dominance ratings for each audio stimulus, recorded from 100 subjects during psycho physical test. The database is formed by different type of audio: from environmental sounds to music, as well as from single sound to complex ones. We start considering the affective dimension of valence within the three categorical classes of low, medium and high pleasure. To investigate this classification task we consider 35 features both in time and frequency domain. With these features, we test three types of classifiers: Bayesian, K Nearest Neighbor and Classification and Regression Tree [3]. We apply a feature selection strategy in order to find the more significant features. Using these features and the Bayesian classifier we have reached an average accuracy of 45%. A similar result is achieved using physiological signals [1]. Starting from our results we believe that dividing each audio files in frames and applying a windowing strategy to evaluate objective features, the final classification performance could significantly increase.
Syntactic information, obtainable through syntactical analysis, plays an important role in many areas of NLP. Researches of Indonesian constituent parser have been very limited with the currently available yields very poor performance of 38.89% and 47.22% using Earley and CYK algorithm respectively. With the availability of the newly introduced Indonesian treebank corpus, we evaluate the performance of Indonesian constituent parser using Trance parser, a language independent constituent parser which employs deep learning. The parser achieved a respectable f-score of 74.91%, a very significant improvement to previous researches.
Abstract Drawing on an analogy between discourse and syntactic trees, this paper chooses 359 Wall Street Journal articles with multiple paragraphs from the Rhetorical Structure Theory (RST) Discourse Treebank, and converts each discourse tree into three additional dependency ones, at discourse, paragraph and sentence levels, with exclusively elementary discourse units of clauses, sentences and paragraphs, respectively. It empirically tests and visually presents the genre-specific “summary+details” or “inverted pyramid” structuring of news discourse. It further extends the idea of inverted pyramid structuring to the paragraph and sentence levels. It proves that the body of the report also has a similar schematic top-down installment organization with macro-propositions on top. It also visually and statistically presents the rhetorical structures at sentence level, which differ to some extent from grammatical structures. Operated in line with the compositionality criterion and hierarchy principle of RST, the converted trees provide unique analytical advantages and constitute new research prospects.
BACKGROUND: Curious parallels between the processes of species and language evolution have been observed by many researchers. Retracing the evolution of Indo-European (IE) languages remains one of the most intriguing intellectual challenges in historical linguistics. Most of the IE language studies use the traditional phylogenetic tree model to represent the evolution of natural languages, thus not taking into account reticulate evolutionary events, such as language hybridization and word borrowing which can be associated with species hybridization and horizontal gene transfer, respectively. More recently, implicit evolutionary networks, such as split graphs and minimal lateral networks, have been used to account for reticulate evolution in linguistics. RESULTS: Striking parallels existing between the evolution of species and natural languages allowed us to apply three computational biology methods for reconstruction of phylogenetic networks to model the evolution of IE languages. We show how the transfer of methods between the two disciplines can be achieved, making necessary methodological adaptations. Considering basic vocabulary data from the well-known Dyen's lexical database, which contains word forms in 84 IE languages for the meanings of a 200-meaning Swadesh list, we adapt a recently developed computational biology algorithm for building explicit hybridization networks to study the evolution of IE languages and compare our findings to the results provided by the split graph and galled network methods. CONCLUSION: We conclude that explicit phylogenetic networks can be successfully used to identify donors and recipients of lexical material as well as the degree of influence of each donor language on the corresponding recipient languages. We show that our algorithm is well suited to detect reticulate relationships among languages, and present some historical and linguistic justification for the results obtained. Our findings could be further refined if relevant syntactic, phonological and morphological data could be analyzed along with the available lexical data.
Inferring implicit discourse relations in natural language text is the most difficult subtask in discourse parsing. Surface features achieve good performance, but they are not readily applicable to other languages without semantic lexicons. Previous neural models require parses, surface features, or a small label set to work well. Here, we propose neural network models that are based on feedforward and long-short term memory architecture without any surface features. To our surprise, our best configured feedforward architecture outperforms LSTM-based model in most cases despite thorough tuning. Under various fine-grained label sets and a cross-linguistic setting, our feedforward models perform consistently better or at least just as well as systems that require hand-crafted surface features. Our models present the first neural Chinese discourse parser in the style of Chinese Discourse Treebank, showing that our results hold cross-linguistically.
We present a novel annotation framework for representing predicate-argument structures, which uses dependency trees to encode the syntactic and semantic roles of a sentence simultaneously. The main contribution is a semantic role transmission model, which eliminates the structural gap between syntax and shallow semantics, making them compatible. A Chinese semantic treebank was built under the proposed framework, and the first release containing about 14K sentences is made freely available. The proposed framework enables semantic role labeling to be solved as a sequence labeling task, and experiments show that standard sequence labelers can give competitive performance on the new treebank compared with state-of-the-art graph structure models.
Morphological segmentation has traditionally been modeled with non-hierarchical models, which yield flat segmentations as output. In many cases, however, proper morphological analysis requires hierarchical structureespecially in the case of derivational morphology. In this work, we introduce a discriminative, joint model of morphological segmentation along with the orthographic changes that occur during word formation. To the best of our knowledge, this is the first attempt to approach discriminative segmentation with a context-free model. Additionally, we release an annotated treebank of 7454 English words with constituency parses, encouraging future research in this area. 1
Abstract Three studies examined gender differences in the effect of storytelling ability on perceptions of a person's attractiveness as a short‐term and long‐term romantic partner. In Study 1, information about a potential partner's storytelling ability was provided. Study 2 participants read a good or poor story supposedly written by a potential partner. Results suggested that only women's attractiveness assessments of men as a long‐term date increased for good storytellers. Storytelling ability did not affect men's ratings of women nor did it affect ratings of short‐term partners. Study 3 suggested that the effect of storytelling ability on long‐term attractiveness for male targets may be mediated by perceived status. Storytelling ability appears to increase perceived status and thus helps men attract long‐term partners.
We propose a framework to model human comprehension of discourse connectives. Following the Bayesian pragmatic paradigm, we advocate that discourse connectives are interpreted based on a simulation of the production process by the speaker, who, in turn, considers the ease of interpretation for the listener when choosing connectives. Evaluation against the sense annotation of the Penn Discourse Treebank confirms the superiority of the model over literal comprehension. A further experiment demonstrates that the proposed model also improves automatic discourse parsing.
The Universal Dependencies (UD) Project seeks to build a cross-lingual studies of treebanks, linguistic structures and parsing. Its goal is to create a set of multilingual harmonized treebanks that are designed according to a universal annotation scheme. In this paper, we report on the conversion of the Uyghur dependency treebank to a UD version of the treebank which we term the Uyghur Universal Dependency Treebank (UyDT). We present the mapping of the Uyghur dependency treebank’s labelling scheme to the UD scheme, along with a clear description of the structural changes required in this conversion.
Term sense disambiguation is very essential for different approaches of NLP, including Internet search engines, information retrieval, Data mining, classification etc. However, the old methods using case frames and semantic primitives are not qualify for solving term ambiguities which needs a lot of information with sentences. This new approach introduces a building structure system of natural language knowledge. In this paper all surface case patterns is classified in advance with the consideration of the meaning of noun. Moreover, this paper introduces an efficient data structure using a trie which define the linkage among leaves and multi-attribute relations. By using this linkage multi-attribute relations, we can get a high frequent access among verbs and noun with an automatic generation of hierarchical relationships. In our experiment a large tagged corpus (Pan Treebank) is used to extract data. In our approach around 11,000 verbs and nouns is used for verifying the new method and made a hierarchy group of its noun. Moreover, the achievement of term disambiguating using our trie structure method and linking trie among leaves is 6% higher than old method.
In accordance with the compositionality criterion and hierarchy principle of Rhetorical Structure Theory (RST), this study reframes each tree in the RST Discourse Treebank into three new dependency trees with ultimate nodes being clauses, sentences, and paragraphs, respectively, which also draw on an analogy between syntactic and discourse trees. Detailed percentages of various RST relations at the three granularity levels are examined, illuminating the discourse processes of organizing units of one granularity level into those of the next upper level and suggesting certain homogeneity and interaction across levels in the Treebank, particularly at the two upper levels. The study demonstrates the applicability of RST analysis between same-level terminal units. With unique analytical advantages, the newly constructed discourse dependency trees provide new research prospects.
We propose a classification framework for semantic type identification of compounds in Sanskrit. We broadly classify the compounds into four different classes namely, Avyayībhāva, Tatpuruṣa, Bahuvrīhi and Dvandva. Our classification is based on the traditional classification system followed by the ancient grammar treatise Adṣṭādhyāyī, proposed by Pāṇini 25 centuries back. We construct an elaborate features space for our system by combining conditional rules from the grammar Adṣṭādhyāyī, semantic relations between the compound components from a lexical database Amarakoṣa and linguistic structures from the data using Adaptor Grammars. Our in-depth analysis of the feature space highlight inadequacy of Adṣṭādhyāyī, a generative grammar, in classifying the data samples. Our experimental results validate the effectiveness of using lexical databases as suggested by Amba Kulkarni and Anil Kumar, and put forward a new research direction by introducing linguistic patterns obtained from Adaptor grammars for effective identification of compound type. We utilise an ensemble based approach, specifically designed for handling skewed datasets and we %and Experimenting with various classification methods, we achieve an overall accuracy of 0.77 using random forest classifiers.
The Internet-scale open source software (OSS) production in various communities are generating abundant reusable resources for software developers. However, how to retrieve and reuse the desired and mature software from huge amounts of candidates is a great challenge: there are usually big gaps between the user application contexts (that often used as queries) and the OSS key words (that often used to match the queries). In this paper, we define the scenario-based query problem for OSS retrieval, and then we propose a novel approach to reformulate the raw query by leveraging the crowd wisdom from millions of developers to improve the retrieval results. We build a software-specific domain lexical database based on the knowledge in open source communities, by which we can expand and optimize the input queries. The experiment results show that, our approach can reformulate the initial query effectively and outperforms other existing search engines significantly at finding mature software.