Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
We present a novel abstractive summarization framework that draws on the recent development of a treebank for the Abstract Meaning Representation (AMR). In this framework, the source text is parsed to a set of AMR graphs, the graphs are transformed into a summary graph, and then text is generated from the summary graph. We focus on the graph-to-graph transformation that reduces the source semantic graph into a summary graph, making use of an existing AMR parser and assuming the eventual availability of an AMR-to-text generator. The framework is data-driven, trainable, and not specifically designed for a particular domain. Experiments on gold-standard AMR annotations and system parses show promising results. Code is available at: https://github.com/summarization
We explore dynamic evaluation, where sequence models are adapted to the recent sequence history using gradient descent, assigning higher probabilities to re-occurring sequential patterns. We develop a dynamic evaluation approach that outperforms existing adaptation approaches in our comparisons. We apply dynamic evaluation to outperform all previous word-level perplexities on the Penn Treebank and WikiText-2 datasets (achieving 51.1 and 44.3 respectively) and all previous character-level cross-entropies on the text8 and Hutter Prize datasets (achieving 1.19 bits/char and 1.08 bits/char respectively).
Chronic pain may alter both affect- and value-related behaviors, which represents a potentially treatable aspect of chronic pain experience. Current understanding of how chronic pain influences the function of brain reward systems, however, is limited. Using a monetary incentive delay task and functional magnetic resonance imaging (fMRI), we measured neural correlates of reward anticipation and outcomes in female participants with the chronic pain condition of fibromyalgia (N = 17) and age-matched, pain-free, female controls (N = 15). We hypothesized that patients would demonstrate lower positive arousal, as well as altered reward anticipation and outcome activity within corticostriatal circuits implicated in reward processing. Patients demonstrated lower arousal ratings as compared with controls, but no group differences were observed for valence, positive arousal, or negative arousal ratings. Group fMRI analyses were conducted to determine predetermined region of interest, nucleus accumbens (NAcc) and medial prefrontal cortex (mPFC), responses to potential gains, potential losses, reward outcomes, and punishment outcomes. Compared with controls, patients demonstrated similar, although slightly reduced, NAcc activity during gain anticipation. Conversely, patients demonstrated dramatically reduced mPFC activity during gain anticipation-possibly related to lower estimated reward probabilities. Further, patients demonstrated normal mPFC activity to reward outcomes, but dramatically heightened mPFC activity to no-loss (nonpunishment) outcomes. In parallel to NAcc and mPFC responses, patients demonstrated slightly reduced activity during reward anticipation in other brain regions, which included the ventral tegmental area, anterior cingulate cortex, and anterior insular cortex. Together, these results implicate altered corticostriatal processing of monetary rewards in chronic pain.
Previous studies have demonstrated differential perception of body expressions between males and females. However, only two recent studies (Kret et al., 2011; Krüger et al., 2013) explored the interaction effect between observer gender and subject gender, and it remains unclear whether this interaction between the two gender factors is gender-congruent (i.e., better recognition of emotions expressed by subjects of the same gender) or gender-incongruent (i.e., better recognition of emotions expressed by subjects of the opposite gender). Here, we used event-related potentials (ERPs) to investigate the recognition of fearful and angry body expressions posed by males and females. Male and female observers also completed an affective rating task (including valence, intensity, and arousal ratings). Behavioral results showed that male observers reported higher arousal rating scores for angry body expressions posed by females than males. ERP data showed that when recognizing angry body expressions, female observers had larger P1 for male than female bodies, while male observers had larger P3 for female than male bodies. These results indicate gender-incongruent effects in early and later stages of body expression processing, which fits well with the evolutionary theory that females mainly play a role in care of offspring while males mainly play a role in family guarding and protection. Furthermore, it is found that in both angry and fearful conditions male observers exhibited a larger N170 for male than female bodies, and female observers showed a larger N170 for female than male bodies. This gender-incongruent effect in the structural encoding stage of processing may be due to the familiarity of the body configural features of the same gender. The current results provide insights into the significant role of gender in body expression processing, helping us understand the issue of gender vulnerability associated with psychiatric disorders characterized by deficits of body language reading.
Recent work on the problem of latent tree learning has made it possible to train neural networks that learn to both parse a sentence and use the resulting parse to interpret the sentence, all without exposure to ground-truth parse trees at training time. Surprisingly, these models often perform better at sentence understanding tasks than models that use parse trees from conventional parsers. This paper aims to investigate what these latent tree learning models learn. We replicate two such models in a shared codebase and find that (i) only one of these models outperforms conventional tree-structured models on sentence classification, (ii) its parsing strategies are not especially consistent across random restarts, (iii) the parses it produces tend to be shallower than standard Penn Treebank (PTB) parses, and (iv) they do not resemble those of PTB or any other semantic or syntactic formalism that the authors are aware of.
Released only a year ago as the outputs of a research project (``Parsing Web 2.0 Sentences'', supported in part by a TÜBİTAK 1001 grant (No. 112E276) and a part of the ICT COST Action PARSEME (IC1207)), IMST and IWT are currently the most comprehensive Turkish dependency treebanks in the literature. This article introduces the final states of our treebanks, as well as a newly integrated hierarchical categorization of the multiheaded dependencies and their organization in an exclusive deep dependency layer in the treebanks. It also presents the adaptation of recent studies on standardizing multiword expression and named entity annotation schemes for the Turkish language and integration of benchmark annotations into the dependency layers of our treebanks and the mapping of the treebanks to the latest Universal Dependencies (v2.0) standard, ensuring further compliance with rising universal annotation trends. In addition to significantly boosting the universal recognition of Turkish treebanks, our recent efforts have shown an improvement in their syntactic parsing performance (up to 77.8{\%}/82.8{\%} LAS and 84.0{\%}/87.9{\%} UAS for IMST/IWT, respectively). The final states of the treebanks are expected to be more suited to different natural language processing tasks, such as named entity recognition, multiword expression detection, transfer-based machine translation, semantic parsing, and semantic role labeling.
This paper presents results from the first statistical dependency parser for Turkish. Turkish is a free-constituent order language with complex agglutinative inflectional and derivational morphology and presents interesting challenges for statistical parsing, as in general, dependency relations are between “portions” of words – called inflectional groups. We have explored statistical models that use different representational units for parsing. We have used the Turkish Dependency Treebank to train and test our parser but have limited this initial exploration to that subset of the treebank sentences with only left-to-right non-crossing dependency links. Our results indicate that the best accuracy in terms of the dependency relations between inflectional groups is obtained when we use inflectional groups as units in parsing, and when contexts around the dependent are employed.
In literature, writers have the liberty to deviate from linguistic norms under a principle known as poetic license. Poetic license allows deviation in favour of making language inspiring. Deviation from linguistic norms often implies that writers can take liberties with word formation, thus neology in literary contexts should be addressed specifically. This article analyses the status of literary coinages in the scope of neology and describes the specific context of children’s literature. The article also offers a typology of nonce formation processes for occasionalisms, with textual analysis, from a corpus of children’s books, using J. Tournier’s matrices of lexicogenesis [2007: 51].
This paper presents a truly full character-level neural dependency parser together with a newly released character-level dependency treebank for Chinese, which has suffered a lot from the dilemma of defining word or not to model character interactions. Integrating full character-level dependencies with character embedding and human annotated character-level part-of-speech and dependency labels for the first time, we show an extra performance enhancement from the evaluation on Chinese Penn Treebank and SJTU (Shanghai Jiao Tong University) Chinese Character Dependency Treebank and the potential of better understanding deeper structure of Chinese sentences.
To approximately parse an unfamiliar language, it helps to have a treebank of a similar language. But what if the closest available treebank still has the wrong word order? We show how to (stochastically) permute the constituents of an existing dependency treebank so that its surface part-of-speech statistics approximately match those of the target language. The parameters of the permutation model can be evaluated for quality by dynamic programming and tuned by gradient descent (up to a local optimum). This optimization procedure yields trees for a new artificial language that resembles the target language. We show that delexicalized parsers for the target language can be successfully trained using such "made to order" artificial languages.
As the amount of unstructured text data that humanity produces overall and on the Internet grows, so does the need to intelligently to process it and extract different types of knowledge from it. Convolutional neural networks (CNNs) and recurrent neural networks (RNNs) have been applied to natural language processing systems with comparative, remarkable results. The CNN is a noble approach to extract higher level features that are invariant to local translation. However, it requires stacking multiple convolutional layers in order to capture long-term dependencies, due to the locality of the convolutional and pooling layers. In this paper, we describe a joint CNN and RNN framework to overcome this problem. Briefly, we use an unsupervised neural language model to train initial word embeddings that are further tuned by our deep learning network, then, the pre-trained parameters of the network are used to initialize the model. At a final stage, the proposed framework combines former information with a set of feature maps learned by a convolutional layer with long-term dependencies learned via long-short-term memory. Empirically, we show that our approach, with slight hyperparameter tuning and static vectors, achieves outstanding results on multiple sentiment analysis benchmarks. Our approach outperforms several existing approaches in term of accuracy; our results are also competitive with the state-of-the-art results on the Stanford Large Movie Review data set with 93.3% accuracy, and the Stanford Sentiment Treebank data set with 48.8% fine-grained and 89.2% binary accuracy, respectively. Our approach has a significant role in reducing the number of parameters and constructing the convolutional layer followed by the recurrent layer as a substitute for the pooling layer. Our results show that we were able to reduce the loss of detailed, local information and capture long-term dependencies with an efficient framework that has fewer parameters and a high level of performance.
Recently, researchers have developed black-box approaches to mine design and interaction data from mobile apps. Although the data captured during this interaction mining is descriptive, it does not expose the design semantics of UIs: what elements on the screen mean and how they are used. This paper introduces an automatic approach for generating semantic annotations for mobile app UIs. Through an iterative open coding of 73k UI elements and 720 screens, we contribute a lexical database of 25 types of UI components, 197 text button concepts, and 135 icon classes shared across apps. We use this labeled data to learn code-based patterns to detect UI components and to train a convolutional neural network that distinguishes between icon classes with 94% accuracy. To demonstrate the efficacy of our approach at scale, we compute semantic annotations for the 72k unique UIs in the Rico dataset, assigning labels for 78% of the total visible, non-redundant elements.
We introduce a novel architecture for dependency parsing: stack-pointer networks (STACKPTR). Combining pointer networks The stack tracks the status of the depthfirst search and the pointer networks select one child for the word at the top of the stack at each step. The STACKPTR parser benefits from the information of the whole sentence and all previously derived subtree structures, and removes the leftto-right restriction in classical transitionbased parsers. Yet, the number of steps for building any (including non-projective) parse tree is linear in the length of the sentence just as other transition-based parsers, yielding an efficient decoding algorithm with O(n 2 ) time complexity. We evaluate our model on 29 treebanks spanning 20 languages and different dependency annotation schemas, and achieve state-of-theart performance on 21 of them.
In this paper, we propose RNN-Capsule, a capsule model based on Recurrent Neural Network (RNN) for sentiment analysis. For a given problem, one capsule is built for each sentiment category e.g., 'positive' and 'negative'. Each capsule has an attribute, a state, and three modules: representation module, probability module, and reconstruction module. The attribute of a capsule is the assigned sentiment category. Given an instance encoded in hidden vectors by a typical RNN, the representation module builds capsule representation by the attention mechanism. Based on capsule representation, the probability module computes the capsule's state probability. A capsule's state is active if its state probability is the largest among all capsules for the given instance, and inactive otherwise. On two benchmark datasets (i.e., Movie Review and Stanford Sentiment Treebank) and one proprietary dataset (i.e., Hospital Feedback), we show that RNN-Capsule achieves state-of-the-art performance on sentiment classification. More importantly, without using any linguistic knowledge, RNN-Capsule is capable of outputting words with sentiment tendencies reflecting capsules' attributes. The words well reflect the domain specificity of the dataset.
We demonstrate that replacing an LSTM encoder with a self-attentive architecture can lead to improvements to a state-ofthe-art discriminative constituency parser. The use of attention makes explicit the manner in which information is propagated between different locations in the sentence, which we use to both analyze our model and propose potential improvements. For example, we find that separating positional and content information in the encoder can lead to improved parsing accuracy. Additionally, we evaluate different approaches for lexical representation. Our parser achieves new state-ofthe-art results for single models trained on the Penn Treebank: 93.55 F1 without the use of any external data, and 95.13 F1 when using pre-trained word representations. Our parser also outperforms the previous best-published accuracy figures on 8 of the 9 languages in the SPMRL dataset.
To fasten treebank construction, it is necessary to design an integrated annotation tool that includes word segmenter, sentence parser for initial tree suggestion, tree visualizer, tree-structure editor, and collaborative functions. In the past, existing tools did not consider an integrated platform that provides preprocessing, automated or semi-automated mechanism for parse tree suggestion, as well as tagged corpus data management. This paper presents a so-called CF Planter, a toolset for semi-automatic Thai treebank construction that consist of word segmenter, part-of-speech tagger, statistical parser, a web-based GUI for syntactic tree refinement and management. Given an input sentence, its most likely syntactic tree is automatically suggested and visualized to an annotator for manual correction before adding into the treebank repository. Whenever a new syntactic tree is appended into the treebank, the treebank repository is iteratively refined by computing a set of newly revised grammar rules based on revised probabilities. Toolset is performed to severally illustrate with grammar frequencies. The toolset facilitates annotators to easily tag tree structure for an input sentence. Finally, the process of automatic suggestion of syntactic tree is evaluated.
This paper presents a methodology for rule based bottom up parsing technique forModern Standard Arabic (MSA) inContext Free Grammar (CFG) formalism in Phrase Structure Grammar (PSG) representation, where the grammar isautomatically extracted from a syntactically annotated corpus.The extracted grammar is used to build an automatic lexicon andgrammar rules module. Furthermore, the extracted CFG is further transformed into Probabilistic Context Free Grammar (PCFG)that could be used in a hybrid approach, which is also calculated automatically. The used corpus is the Penn ArabicTreebank(PATB)and algorithm implementation is performed with Natural Language Processing Toolkit (NLTK).The parsershowed that automatic extraction of grammar improved the grammar building phase in both coverage of structures and timeneeded, but still needs further manual constrains addition. Automatic extraction of grammar is able to enhance rule basedgrammar parsers and it will enable a new paradigm of statistically directed symbolic parsing.
Detection and correction of errors and inconsistencies in "gold treebanks" are becoming more and more central topics of corpus annotation. The paper illustrates a new incremental method for enhancing treebanks, with particular emphasis on the extension of error patterns across different textual genres and registers. Impact and role of corrections have been assessed in a dependency parsing experiment carried out with four different parsers, whose results are promising. For both evaluation datasets, the performance of parsers increases, in terms of the standard LAS and UAS measures and of a more focused measure taking into account only relations involved in error patterns, and at the level of individual dependencies.
Recurrent neural nets (RNN) and convolutional neural nets (CNN) are widely used on NLP tasks to capture the long-term and local dependencies, respectively. Attention mechanisms have recently attracted enormous interest due to their highly parallelizable computation, significantly less training time, and flexibility in modeling dependencies. We propose a novel attention mechanism in which the attention between elements from input sequence(s) is directional and multi-dimensional (i.e., feature-wise). A light-weight neural net, "Directional Self-Attention Network (DiSAN)," is then proposed to learn sentence embedding, based solely on the proposed attention without any RNN/CNN structure. DiSAN is only composed of a directional self-attention with temporal order encoded, followed by a multi-dimensional attention that compresses the sequence into a vector representation. Despite its simple form, DiSAN outperforms complicated RNN models on both prediction quality and time efficiency. It achieves the best test accuracy among all sentence encoding methods and improves the most recent best result by 1.02% on the Stanford Natural Language Inference (SNLI) dataset, and shows state-of-the-art test accuracy on the Stanford Sentiment Treebank (SST), Multi-Genre natural language inference (MultiNLI), Sentences Involving Compositional Knowledge (SICK), Customer Review, MPQA, TREC question-type classification and Subjectivity (SUBJ) datasets.
We present PAWS, a multi-lingual parallel treebank with coreference annotation. It consists of English texts from the Wall Street Journal translated into Czech, Russian and Polish. In addition, the texts are syntactically parsed and word-aligned. PAWS is based on PCEDT 2.0 and continues the tradition of multilingual treebanks with coreference annotation. The paper focuses on the coreference annotation in PAWS and its language-specific differences. PAWS offers linguistic material that can be further leveraged in cross-lingual studies, especially on coreference.
We propose Efficient Neural Architecture Search (ENAS), a fast and inexpensive approach for automatic model design. In ENAS, a controller learns to discover neural network architectures by searching for an optimal subgraph within a large computational graph. The controller is trained with policy gradient to select a subgraph that maximizes the expected reward on the validation set. Meanwhile the model corresponding to the selected subgraph is trained to minimize a canonical cross entropy loss. Thanks to parameter sharing between child models, ENAS is fast: it delivers strong empirical performances using much fewer GPU-hours than all existing automatic model design approaches, and notably, 1000x less expensive than standard Neural Architecture Search. On the Penn Treebank dataset, ENAS discovers a novel architecture that achieves a test perplexity of 55.8, establishing a new state-of-the-art among all methods without post-training processing. On the CIFAR-10 dataset, ENAS designs novel architectures that achieve a test error of 2.89%, which is on par with NASNet (Zoph et al., 2018), whose test error is 2.65%.
This paper describes Stanford's system at the CoNLL 2018 UD Shared Task. We introduce a complete neural pipeline system that takes raw text as input, and performs all tasks required by the shared task, ranging from tokenization and sentence segmentation, to POS tagging and dependency parsing. Our single system submission achieved very competitive performance on big treebanks. Moreover, after fixing an unfortunate bug, our corrected system would have placed the 2 nd, 1 st, and 3 rd on the official evaluation metrics LAS, MLAS, and BLEX, and would have outperformed all submission systems on lowresource treebank categories on all metrics by a large margin. We further show the effectiveness of different model components through extensive ablation studies. * These authors contributed roughly equally.
When adding enhanced dependencies to an existing UD treebank, one can opt for heuristics that predict the enhanced dependencies on the basis of the UD annotation only. If the treebank is the result of conversion from an underlying treebank, an alternative is to produce the enhanced dependencies directly on the basis of this underlying annotation. Here we present a method for doing the latter for the Dutch UD treebanks. We compare our method with the UD -based approach of Schuster et al. (2018). While there are a number of systematic differences in the output of both methods, it appears these are the result of insufficient detail in the annotation guidelines and it is not the case that one approach is superior over the other in principle.
The performance of a machine translation system depends on the availability of bilingual lexical dictionary and completion of its word sense disambiguation performance. Word sense disambiguation plays a vital role in several applications such as machine translation, information retrieval and many other Natural Language Processing (NLP). In oder to construct a reliable machine translation system, not only consistent lexical database like WordNet (WN), but also word sense distinguished database, and bilingual lexical machine readable dictionary (MRD) are required to achieve an accurate translation. Since WN is for the English language, there have been several approcaches building WN like database for other languages, also known as multilingual WN, by linking and extending bilingual MRDs with synsets from WN. Lexical database like WN for Myanmar languae have been proposed in recent years which are bilingual lexical databases as the synsets from WN are merged with bilingual MRD. However, there are few limitations in constructing bilingual database based on MRD where thesauries are derived from WN yet there are no directed alignemnt of meanings between English and Myanmar language. In order to design a complete lexical database with correct parallel meanings for both languages, our proposed method composes a model which accurately matches the similar meanings between English words and Myanmar words by translating and comparing the word senses for both languages. The experimental result shows that the meanings for both languages are aligned as accurate as the manual alignment of the word senses for lexical database.
This paper describes our system (HIT-SCIR) submitted to the CoNLL 2018 shared task on Multilingual Parsing from Raw Text to Universal Dependencies. We base our submission on Stanford's winning system for the CoNLL 2017 shared task and make two effective extensions: 1) incorporating deep contextualized word embeddings into both the part of speech tagger and dependency parser; 2) ensembling parsers trained with different initialization. We also explore different ways of concatenating treebanks for further improvements. Experimental results on the development data show the effectiveness of our methods. In the final evaluation, our system was ranked first according to LAS (75.84%) and outperformed the other systems by a large margin.
We know very little about how neural language models (LM) use prior linguistic context.In this paper, we investigate the role of context in an LSTM LM, through ablation studies.Specifically, we analyze the increase in perplexity when prior context words are shuffled, replaced, or dropped.On two standard datasets, Penn Treebank and WikiText-2, we find that the model is capable of using about 200 tokens of context on average, but sharply distinguishes nearby context (recent 50 tokens) from the distant history.The model is highly sensitive to the order of words within the most recent sentence, but ignores word order in the long-range context (beyond 50 tokens), suggesting the distant past is modeled only as a rough semantic field or topic.We further find that the neural caching model (Grave et al., 2017b) especially helps the LSTM to copy words from within this distant context.Overall, our analysis not only provides a better understanding of how neural LMs use their context, but also sheds light on recent success from cache-based models.
We present the Uppsala system for the CoNLL 2018 Shared Task on universal dependency parsing. Our system is a pipeline consisting of three components: the first performs joint word and sentence segmentation; the second predicts part-ofspeech tags and morphological features; the third predicts dependency trees from words and tags. Instead of training a single parsing model for each treebank, we trained models with multiple treebanks for one language or closely related languages, greatly reducing the number of models. On the official test run, we ranked 7th of 27 teams for the LAS and MLAS metrics. Our system obtained the best scores overall for word segmentation, universal POS tagging, and morphological features.
The significance of tourism within the ASEAN region is recognised by multiple stakeholders. Presenting and promoting a distinctive image of tourism is a common agenda across the ten ASEAN nations. The aim of this study is to document and interpret the emotional connotations of ASEAN tourism slogans. Arguably, such messages provide an initial guide to the appeal and competitive advantages of each individual country. The study is underpinned by considering key ideas on destination positioning and the lexical analysis of emotions. By mining archival resources about word frequencies, synonyms and meanings, the positions of the slogans in an emotion space originally developed by Plutchik were compared and plotted. Joy, admiration and ecstasy were the dominant emotional connotations of most slogans. Thailand and Malaysia have the most distinctive tourism slogans, followed by Vietnam and Laos. Expressions used in the slogans for these four nations overlapped less with other countries across the families of emotion words.
Developed by the CLLD project with support from the Department of Linguistic and Cultural Evolution of the Max Planck Institute for the Science of Human History.
In the central nervous system the neuropeptide oxytocin mediates a range of behaviors related primarily to emotionality. One factor that influences oxytocinergic communication in the human brain and correlates with emotional behaviors is the single nucleotide polymorphism rs53576 on the oxytocin receptor gene (OXTR). For example, variations in this OXTR genotype are related to parental, altruistic, and other prosocial behaviors. Electroencephalographic waveforms of visually evoked response potentials recorded at the midline parietal electrode site display a prominent component putatively involved with attention allocation called the late positive potential. The magnitude of the late positive potential was found to be significantly higher in homozygous G allele individuals compared with A allele carriers when viewing negative emotionally charged images. Inversely, A allele carriers rated these negative images as more arousing, when measured by the Self-Assessment Manikin rating scale. These data suggest that OXTR functioning contributes to visual processing and subjective experience of negative stimuli.
This paper addresses the scalability challenge of architecture search by formulating the task in a differentiable manner. Unlike conventional approaches of applying evolution or reinforcement learning over a discrete and non-differentiable search space, our method is based on the continuous relaxation of the architecture representation, allowing efficient search of the architecture using gradient descent. Extensive experiments on CIFAR-10, ImageNet, Penn Treebank and WikiText-2 show that our algorithm excels in discovering high-performance convolutional architectures for image classification and recurrent architectures for language modeling, while being orders of magnitude faster than state-of-the-art non-differentiable techniques. Our implementation has been made publicly available to facilitate further research on efficient architecture search algorithms.
This lexical cognate data was exported from the Indo-European Lexical Database (IELex) supporting the publication: Verkerk, Annemarie. (2018). Detecting non-tree-like signal using multiple tree topologies. Journal of Historical Linguistics. If you are looking to use the IELex data in your own research, we recommend you use the latest version, which is available along with documentation, phylogenetic tree samples and example BEAST control scripts.
The present paper shows how the current Universal Dependency treebanks can be used for language typology studies and can reveal structural syntactic features of languages. Two methods, one existing method and one newly proposed method, based on dependency treebanks as typological measurements, are presented and tested in order to assess both the coherence of the underlying syntactic data and the validity of the methods themselves. The results show that both methods are valid for positioning a language in the typological continuum, although they probably reveal different typological features of languages.
Cilin is one of the most popular semantic knowledge bases in Chinese information processing. Due to its coding scheme and taxonomical arrangement, some semantic relations among words are not explicitly shown. Our work aims to characterize its semantic relations by adding tags and compound codes to optimize the taxonomy and hierarchy of Cilin. Experiment results show that using the tag-augmented Cilin as a knowledge base improves the performance in the task of semantic similarity measurement.
Many recent studies, academic and non-academic alike, have argued that the use of Urdu in Bollywood has started to decline. These studies, important as they are, however, suffer from some limitations. They are either impressionistic or based on non-representative data. Furthermore, they do not specify the object of the study or the site of the assumed decline of Urdu. Therefore, it remains vague which element of Urdu, for example sounds, words, syntax or script is under investigation. Similarly, it is not clear which component of film for example titles, dialogues or songs are experiencing the decline. Fulfilling this research gap this paper makes two contributions. Analyzing songs from 1959 to 2010’s, it empirically demonstrates the decline by documenting the shift in the pronunciation of the sounds /kh̲/, /gh̲/, and /q/ from the Urdu to Hindi phonetic norms. Singers from the 1990’s, unlike those from the previous generations, merge them with the sounds /kh/, /gh/, and /q/. The paper also makes a methodological contribution in that it shows how language in cinema can be studied empirically using a corpus.
We conducted a semantic similarity study of semantic concepts in the context of the Holy Book Quran. Semantic similarity examines the degree of likeness and shared common properties of two concepts. For example, the Quranic concept of Allah and God will result in a high score of semantic similarity, whereas hell and paradise will yield in a low score because of its extremely different attributes and semantic features. Apart from that, we also delivered the Quranic concept semantic similarity standard dataset which consists of some pairs of Quranic concept along with its similarity score, which was manually annotated by human raters. This dataset resulted in the score of inter-annotator agreement 0.63, not far from the the ones yielded by some well-known datasets such as WordSim and Simlex. Furthermore, to measure the semantic similarity score, we chose the knowledge-based approach by utilizing lexical database properties such as the length and depth of a synonym set (synset). We then applied it to Yuhua Li equation, which has been considered to be the baseline among researchers within the problem of semantic similarity. In terms of the result, our system gained Pearson's correlation 0.33 and Spearman's 0.19. By considering inter-annotator agreement 0.63 that our Quranic standard dataset has as the upper bound score, there are still quite large room for improvement to better mimicking Muslim's intuition to measure the degree of similarity of concepts within the domain of Quran.
Dependency distance minimization (DDM) is found as a universal quantitative property of natural languages. To investigate whether second language learners develop their interlanguage system under the pressure of DDM, we selected 367 Chinese EFL learners of nine consecutive grades, built one second language dependency treebank and two corresponding random treebanks and fitted different probability distribution models to dependency distances. It was found that: (1) The mean dependency distance (MDD) of interlanguage increases significantly across nine grades and the MDD of high-level learners doesn't reach the level of English native speakers. (2) The MDDs of interlanguage at different learning phases are significantly lower than their corresponding random languages (RL1 and RL2), indicating that learners develop their English proficiency under the pressure of DDM. (3) The distribution of dependency distances of RL1 cannot fit the Zipf-Alekseev distribution, but that of RL2 can. The parameters in the Zipf-Alekseev distribution of RL2 have no correlation with learners' language proficiency.
L’article présente plusieurs normes utilisées pour la représentation des données dans les dictionnaires électroniques et lexiques destinés aux outils de Traitement automatique des Langues (TAL). Les normes présentées préconisent la représentation des informations linguistiques dans des dictionnaires électroniques selon le modèle TEI (Text Encoding Initiative) et le modèle LMF (Lexical Markup Framework). Nous nous intéressons en particulier aux dictionnaires de collocations à l’adaptation du modèle LMF pour la représentation de ce type de données.
It has been quite a challenge to diagnose Mild Cognitive Impairment due to Alzheimer’s disease (MCI) and Alzheimer-type dementia (AD-type dementia) using the currently available clinical diagnostic criteria and neuropsychological examinations. As such we propose an automated diagnostic technique using a variant of deep neural networks language models (DNNLM) on the verbal utterances of affected individuals. Motivated by the success of DNNLM on natural language tasks, we propose a combination of deep neural network and deep language models (D2NNLM) for classifying the disease. Results on the DementiaBank language transcript clinical dataset show that D2NNLM sufficiently learned several linguistic biomarkers in the form of higher order n-grams to distinguish the affected group from the healthy group with reasonable accuracy on very sparse clinical datasets. [ABSTRACT FROM AUTHOR], Copyright of PLoS ONE is the property of Public Library of Science and its content may not be copied or ema)
In this article the associations of the alla lexeme with the national-cultural seme in Uzbek language were studied on the basis of associative experiment and analyzed linguistic. The understanding of the Turkic language about the alla lexeme, the linguistic memory, the reserve and the knowledge of the lexical units were clarified.This testifies on the existence of an associative-conceptual principle in lexical system of the Uzbek language. The observed standardization in the lexical association of native speakers of the Uzbek language has much in common with the processes of lexical association in other languages. This proves the existence of associative universals among speakers of different languages and about the similar structure of the world reflected in the presentation of speakers of these languages. In the associative lexical system of the Uzbek language, a special place is occupied by specific reactions that reflect different realities Uzbek life. This national specificity of association distinguishes one nation from another and is like a symbol of the national culture. There had been compiled the first ever dictionary of the Uzbek language associative norms. Typically Uzbek standards are normally reflecting national picture of the world existing in the consciousness of Uzbek language speakers and demonstrate the national spirit of the language and the people.
Under creative speaking grammatical and extra-linguistic factors are interplaying. Therefor classification of such processes is still relevant. Potential words as the system determined type of ad-hock creativity can be classified according to the grade of deviation, concerning as derivation model components, as non-linguistic factors action, influence of social context. Each model has more or less strict components parameters: formant meaning and definite part of speech for stem. During communication speaker can actualize any of the forms, which one lexical group paradigms have. Created form can appear despite the model maturity and this is the way of its occasionality rising. If there is no lacuna in a model paradigm so some form, duplicating another one, reveals special pragmatic need of the speaker. Considering distance between system norms and real speech circumstances, we are able to fix difference among potential words as more or less occasional. So we get a scale with conventional numerical indication of the occasionality growth – from zero to one with a half index for cases of incomplete characteristic. The last one means light deviation of the model: formant meaning variation or part of speech shifting for stem, or weak perlocution in the structure of intention. Zero index marks model safety, available system gap for the potential form and absence in the speech act any perlocution or social ranging. The index “one” fixes opposite characteristic: a deep model deviation (e.g. stem structure breach), system gap absence for the potential form or duplicating paradigm fragment and explicit under the speech act perlocution or social ranging. The scale can be used as the base for forecasting of the usage “future” for each potential word, if the absolute zero and absolute one will be considered as the indexes of system integration and social demand.
A survey of reported comparative constructions in the Koyukon, Ahtna and Tanana Athabascan languages of Alaska shows that many fall into A dimensional verb is accompanied by a modifying postpositional phrase, with the standard being the object of the postposition. Superlatives are not as well represented in lexical documentation as comparatives, which are themselves rare in texts and difficult to elicit. Structured elicitation of comparatives and superlatives in Ahtna and Koyukon supports observations that this rarity is related to cultural norms in Athabascan communities, where comparison (especially of people) can be considered rude, and superlatives evidence of inappropriate pride.
The Trail Making Test (TMT) is used in neuropsychological clinical practice to assess aspects of attention and executive function. The test consists of two parts (A and B) and requires drawing a trail between elements. Many patients are assessed with their non-dominant hand because of motor dysfunction that prevents them from using their dominant hand. Since drawing with the non-dominant hand is not an automatic task for many people, we explored the effect of hand use on TMT performance. The TMT was administered digitally in order to analyze new outcome measures in addition to total completion time. In a sample of 82 healthy participants, we found that non-dominant hand use increased completion times on the TMT B but not on the TMT A. The average completion time increased by almost 5 seconds, which may be clinically relevant. A substantial number of participants who performed the TMT with their non-dominant hand had a B/A ratio score of 2.5 or higher. In clinical practice, an abnormally high B/A ratio score may be falsely attributed to cognitive dysfunction. With our digitized pen data, we further explored the causes of the reduced TMT B performance by using new outcome measures, including individual element completion times and interelement variability. These measures indicated selective interference between non-dominant hand use and executive functions. Both non-dominant hand use and performance of the TMT B seem to draw on the same, limited higher-order cognitive resources.
Linguistic register reflects changes in speech that depend on the situation, especially the status of listeners and listener-speaker relationships. Following the sociolinguistic rules of register is essential in establishing and maintaining social interactions. Recent research suggests that children over 3 years of age can understand appropriate register-listener relationships as well as the fact that people change register depending on their listeners. However, given previous findings that infants under 2 years of age have already formed both social and speech categories, it may be possible that even younger children can also understand appropriate register-listener relationships. The present study used Infant-Directed Speech (IDS) and formal Adult-Directed Speech (ADS) to examine whether 20-month-old toddlers can understand register-listener relationships. In Experiment 1, we used a violation-of-expectation method to examine whether 20-month-olds understand the individual associatio)