Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
This chapter takes up the issue of authenticity in language pedagogy. Traditional views of authenticity take the native speaker to be the primary authority for linguistic norms. Written standard language is especially highly valued here. It is argued herein that TELL environments are equally valid as learning environments, and that students can use the freedom they provide to develop their own locally negotiated cultural and linguistic norms. Evidence is provided that students on a net-based MA program develop their own norms for reducing language, and use them and other means to mark membership of a local TELL community. Thus, TELL is a rich and authentic environment for learners of English to become what is referred to as “language practitioners.”
Abstract In the context of the Index Thomisticus Treebank project, we have enhanced the full text of Bellum Catilinae by Sallust with semantic annotation. The annotation style resembles the one used for the so called “tectogrammatical” layer of the Prague Dependency Treebank. By exploiting the results of semantic role labeling, ellipsis resolution and coreference analysis, this paper presents a network-based study of the main Actors and Actions (and their relations) in Bellum Catilinae.
並列構造解析の主たるタスクは並列する句の範囲を同定することである.並列構造は文の構文・意味の解析において有用な特徴となるが,これまで決定的な解析手法が確立されておらず,現在の最高精度の構文解析器においても誤りを生じさせる主たる要因となっている.既存の並列句範囲の曖昧性解消手法は並列構造の類似性のみの特性や構文解析器の結果に強く依存しているという問題があった.本研究では,近年自然言語解析に広く使用されているリカレントニューラルネットワークを用いて,構文解析の結果を用いずに単語の表層形と品詞情報のみから並列句の類似性と可換性の特徴ベクトルを計算し,並列構造の範囲を予測する手法を提案する.Penn Treebank と GENIA コーパスを用いた実験の結果,提案手法によって先行研究を上回る解析精度を得た.
Accurate natural language processing systems rely heavily on annotated datasets. In the absence of such datasets, transfer methods can help to develop a model by transferring annotations from one or more rich-resource languages to the target language of interest. These methods are generally divided into two approaches: 1) annotation projection from translation data, aka parallel data, using supervised models in rich-resource languages, and 2) direct model transfer from annotated datasets in rich-resource languages. In this thesis, we demonstrate different methods for transfer of dependency parsers and sentiment analysis systems. We propose an annotation projection method that performs well in the scenarios for which a large amount of in-domain parallel data is available. We also propose a method which is a combination of annotation projection and direct transfer that can leverage a minimal amount of information from a small out-of-domain parallel dataset to develop highly accurate transfer models. Furthermore, we propose an unsupervised syntactic reordering model to improve the accuracy of dependency parser transfer for non-European languages. Finally, we conduct a diverse set of experiments for the transfer of sentiment analysis systems in different data settings. A summary of our contributions are as follows: * We develop accurate dependency parsers using parallel text in an annotation projection framework. We make use of the fact that the density of word alignments is a valuable indicator of reliability in annotation projection. * We develop accurate dependency parsers in the absence of a large amount of parallel data. We use the Bible data, which is in orders of magnitude smaller than a conventional parallel dataset, to provide minimal cues for creating cross-lingual word representations. Our model is also capable of boosting the performance of annotation projection with a large amount of parallel data. Our model develops cross-lingual word representations for going beyond the traditional delexicalized direct transfer methods. Moreover, we propose a simple but effective word translation approach that brings in explicit lexical features from the target language in our direct transfer method. * We develop different syntactic reordering models that can change the source treebanks in rich-resource languages, thus preventing learning a wrong model for a non-related language. Our experimental results show substantial improvements over non-European languages. * We develop transfer methods for sentiment analysis in different data availability scenarios. We show that we can leverage cross-lingual word embeddings to create accurate sentiment analysis systems in the absence of annotated data in the target language of interest. We believe that the novelties that we introduce in this thesis indicate the usefulness of transfer methods. This is appealing in practice, especially since we suggest eliminating the requirement for annotating new datasets for low-resource languages which is expensive, if not impossible, to obtain.
The contents and structure of semantic memory have been the focus of much recent research, with major advances in the development of distributional models, which use word co-occurrence information as a window into the semantics of language. In parallel, connectionist modeling has extended our knowledge of the processes engaged in semantic activation. However, these two lines of investigation have rarely been brought together. Here, we describe a processing model based on distributional semantics in which activation spreads throughout a semantic network, as dictated by the patterns of semantic similarity between words. We show that the activation profile of the network, measured at various time points, can successfully account for response times in lexical and semantic decision tasks, as well as for subjective concreteness and imageability ratings. We also show that the dynamics of the network is predictive of performance in relational semantic tasks, such as similarity/relatedness rating. Our results indicate that bringing together distributional semantic networks and spreading of activation provides a good fit to both automatic lexical processing (as indexed by lexical and semantic decisions) as well as more deliberate processing (as indexed by ratings), above and beyond what has been reported for previous models that take into account only similarity resulting from network structure.
Temporal models based on recurrent neural networks have proven to be quite\npowerful in a wide variety of applications. However, training these models\noften relies on back-propagation through time, which entails unfolding the\nnetwork over many time steps, making the process of conducting credit\nassignment considerably more challenging. Furthermore, the nature of\nback-propagation itself does not permit the use of non-differentiable\nactivation functions and is inherently sequential, making parallelization of\nthe underlying training process difficult. Here, we propose the Parallel\nTemporal Neural Coding Network (P-TNCN), a biologically inspired model trained\nby the learning algorithm we call Local Representation Alignment. It aims to\nresolve the difficulties and problems that plague recurrent networks trained by\nback-propagation through time. The architecture requires neither unrolling in\ntime nor the derivatives of its internal activation functions. We compare our\nmodel and learning procedure to other back-propagation through time\nalternatives (which also tend to be computationally expensive), including\nreal-time recurrent learning, echo state networks, and unbiased online\nrecurrent optimization. We show that it outperforms these on sequence modeling\nbenchmarks such as Bouncing MNIST, a new benchmark we denote as Bouncing\nNotMNIST, and Penn Treebank. Notably, our approach can in some instances\noutperform full back-propagation through time as well as variants such as\nsparse attentive back-tracking. Significantly, the hidden unit correction phase\nof P-TNCN allows it to adapt to new datasets even if its synaptic weights are\nheld fixed (zero-shot adaptation) and facilitates retention of prior generative\nknowledge when faced with a task sequence. We present results that show the\nP-TNCN's ability to conduct zero-shot adaptation and online continual sequence\nmodeling.\n
This thesis aims to examine metapragmatic discourses on linguistic politeness illustrated in Korean language how-to literature. The primary task lies in contextualizing the native awareness of ene yeycel (linguistic politeness in Korean) within the interests or values of certain social groups. The first group, South Korean government-sanctioned agencies, led a linguistic campaign promoting a new standard speech model in 1992. Language professionals, the second group of social actors, produced popular language how-to literature, especially after the establishment of the hegemonic standard speech model. Both language standardizing policy and the participants in the how-to industry represent the cultural process of constructing language and social conventions. The “normative” culture of ene yeycel can be empowered and widely circulated, gaining wider social practice. Standardization of honorification came to the surface as a public issue along with a new “cultural policy” of the Ministry of Cultural Affairs in 1990. In this cultural-political circumstance, the social meaning of standardized honorification was rediscovered as indigenous culture, a group identity shared by Korean speakers. Positively valorizing honorification as linguistic and cultural tradition, the standardized model preserves the sophisticated use of honorifics and reinforces superior-inferior relationships. However, the standard model of ene yeycel can be subjective and arbitrary. Moreover, different styles are too easily proscribed as errors made by sloppy speakers. Language how-to literature produces more diversified interpretations than the standard speech manual. As language users are confronted with the challenges of finding the proper level of honorification, language how-to manuals provide justifications to help speakers prioritize linguistic norms when internalizing social relationships. Positive valorizations of honorification derive from a speaker's respect for the interlocutor's social status or personality. Negative valorizations of honorification view deferential politeness as a kind of discriminatory behaviour indexing power-difference. The positive or negative values of honorification are based on different concepts of ene yeycel and on different identifications of social relationships. Such conceptualizations rationalize whether speakers should support honorification or not, and lead them to discuss language use in current society.
This paper is an analysis of the Hill‟s Strategy Development Framework and application of the framework to a Fast Food Restaurant business which is operating in Brunei Darussalam.The study of this article concentrates on corporate objectives, marketing strategy, order qualifiers, order winners, and the operations strategy within the company.A review of the relevant literature conducted on corporate goals, competitive priorities, Order Qualifier and Order Winner. The methodology based on a desk review of secondary data and non-participant observations research approach.The findings demonstrated that Fast Food Restaurant Business in Brunei Darussalam required more considerable attention to focus on improving the Customer Service Relationship (CSR) value.It is crucial to concentrate on CSR that would enhance the business brand value, a better image rating and thereby contributing to company‟s sales and gaining a new customer.The study recommended that Fast Food Restaurant Business create a new market such as selling the frozen product.Fast Food Restaurant Business potentially can sell this in their store or export their patent right on the product internationally.The contribution of this paper is to provide provides a review of the application of the strategic framework to uncover the potential customer benefits package of the strategy.
Introduction:Iliac artery endofibrosis (IAE) is an uncommon disease, poorly studied pathology with devastating effects and different therapeutic approaches affecting young people who practise intensive sports, especially cyclists. The evolution of the process not only depends on the diagnosis and therapeutic action, but also on the acceptance and attitude of the patient and subsequent professional guidance. Case description:This is the case description of a professional triathlon athlete that had one previous iliac surgical revascularization for an IAE Iliac and was admitted in our department five times with subacute lower limb ischemia affecting both legs between 2013 and 2016. Clinical findings and image tests are reported, as well as medical procedures performed. Indications based on clinical, functional and imaging ratings were clear, but his professional activity was not completely abandoned. Finally, after four endovascular procedures with good immediate results, he was warned of the seriousness of the process since the etiopathogenic reason. At the present moment patient is asymptomatic, under routine controls, working as successful triathlon coach. Discussion and conclusion:The fact that an external mechanical stress is the reason of repeated iliac artery injury suggests that an open surgical approach correcting the external muscular compression or arterial deformation should be a definitive but also aggressive solution according to literature. However, endovascular procedures and new endovascular devices are an increasingly promising option with a very low surgical risk. No matter the revascularization performed, the persistence of sports intensive practice carries a high risk of recurrence. Sport practise cessation is mandatory in some cases in order to assure revascularization long-term patency, but also a well conducted professional orientation is needed to complete the therapeutic action.
Temporal models based on recurrent neural networks have proven to be quite powerful in a wide variety of applications, including language modeling and speech processing. However, to train these models, one relies on back-propagation through time, which entails unfolding the network over many time steps, making the process of conducting credit assignment considerably more challenging. Furthermore, the nature of back-propagation itself does not permit the use of non-differentiable activation functions and is inherently sequential, making parallelization of the underlying training process very difficult. In this work, we propose the Parallel Temporal Neural Coding Network, a biologically inspired model trained by the local learning algorithm known as Local Representation Alignment, that aims to resolve the difficulties and problems that plague recurrent networks trained by back-propagation through time. Most notably, this architecture requires neither unrolling nor the derivatives of its internal activation functions. We compare our model and learning procedure to other online back-propagation-through-time alternatives (which also tend to be computationally expensive), including real-time recurrent learning, echo state networks, and unbiased online recurrent optimization, and show that it outperforms them on sequence modeling benchmarks such as Bouncing MNIST, a new benchmark we call Bouncing NotMNIST, and Penn Treebank. Notably, our approach can, in some instances, even outperform full back-propagation through time itself as well as variants such as sparse attentive back-tracking. Furthermore, we present promising experimental results that demonstrate our model's ability to conduct zero-shot adaptation.
The paper tackles the question of what the dynamics of wordplay mean for Early Modern language philosophy and what function wordplay fulfills at a time when linguistic norms and cultural values of a particular language are being sought. In Part 1, the current definition of wordplay suggested in In Part 2, we give a brief sketch of the main features of Early Modern linguistic thought with a particular focus on the concepts of play and wordplay. As one of the language theorists of 17th century Germany, Georg Philipp Harsdrffer (1607-1658) is widely known for the sophisticated integration of these concepts into his "linguistic" oeuvre, and this will determine the main focus of the current article. Two of Harsdrffer's works will be the center of attention: the Frauenzimmer Gesprchspiele (FZG), published 1643-1649 in Nuremberg, an eight-volume series of dialogues about social, poetic and scientific matters, which incorporates much of Harsdrffer's thoughts on language and one of the best-sellers of the 17th century, and the Delitiae Mathematicae et Physicae (DMP), a three-volume scientific work, to which Harsdrffer added the last two of the three volumes (1651( -1653, Nuremberg), Nuremberg). Based on the study of various subtypes of wordplay with letters in Part 3, we shall argue that in the context of baroque linguistic ideas wordplay should be defined in a broader sense. It is deeply rooted in a particular view of language peculiar to European baroque culture that provided a conceptual background not only for language "theories", poetry, education and standards of knowledge but also for the role and functions of wordplay. As Harsdrffer found his inspiration in and was strongly influenced by similar ideas of other scientists, particularly in Italy and France, the results of the analysis of the German baroque sources allow for more general assumptions that are not restricted to one language only.
This paper explains transition dependency parsing approaches to build a dependency parser for Telugu language. Telugu treebank is given as an input to transition dependency parsers. One of the best transition dependency parser is the Malt parser. It is an independent system and it has nine methods to parse a sentence of any language. We have applied the treebank on all the methods of a Malt parser among which Arc-eager parser produces state-of-art results for Telugu language. Arc-eager method was produced LA (Label Accuracy) of 63%, UAS (Unlabeled Attachment Score) of 88.1% and LAS (Label Attachment Score) of 62.3%. In this paper we discuss a brief introduction of all Malt Parsing methods and an in detail explanation of Arc-eager dependency parsing.
Introduction. Nowadays the language of modern television broadcasts’ speechesis more and more in the focus of linguistic study. Special interest of given paper is comprisedby the variability in gender categorization of nouns as it occurs in speeches of Ukrainian TVprograms anchorpersons. The object of the paper is the choice of nouns in modern TV speech,that is distinguished due to the variability of grammatical category of the gender of nouns waysof realization.Purpose of the article is to analyze the nouns that have suffered changes in thegrammatical category of gender within the current language trends being implemented inthe broadcast of Ukrainian television. In addition, the aim is to outline some of the reasonsfor the emergence of such tendencies, the relative use of these nouns, the degree of codificationin modern lexicographic sources.Methods of research. The research is grounded on descriptive method, the method ofempiric analysis, immediate constituents’ analysis, and contextual analysis.Results. Studying the language used in contemporary Ukrainian TV speeches convincinglydemonstrated how formal-grammatical indicators of the category of gender of nouns in moderntelecommunication vary (shift), implemented in the following modifications: male genius femalegenus, female genus male genus. In addition, the review of codification in the dictionariesof the analyzed noun units shows that variational changes in the morphology of the noun andin particular, in its morphological and grammatical categories nowadays cause not onlythe dislocation of the current linguistic norm, but also tend to change the morphological norms. Conclusion. Analyzing the language used in speeches of informative and entertainingUkrainian TV programs’ anchorpersons, we have concluded that they prefer to choose differentgrammatical variations in favor of a specific counterpart to the grammatical category of the gender,often a revitalized or dialectal word used.
There exist distinctive words that are used to express same semantics and as a result of this it has become hard to quantify the exact matching of words. To deal with this issue, past investigations endeavored to ascertain a likeness between distinctive pair of words. Conventional methodologies for computing word similarity are based on repositories like WordNet. It is a manually created lexical database and it processes semantic connection between various words. However, WordNet is a universally useful asset but wide range of words are not present in it and furthermore there exist an issue of identifying the meaning of words. Implication of words are diverse in WordNet when we utilize it in a textual framework. There exists a need of the refined approach that can gauge words resemblance in light of their co-occurrence. In this examination, we proposed an approach that registers likeness in text particular words, with the assistance of literary substance of various posts on StackOverflow. Our proposed strategy figures out word similarities in text by ascertaining the weighted co-occurrence in view of Computing Term Cooccurrence (CTC) and SentiWordNet. The exploratory outcome demonstrates that our system proposed an arrangement of words that are identified with text data is exceptional. Moreover, when it was compared with WordNet-based strategy named as WordNetres, it results with better outcomes.
Abstract Facial expressions are fundamental to interpersonal communication, including social interaction, and allow people of different ages, cultures, and languages to quickly and reliably convey emotional information. Historically, facial expression research has followed from discrete emotion theories, which posit a limited number of distinct affective states that are represented with specific patterns of facial action. Much less work has focused on dimensional features of emotion, particularly positive and negative affect intensity. This is likely, in part, because achieving inter-rater reliability for facial action and affect intensity ratings is painstaking and labor-intensive. We use computer-vision and machine learning (CVML) to identify patterns of facial actions in 4,648 video recordings of 125 human participants, which show strong correspondences to positive and negative affect intensity ratings obtained from highly trained coders. Our results show that CVML can both (1) determine the importance of different facial actions that human coders use to derive positive and negative affective ratings, and (2) efficiently automate positive and negative affect intensity coding on large facial expression databases. Further, we show that CVML can be applied to individual human judges to infer which facial actions they use to generate perceptual emotion ratings from facial expressions.
This paper carries out an empirical analysis of various dropout techniques for language modelling, such as Bernoulli dropout, Gaussian dropout, Curriculum Dropout, Variational Dropout and Concrete Dropout. Moreover, we propose an extension of variational dropout to concrete dropout and curriculum dropout with varying schedules. We find these extensions to perform well when compared to standard dropout approaches, particularly variational curriculum dropout with a linear schedule. Largest performance increases are made when applying dropout on the decoder layer. Lastly, we analyze where most of the errors occur at test time as a post-analysis step to determine if the well-known problem of compounding errors is apparent and to what end do the proposed methods mitigate this issue for each dataset. We report results on a 2-hidden layer LSTM, GRU and Highway network with embedding dropout, dropout on the gated hidden layers and the output projection layer for each model. We report our results on Penn-TreeBank and WikiText-2 word-level language modelling datasets, where the former reduces the long-tail distribution through preprocessing and one which preserves rare words in the training and test set.
Abstract In this article we provide a practical demonstration of how syntactically annotated corpora (treebanks), particularly the English Historical Parsed Corpora Series, can be used to investigate research questions with a diachronic depth and synchronic breadth that would not otherwise be possible. The phenomenon under investigation is split coordination, in which two parts of a conjoined constituent appear separated in the clause (e.g., and this is where my aunt lives and my uncle ). It affects every type of coordinated constituent (subject/object DPs, predicate and attributive ADJPs, ADVPs, PPs and DP objects of P) in Old English (OE); and it, or a superficially similar construction, occurs continuously throughout the attested period from approximately 800 to the present day. Despite its synchronic range and diachronic persistence, split coordination has received surprisingly little attention in the diachronic literature, with the exception of Perez Lorido’s (2009) limited study of split subjects in eight OE texts. Its modern counterpart is most frequently analysed as Bare Argument Ellipsis (BAE). Although the OE and Present-Day English constructions appear superficially similar, we show that not all of the OE data is amenable to a BAE analysis. We bring to bear different types of evidence (structural, discourse/performance effects, rate of change, etc.) to argue that split coordination in fact represents two different constructions, one of which remains stable over time while the other is lost in the post-Middle English period.
NLTK toolkit is an API platform built with Python language to interact with humans through natural language. The very first version of NLTK was released in 2005 (1.4.3), which was compatible with Python 2.4. The latest version was in September 2017 NLTK (3.2.5), which incorporated features like Arabic stemmers, NIST evaluation, MOSES tokenizer, Stanford segmenter, treebank detokenizer, verbnet, and vader, etc. NLTK was created in 2001 as a part of Computational Linguistic Department at the University of Pennsylvania. Since then it has been tested and developed. The important packages of this system are 1) corpus builder, 2) tokenizer, 3) collocation, 4) tagging, 5) parsing, 6) metrics, and 7) probability distribution system. Toolbox NLTK was built to meet four primary requirements: 1) Simplicity: An substantive framework for building blocks; 2) Consistency: Consistent interface; 3) Extensibility: Which can be easily scaled; and 4) Modularity: All modules are independent of each other.
Language models, which are used in various tasks including speech recognition and sentence completion, are usually used with texts covering various domains. Therefore, domain adaptation has been a long-ongoing challenge in language model research. Conventional methods mainly work by the addition of a domain dependent bias. In this paper, we propose a novel way to adapt neural network-based language models. Our proposed approach relies on a linear combination of factorised hidden layers, which are learnt by the network. For domain adaptation, we use topic features from latent Dirichlet allocation. These features are input into an auxiliary network, and the output of this network is used to calculate the hidden layer weights. Both the auxiliary network and the main network can be trained jointly by error backpropagation. This makes our proposed approach completely unsupervised. To evaluate our method, we show results for the well-known Penn Treebank and the TED-LIUM dataset.
Easy-first parsing relies on subtree re-ranking to build the complete parse tree. Whereas the intermediate state of parsing processing is represented by various subtrees, whose internal structural information is the key lead for later parsing action decisions, we explore a better representation for such subtrees. In detail, this work introduces a bottom-up subtree encoding method based on the child-sum tree-LSTM. Starting from an easy-first dependency parser without other handcraft features, we show that the effective subtree encoder does promote the parsing process, and can make a greedy search easy-first parser achieve promising results on benchmark treebanks compared to state-of-the-art baselines. Furthermore, with the help of the current pre-training language model, we further improve the state-of-the-art results of the easy-first approach.
espanolEn este articulo se lleva a cabo un analisis contrastivo de las cuestiones normativas, referentes a las categorias gramaticales de los determinantes y de los pronombres, incluidas en las gramaticas y en los diccionarios de la Real Academia Espanola. El corpus lo integran las diferentes ediciones de su obra gramatical (1771, 1796, 1854, 1870, 1883, 1911, 1917, el Esbozo de 1973 y la NGLE de 2009), asi como las veintitres ediciones de su obra lexicografica (desde 1780 hasta 2014). Se han examinado, clasificado y descrito los asuntos referentes a la norma, extraidos tras el cotejo de ambas obras a lo largo de la historia. Las principales conclusiones de este estudio revelan los siguientes datos: por una parte, son las gramaticas las que dedican una mayor atencion a los temas prescriptivos; por otra parte, en ciertas ocasiones, se observa una falta de coherencia entre las gramaticas y los diccionarios tanto en el seguimiento de la norma como en los criterios de correccion empleados; finalmente, con respecto a los diccionarios, sobresalen las ediciones manuales y la vigesima tercera edicion, dado su interes por recoger un mayor numero de alusiones al buen uso linguistico. EnglishIn this paper, a contrastive analysis of normative issues concerning the grammatical categories of determinants and pronouns included in the Spanish Royal Academy Grammars and Dictionaries is carried out. The corpus comprises the different editions of its grammatical work (1771, 1796, 1854, 1870, 1883, 1911, 1917, the 1973 Sketch and the NGLE of 2009), as well as the twenty-three editions of its lexicographical work (from 1780 to 2014). Issues related to the linguistic norm, which have been extracted from a comparison between the different editions throughout history, have been examined, classified and described. Results from this study reveal the following data: on the one hand, Grammars pay closer attention to prescriptive issues; on the other hand, a lack of coherence between Grammars and Dictionaries both in the follow-up of the linguistic norm and in the correction criteria used can be observed. In addition to that and with regard to dictionaries, the manual editions along with the 23rd edition are exceptional provided their interest in collecting a greater number of allusions to proper linguistic use.
Combination of customer needs and quantitative data is an idea to produce a various value. Therefore, this research proposed method for multidisciplinary design optimization of hub airport which optimized customer needs and quantitative data simultaneously. For customer needs, target of tourism and season were assigned into design variables and optimal weight was calculated by using SGD method. Design variables of quantitative data were number of transit, transportation fee and time from airport to World Heritage, and calculated value by using AHP which Monte Carlo simulation was applied for updating optimal value. To obtain optimal value for customer needs, review of tourism was extracted by text mining and count number of review in each World Heritage, word rating point was decided to calculate the weight of each World Heritage. Next, results of customer needs and quantitative data were normalized for collaborative optimization. Finally, the results showed a various optimal solution obtained for each design variable.
Cross-entropy loss is a common choice when it comes to multiclass classification tasks and language modeling in particular. Minimizing this loss results in language models of very good quality. We show that it is possible to fine-tune these models and make them perform even better if they are fine-tuned with sum of cross-entropy loss and reverse Kullback-Leibler divergence. The latter is estimated using discriminator network that we train in advance. During fine-tuning probabilities of rare words that are usually underestimated by language models become bigger. The novel approach that we propose allows us to reach state-of-the-art quality on Penn Treebank: perplexity decreases from 52.4 to 52.1. Our fine-tuning algorithm is rather fast, scales well to different architectures and datasets and requires almost no hyperparameter tuning: the only hyperparameter that needs to be tuned is learning rate.
Ratings are important in attracting foreign capital so they play a great role in the financial system of a country. The aim of the study is to investigate the impact of macroeconomic indicators on sovereign credit ratings assigned by Fitch. For this aim Panel ordered probit model was applied to the annual data from 2000 to 2011. The analysis rests on panel of 44 countries. According to the results obtained it can be concluded that gross domestic product growth rate, per capita gross domestic product, unemployment, export, default history and the level of economic development significantly affect ratings.
Shi, Huang, and Lee (2017) obtained state-of-the-art results for English and Chinese dependency parsing by combining dynamic-programming implementations of transition-based dependency parsers with a minimal set of bidirectional LSTM features. However, their results were limited to projective parsing. In this paper, we extend their approach to support non-projectivity by providing the first practical implementation of the MH_4 algorithm, an $O(n^4)$ mildly nonprojective dynamic-programming parser with very high coverage on non-projective treebanks. To make MH_4 compatible with minimal transition-based feature sets, we introduce a transition-based interpretation of it in which parser items are mapped to sequences of transitions. We thus obtain the first implementation of global decoding for non-projective transition-based parsing, and demonstrate empirically that it is more effective than its projective counterpart in parsing a number of highly non-projective languages
In this work we describe the system built for the three English subtasks of\nthe SemEval 2016 Task 3 by the Department of Computer Science of the University\nof Houston (UH) and the Pattern Recognition and Human Language Technology\n(PRHLT) research center - Universitat Polit`ecnica de Val`encia: UH-PRHLT. Our\nsystem represents instances by using both lexical and semantic-based similarity\nmeasures between text pairs. Our semantic features include the use of\ndistributed representations of words, knowledge graphs generated with the\nBabelNet multilingual semantic network, and the FrameNet lexical database.\nExperimental results outperform the random and Google search engine baselines\nin the three English subtasks. Our approach obtained the highest results of\nsubtask B compared to the other task participants.\n
Abstract. Patterns of facial reactivity and attentional allocation to emotional facial expressions, and how these are moderated by gaze direction, are not clearly established. Among a sample of undergraduate university students, aged between 17 and 22 years (76% female), corrugator and zygomatic reactivity, as measured by facial electromyography, and attention allocation, as measured by the startle reflex and startle-elicited N100, was examined while viewing happy, neutral, angry and fearful facial expressions, which were presented at either 0- or 30-degree gaze. Results indicated typically observed facial mimicry to happy faces but, unexpectedly, “smiling” facial responses to fearful, and to a lesser extent, angry faces. This facial reactivity was not influenced by gaze direction. Furthermore, emotional facial expressions did not elicit increased attentional allocation. Likewise, matched facial expressions did not elicit increased attentional allocation. Rather, happy and fearful faces with direct (0°) gaze elicited increased controlled attentional allocation, and averted (30°) gaze faces, regardless of emotional expression, elicited preferential, early cortical processing. These findings suggest typical facial mimicry to happy faces, but unexpected facial reactivity to angry and fearful faces, perhaps due to an attempt to regulate social bonds during threat perception. Findings also suggest a divergence in controlled versus preferential, early cortical attentional processing for direct compared to averted gaze faces. These findings relate to young, mostly female, adults attending university. The experiment should be repeated with a larger sample drawn from the general community, with a broader age range and gender balance, and with a stimulus set with validated subjective valence and arousal ratings. This can reduce Type II error and establish normative patterns of facial reactivity and attentional processing of emotional facial expressions with different gaze directions.
Requirement is a formal expression of user’s need. It is the main foundation of any software development project. Natural language (NL) is often used to express and write system requirements specifications as well as user requirements. However, there is a very high probability that more than half natural language requirements can be ambiguous, incomplete and inaccurate. A software engineer can miss-interpret the natural language requirements and can generate an erroneous software model, which finally will lead to project failure. Earlier, we have introduced a prototype tool that provides natural language requirements authoring facilities and consistency checking to assist requirement engineers when working with informal and semi-formal requirements. However, the tool has pattern limitation to support the extraction of the essential requirements from the NL requirements. Therefore this study is aimed to enhance the accuracy and scalability of the tool to capture the essential requirements from the NL requirements. Our approach is to implement lexical analysis and embed an English lexical database where it will serve as a thesaurus in the tool. This tool is expected to be able to find the synonym of the extracted phrases (essential requirements) in the database to match it to the essential interaction pattern (phrases and expressions) in the library. Our future work will focus on the next phase of requirements engineering, which is requirements validation.
Artikkeli käsittelee suomentamiseen liittyviä ideologioita ja normeja 1800-luvun tietokirjallisuudessa. Tapaustutkimuksena on Werner Söderström Osakeyhtiön tietokirjojen suomennostoiminta 1800-luvun lopulla. Tutkimus kytkeytyy kääntämisen sosiologiaan ja historiaan, ja siinä arvioidaan myös, miten ja missä määrin historiallisia käännösprosesseja voidaan rekonstruoida. Käännösprosesseja lähestytään tarkastelemalla eri toimijoiden − kustantaja, kääntäjä, kieliasiantuntija, tekstin arvioija − osuutta käännösprosessissa. Tutkimuksen aineistona on kustantajan ja kääntäjän kääntämistä ja kielellisiä valintoja käsittelevä kirjeenvaihto, jonka avulla on mahdollista valottaa eri suunnista kääntäjän arkea, yhteisöllisiä arvoja ja normeja käännösvalintojen taustalla sekä niitä henkilökohtaisia asenteita, jotka ohjaavat kääntäjiä erilaisiin valintoihin.
 Analyysin tuloksena voi päätellä, että ammattikirjoittajina kääntäjät olivat hyvin tietoisia erilaisista kielellisistä ja kääntämiseen liittyvistä normeista. Käytännön työssä kääntäjät toimivat kuitenkin usein erilaisten normien ristipaineessa, jolloin vastakkain asettuivat esimerkiksi alkuteoksen luonteen säilyttäminen ja toisaalta sen kotouttaminen. Kääntäjät olivat myös tietoisia kielen vaihtelevista normeista, tunsivat käynnissä olevat kielikeskustelut ja mukauttivat herkästi kielenkäyttöään kulloinkin vallitsevien kirjakielen normien mukaiseksi.
 
 Norms and ideologies of translation in light of correspondence between publisher and translator in 19th-century Finland
 This article analyses the ideologies and norms that guided the translation of works of non-fiction in 19th-century Finland. As a case study the article analyses the processes involved in the publication of non-fiction at the Werner Söderström Ltd publishing house at the end of the 19th century. The research takes as its base theories examining the sociology and history of translation. It also aims to evaluate how and to what extent historical translation processes can be reconstructed. Translation is approached as a collaborative process involving various actors: publisher, translator, language editor, and expert reader. The data consists of correspondence between publisher and translator that deals with matters of translation or language. This correspondence sheds light on the everyday life of the translator and the socially accepted norms and ideologies that guide the translation process. It also reveals the stance of publishers concerning the choice of translator, a factor that can lead to very different end products.
 The analysis shows that, as professional writers, translators at the end of the 19th century were well aware of contemporary translational norms. In practice, translators were caught between various conflicting pressures – regarding, for instance, questions such as whether one should follow the original text as close as possible to preserve its unique style or assimilate the text to a Finnish context to help the reader. The data also shows that translators were well aware of linguistic norms; they were acquainted with current and past debates, and in assimilating their use of language they remained sensitive to prevailing norms.
Estimating the entropy based on data is one of the prototypical problems in distribution property testing and estimation. For estimating the Shannon entropy of a distribution on $S$ elements with independent samples, [Paninski2004] showed that the sample complexity is sublinear in $S$, and [Valiant--Valiant2011] showed that consistent estimation of Shannon entropy is possible if and only if the sample size $n$ far exceeds $\frac{S}{\log S}$. In this paper we consider the problem of estimating the entropy rate of a stationary reversible Markov chain with $S$ states from a sample path of $n$ observations. We show that: (1) As long as the Markov chain mixes not too slowly, i.e., the relaxation time is at most $O(\frac{S}{\ln^3 S})$, consistent estimation is achievable when $n \gg \frac{S^2}{\log S}$. (2) As long as the Markov chain has some slight dependency, i.e., the relaxation time is at least $1+Ω(\frac{\ln^2 S}{\sqrt{S}})$, consistent estimation is impossible when $n \lesssim \frac{S^2}{\log S}$. Under both assumptions, the optimal estimation accuracy is shown to be $Θ(\frac{S^2}{n \log S})$. In comparison, the empirical entropy rate requires at least $Ω(S^2)$ samples to be consistent, even when the Markov chain is memoryless. In addition to synthetic experiments, we also apply the estimators that achieve the optimal sample complexity to estimate the entropy rate of the English language in the Penn Treebank and the Google One Billion Words corpora, which provides a natural benchmark for language modeling and relates it directly to the widely used perplexity measure.
The article presents the results of word-formative and semantic analysis of Middle Czech verbs consisting of the prefix roz(e)- contained in Lexical database of humanistic and baroque Czech (https://madla.ujc.cas.cz). The analysis partly confirms, partly corrects the results of earlier analyses, above all, carried out by D. Šlosar (1981).
This paper aims to present the theoretical considerations and methodology used in elaborating a formal characterization of a fuzzy grammar in natural language grammars. It specifically focuses on the syntax of Spanish. However, we suggest that this methodology could be used to define a universal model for describing any kind of natural language grammar that takes into account fuzziness. Objective data based on frequencies were extracted using the Spanish Universal Dependencies Corpus Treebank and the Marsagram tool. These data allow us to describe Spanish Natural Language Grammar in terms of its constraints using the Property Grammars Theory. The work presented here could be applied in the form of an algorithm for parsing which could be of benefit to various areas of language and technology such as self-taught language learning software (in which violations and degree of violation could be tagged), data mining or human-machine interfaces.
We explore whether it is possible to build lighter parsers, that are statistically equivalent to their corresponding standard version, for a wide set of languages showing different structures and morphologies. As testbed, we use the Universal Dependencies and transition-based dependency parsers trained on feed-forward networks. For these, most existing research assumes de facto standard embedded features and relies on pre-computation tricks to obtain speed-ups. We explore how these features and their size can be reduced and whether this translates into speed-ups with a negligible impact on accuracy. The experiments show that grand-daughter features can be removed for the majority of treebanks without a significant (negative or positive) LAS difference. They also show how the size of the embeddings can be notably reduced.
We introduce a method to reduce constituent parsing to sequence labeling. For each word w_t, it generates a label that encodes: (1) the number of ancestors in the tree that the words w_t and w_{t+1} have in common, and (2) the nonterminal symbol at the lowest common ancestor. We first prove that the proposed encoding function is injective for any tree without unary branches. In practice, the approach is made extensible to all constituency trees by collapsing unary branches. We then use the PTB and CTB treebanks as testbeds and propose a set of fast baselines. We achieve 90.7% F-score on the PTB test set, outperforming the Vinyals et al. (2015) sequence-to-sequence parser. In addition, sacrificing some accuracy, our approach achieves the fastest constituent parsing speeds reported to date on PTB by a wide margin.
In recent years, with increase in the use of internet the multimedia contents on it have rapidly increased. Users may need to go through a video in a top down manner i.e. browsing the videos, or in bottom up manner i.e. retrieving specific information from videos. They may also want to go through the summary or through the highlights of the videos. This has necessitated the need to handle multimedia resources effectively. This paper proposes an automatic method for aligning scripts of lecture videos with captions. Alignment is needed to extract time information from captions and insert it in the scripts, to create index of the videos. No alignment work has been previously done in lecture videos domain. Alignment methods proposed for other type of videos are not applicable for lecture videos because, different similarity techniques behave differently on different types of datasets. The proposed method uses transcripts of lecture videos, SRT file of captions available along with lecture videos and captions generated from auto-caption generation feature of YouTube. The captions and scripts are then aligned using a dynamic programming technique. No such work has been previously done for lecture videos. Most important aspect of alignment is similarity measure. In the proposed work we have used three similarity measures cosine, jaccard, and dice. A comparative analysis of these measures is given in the paper. We also use a large lexical database of English words known as WordNet for word-to-word similarity. The experimental result shows comparison of various similarity techniques and YouTube captions.
Hoarding is a mental and public health problem stemming from difficulty associated with discarding one's possessions and resulting clutter. In the last decade, a visual method, called "Clutter Image Rating" (CIR), has been developed for the assessment of hoarding severity. It involves rating clutter in patient's home on the CIR scale from 1 to 9 using a set of reference images. Such assessment, however, is time-consuming, subjective, and may be non-repeatable. In this paper, we propose a new automatic clutter assessment method from images, according to the CIR scale, based on deep learning. While, ideally, the goal is to perfectly classify clutter, trained professionals admit assigning CIR values within ±1. Therefore, we study two loss functions for our network: one that aims to precisely assign a CIR value and one that aims to do so within ±1. We also propose a weighted combination of these loss functions that, as a byproduct, allows us to control the CIR mean absolute error (MAE). On a recently-collected dataset, we achieved ±1 accuracy of 82% and MAE of 0.88, significantly outperforming our previous results of 60% and 1.58, respectively.
One of the most demanded types of text in higher education is argumentation, present in different discursive genres such as the academic essay. Therefore, becoming familiarized with the features and structures of argumentative texts and the different kinds of arguments is essential to perform in the different disciplines. On the other hand, research shows that collaborative writing, peer review and the use of rubrics improve the quality of written productions in different educational levels.\nThe objective of this paper is to present the impact of a rubric for the evaluation of argumentative texts in higher education. We will report on the process of elaboration of the rubric (creation, validation and pilot experiment) and its dimensions (content, argumentation and persuasion, coherence and cohesion, and use of the linguistic norm), as well as the results about the students’ perception of its use. Results show the students’ positive response and the tool’s potential. We find it convenient to elaborate rubrics for the different types of text and discursive academic genres in higher education.
One of the goals of the Russian language course in the primary school is the formation of the communicative literacy. The content of the course should be aimed at understanding the wealth of linguistic means by primary school children; the formation of the ability to detect a violation of linguistic norms and the inadequacy of the linguistic means used in the speech situation; the accumulation of the experience in choosing of linguistic means in accordance with the peculiarities of the speech situation; the creation of oral and written texts that meet the criteria of content, connectivity, compliance with the norms of the Russian literary language. The article considers the classification of exercises that contribute to the formation of communicative literacy. The author gives the examples of exercises where the student acts in different roles: the student is an observer of the speech situation and analyzes the adequacy of the choice of linguistic means; the student is a direct participant in the given speech situation and makes a choice of language facilities; the student is offered to create the speech situation himself, to independently construct an oral and written text.
The article is devoted to two approaches in teaching one of the so-called "polynational" languages -the Spanish language. The national-variant differentiation and the internal variability of the Spanish language (on the Iberian Peninsula), which has developed to the present moment, is reflected in the educational process. In different countries and at various stages of education, a monocentric or polycentric approach is adopted in teaching Spanish as a foreign language. With a monocentric approach, the question arises of the linguistic norm of the Spanish language as an instrument of instruction. In the polycentric approach, it is necessary to decide which national variants to include in the training course and to what extent, on which features to pay more attention (phonetics, grammar, vocabulary). The experience of teaching Spanish in the system of compulsory education of Russia was considered on the example of three Russian universities: Peoples Friendship University of Russia, Moscow Region State University, Moscow State Linguistic University.
The article analyzes the actual concept of linguistic expertise of translation and determines its place in linguistic expertology. The substantiation of the difference between the expert opinion and the usual translation quality assessment is a fundamentally new approach to this problem. In connection with this, modern concepts of equivalence, types and methods of determining translation errors in correlation with the objectives of the expert opinion are touched upon. An attempt is made to systematize linguistic (interlingual correspondence, directed equivalence, linguistic norms, etc.), communicative-psychological (communicative intention, translation creativity, emotive evaluation system, recipient’s response), information (text depth, static and dynamic information, etc.) and logical (opinion, judgment, statement) tools of linguistic expertise of translation. Preliminary conclusions based on the conducted research can contribute to increasing the effectiveness of forensic linguistic expertise of translation, and also prove useful in the process of preparing translators for expert activity.
Users review about an app is a crucial component for open mobile application market, such as the AppStore and the Google play. Analyzing these reviews can reveal user's sentiment towards a feature in the app. There exist several analytical tools to summarize user reviews and extract meaningful sense out of them. However, these tools are still limited in terms of expressiveness and accurately classifying the reviews into more than a positive and a negative review. There is a need to get more insights from user app reviews and direct it to future app development. In this paper, we present our result of analyzing user reviews of 20 food journaling and health tracking apps. We gathered and analyzed reviews per app and classified them into three distinct categories using the sentiment treebank with recursive neural tensor network. We then analyzed the vocabulary frequency per category using the Gensim implementation of Word2Vec model. The analysis result clustered the reviews into good, bad and ugly feature reviews. Different usage patterns were detected from users review. We identified major reasons why users express a certain sentiment towards an app and learned how users' satisfaction or complaints was related to a specific feature. This research could be a guideline for app developers to follow when developing an app to refrain from adopting techniques that might demotivate (hinder) the application use or adopt those perceived positively by the users.
The path towards electronic healthcare records is nonetheless not free of challenges including the large amount of clinical information buried in narrative content. Medication information is one of the most important types of clinical data in electronic healthcare records. It is critical for healthcare safety and quality, as well as for clinical research to have such information identified correctly. Natural language processing (NLP) is essential to phenotyping the medication data. However, recognizing medication patterns based on general NLP techniques fail short to identify such patterns with great accuracy even if they were trained with relevant clinical treebanks or corpuses. This article describes how Clojure and OpenNLP API can be used to identify medication patterns and train the clinical narrative chuncker to accurately identify given medication patterns.
Images and language convey meaning that depend on the viewpoints and contextual background of those perceiving them. Taxonomies help order meaning such that information granules, image elements and words, make sense in relation to one another and to their mutual global context. In linguistics, hypernyms cover semantically broader context then their subordinate hyponym. In images, superordinate spatial-taxons (object groups or foreground) cover more abstract regions then their child subordinate spatial-taxons (objects or salient object parts). In this paper I use fuzzy granularization and fuzzy perceptualization as proposed by Zadeh 2002 to explore image annotation by using Zadeh's Restriction-centered Theory of Truth and Meaning as proposed in 2013. The approach uses human annotated image data, search engine queries and data collected from WordNet (A Lexical Database for English maintained by Princeton University). I discuss implications for Shannon, Integrated, and Zadeh Information Theory.
This article explores whether and how network visualization can benefit philological and historical-linguistic study. This is illustrated with a corpus-based investigation of scribes' language use in a lemmatized and morphologically annotated corpus of documentary Latin (Late Latin Charter Treebank, LLCT2). We extract four continuous linguistic variables from LLCT2 and utilize a gradient colour palette in Gephi to visualize the variable values as node attributes in a trimodal network which consists of the documents, writers, and writing locations underlying the same corpus. We call this network the "LLCT2 network". The geographical coordinates of the location nodes form an approximate map, which allows for drawing geographical conclusions. The linguistic variables are examined both separately and as a sum variable, and the visualizations presented as static images and as interactive Sigma.js visualizations. The variables represent different domains of language competence of scribes who learnt written Latin practically as a second-language. The results show that the network visualization of linguistic features helps in observing patterns which support linguistic-philological argumentation and which risk passing unnoticed with traditional methods. However, the approach is subject to the same limitations as all visualization techniques: the human eye can only perceive a certain, relatively small amount of information at a time.
Word vectors are at the core of many natural language processing tasks. Recently, there has been interest in post-processing word vectors to enrich their semantic information. In this paper, we introduce a novel word vector post-processing technique based on matrix conceptors (Jaeger2014), a family of regularized identity maps. More concretely, we propose to use conceptors to suppress those latent features of word vectors having high variances. The proposed method is purely unsupervised: it does not rely on any corpus or external linguistic database. We evaluate the post-processed word vectors on a battery of intrinsic lexical evaluation tasks, showing that the proposed method consistently outperforms existing state-of-the-art alternatives. We also show that post-processed word vectors can be used for the downstream natural language processing task of dialogue state tracking, yielding improved results in different dialogue domains.