Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Introduction. Nowadays the language of modern television broadcasts’ speechesis more and more in the focus of linguistic study. Special interest of given paper is comprisedby the variability in gender categorization of nouns as it occurs in speeches of Ukrainian TVprograms anchorpersons. The object of the paper is the choice of nouns in modern TV speech,that is distinguished due to the variability of grammatical category of the gender of nouns waysof realization.Purpose of the article is to analyze the nouns that have suffered changes in thegrammatical category of gender within the current language trends being implemented inthe broadcast of Ukrainian television. In addition, the aim is to outline some of the reasonsfor the emergence of such tendencies, the relative use of these nouns, the degree of codificationin modern lexicographic sources.Methods of research. The research is grounded on descriptive method, the method ofempiric analysis, immediate constituents’ analysis, and contextual analysis.Results. Studying the language used in contemporary Ukrainian TV speeches convincinglydemonstrated how formal-grammatical indicators of the category of gender of nouns in moderntelecommunication vary (shift), implemented in the following modifications: male genius femalegenus, female genus male genus. In addition, the review of codification in the dictionariesof the analyzed noun units shows that variational changes in the morphology of the noun andin particular, in its morphological and grammatical categories nowadays cause not onlythe dislocation of the current linguistic norm, but also tend to change the morphological norms. Conclusion. Analyzing the language used in speeches of informative and entertainingUkrainian TV programs’ anchorpersons, we have concluded that they prefer to choose differentgrammatical variations in favor of a specific counterpart to the grammatical category of the gender,often a revitalized or dialectal word used.
This paper explains transition dependency parsing approaches to build a dependency parser for Telugu language. Telugu treebank is given as an input to transition dependency parsers. One of the best transition dependency parser is the Malt parser. It is an independent system and it has nine methods to parse a sentence of any language. We have applied the treebank on all the methods of a Malt parser among which Arc-eager parser produces state-of-art results for Telugu language. Arc-eager method was produced LA (Label Accuracy) of 63%, UAS (Unlabeled Attachment Score) of 88.1% and LAS (Label Attachment Score) of 62.3%. In this paper we discuss a brief introduction of all Malt Parsing methods and an in detail explanation of Arc-eager dependency parsing.
Temporal models based on recurrent neural networks have proven to be quite powerful in a wide variety of applications, including language modeling and speech processing. However, to train these models, one relies on back-propagation through time, which entails unfolding the network over many time steps, making the process of conducting credit assignment considerably more challenging. Furthermore, the nature of back-propagation itself does not permit the use of non-differentiable activation functions and is inherently sequential, making parallelization of the underlying training process very difficult. In this work, we propose the Parallel Temporal Neural Coding Network, a biologically inspired model trained by the local learning algorithm known as Local Representation Alignment, that aims to resolve the difficulties and problems that plague recurrent networks trained by back-propagation through time. Most notably, this architecture requires neither unrolling nor the derivatives of its internal activation functions. We compare our model and learning procedure to other online back-propagation-through-time alternatives (which also tend to be computationally expensive), including real-time recurrent learning, echo state networks, and unbiased online recurrent optimization, and show that it outperforms them on sequence modeling benchmarks such as Bouncing MNIST, a new benchmark we call Bouncing NotMNIST, and Penn Treebank. Notably, our approach can, in some instances, even outperform full back-propagation through time itself as well as variants such as sparse attentive back-tracking. Furthermore, we present promising experimental results that demonstrate our model's ability to conduct zero-shot adaptation.
Introduction:Iliac artery endofibrosis (IAE) is an uncommon disease, poorly studied pathology with devastating effects and different therapeutic approaches affecting young people who practise intensive sports, especially cyclists. The evolution of the process not only depends on the diagnosis and therapeutic action, but also on the acceptance and attitude of the patient and subsequent professional guidance. Case description:This is the case description of a professional triathlon athlete that had one previous iliac surgical revascularization for an IAE Iliac and was admitted in our department five times with subacute lower limb ischemia affecting both legs between 2013 and 2016. Clinical findings and image tests are reported, as well as medical procedures performed. Indications based on clinical, functional and imaging ratings were clear, but his professional activity was not completely abandoned. Finally, after four endovascular procedures with good immediate results, he was warned of the seriousness of the process since the etiopathogenic reason. At the present moment patient is asymptomatic, under routine controls, working as successful triathlon coach. Discussion and conclusion:The fact that an external mechanical stress is the reason of repeated iliac artery injury suggests that an open surgical approach correcting the external muscular compression or arterial deformation should be a definitive but also aggressive solution according to literature. However, endovascular procedures and new endovascular devices are an increasingly promising option with a very low surgical risk. No matter the revascularization performed, the persistence of sports intensive practice carries a high risk of recurrence. Sport practise cessation is mandatory in some cases in order to assure revascularization long-term patency, but also a well conducted professional orientation is needed to complete the therapeutic action.
This thesis aims to examine metapragmatic discourses on linguistic politeness illustrated in Korean language how-to literature. The primary task lies in contextualizing the native awareness of ene yeycel (linguistic politeness in Korean) within the interests or values of certain social groups. The first group, South Korean government-sanctioned agencies, led a linguistic campaign promoting a new standard speech model in 1992. Language professionals, the second group of social actors, produced popular language how-to literature, especially after the establishment of the hegemonic standard speech model. Both language standardizing policy and the participants in the how-to industry represent the cultural process of constructing language and social conventions. The “normative” culture of ene yeycel can be empowered and widely circulated, gaining wider social practice. Standardization of honorification came to the surface as a public issue along with a new “cultural policy” of the Ministry of Cultural Affairs in 1990. In this cultural-political circumstance, the social meaning of standardized honorification was rediscovered as indigenous culture, a group identity shared by Korean speakers. Positively valorizing honorification as linguistic and cultural tradition, the standardized model preserves the sophisticated use of honorifics and reinforces superior-inferior relationships. However, the standard model of ene yeycel can be subjective and arbitrary. Moreover, different styles are too easily proscribed as errors made by sloppy speakers. Language how-to literature produces more diversified interpretations than the standard speech manual. As language users are confronted with the challenges of finding the proper level of honorification, language how-to manuals provide justifications to help speakers prioritize linguistic norms when internalizing social relationships. Positive valorizations of honorification derive from a speaker's respect for the interlocutor's social status or personality. Negative valorizations of honorification view deferential politeness as a kind of discriminatory behaviour indexing power-difference. The positive or negative values of honorification are based on different concepts of ene yeycel and on different identifications of social relationships. Such conceptualizations rationalize whether speakers should support honorification or not, and lead them to discuss language use in current society.
Temporal models based on recurrent neural networks have proven to be quite\npowerful in a wide variety of applications. However, training these models\noften relies on back-propagation through time, which entails unfolding the\nnetwork over many time steps, making the process of conducting credit\nassignment considerably more challenging. Furthermore, the nature of\nback-propagation itself does not permit the use of non-differentiable\nactivation functions and is inherently sequential, making parallelization of\nthe underlying training process difficult. Here, we propose the Parallel\nTemporal Neural Coding Network (P-TNCN), a biologically inspired model trained\nby the learning algorithm we call Local Representation Alignment. It aims to\nresolve the difficulties and problems that plague recurrent networks trained by\nback-propagation through time. The architecture requires neither unrolling in\ntime nor the derivatives of its internal activation functions. We compare our\nmodel and learning procedure to other back-propagation through time\nalternatives (which also tend to be computationally expensive), including\nreal-time recurrent learning, echo state networks, and unbiased online\nrecurrent optimization. We show that it outperforms these on sequence modeling\nbenchmarks such as Bouncing MNIST, a new benchmark we denote as Bouncing\nNotMNIST, and Penn Treebank. Notably, our approach can in some instances\noutperform full back-propagation through time as well as variants such as\nsparse attentive back-tracking. Significantly, the hidden unit correction phase\nof P-TNCN allows it to adapt to new datasets even if its synaptic weights are\nheld fixed (zero-shot adaptation) and facilitates retention of prior generative\nknowledge when faced with a task sequence. We present results that show the\nP-TNCN's ability to conduct zero-shot adaptation and online continual sequence\nmodeling.\n
Functional alterations of the default mode network (DMN) are frequently reported in psychotic disorders, but the functional role of these alterations remains poorly known. In addition to previous studies that have applied different types of tasks or recorded resting-state neuroimaging data, there has recently been more interest in the use of movie stimuli in studying brain functioning in patient populations, because this could provide a more naturalistic account of brain functioning in real life-like situations. Seventy-one first-episode psychosis (FEP) patients (mean age = 26.0 yrs, 47 (66%) males) and 57 controls (mean age = 26.86 yrs, 24 (42%) males) from the Helsinki Early Psychosis Study watched scenes from the movie Alice in Wonderland (Tim Burton, 2010) during 3 T fMRI-BOLD imaging. We used intersubject correlation (ISC) analysis, in which the correlation between voxel-wise BOLD time series in every within-group pair of subjects is calculated. In this study, time-windowed ISC was calculated with a 10-TR (time of repetition, 1.8 s) window with 1-TR steps over the fMRI time series. In each ISC window, a two-sample t test was performed to obtain a t-statistic time series of differences between the groups. An independent group of control subjects (n = 17, 10 males, mean age 26.5 yrs) rated how emotionally arousing the currently seen events of the stimulus are, producing a time-varying rating used as a regressor. General linear model was used to identify brain regions where the t-statistic time series covaries with the arousal rating. To make the interpretation of results less ambiguous, the arousal rating was divided into high and low arousal regressor by z scoring the rating and taking only the positive and negative values, respectively. Nonparametric clusterwise permutation test was used for statistical inference (cluster-defining threshold of p = 0.05, familywise error corrected threshold of p = 0.05, number of permutations = 5000). Furthermore, by using an experience-sampling setup during the same brain-scanning session, a partially overlapping sample of participants reported how emotionally aroused they were feeling during scanning. The results show significant correlation between the t-statistic time series and low arousal regressor, especially in the DMN including the anterior and posterior cingulate cortex, medial prefrontal cortex, precuneus, and bilateral lateral temporoparietal regions. Closer inspection reveals that during moments of low arousal in the movie stimulus, the ISC of healthy controls goes up but the ISC of patients does not. In the experience-sampling portion of the study, the patients reported more arousal than the control subjects. Intersubject correlation in the DMN depended differentially on arousal in FEP patients and control subjects. More specifically, during moments when the stimulus was rated less emotionally arousing, control subjects’ DMN functioning synchronized more while the patients’ did not. In connection with the difference in reported arousal during the same imaging session, our findings provide preliminary evidence for a contribution of arousal on the functional alterations of the DMN and suggest that this may be related to higher baseline arousal in the patients. Higher arousal and the related distortion of high order integrative functioning that characterizes DMN could contribute to the pathogenesis of psychosis.
Abstract The aim of the contribution is to introduce a database of linguistic forms and their functions built with the use of the multi-layer annotated corpora of Czech, the Prague Dependency Treebanks. The purpose of the Prague Database of Forms and Functions (ForFun) is to help the linguists to study the form-function relation, which we assume to be one of the principal tasks of both theoretical linguistics and natural language processing. We demonstrate possibilities of the exploitation of the ForFun database. This article is largely based on a paper presented at the 16th International Workshop on Treebanks and Linguistic Theories in Prague (Bejček et al., 2017).
The linguistic database is also positioned as an actual way of formalizing and organizing phraseological units, terms for designating types of phraseological units. The main principle of systematization of the latter in the study is the thesaurus principle, that is the filling of the paradigm «terminological system – terminological microsystem – terminological subsystem – term», represented by a linguistic database.
The article discusses the terms that nominate the language of written artifacts documented by the Cyrillic on the Ukrainian-Byelorussian lands in the XIV-XVI centuries; the expediency of using the notion “literary language” as to the Ukrainian literary written tradition of the XIV–XVI centuries is clarified; the content of the term “linguistic norm” is outlined, its characteristics in the investigated period are determined.
This chapter discusses the new and changing conditions for linguistic norms in literary fiction of the post-Soviet era. In particular, it looks at how the interrelationship between the language of literature (<italic>iazyk literatury</italic>) and the standard language (<italic>literaturnyi iazyk</italic>) has been challenged by several processes of sociolinguistic change, including initiatives in language policy.
The article deals with the issue of translation as an important means of communication between individuals who speak different languages and belong to different cultures.The article analyzes the translation as interlingual communicative phenomenon.The translation process is determined by the linguistic norms, communicative situations, functional parameters of the original text and translation norms.The role and tasks of an interpreter in the process of intercultural communication of individuals are defined.
The short note describes the chart parser for multimodal type-logical grammars which has been developed in conjunction with the type-logical treebank for French. The chart parser presents an incomplete but fast implementation of proof search for multimodal type-logical grammars using the "deductive parsing" framework. Proofs found can be transformed to natural deduction proofs.
Polycentric Spanish Norm Towards the Polish‑Spanish Legal Translation The Spanish, being the official language of Spain and many other countries, is characterized by an important dialectal diversity that is reflected in the differences at all linguistic levels: phonetic, morphological, syntactic and lexico‑semantic, etc. All these differences raise controversies and discussions about the existence of a linguistic norm depending on the perspective that can have a monocentric or polycentric character. In this contribution we present some arguments for the second one. To this end, we rely on translations, starting simultaneously from the semasiological and onomasiological perspective, of some Polish‑Spanish legal terms in which it is essential to take into account, the diatopic variation as well as the norm whose character is polycentric.
Because the most common transition systems are projective, training a transition-based dependency parser often implies to either ignore or rewrite the non-projective training examples, which has anadverse impact on accuracy. In this work, we propose a simple modification of dynamic oracles, which enables the use of non-projective data when training projective parsers. Evaluation on 73~treebanks shows that our method achieves significant gains (+2 to +7 UAS for the most non-projective languages) and consistently outperforms traditional projectivization and pseudo-projectivizationapproaches.
The Universal Dependencies project is currently comprised of 71 languages and 122 treebanks, and aims to find morphological and syntactic characteristics that can be applied to multiple languages for parallel language processing. In this paper, we introduce Universal POS, which is a morphological tagset for UD, and propose a method to automatically convert existing Korean morphological tagset into UPOS. In order to apply the UPOS tagset, which is based on refraction words such as English, to the Korean language, it is necessary to try a one-to-many mapping between the UPOS individual tag and the 21st century Sejong tag combination. (Yonsei University)
This study examines how the acoustic input (the surface form) and the abstract linguistic representation (the underlying representation) interact during spoken word recognition by investigating left-dominant tone sandhi, a tonal alternation in which the underlying tone of the first syllable spreads to the sandhi domain. We conducted an auditory-auditory priming lexical decision experiment on Shanghai left-dominant sandhi words, in which each disyllabic target ([tɕi55 dɛ31] “egg”) was preceded by monosyllabic primes either sharing the same underlying tone ([tɕi55]), surface tone ([tɕi53] “machine”), or being unrelated to the tone of the first syllable of the sandhi targets ([tɕi24] “to remember”). Results showed a surface priming effect, but not an underlying priming effect. Moreover, the surface priming did not interact with speakers’ familiarity ratings to the sandhi targets. The results are discussed in the context of how phonological opacity, productivity, and the directionality of tone sandhi patterns influence the representation of tone sandhi words as well as how the lexicality of the primes and the participants’ usage pattern of Shanghai may have influenced the results.
The repertoire of forms of address can be considered as one of the determinants of the discourse genre, which makes it possible to capture its evolution and cultural variations. From such comparative, intra- and intercultural perspective, adopting an interactive approach in the analysis of political discourse, we will look at the practice of addressing one another in the French and Polish politicalmedia discourse. While in both languages the linguistic norm recommends the use of the polite forms of address in official situations, the cases of the use of the familiar pronoun tu / ty in media interactions between politicians are not rare at all. Whether it is an informal talk of politicians caught by the media, a television pre-election debate, or a meeting of the heads of state, addressing the other person by the familiar forms is a manifestation of a deliberate blurring of the boundaries between the front-stage and backstage in political discourse in order to create the impression of intimacy andequality between the interlocutors.
The first edition of one of the most important and mysterious novels of the 20th century appeared more than fifty years ago. Despite the passage of time The Master and Margarita still enjoys popularity; it also intrigues and inspires. Until now five Polish translations of Bulgakov’s novel have appeared. It is known that the interpretation of the original might be expressed in the form of many potential texts that are communicatively equivalent. There is no doubt that it is the translator who plays a vital role in any translation; her/his personality, life experience, knowledge, skills, and also the times s/he lives in regulate the target text. That is why, no matter how many times a text is translated, the final product will always be different. Taking this into consideration, the author will compare the three Polish translations of Bulgakov’s Master and Margarita, paying attention to the diachronic perspective as far as linguistic norms are concerned, the modernity of language, and the way the anthroponyms are expressed.
The specification of concept human in the phraseology of East Steppe Ukrainian dialects based on the code system of culture and formation of the correspondent linguistic database have been found and described in the work.The concept human has been positioned as one of the main segments of the conceptual picture of Ukrainian’s world and structured into four components: 1) conceptual that involves the review of human’s definitions in the lexicographical works; 2) descriptive that is represented by secondary sign system; 3) semantic that is represented by phraseoideographic paradigm of areal phraseological units and 4) estimative that is formed by positive evaluative, negative evaluative and neutral evaluative areal phraseological units.The code system of culture including gradual separation of somatic, subject, zoomorphic, anthropic, phytomorphic, natural, gastronomic, spiritual, quantitative, temporal, mythological, actional, sensor, spatial, qualificative, speech codes and two-component and three-component intercode passages of the concept human in the phraseology of East Steppe Ukrainian dialects with following quantitative representation has been identified.The step-by-step algorithm of design cycle of linguistic database «Concept human in the phraseology of East Steppe Ukrainian proverbs» with the compulsory parametrization of areal phraseological units according to ideographic, axiological, structural qualifications has been made.
We introduce the task of predicting adverbial presupposition triggers such as also and again. Solving such a task requires detecting recurring or similar events in the discourse context, and has applications in natural language generation tasks such as summarization and dialogue systems. We create two new datasets for the task, derived from the Penn Treebank and the Annotated English Gigaword corpora, as well as a novel attention mechanism tailored to this task. Our attention mechanism augments a baseline recurrent neural network without the need for additional trainable parameters, minimizing the added computational cost of our mechanism. We demonstrate that our model statistically outperforms a number of baselines, including an LSTM-based language model.
The Norm in Translation and the Feeling of Lack – an Attempt of Psychoanalytical Reflection on the Experience of Translation This article is an attempt to use psychoanalytical language for describing the experience of translation and the discourse on translation. While considering the recognitions and method proposed by Tadeusz Sławek in an article titled, „Kalibanizm. Filozoficzne dylematy tłumaczenia“ [Kalibanism. The philosophic dilemma of translation], I try to describe them and develop, to say a few words about the translation and linguistic norm in the context of feeling of lack widely spread in the discourse on translation. Thinking about translation process as one affected by the melancholy I try to point out the indelible contradiction between the existence of translation norm and the imaginative concept of semantic plenitude of the original text. Repeating after Sławek I make an attempt to indicate a “crypt” that is being build by the target language around the original work, a crpyt existence of which was omitted with silence, although it have had a grate influence on the way that translation funcion in our culture.
We describe the first automatic approach for merging coreference annotations obtained from multiple annotators into a single gold standard. This merging is subject to certain linguistic hard constraints and optimization criteria that prefer solutions with minimal divergence from annotators. The representation involves an equivalence relation over a large number of elements. We use Answer Set Programming to describe two representations of the problem and four objective functions suitable for different datasets. We provide two structurally different real-world benchmark datasets based on the METU-Sabanci Turkish Treebank and we report our experiences in using the Gringo, Clasp, and Wasp tools for computing optimal adjudication results on these datasets.
I participated in the CMLD (Computational Methods for Endangered Languages) conference last week at ENS in Paris. Laurent Besacier was also attending. Here is a short report.-About fifty participants, several were also present at the workshop on Uralic languages in Saint Petersburg a year ago-Proposal to use facial recognition for metadata on speakers' names (Niko Partanen, one of the organizers).-Joachim Nivre presented the international effort on "Universal Dependencies (Treebanks)". Already 60 languages. I ask him what he plans to do for the remaining 6840... He says it's a funding problem. He agrees that they should be treated by family of languages to save money.-Presentation by Jargal Badagarov (Mongolia:
Based on Chinese dependency treebank PMT 1.0, the present study investigates the positional aspects of dependency distance (DD) quantitatively. Results show that (1) as the word position in the sentence increases, the tendency of mean dependency distance (MDD) in different sentence length groups shows striking similarity. The two longest MDDs generally are in the sentence-initial and sentence-final positions; (2) The consensus string (CS) and weighted consensus string (WCS) show some characteristics, which demonstrates again that human cognition plays an important role in affecting DD and dependency distance minimization is a universal tendency; (3) The distribution of DD in each sentential position can be captured by power law, which implies that something like a vertical structure of texts exists.
Democratization in speech have not only broadened the ways of language expression, manifestations of linguistic individuality, but have led to many negative phenomena. This is typical not only for marginal communication, but also for political discourse, especially � for the media, which has a huge impact on the speech behavior of society. Nowadays, the concept of ethical and linguistic standard have been actualized, it is developed not only in the framework of ecological linguistics, but also in legal linguistics. In the context of ethical and speech norms, it is important to note the words usage is inseparable from the categories of ethics. These new phenomena are due to the combination of all the circumstances of socio-political and cultural life. It is impossible to give any recommendations in the field of regulation in general and ethical and linguistic norms in particular without taking them into account. The methodology of the work is based on a combination of panchronic and diachronic approaches to the language. The leading method is extrapolation of language theories, which arose in the same historical conditions, to the conditions of different historical reality, synthesis of interpretative and comparative approaches to the material, component-semantic and contextual analysis, composite analysis.
The article is devoted to the problem of identifying the stylistic functions of addresses, which are used in Internet communication. To achieve this goal, the author has solved several problems. For the first, the main features of Internet communication are anonymity, mediation, distance and frequent violation of linguistic norms. The last attribute refers to the adresses, which are used in Internet messages. For the second, official and household addreses are used in the Internet communication,. The stylistic function of the official addresses is the indication of the status and social role of the addressee. The stylistic function of household appeals is the indication of proximity between communication participants, the emphasis on positive or negative connotations of addresses.In addition, household addresses reflect the new social trends associated with the using of e-mail, aliases etc. The main conclusions of this research: the greatest number of the addresses in Internet communication is recorded in the materials of business correspondence. The users of the network very rarely use the household addresses. The author believes that the main reason for such quantitative dynamics is the avoidance or inability of addressees to show an emotional attitude to their interlocutor.
We propose the dense RNN, which has the fully connections from each hidden state to multiple preceding hidden states of all layers directly. As the density of the connection increases, the number of paths through which the gradient flows can be increased. It increases the magnitude of gradients, which help to prevent the vanishing gradient problem in time. Larger gradients, however, can also cause exploding gradient problem. To complement the trade-off between two problems, we propose an attention gate, which controls the amounts of gradient flows. We describe the relation between the attention gate and the gradient flows by approximation. The experiment on the language modeling using Penn Treebank corpus shows dense connections with the attention gate improve the model’s performance.
This study aims at exploring new norms as to the textual additions in parentheses (=TAiPs) in the translation of a Quranic text as writer-oriented devices of textuality. Coding for this sort of information could be useful in establishing an impact on any decision-making process on the TL version; such TAiPs can give a translated text of the Quran unity and purpose and distinguish it from a disconnected sequence of sentences. Six small-sized chapters of the Quran were selected as a research sample including a number of four handred forty two (442) TAiPs. Two writer-oriented kinds of textuality were found: cohesivity at the levels of grammar and lexis to be in form of recurrence, reference, substitution, ellipsis and conjunction; and relationality by coherence and intentionality to be in form of reiteration, collocation, connotation, evocation and interpretation. The study is a detailed analysis of such a severely criticized yet officially approved English interpretation of the Quran as the Hilali and Khan Translation (=HKT) against a predetermined set of text-linguistic norms. The strength or weakness of TAiPs as to how they might alleviate or aggravate the TL version is eventually identified for sake of improvement.
Abstract People remember events and materials better when these are congruent with their mood at retrieval; this is known as the mood-congruent memory bias. This effect is largest when the materials are self-referential and this is known as the self-reference effect. We present two word rating studies, to create a list of self-referential valenced words that may be used as stimuli to investigate the influence of valence on cognitive processing in depressive ruminators. Words selected from the Affective Norms for English Words pool were rated by an unselected sample for self-referentiality (Study 1) and validated with ratings provided by depressive ruminators. As hypothesized, depressive ruminators rated negative words as more self-referential than an unselected sample. Using this list, valence differentiated performance between depressive ruminators and healthy controls in a working memory updating task. We thus created a list of self-referential valenced words matched on factors that influence word processing.
The statistical parsing of morphologically rich languages is hindered by the inability of parsers to collect solid statistics because of the large number of word types in such languages. There are however two separate but connected problems, reducing data sparsity of known words and handling rare and unknown words. Methods for tackling one problem may inadvertently negatively impact methods to handle the other. We perform a tightly controlled set of experiments to reduce data sparsity through class-based representations in combination with unknown word signatures with two PCFG-LA parsers that handle rare and unknown words differently on the German TiGer treebank. We demonstrate that methods that have improved results for other languages do not transfer directly to German, and that we can obtain better results using a simplistic model rather than a more generalized model for rare and unknown word handling.
Detecting lexical entailment plays a fundamental role in a variety of natural language processing tasks and is key to language understanding. Unsupervised methods still play an important role due to the lack of coverage of lexical databases in some domains and languages. Most of the previous approaches were either based on statistical hypothesis of specific entailment relations or tried to encode word relations in low-dimensional vector embeddings. This thesis builds upon one of the few approaches which intrinsically model entailment in a vector space. We then further generalize this model by introducing an alternative, distributional representations for words which harnesses tools from optimal transport to define distance or entailment measures between such representations. We evaluated the models on hypernymy detection where our distributional estimate significantly improves over the underlying model and even outperforms state-of-the-art on some datasets.
In this paper we present the linguistic databases developed during our 8-year lexicographic research on the Modern Greek Standard (MGS) verbal system. Apart from the intermediate databases presented, the main products are (a) a new conjugation system of 385 paradigmatic models, which allows for the automatic generation of all verbal lexical morphemes and monolexical forms (b) a statistically established database of 151,536 distinctive verb-final grapheme sequences which allow for the automatic tagging of all monolexical verbal tokens without the traditional intervention of any built-in lexicon, and (c) a linear Iemmatisation morphophonological rule system accessed on the basis of the distinctive grapheme sequences identified.
Tree-structured neural network architectures for sentence encoding draw inspiration from the approach to semantic composition generally seen in formal linguistics, and have shown empirical improvements over comparable sequence models by doing so. Moreover, adding multiplicative interaction terms to the composition functions in these models can yield significant further improvements. However, existing compositional approaches that adopt such a powerful composition function scale poorly, with parameter counts exploding as model dimension or vocabulary size grows. We introduce the Lifted Matrix-Space model, which uses a global transformation to map vector word embeddings to matrices, which can then be composed via an operation based on matrix-matrix multiplication. Its composition function effectively transmits a larger number of activations across layers with relatively few model parameters. We evaluate our model on the Stanford NLI corpus, the Multi-Genre NLI corpus, and the Stanford Sentiment Treebank and find that it consistently outperforms TreeLSTM
This paper describes our system (SLT-Interactions) for the CoNLL 2018 shared task: Multilingual Parsing from Raw Text to Universal Dependencies. Our system performs three main tasks: word segmentation (only for few treebanks), POS tagging and parsing. While segmentation is learned separately, we use neural stacking for joint learning of POS tagging and parsing tasks. For all the tasks, we employ simple neural network architectures that rely on long short-term memory (LSTM) networks for learning task-dependent features. At the basis of our parser, we use an arc-standard algorithm with Swap action for general non-projective parsing. Additionally, we use neural stacking as a knowledge transfer mechanism for cross-domain parsing of low resource domains. Our system shows substantial gains against the UDPipe baseline, with an average improvement of 4.18% in LAS across all languages. Overall, we are placed at the 12 th position on the official test sets.
Despite the advances in information processing systems, word-sense disambiguation tasks are far to be satisfactory as testified by numerous limitations of current translation systems and text inference systems. This paper attempts to investigate new techniques in knowledge based word-sense disambiguation field. First, by exploring the WordNet lexical database and part-of-speech conversion through the established CatVar database that translates all non-noun words into their noun counterparts, and following the spirit of Lesk's disambiguation algorithm, a new disambiguation algorithm that maximizes the overall semantic similarity in the sense of Wu and Palmer measure between each sense of the target word and synsets of words of the context, is established. Second, motivated by the existence of WordNet domains for individual synsets, an overlapping based approach that quantifies the set intersection of synset domains, if not empty, or the hierarchy structure of the domains links through a simple path-length measure is put forward. Third, instead of exploring the whole set of words involved in the context, a selective approach that uses syntactic feature as outputted by Stanford Parser and a fixed length windowing is developed. The developed algorithms are evaluated according to two commonly employed dataset where a clear improvement to the baseline algorithm has been acknowledged.
Rapid advances in information technology and proliferation of social media services have caused a radical transformation of human communication. Having created a social media presence people engage in computer-mediated communication, set their own goals as well as perfect their knowledge of English as a global language. The richness and diversity of computer-mediated discourse is concentrated in multiple online experiences and therefore enables to study a great number of linguistic changes. Various studies of computer-mediated discourse analyze socio-psychological characteristics in coherent sequences of sentences, propositions, speech or turns-at-talk. The given article aims at presenting vocabulary teaching strategies to new computer-mediated language and their influence on students' acquisition. Our contribution provides an overview of the recent new entries of computer-mediated vocabulary in online crowdsourced dictionaries. The main ways of forming new words as well as wide-spread semantic changes are viewed (acronyms, compounds, suffixes, blended words, conversion, etc.) Computer-mediated vocabulary teaching in the classroom covers change, diversity, disputes in economical, political, social spheres (such as narcissistic tendencies, emotional correctness, excessive use of social media and increasing reliance on technology, equal rights movement, task-based employment, etc). As the social media universe strives for a thorough integration with the user's life language learners are expected to embrace the latest changes in linguistic norms and devise an appropriate philosophy of language management.
This study aims to electronically assessed (e-assessment) students’ replies in response to teachers’ question. It can be useful to systematize the question answering context regarding matching text semantically through WordNet semantic similarity techniques. WordNet is a lexical database of words’ synonyms. It uses group of synonyms called synsets for semantical operation of English text. For this purpose, a new methodology is proposed to automate e-assessment in the field of education. The collected dataset contains 210 pairs of words extracted from different undergraduate students’ replies in contradiction of teacher’s question statement. Further WordNet similarity measures i.e. Path Length, Lin, Wu &Palmer and Hirst & Onge are used to compute the semantic relatedness score. In the pilot study 42 pair of words were extracted from 8 students’ replies, which are marked using semantic similarity measures and equated with teacher’s marks. Teachers are provided with four boxes of the mark while our developed method provides a precise measure of marks. The experiment is shown with comprehensive dataset resulting with words’ frequencies in similarity measures.
Introduction: The purpose of this investigation was to examine the independent and combined effects of caffeine (CAF) alone or as a part of a multi-ingredient pre-workout supplement (PWS) on resistance exercise performance in recreationally active males. Methods: In a single-blind, randomized, placebo (PLA) controlled, crossover design; 10 recreationally active males (20.5 ± 0.9; 178.9 ± 7.7 cm; 81.8 ± 11.5 kg) completed three laboratory visits, after determination of one repetition-maximum (1-RM) on the bench press and leg press, where they performed bench press and leg press to failure at a load of 70% 1-RM. Subjects were randomly assigned to ingest either one serving of a commercially available PWS (C4 Original, Cellucor, Bryan, TX, United States), a dosage-matched anhydrous CAF beverage (150 mg), or a taste- matched PLA beverage. Heart rate (HR), affect, rating of perceived exertion (RPE), and mood state was assessed 20 minutes pre and post-substance ingestion, and immediately after exercise.Results: Participants completed significantly more repetitions to failure (p = 0.006) and lifted significantly greater weight (p = 0.009) during Leg Press in the PWS and CAF conditions compared to the PLA condition. There was not a significant difference found between CAF and PWS trials (p > 0.05).Conclusions: This data suggests that both CAF and PWS may have a positive effect on exercise performance in leg press but was not effective in increasing muscular endurance in bench press. The commercially available PWS offered no additional ergogenic effects when compared to the CAF.
The article deals with the variant terms in normative aspect codified in Ukrainian art lexicography of the 21st century. Dictionary codification of variant terms indicates changing in the language and deliberate influence of the society on the development of terminological norm. Variation is a existence form of objects of the surrounding reality, in particular, of scientific concepts, which defines the laws of their function and interaction. The choice of sources is due to the fact that the selected dictionaries are represented modern art knowledge. Dictionaries play a significant role in the normalization of language, the spread of linguistic norms, and therefore they are a grateful and relevant material for the analysis of variation in the Ukrainian art terminology. The article focuses on the importance of the scientific philological study of art terminology – the field of knowledge, which is rapidly developing in modern conditions, acquiring new meanings and forms. The variant terms of the art terminology, codified in Ukrainian special vocabulary, are analyzed. Three types of variant terms, phonemic, derivational and morphological-phonemic units, are fixed in the Ukrainian art terminology. It was found out that among the reasons for the occurrence of phonemic variant terms of the analyzed terminology tends to facilitate articulation of the learned term; the appearance of derivation of variant terms is conditioned by the presence of various derivative models in the Ukrainian language and the search for forms of terms that correspond most closely to modern productive models of term derivation; functioning of morphological-phonemic variant terms is explained by different degrees of grammatical adaptation of foreign-language art terms. It also traces the effect of an analogy inherent to all three of the varieties mentioned. In general, the article discusses the essence of the problem of terminological variation as one of the most relevant processes in the regulation and standardization of the Ukrainian art terminology.
Comparative vocabulary of Lolo-Burmese languages, with glosses and word lists merged from Shintani (2001) and Lama (2012). This spreadsheet is a work in progress.
Abstract: Even though it has been noted that comparative concepts for typology are merely instruments for research and may therefore differ across researchers (Haspelmath 2010), we want different databases to be comparable. For example, we would like to compare data from WALS (Haspelmath et al. 2005/2013) to be comparable with data from SAILS Online (Muysken et al.) The task is thus similar to the task of lexical comparison across languages by means of a set of comparison meanings. For the latter, a semi-standard ontology now exists: The Concepticon (List et al. 2016, concepticon.clld.org), which has over 2500 comparison meanings which bring together comparison meanings from diverse lexical databases. This allows automatic comparison of lexical forms from diverse databases- The present talk explores the possibility of a counterpart of this for grammatical comparative concepts, called Grammaticon, which would facilitate the comparison of different grammatical datasets. The goal of allowing comparability, also for machines (i.e. interoperability), is similar to that of the GOLD ontology (General Ontology for Linguistic Description, Farrar & Langendoen 2003), but the latter attempts to come up with a general set of concepts for description, which is impossible, because each language has its own system. The GOLD’s goal thus resembles the idea of Wierzbicka’s natural semantic metalanguage, which can be used to describe any language, but which has not proved practical. The Grammaticon’s goal is more modest, in that it proposes a set of highly ecumenical comparative concepts with associated terms that can be used in comparison. No claim is made that these concepts should be useful for describing individual languages. References Farrar, Scott & D. Terence Langendoen. 2003. A linguistic ontology for the Semantic Web. GLOT International 7(3). 97-100. (Available at http://www.linguistics-ontology.org/gold.html) Haspelmath, Martin. 2010. Comparative concepts and descriptive categories in crosslinguistic studies. Language 86(3). 663–687. Haspelmath, Martin, Matthew Dryer, David Gil & Bernard Comrie (eds.). 2005. The world atlas of language structures. Oxford: Oxford University Press. (2013 online version available at http://wals.info) Kortmann, Bernd & Lunkenheimer, Kerstin (eds.) 2013. The Electronic World Atlas of Varieties of English. Leipzig: Max Planck Institute for Evolutionary Anthropology. (Available at http://ewave-atlas.org) List, Johann-Mattis & Cysouw, Michael & Forkel, Robert (eds.) 2016. Concepticon. Jena: Max Planck Institute for the Science of Human History. Muysken, Pieter, Harald Hammarström, Olga Krasnoukhova, Neele Müller, Joshua Birchall, Simon van de Kerke, Loretta O'Connor, Swintha Danielsen, Rik van Gijn & George Saad. 2016. South American Indigenous Language Structures (SAILS) Online. Leipzig: Max Planck Institute for Evolutionary Anthropology. (Available at http://sails.clld.org)
Knowledge Organization Systems (KOS), in the form of classification systems, thesauri, lexical databases, ontologies, and taxonomies, play a crucial role in digital information management and applications generally. Carrying semantics in a well-controlled and documented way, Knowledge Organisation Systems serve a variety of important functions: tools for representation and indexing of information and documents, knowledge-based support to information searchers, semantic road maps to domains and disciplines, communication tool by providing conceptual framework, and conceptual basis for knowledge based systems, e.g. automated classification systems. New networked KOS (NKOS) services and applications are emerging, and we have reached a stage where many KOS standards exist and the integration of linked services is no longer just a future scenario. This editorial describes the workshop outline and overview of presented papers at the 15th European Networked Knowledge Organization Systems Workshop (NKOS 2016) in Hannover, Germany.
The African wordnet (AWN) provides South African indigenous languages with a platform to access a machine-readable lexical database organised by meaning. The creation of the African wordnet was based on the Princeton wordnet. As in the case of the Princeton wordnet, the African wordnet groups African language words into sets of synonyms along with short definitions and usage examples, as well as records relations between synonyms. This article examines a number of synsets in order to identify the word-formation processes used by various linguists in constructing the AWN. Since the English Princeton wordnet was used as the basis for the lexical database in the creation of the African wordnet, various word-formation strategies had to be used to account for lexical items that are not lexicalised in the African languages. Access to the created synsets was gained via a web browser, which is an automated text analysis application.