Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
The Web has accumulated a rich source of information, such as text, image, rating, etc, which represent different aspects of user preferences. However, the heterogeneous nature of this information makes it difficult for recommender systems to leverage in a unified framework to boost the performance. Recently, the rapid development of representation learning techniques provides an approach to this problem. By translating the various information sources into a unified representation space, it becomes possible to integrate heterogeneous information for informed recommendation.
Comunicacio presentada a: the Fourth International Conference on Dependency Linguistics (Depling 2017), celebrada a Pisa, Italia, del 18 al 20 de setembre 2017.
Virtual reality (VR) has been proposed as a methodological tool to study the basic science of psychology and other fields. One key advantage of VR is that sharing of virtual content can lead to more robust replication and representative sampling. A database of standardized content will help fulfill this vision. There are two objectives to this study. First, we seek to establish and allow public access to a database of immersive VR video clips that can act as a potential resource for studies on emotion induction using virtual reality. Second, given the large sample size of participants needed to get reliable valence and arousal ratings for our video, we were able to explore the possible links between the head movements of the observer and the emotions he or she feels while viewing immersive VR. To accomplish our goals, we sourced for and tested 73 immersive VR clips which participants rated on valence and arousal dimensions using self-assessment manikins. We also tracked participants' rotational head movements as they watched the clips, allowing us to correlate head movements and affect. Based on past research, we predicted relationships between the standard deviation of head yaw and valence and arousal ratings. Results showed that the stimuli varied reasonably well along the dimensions of valence and arousal, with a slight underrepresentation of clips that are of negative valence and highly arousing. The standard deviation of yaw positively correlated with valence, while a significant positive relationship was found between head pitch and arousal. The immersive VR clips tested are available online as supplemental material.
We describe a variant of Child-Sum Tree-LSTM deep neural network (Tai et al, 2015) fine-tuned for working with dependency trees and morphologically rich languages using the example of Polish. Fine-tuning included applying a custom regularization technique (zoneout, described by (Krueger et al., 2016), and further adapted for Tree-LSTMs) as well as using pre-trained word embeddings enhanced with sub-word information (Bojanowski et al., 2016). The system was implemented in PyTorch and evaluated on phrase-level sentiment labeling task as part of the PolEval competition.
We systematically explore regularizing neural networks by penalizing low\nentropy output distributions. We show that penalizing low entropy output\ndistributions, which has been shown to improve exploration in reinforcement\nlearning, acts as a strong regularizer in supervised learning. Furthermore, we\nconnect a maximum entropy based confidence penalty to label smoothing through\nthe direction of the KL divergence. We exhaustively evaluate the proposed\nconfidence penalty and label smoothing on 6 common benchmarks: image\nclassification (MNIST and Cifar-10), language modeling (Penn Treebank), machine\ntranslation (WMT'14 English-to-German), and speech recognition (TIMIT and WSJ).\nWe find that both label smoothing and the confidence penalty improve\nstate-of-the-art models across benchmarks without modifying existing\nhyperparameters, suggesting the wide applicability of these regularizers.\n
In this paper, I will describe the methodology to develop the first sample of a dependency treebank collecting German aesthetic writings of the late 18th century. A gold standard of the target data was annotated in order to evaluate some datadriven tools, trained on contemporary web news. Results are reported and discussed.
Multiword expressions (MWEs) are linguistic objects containing two or more words and showing idiosyncratic behavior at different levels. Treebanks with annotated MWEs enable studies of such properties, as well as training and evaluation of MWE-aware parsers. However, few treebanks contain full-fledged MWE annotations. We show how this gap can be bridged in Polish by projecting 3 MWE resources on a constituency treebank.
Abstract This paper addresses the feasibility of cross-lingual parsing with Universal Dependencies (UD) between Romance languages, analyzing its performance when compared to the use of manually annotated resources of the target languages. Several experiments take into account factors such as the lexical distance between the source and target varieties, the impact of delexicalization, the combination of different source treebanks or the adaptation of resources to the target language, among others. The results of these evaluations show that the direct application of a parser from one Romance language to another reaches similar labeled attachment score (LAS) values to those obtained with a manual annotation of about 3,000 tokens in the target language, and unlabeled attachment score (UAS) results equivalent to the use of around 7,000 tokens, depending on the case. These numbers can noticeably increase by performing a focused selection of the source treebanks. Furthermore, the removal of the words in the training corpus (delexicalization) is not useful in most cases of cross-lingual parsing of Romance languages. The lessons learned with the performed experiments were used to build a new UD treebank for Galician, with 1,000 sentences manually corrected after an automatic cross-lingual annotation. Several evaluations in this new resource show that a cross-lingual parser built with the best combination and adaptation of the source treebanks performs better (77 percent LAS and 82 percent UAS) than using more than 16,000 (for LAS results) and more than 20,000 (UAS) manually labeled tokens of Galician.
In this paper, we describe a method for mapping the phonological feature location of Swedish Sign Language (SSL) signs to the meanings in the Swedish semantic dictionary SALDO.By doing so, we observe clear differences in the distribution of meanings associated with different locations on the body.The prominence of certain locations for specific meanings clearly point to iconic mappings between form and meaning in the lexicon of SSL, which pinpoints modalityspecific properties of the visual modality.
The paper contains the research of noun-compounds from modern Tibetan corpus with the use of a relational lexical database. The lexical database represents a consistent classification of meanings of Tibetan lexical units with different relations between them. The paper describes the structure of the database; principles of process work with Tibetan compounds; recognized types of compounds semantic structure.
In this paper, we propose using a "bootstrapping" method for constructing a dependency treebank of Arabic tweets. This method uses a rule-based parser to create a small treebank of one thousand Arabic tweets and a data-driven parser to create a larger treebank by using the small treebank as a seed training set. We are able to create a dependency treebank from unlabelled tweets without any manual intervention. Experiments results show that this method can improve the speed of training the parser and the accuracy of the resulting parsers.
In this paper, we present evaluation of URDU.KON-TB in the dependency parsing domain. The URDU.KON-TB treebank is developed on the bases of the phrase structure and hyper dependency structure which are only functional constituent's label. Treebank was annotated with three levels of annotation tagset, the semi-semantic POS (SSP), semi-semantic Syntactic (SSS) and Functional (F) tagset and was checked for the Phrase Structure Parsing domain. To evaluate this treebank in the Dependency Parsing domain we have selected MaltParser. To use data in the parser, we have converted the URDU.KON-TB treebank annotated data according to the CONLL format. The compatibility of data to CoNLL is also measured along with usability of data in the dependency parsing domain. To make the data compatible, few assumptions are taken. The converted data is used to evaluate the system by dividing 80% data as training data and 20% data as testing data. We have performed eight experiments. Four experiments are conducted with six different feature models with converted data. The experiments results show URDU.KON-TB treebank is not suitable for the dependency parsing as dependency relation because Head information was missing in the treebank. We then performed four experiments with an assumption based enhancement by adding Head information. The algorithm used to train and test data is Nivre arc-agear algorithm. The new experiments show this treebank data can be used to develop new dependency treebank for Urdu.
This paper describes the use of GrETEL for linguistic research. GrETEL is a linguistic search tool that enables users to look up constructions in syntactically annotated corpora or <i>treebanks</i>. It provides online access to the data, allowing users to query a treebank using either an example sentence or an XPath expression in order to look for similar constructions. A major asset of GrETEL is that it enables non-technical users to consult treebanks in a user-friendly way, which is also in line with the main CLARIN goal of applying the results of speech and language technology to research in the humanities and the social sciences. Besides a description of the querying procedure in GrETEL, this paper presents a selection of research in Dutch syntax and semantics that has been carried out using GrETEL. Furthermore, an overview is given of further developments.
The paper illustrates an effective and innovative method for detecting erroneously annotated arcs in gold dependency treebanks based on an algorithm originally developed to measure the reliability of automatically produced dependency relations. The method permits to significantly restrict the error search space and, more importantly, to reliably identify patterns of systematic recurrent errors which represent dangerous evidence to a parser which tendentially will replicate them. Achieved results demonstrate effectiveness and reliability of the method.
Public textual cyberbullying has become one of the most prevalent issues associated with\nonline safety of young people, particularly on social networks. To address this issue, we\nargue that the boundaries of what constitutes public textual cyberbullying needs to be first\nidentified and a corresponding linguistically motivated definition needs to be advanced.\nThus, we propose a definition of public textual cyberbullying that contains three necessary\nand sufficient elements: the personal marker, the dysphemistic element and the\ncyberbullying link between the previous two elements. Subsequently, we argue that one of\nthe cornerstones in the overall process of mitigating the effects of cyberbullying is the\ndesign of a cyberbullying lexical database that specifies what linguistic and cyberbullying\nspecific information is relevant to the detection process. In this vein, we propose a novel\ncyberbullying lexical database based on the definition of public textual cyberbullying. The\noverall architecture of our cyberbullying lexical database is determined semantically, and, in\norder to facilitate cyberbullying detection, the lexical entry encapsulates two new semantic\ndimensions that are derived from our definition: cyberbullying function and cyberbullying\nreferential domain. In addition, the lexical entry encapsulates other semantic and syntactic \ninformation, such as sense and syntactic category, information that, not only aids the\nprocess of detection, but also allows us to expand the cyberbullying database using\nWordNet (Miller, 1993).
Similarity calculation between business process models has an important role in managing repository of business process model. One of its uses is to facilitate the searching process of models in the repository. Business process similarity is closely related to semantic string similarity. Semantic string similarity is usually performed by utilizing a lexical database such as WordNet to find the semantic meaning of the word. The activity name of the business process uses terms that specifically related to the business field. However, most of the terms in business domain are not available in WordNet. This case would decrease the semantic analysis quality of business process model. Therefore, this study would try to improve semantic analysis of business process model. We present a new lexical database called B-BabelNet. B-BabelNet is a lexical database built by using the same method in BabelNet. We attempt to map the Wikipedia page to WordNet database but only focus on the word related to the domain of business. Also, to enrich the vocabulary in the business domain, we also use terms in the business-specific online dictionary (businessdictionary.com). We utilize this database to do word sense disambiguation process on business process model activity’s terms. The result from this study shows that the database can increase the accuracy of the word sense disambiguation process especially in particular terms related to the business and industrial domains.
Public textual cyberbullying has become one of the most prevalent issues associated with online safety of young people, particularly on social networks. To address this issue, we argue that the boundaries of what constitutes public textual cyberbullying needs to be first identified and a corresponding linguistically motivated definition needs to be advanced. Thus, we propose a definition of public textual cyberbullying that contains three necessary and sufficient elements: the personal marker, the dysphemistic element and the cyberbullying link between the previous two elements. Subsequently, we argue that one of the cornerstones in the overall process of mitigating the effects of cyberbullying is the design of a cyberbullying lexical database that specifies what linguistic and cyberbullying specific information is relevant to the detection process. In this vein, we propose a novel cyberbullying lexical database based on the definition of public textual cyberbullying. The overall architecture of our cyberbullying lexical database is determined semantically, and, in order to facilitate cyberbullying detection, the lexical entry encapsulates two new semantic dimensions that are derived from our definition: cyberbullying function and cyberbullying referential domain. In addition, the lexical entry encapsulates other semantic and syntactic
We release Galactic Dependencies 1.0---a large set of synthetic languages not found on Earth, but annotated in Universal Dependencies format. This new resource aims to provide training and development data for NLP methods that aim to adapt to unfamiliar languages. Each synthetic treebank is produced from a real treebank by stochastically permuting the dependents of nouns and/or verbs to match the word order of other real languages. We discuss the usefulness, realism, parsability, perplexity, and diversity of the synthetic languages. As a simple demonstration of the use of Galactic Dependencies, we consider single-source transfer, which attempts to parse a real target language using a parser trained on a "nearby" source language. We find that including synthetic source languages somewhat increases the diversity of the source pool, which significantly improves results for most target languages.
We explore the properties of byte-level recurrent language models. When given sufficient amounts of capacity, training data, and compute time, the representations learned by these models include disentangled features corresponding to high-level concepts. Specifically, we find a single unit which performs sentiment analysis. These representations, learned in an unsupervised manner, achieve state of the art on the binary subset of the Stanford Sentiment Treebank. They are also very data efficient. When using only a handful of labeled examples, our approach matches the performance of strong baselines trained on full datasets. We also demonstrate the sentiment unit has a direct influence on the generative process of the model. Simply fixing its value to be positive or negative generates samples with the corresponding positive or negative sentiment.
In this work, we present a minimal neural model for constituency parsing based on independent scoring of labels and spans. We show that this model is not only compatible with classical dynamic programming techniques, but also admits a novel greedy top-down inference algorithm based on recursive partitioning of the input. We demonstrate empirically that both prediction schemes are competitive with recent work, and when combined with basic extensions to the scoring model are capable of achieving state-of-the-art single-model performance on the Penn Treebank (91.79 F1) and strong performance on the French Treebank (82.23 F1).
We first present a minimal feature set for transition-based dependency parsing, continuing a recent trend started by Kiperwasser and Goldberg (2016a) and Cross and Huang (2016a) of using bi-directional LSTM features. We plug our minimal feature set into the dynamic-programming framework of Huang and Sagae (2010) and With our minimal features, we also present Opn 3 q global training methods. Finally, using ensembles including our new parsers, we achieve the best unlabeled attachment score reported (to our knowledge) on the Chinese Treebank and the "second-best-in-class" result on the English Penn Treebank.
The article is devoted to a lexical analysis of the New Testament translation into Russian performed by an established statespersonKonstantin Pobedonostsev (1827Pobedonostsev ( -1907) ) at the beginning of the 20 th century. The lexical particularities of it have been revealed by means of its lexical comparison with the Synodal translation and the Church-Slavonic liturgical version. According to academic interpretation of the data collected the author stated that Konstantin Pobedonostsev's translationpreservesmore resemblance to the Church-Slavonic liturgical version, as there are 185 wordsinhis work, which were not found in the Synodal translation, but were used in Church-Slavonic liturgical version. The major part of the words is registered in lexicographical sources and reflects language norms of Pobedonostsev's lifetime. However, the smaller part is registered in the Russian Language National Corpus as being presented in the texts of the 1820-1920 period. The vocabulary specificity of the translation version under study lies in vast references to the Church-Slavonic liturgical version word pool. This fact is explained by stylistic preferences of the Gospel translators as well as by the target addressee image (people who are well informed about the orthodox liturgical tradition). In conclusion the author states that in his translation Konstantin Pobedonostsev never took words from the Church-Slavonic texts without prolonged meditation, the replacement cases seem to be an intentional choice of the words that were frequently used in the speech of his epoch.
Background: Gestures can provide an excellent natural alternative to verbal communication in people with aphasia (PWA). However, despite numerous studies focusing on gesture production in aphasia, it is still a matter of debate whether the gesture system remains intact after language impairment and how PWA use gestures to improve communication. A likely source for the contradicting results is that many studies were conducted on individual cases or in heterogeneous groups of individuals with additional cognitive deficits such as conceptual impairment and comorbid conditions such as limb apraxia.Aims: The goal of the current study was to evaluate the integrity and function of gestures in PWA in light of cognitive theories of language–gesture relationship. Since all such theories presuppose the integrity of the conceptual system, and the absence of comorbid conditions that selectively impair gesturing (i.e., limb apraxia), our sample was selected to fulfill these assumptions.Methods & Procedures: We examined gesture production in eight PWA with preserved auditory comprehension, no comorbidities, and various degrees of expressive deficit, as well as 11 age- and education-matched controls, while they described events in 20 normed video clips. Both speech and gesture data were coded for quantitative measures of informativeness, and gestures were grouped into several functional categories (matching, complementary, compensatory, social cueing, and facilitating lexical retrieval) based on correspondence to the accompanying speech. Using rigorous group analyses, individual-case analyses, and analyses of individual differences, we provide converging evidence for the integrity and type of function(s) served by gesturing in PWA.Outcomes & Results: Our results indicate that the gesture system can remain functional even when language production is severely impaired. Our PWA heavily relied on iconic gestures to compensate for their language impairment, and the degree of such compensation was correlated with the extent of language impairment. In addition, we found evidence that producing iconic gestures was related to higher success rates in resolving lexical retrieval difficulties.Conclusions: When comprehension and comorbidities are controlled for, impairment of language and gesture systems is dissociable. In PWA with good comprehension, gesturing can provide an excellent means to both compensate for the impaired language and act as a retrieval cue. Implications for cognitive theories of language–gesture relationship and therapy are discussed.
The all-Russian lexeme время and its derivatives in the dialects of the Russian language are considered. The author believes that the semantic volume of the word, well-known to the literary language, and its dialectal counterparts may not be identical due to different discursive conditions generated by the culture. The relevance of the study is determined by the increased attention in modern linguistics to the problems of reflection of traditional culture in the language, as well as the issues of diachronic description of the vocabulary of the Russian language and the history of particular words. Based on the analysis of lexical-semantic variants of the word время and meanings of the words with - врем - root in Russian dialects the understanding of time in traditional culture is refined. It is reported that the word время in the traditional sense names not the whole period of human life from birth to death, but only the period of biological maturity associated with the ability to procreate. It is proved that the period of maturity in the people’s culture and language is assessed as a period of prosperity, which becomes the basis for the submission of the norm in human life and a landmark in the awareness of the life space of a person. It is established that the semantics of the Russian word время has accumulated the most ancient etymological meanings of the words год and пора.
<em>The article deals with the basic aspects of the speech genre of condemnation on the materials from journalistic texts (in particular online publications). It is outlined the main features of journalistic style that lead to the need to study the problem of interaction between speakers in terms of social relations, and the research of verbal reactions of society to the actions of some of its representatives who violate certain moral and social norms. The basic language units, typical of this genre, and in particular for texts on political issues are analyzed in the article. The current works of the researchers in the field of the study of speech genres that operate within the discourse of confrontation are analyzed. The relevance of the problem is caused by the necessity of detailed study of evaluative speech genres in Ukrainian press, providing the identification of pragmatic characteristics of speech acts and linguistic units of different levels, which organize the communicative intention of the speaker. The aim of the article is to highlight the basic aspects of the speech genre of condemnation in Ukrainian journalism and the main features of its verbal expression. Linguistic units, typical for expression of condemnation are dominated by those the semantics of which contain negative evaluation and structure-cliches with negative semantics. However, there are cases when the text does not contain any lexical units with semantics of negative evaluation, but there is general condemnation pragmatics.</em>
The Macedonian Recension of the Church Slavonic language from its gradual beginning during the 11th century, especially in the second half, reaches its full development in the 12th century, when there is a consolidation of the basic norms of the Macedonian Church Slavonic literacy. The consolidation of these norms is connected and in continuity with the Old Slavonic Glagolitic period in the work of the Ohrid literary center. The paper presents representative examples that characterize the language of the Macedonian Old Church Slavonic literacy on orthographic, phonological, morpho-syntactic and lexical level.
The aim of this thesis is to contrast the verbal and non-verbal persuasive techniques used in display hoardings in the Czech Republic and Germany. The mode and frequency of the features employed to attract consumers in both countries are compared. The specific goal of the thesis is to find out whether such advertisements can possibly mirror the values and norms of the society. The thesis is divided into three chapters. The first discusses the merits and demerits of the billboard as an advertising mechanism. In this section, the relevant laws of the Czech Republic and Germany are surveyed. In addition, the current situation in both countries with regard to this type of open-air promotion is examined. The second chapter summarises what is known about persuasion in general and the manipulative techniques employed to influence people. In this context, linguistic devices in Czech and German to achieve this end are evaluated. The final chapter provides an analysis of the material gathered for the thesis. The database consists of posters from 30 Czech and 30 German billboards which were photographed in the regions of South Bohemia and Lower Bavaria. The objective was to create a representative sample of contemporary outdoor advertisements in both countries. From the outset, close attention is paid to verbal features. These are divided into phonetic, word-forming, syntactic, and lexical devices. In terms of non-verbal features, typography, colour, and visual style are scrutinized. Since an advertisement is a complex means of communication, the interaction between the visual format and the text is also discussed.
Our study examines the extent to which French immersion students use lax /ɪ/ in the same linguistic context as native speakers of Canadian French. Our results show that the lax variant is vanishingly rare in the speech of immersion students and is used by only a small minority of individuals. This is interpreted as a limitation of French immersion students’ sociolinguistic competence. Within the group of students who do use both variants, we document a positive correlation between female and middle-class students and use of the lax variant and suggest these speakers are generally more sensitive to sociolinguistic variation. A reverse correlation between English cognates and laxing was found. This is taken as evidence that the learning of laxing is lexically mediated.
Berkeley FrameNet is a lexico-semantic resource for English based on the theory of frame semantics. It has been exploited in a range of natural language processing applications and has inspired the development of framenets for many languages. We present a methodological approach to the extraction and generation of a computational multilingual FrameNet-based grammar and lexicon. The approach leverages FrameNet-annotated corpora to automatically extract a set of cross-lingual semantico-syntactic valence patterns. Based on data from Berkeley FrameNet and Swedish FrameNet, the proposed approach has been implemented in Grammatical Framework (GF), a categorial grammar formalism specialized for multilingual grammars. The implementation of the grammar and lexicon is supported by the design of FrameNet, providing a frame semantic abstraction layer, an interlingual semantic application programming interface (API), over the interlingual syntactic API already provided by GF Resource Grammar Library. The evaluation of the acquired grammar and lexicon shows the feasibility of the approach. Additionally, we illustrate how the FrameNet-based grammar and lexicon are exploited in two distinct multilingual controlled natural language applications. The produced resources are available under an open source license.
The article reviews English words and expressions recorded in Word Spy online dictionary of neologisms within the last three decades and conceptualized around the notion of social capital viewed as a civil society attribute and a valuable resource for the sustainable economic development. The meaning of the newly coined words gets the onomasiological coverage within the framework of neology and the social capital theory. The lexical units are analyzed through extra- and intralinguistic motivators of their emergence in the language inventory as well as the formal and semantic composition. The study reveals that the connotatively marked transnominations used to indicate the internal corporate communications outnumber the proper neologisms that refer to new policies and practices developed by a company to operate in its business environment. As the majority of neologisms possess the metaphorical potential, their intensive use in modern business communication results in violating its traditional norms. Thus, English professional discourse tends to experience the loss of its conventionality in favour of increased efficiency of every single communicative act.
Over the past decade Ukrainian terminology decade has taken a significant step forward, due to many factors: 1) many universities introduced special courses on problems of terminology and professional terminology; 2) familiar with the term and professional terminosystem, included in Ukrainian language for professional purposes; 3) issued a lot of educational literature (textbooks, manuals, workshops, etc), which help the students in learning process; 4) published a significant number of terminological dictionaries (explanatory and translated), and the thesaurus; 4) defended a number of dissertations in terminology; 5) There are considered and analyzed the theses of the Ukrainian terminology in different scientific areas, published during the years 2000-2016, revealed the specifics of each work, made general conclusion about the condition of Ukrainian terminology at the time of 2016. The main focus is made on the branch termsystems, peculiarities of their description, and new directions in Ukrainian terminology. conducted regional and international conferences on terminology (Kyiv, Rivne, Lviv). In the first decades of the XXI century, a number of scientific articles in Ukrainian scientific terminology has been published. This article covers abstracts, which describe certain scientific fields, conducted a detailed review of the problems, positions of analysis, and conclusions. As we know, every scientific field has its own terminology which is a specific system, it has its own logical organization, well-known to specialists, and linguistic specificity that needs to be found by linguists. Only since 2000 year till now, there has been defended more than 100 dissertations which describe about 50 scientific branches and give recommendations for improving in them terms and nomens due to their terminological and common language norms. Today in terminology a number of independent branches of research has been singled out: theoretical, applied, historical comparative functional etc. Recently works of cognitive (epistemological) directions, terminological theory of the text, and the like have appeared. A separate branch is the consideration of special vocabulary in the dialects of Ukrainian language. The attention of researchers to terminology has been drawn from different points of view: thematic and lexical-semantic content, structure, origin, formation, development, functioning, systemic organization, paradymatics and syntagmatic, normalization and codification. On the one hand, this is good, because it has been mainly used some analogical schemes of analysis, but from the other hand there is not enough work that would summarize all these studies (each direction and each branch in particular), define the specificity in each of the, language-structure, respect, lexical, word forming systemic etc. Even a brief review of abstracts’ dissertations in terminology permits with certain to speak about a number of still not investigated fields of science. And it gives the opportunity to young scientists to choose in their scientific work the properniche that will allow not only to get a scientific degree, but also to help the Ukrainian science in the future, some approval in it of Ukrainian language.
Abstract English datives show two syntactic patterns, the double object dative (DOD) and the prepositional dative (PD). The alternation between DOD and PD is influenced by three contextual factors: lexical verbs, syntactic weights, and information structures. However, it has been observed that English dative alternation by second language (L2) learners significantly deviates from the native norm. Accordingly, this study examines whether the three factors are influential when L2 learners produce dative sentences, by analyzing a learner corpus and a native speaker corpus. Results show that the learners produced PD significantly more frequently than the native speakers did. Even when DOD should be contextually preferred, the learners produced many PD sentences. These results suggest that L2 learners have trouble noticing the contextual factors when structuring English datives. The finding is further discussed as it relates to the major tenets of L2 acquisition such as cross-linguistic transfer, constructional knowledge, and language processing.
Stabilometry is a technique that aims to study the body sway of human subjects, employing a force platform. The signal obtained from this technique refers to the position of the foot base ground-reaction vector, known as the center of pressure (CoP). The parameters calculated from the signal are used to quantify the displacement of the CoP over time; there is a large variability, both between and within subjects, which prevents the definition of normative values. The intersubject variability is related to differences between subjects in terms of their anthropometry, in conjunction with their muscle activation patterns (biomechanics); and the intrasubject variability can be caused by a learning effect or fatigue. Age and foot placement on the platform are also known to influence variability. Normalization is the main method used to decrease this variability and to bring distributions of adjusted values into alignment. In 1996, O’Malley proposed three normalization techniques to eliminate the effect of age and anthropometric factors from temporal-distance parameters of gait. These techniques were adopted to normalize the stabilometric signal by some authors. This paper proposes a new method of normalization of stabilometric signals to be applied in balance studies. The method was applied to a data set collected in a previous study, and the results of normalized and nonnormalized signals were compared. The results showed that the new method, if used in a well-designed experiment, can eliminate undesirable correlations between the analyzed parameters and the subjects’ characteristics and show only the experimental conditions’ effects.
The research into Bazhov’s tales has revealed frequent occurrence of the dialect and colloquial language, which is a deviation from the literary norm. Such deviations are caused by the author’s intention to preserve the local Ural language, to make the speech of the characters more colorful and authentic, and to make the text more expressive. As a result, a translator meets certain challenges – to follow the style and preserve the specific character of the original tale, to recreate it in a foreign language in such a way that the target text could have an effect on a foreign reader similar to that produced by the source text on a native reader. It is especially difficult to convey the Russian realities of the times of the Old Urals reflected in the author’s socio-cultural comments. Looking for the best solutions to these translation problems, it is necessary to rely on recommendations of reputable translation experts. The authors also give a number of recommendations for translation of such texts. The research has both theoretical and practical orientation. The research object of the article are Pavel Bazhov’s tales, the subject is the analysis of their lexical special features. The relevance of this research can be explained by the scientific interest in folklore and its language as a significant part of any culture and mentality. Bazhov’s tales have been studied by Russian philologists but not so deeply in the aspect of translation into other languages. Bazhov’s collection of tales is a good example of using a living poetic language of the Ural region, with its specific phraseology and local dialect features, to create an authentic atmosphere in literary works. This causes difficulties in translating, which can be overcome by means of pre-translation stylistic analysis of texts. Key words: genre of tales; pre-translation analysis; problem of folk tales translation; national and cultural specific character; lexical special features; translatability.
Cross-cultural communication affects not only the translations per se, but also target culture and thinking in general. Globalization, migration, tourism, student exchanges, international trade and business, and first of all the openness of media brings numerous new concepts and terms into languages. Yet, the direct lexical impact is only part of the process; there is also a broad effect on target language composition/corpus, conventions, norms and even deep structures. Most ‘original’ texts today carry many of the same traits as translations. Interference has long ceased to be characteristic of translated texts only. Translations in many languages constitute more than half of the texts that an average citizen ‘consumes’. We cannot speak anymore of a clear dichotomy of ‘translation language’ versus the real language – there is no isolation in the modern world. One can view this asymmetrical phenomenon as a deplorable interference, as linguistic and cultural imperialism or as a general standardization of languages with a consequent potential loss of cultural uniqueness. Yet it can hardly be affected, as language change is inevitable, and in the modern world translation functions as a major vehicle of change. It also calls for a review of some of the traditional approaches to translation theory issues within the framework of the new globalized, international and multilingual communication.
Technological advancements in combination with significant reductions in price have made it practically feasible to run experiments with multiple eye trackers. This enables new types of experiments with simultaneous recordings of eye movement data from several participants, which is of interest for researchers in, e.g., social and educational psychology. The Lund University Humanities Laboratory recently acquired 25 remote eye trackers, which are connected over a local wireless network. As a first step toward running experiments with this setup, demanding situations with real time sharing of gaze data were investigated in terms of network performance as well as clock and screen synchronization. Results show that data can be shared with a sufficiently low packet loss (0.1 %) and latency (M = 3 ms, M A D = 2 ms) across 8 eye trackers at a rate of 60 Hz. For a similar performance using 24 computers, the send rate needs to be reduced to 20 Hz. To help researchers conduct similar measurements on their own multi-eye-tracker setup, open source software written in Python and PsychoPy are provided. Part of the software contains a minimal working example to help researchers kick-start experiments with two or more eye trackers.
This study introduces the Sentiment Analysis and Cognition Engine (SEANCE), a freely available text analysis tool that is easy to use, works on most operating systems (Windows, Mac, Linux), is housed on a user’s hard drive (as compared to being accessed via an Internet interface), allows for batch processing of text files, includes negation and part-of-speech (POS) features, and reports on thousands of lexical categories and 20 component scores related to sentiment, social cognition, and social order. In the study, we validated SEANCE by investigating whether its indices and related component scores can be used to classify positive and negative reviews in two well-known sentiment analysis test corpora. We contrasted the results of SEANCE with those from Linguistic Inquiry and Word Count (LIWC), a similar tool that is popular in sentiment analysis, but is pay-to-use and does not include negation or POS features. The results demonstrated that both the SEANCE indices and component scores outperformed LIWC on the categorization tasks.
Cet article traite de la mise en place du contenu lexical des manuels d’enseignement de l’arabe, langue étrangère, à l’Institut supérieur des langues de l’université de Damas (ISLUD). Bien que ce contenu ne s’éloigne pas des normes habituelles de création des manuels d’enseignement (centrés sur l’adaptation d’un lexique arabe littéraire classique ou moderne), il se rapproche autant que possible de la pratique quotidienne du langage, très riche en lexique dialectal. C’est donc la problématique de la polyglossie de l’arabe qui est ici abordée. Par l’analyse didactique et linguistique d’un corpus télévisuel syrien (Marāyā 2013), on propose un lexique plus adapté au besoin des apprenants, dans le sens où il ne permettrait pas seulement de comprendre les ressources textuelles de l’arabe, mais aussi de renforcer la capacité de comprendre et d’échanger dans des situations de communication quotidiennes.
Prism adaptation induces rapid recalibration of visuomotor coordination. The neural mechanisms of prism adaptation have come under scrutiny since the observations that the technique can alleviate hemispatial neglect following stroke, and can alter spatial cognition in healthy controls. Relative to non-imaging behavioral studies, fMRI investigations of prism adaptation face several challenges arising from the confined physical environment of the scanner and the supine position of the participants. Any researcher who wishes to administer prism adaptation in an fMRI environment must adjust their procedures enough to enable the experiment to be performed, but not so much that the behavioral task departs too much from true prism adaptation. Furthermore, the specific temporal dynamics of behavioral components of prism adaptation present additional challenges for measuring their neural correlates. We developed a system for measuring the key features of prism adaptation behavior within an fMRI environment. To validate our configuration, we present behavioral (pointing) and head movement data from 11 right-hemisphere lesioned patients and 17 older controls who underwent sham and real prism adaptation in an MRI scanner. Most participants could adapt to prismatic displacement with minimal head movements, and the procedure was well tolerated. We propose recommendations for fMRI studies of prism adaptation based on the design-specific constraints and our results.
In this article we report a computational semantic analysis of the presidential candidates’ speeches in the two major political parties in the USA. In Study One, we modeled the political semantic spaces as a function of party, candidate, and time of election, and findings revealed patterns of differences in the semantic representation of key political concepts and the changing landscapes in which the presidential candidates align or misalign with their parties in terms of the representation and organization of politically central concepts. Our models further showed that the 2016 US presidential nominees had distinct conceptual representations from those of previous election years, and these patterns did not necessarily align with their respective political parties’ average representation of the key political concepts. In Study Two, structural equation modeling demonstrated that reported political engagement among voters differentially predicted reported likelihoods of voting for Clinton versus Trump in the 2016 presidential election. Study Three indicated that Republicans and Democrats showed distinct, systematic word association patterns for the same concepts/terms, which could be reliably distinguished using machine learning methods. These studies suggest that given an individual’s political beliefs, we can make reliable predictions about how they understand words, and given how an individual understands those same words, we can also predict an individual’s political beliefs. Our study provides a bridge between semantic space models and abstract representations of political concepts on the one hand, and the representations of political concepts and citizens’ voting behavior on the other.
The means of verbalization of the concept of CHALLENGE and the functioning of its lexical representatives in the texts of glossy magazines are considered. The conceptual, axiological and figurative components of the concept are studied by the method of conceptual analysis. It is noted that the concept under study has a large number of verbal representatives, characterized by the originality of the semantic and structural peculiarities in texts about fashion. A comprehensive study of the concept of CHALLENGE is necessary in order to ensure the effectiveness of the impact on potential buyers - recipients of “fashionable” product. The analysis of lexical means and ways of presentation of this concept in the fashion discourse creates the basis to explore means of speech manipulation by the consciousness of the target audience to change and correct beliefs and attitudes of its representatives. Axiological component of the studied concept is commented. The author argues that the main structural elements of the conceptual component of the concept of CHALLENGE in texts of glossy magazines are “a challenge to the norms of behavior in society,” “challenge to the fashion,” “challenge to myself.” These data make a presentation about one of the most popular concepts prevalent in the minds of readers of glossy magazines.
We propose a question answering (QA) approach for standardized science exams that both identifies correct answers and produces compelling human-readable justifications for why those answers are correct. Our method first identifies the actual information needed in a question using psycholinguistic concreteness norms, then uses this information need to construct answer justifications by aggregating multiple sentences from different knowledge bases using syntactic and lexical information. We then jointly rank answers and their justifications using a reranking perceptron that treats justification quality as a latent variable. We evaluate our method on 1,000 multiple-choice questions from elementary school science exams, and empirically demonstrate that it performs better than several strong baselines, including neural network approaches. Our best configuration answers 44% of the questions correctly, where the top justifications for 57% of these correct answers contain a compelling human-readable justification that explains the inference required to arrive at the correct answer. We include a detailed characterization of the justification quality for both our method and a strong baseline, and show that information aggregation is key to addressing the information need in complex questions.