Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
An important research field in the area of text mining is text categorization. Most of the real world documents are multi-label in nature. In this paper we have proposed a novel method for automated and effective categorization of multi-label text documents. The proposed method is based on lexical and semantics concepts. Tokens are identified in the text documents using standard IEEE taxonomy. To analyze the semantic relationships between tokens, standard lexical database WordNet is used. The proposed method is tested on a dataset of 150 research articles of computer science domain from IEEE Xplore digital library. It has shown a significantly good performance with an accuracy of 75%.
The linguistic database is also positioned as an actual way of formalizing and organizing phraseological units, terms for designating types of phraseological units. The main principle of systematization of the latter in the study is the thesaurus principle, that is the filling of the paradigm «terminological system – terminological microsystem – terminological subsystem – term», represented by a linguistic database.
We present a complete, automated, and efficient approach for utilizing valency analysis in making dependency parsing decisions. It includes extraction of valency patterns, a probabilistic model for tagging these patterns, and a joint decoding process that explicitly considers the number and types of each token's syntactic dependents. On 53 treebanks representing 41 languages in the Universal Dependencies data, we find that incorporating valency information yields higher precision and F1 scores on the core arguments (subjects and complements) and functional relations (e.g., auxiliaries) that we employ for valency analysis. Precision on core arguments improves from 80.87 to 85.43. We further show that our approach can be applied to an ostensibly different formalism and dataset, Tree Adjoining Grammar as extracted from the Penn Treebank; there, we outperform the previous stateof-the-art labeled attachment score by 0.7. Finally, we explore the potential of extending valency patterns beyond their traditional domain by confirming their helpfulness in improving PP attachment decisions. 1
Polycentric Spanish Norm Towards the Polish‑Spanish Legal Translation The Spanish, being the official language of Spain and many other countries, is characterized by an important dialectal diversity that is reflected in the differences at all linguistic levels: phonetic, morphological, syntactic and lexico‑semantic, etc. All these differences raise controversies and discussions about the existence of a linguistic norm depending on the perspective that can have a monocentric or polycentric character. In this contribution we present some arguments for the second one. To this end, we rely on translations, starting simultaneously from the semasiological and onomasiological perspective, of some Polish‑Spanish legal terms in which it is essential to take into account, the diatopic variation as well as the norm whose character is polycentric.
Standard neural network architectures are non-linear only by virtue of a simple element-wise activation function, making them both brittle and excessively large. In this paper, we consider methods for making the feed-forward layer more flexible while preserving its basic structure. We develop simple drop-in replacements that learn to adapt their parameterization conditional on the input, thereby increasing statistical efficiency significantly. We present an adaptive LSTM that advances the state of the art for the Penn Treebank and WikiText-2 word-modeling tasks while using fewer parameters and converging in less than half as many iterations.
This paper describes our system (HIT-SCIR) submitted to the CoNLL 2018 shared task on Multilingual Parsing from Raw Text to Universal Dependencies. We base our submission on Stanford's winning system for the CoNLL 2017 shared task and make two effective extensions: 1) incorporating deep contextualized word embeddings into both the part of speech tagger and dependency parser; 2) ensembling parsers trained with different initialization. We also explore different ways of concatenating treebanks for further improvements. Experimental results on the development data show the effectiveness of our methods. In the final evaluation, our system was ranked first according to LAS (75.84%) and outperformed the other systems by a large margin.
We present an architecture based on neural networks to generate natural language from unordered dependency trees. The task is split into the two subproblems of word order prediction and morphology inflection. We test our model gold corpus (the Italian portion of the Universal Dependency treebanks) and an automatically parsed corpus from the Web.
We know very little about how neural language models (LM) use prior linguistic context.In this paper, we investigate the role of context in an LSTM LM, through ablation studies.Specifically, we analyze the increase in perplexity when prior context words are shuffled, replaced, or dropped.On two standard datasets, Penn Treebank and WikiText-2, we find that the model is capable of using about 200 tokens of context on average, but sharply distinguishes nearby context (recent 50 tokens) from the distant history.The model is highly sensitive to the order of words within the most recent sentence, but ignores word order in the long-range context (beyond 50 tokens), suggesting the distant past is modeled only as a rough semantic field or topic.We further find that the neural caching model (Grave et al., 2017b) especially helps the LSTM to copy words from within this distant context.Overall, our analysis not only provides a better understanding of how neural LMs use their context, but also sheds light on recent success from cache-based models.
We present the Uppsala system for the CoNLL 2018 Shared Task on universal dependency parsing. Our system is a pipeline consisting of three components: the first performs joint word and sentence segmentation; the second predicts part-ofspeech tags and morphological features; the third predicts dependency trees from words and tags. Instead of training a single parsing model for each treebank, we trained models with multiple treebanks for one language or closely related languages, greatly reducing the number of models. On the official test run, we ranked 7th of 27 teams for the LAS and MLAS metrics. Our system obtained the best scores overall for word segmentation, universal POS tagging, and morphological features.
The significance of tourism within the ASEAN region is recognised by multiple stakeholders. Presenting and promoting a distinctive image of tourism is a common agenda across the ten ASEAN nations. The aim of this study is to document and interpret the emotional connotations of ASEAN tourism slogans. Arguably, such messages provide an initial guide to the appeal and competitive advantages of each individual country. The study is underpinned by considering key ideas on destination positioning and the lexical analysis of emotions. By mining archival resources about word frequencies, synonyms and meanings, the positions of the slogans in an emotion space originally developed by Plutchik were compared and plotted. Joy, admiration and ecstasy were the dominant emotional connotations of most slogans. Thailand and Malaysia have the most distinctive tourism slogans, followed by Vietnam and Laos. Expressions used in the slogans for these four nations overlapped less with other countries across the families of emotion words.
The short note describes the chart parser for multimodal type-logical grammars which has been developed in conjunction with the type-logical treebank for French. The chart parser presents an incomplete but fast implementation of proof search for multimodal type-logical grammars using the "deductive parsing" framework. Proofs found can be transformed to natural deduction proofs.
We propose a novel approach to Vietnamese word segmentation. Our approach is based on the Single Classification Ripple Down Rules methodology (Compton and Jansen, 1990), where rules are stored in an exception structure and new rules are only added to correct segmentation errors given by existing rules. Experimental results on the benchmark Vietnamese treebank show that our approach outperforms previous state-of-the-art approaches JVnSegmenter, vnTokenizer, DongDu and UETsegmenter in terms of both accuracy and performance speed. Our code is open-source and available at: https://github.com/datquocnguyen/RDRsegmenter.
Developed by the CLLD project with support from the Department of Linguistic and Cultural Evolution of the Max Planck Institute for the Science of Human History.
Anger is considered a unique high-arousal and approach-related negative emotion. The influence of individual differences in trait anger on the processing of visual stimuli is relevant to questions about emotional processing and remains to be explored. Using functional magnetic resonance imaging (fMRI), we explored the neural responses to standardized images, selected based on valence and arousal ratings in a group of men with high trait anger compared to those with normative to low anger scores (controls). Results show increased activation in the left-lateralized ventral fronto-parietal attention network to unpleasant images by individuals with high trait anger. There was also a group by arousal interaction in the left thalamus/pulvinar such that individuals with high trait anger had increased pulvinar activation to the high-arousal (versus low arousal) unpleasant images as compared to controls. Thus, individual differences in trait anger in men are associated with brain regions subserving executive attentional and sensory integration during the processing of unpleasant emotional stimuli, particularly to high arousal images.
In the central nervous system the neuropeptide oxytocin mediates a range of behaviors related primarily to emotionality. One factor that influences oxytocinergic communication in the human brain and correlates with emotional behaviors is the single nucleotide polymorphism rs53576 on the oxytocin receptor gene (OXTR). For example, variations in this OXTR genotype are related to parental, altruistic, and other prosocial behaviors. Electroencephalographic waveforms of visually evoked response potentials recorded at the midline parietal electrode site display a prominent component putatively involved with attention allocation called the late positive potential. The magnitude of the late positive potential was found to be significantly higher in homozygous G allele individuals compared with A allele carriers when viewing negative emotionally charged images. Inversely, A allele carriers rated these negative images as more arousing, when measured by the Self-Assessment Manikin rating scale. These data suggest that OXTR functioning contributes to visual processing and subjective experience of negative stimuli.
This paper chooses two news-genre dependency treebanks, one in Chinese and one in English and examines the synergetics or the interrelations among dependency tree widths, heights and sentence lengths in the framework of dependency grammar. When sentences grow longer, the dependency trees grow both taller and higher. The growths of heights and widths are competing with each other, resulting in a minimization of dependency distance. The product of the widest layer and its width corresponds to the sentence length. These correlations are integrated into a tentative synergetic syntactic model based on the framework of dependency grammar.
Ambivalence is a common experience that permeates a broad range of research. Unfortunately, quantifying ambivalence has proven a daunting task, with researchers limited to studying vacillating ambivalence, VA (i.e., temporal oscillations between favor/disfavor evaluations of an attitude object). Here, we demonstrate the use of the density matrix to measure both VA and what we term “simultaneous ambivalence” (SA): ambivalence that manifests itself as “in the moment” concurrent favor/disfavor evaluations. In a methodological study we gave participants the option of either single-responding or double-responding to questionnaire items regarding a controversial topic (i.e., affirmative action). Since standard statistical procedures provide no means for analyzing double responses, such data are routinely treated as “bad.” As demonstrated here, the density matrix provides an unambiguous and relatively easy means of accounting for double responses, which is our indicator of SA. Our data are well explained by a mixture model, with participants divided into two nearly equal groups of SA and non-SA participants, and provide evidence that the general phenomenon of SA transcends differences of gender and ethnicity. Further, the density matrix data are consistent with viewing SA and VA as distinct ambivalence constructs.
Up to now, the potential of eye tracking in science as well as in everyday life has not been fully realized because of the high acquisition cost of trackers. Recently, manufacturers have introduced low-cost devices, preparing the way for wider use of this underutilized technology. As soon as scientists show independently of the manufacturers that low-cost devices are accurate enough for application and research, the real advent of eye trackers will have arrived. To facilitate this development, we propose a simple approach for comparing two eye trackers by adopting a method that psychologists have been practicing in diagnostics for decades: correlating constructs to show reliability and validity. In a laboratory study, we ran the newer, low-cost EyeTribe eye tracker and an established SensoMotoric Instruments eye tracker at the same time, positioning one above the other. This design allowed us to directly correlate the eye-tracking metrics of the two devices over time. The experiment was embedded in a research project on memory where 26 participants viewed pictures or words and had to make cognitive judgments afterwards. The outputs of both trackers, that is, the pupil size and point of regard, were highly correlated, as estimated in a mixed effects model. Furthermore, calibration quality explained a substantial amount of individual differences for gaze, but not pupil size. Since data quality is not compromised, we conclude that low-cost eye trackers, in many cases, may be reliable alternatives to established devices.
Crowdsourcing services, such as MTurk, have opened a large pool of participants to researchers. Unfortunately, it can be difficult to confidently acquire a sample that matches a given demographic, psychographic, or behavioral dimension. This problem exists because little information is known about individual participants and because some participants are motivated to misrepresent their identity with the goal of financial reward. Despite the fact that online workers do not typically display a greater than average level of dishonesty, when researchers overtly request that only a certain population take part in an online study, a nontrivial portion misrepresent their identity. In this study, a proposed system is tested that researchers can use to quickly, fairly, and easily screen participants on any dimension. In contrast to an overt request, the reported system results in significantly fewer (near zero) instances of participant misrepresentation. Tests for misrepresentations were conducted by using a large database of past participant records (~45,000 unique workers). This research presents and tests an important tool for the increasingly prevalent practice of online data collection.
This paper focuses on a novel methodology of subjective speech quality measurement and repeatability of its results between laboratory conditions and simulated environmental conditions. A single set of speech samples was distorted by various background noises and low bit-rate coding techniques. This study aimed to compare results of subjective speech quality tests with and without a parallel task deploying the ITU-T P.835 methodology. Afterward, tests results performed with and without a parallel task were compared using Pearson correlation, CI95, and numbers of opposite pair-wise comparisons. The tests show differences in results in the case of a parallel task. [ABSTRACT FROM AUTHOR], Copyright of PLoS ONE is the property of Public Library of Science and its content may not be copied or emailed to multiple sites or posted to a listserv without the copyright holder's express written permission. However, users may print, download, or email articles for individual use. This abstract may)
Introduction: Natural resource management uses expert judgement to estimate facts that inform important decisions. Unfortunately, expert judgement is often derived by informal and largely untested protocols, despite evidence that the quality of judgements can be improved with structured approaches. We attribute the lack of uptake of structured protocols to the dearth of illustrative examples that demonstrate how they can be applied within pressing time and resource constraints, while also improving judgements. Aims and methods: In this paper, we demonstrate how the IDEA protocol for structured expert elicitation may be deployed to overcome operational challenges while improving the quality of judgements. The protocol was applied to the estimation of 14 future abiotic and biotic events on the Great Barrier Reef, Australia. Seventy-six participants with varying levels of expertise related to the Great Barrier Reef were recruited and allocated randomly to eight groups. Each participant)
English in Singapore has always presented a balancing act for its founders. The colonial era saw a distinct role for English, i.e. to produce English-speaking officers for the British administration, while modern Singapore sees English being used as both a national and international lingua franca and as a major language that connects the island city-state to the world. ‘English-knowing bilingualism’ has gained ascendancy in Singapore and may become a core competency for the 21st-century world with the rise in status of English as a global language. However, the path to English-knowing bilingualism in the pluri-lingual and heterogeneous country was often marked by paradoxical debates surrounding the issues of language maintenance and shift, identity and the transmission of values, equity and meritocracy, as well as balancing between local versus global linguistic norms and standards. This paper focuses on the continuing debates, from the past to the present, as new challenges arise and argues how a new balance has to be achieved in the language strategy, policy and management for future-readiness in Singapore.
Over the past century, personality theory and research has successfully identified core sets of characteristics that consistently describe and explain fundamental differences in the way people think, feel and behave. Such characteristics were derived through theory, dictionary analyses, and survey research using explicit self-reports. The availability of social media data spanning millions of users now makes it possible to automatically derive characteristics from behavioral data—language use—at large scale. Taking advantage of linguistic information available through Facebook, we study the process of inferring a new set of potential human traits based on unprompted language use. We subject these new traits to a comprehensive set of evaluations and compare them with a popular five factor model of personality. We find that our language-based trait construct is often more generalizable in that it often predicts non-questionnaire-based outcomes better than questionnaire-based traits (e.g)
Scholarly studies and common accounts of national politics enjoy pointing out the resilience of ideological divides among populations. Building on the image of political cleavages and geographic polarization, the regionalization of politics has become a truism across Northern democracies. Left unquestioned, this geography plays a central role in shaping electoral and referendum campaigns. In Europe and North America, observers identify recurring patterns dividing local populations during national votes. While much research describes those patterns in relation to ethnicity, religious affiliation, historic legacy and party affiliation, current approaches in political research lack the capacity to measure their evolution over time or other vote subsets. This article introduces “Dyadic Agreement Modeling” (DyAM), a transdisciplinary method to assess the evolution of geographic cleavages in vote outcomes by implementing a metric of agreement/disagreement through Network Analysis. Unlike ex)
The Newest Vital Sign (NVS) is a simple, quick and accurate screening test for health literacy (HL). It has been validated for different languages but, to date, not for the Croatian language. The aim of this study was to develop a linguistically validated Croatian version of the NVS and to use it at a later stage in a pilot study of health literacy assessment of hospital patients in Croatia. A full linguistic validation procedure was applied, including forward and backward translation, expert panel review, cognitive interview with 10 respondents from general population, and full involvement in the procedure of one of the screening test developers, the lead author of the NVS-UK version. HL testing on 100 hospital patients (55% women, median age 63.5 years) revealed 58% of patients had less than adequate HL level (scores less than 4), and mean NVS total score was 3.34. A positive significant association was observed between HL and educational level (p = 0.002). A high percentage of pati)
Introduction: Oral Anticoagulation therapy (OAC) is highly effective in the management of thromboembolic disorders. An adequate level of knowledge is important for self-management and optimizing clinical outcomes. The Anticoagulation Knowledge Tool (AKT) was developed to assess OAC knowledge and caters for both patients prescribed direct oral anticoagulants or vitamin K antagonist (VKA). However, evidence regarding its psychometric proprieties, validity and reliability are unavailable in non-English speaking settings. For this reason, the aim of this study is to provide further evidence of validity for AKT and also developing an Italian AKT version (I-AKT) supported by evidence of validity and reliability. Methods: A multiphase study was conducted which included the following: cultural and linguistic validity; i.e. content validity; construct validity; reliability assessment. The Construct validity was performed using the contrasted group approach using three groups comprised of hea)
Background: In practical research, it was found that most people made health-related decisions not based on numerical data but on perceptions. Examples include the perceptions and their corresponding linguistic values of health risks such as, smoking, syringe sharing, eating energy-dense food, drinking sugar-sweetened beverages etc. For the sake of understanding the mechanisms that affect the implementations of health-related interventions, we employ fuzzy variables to quantify linguistic variable in healthcare modeling where we employ an integrated system dynamics and agent-based model. Methodology: In a nonlinear causal-driven simulation environment driven by feedback loops, we mathematically demonstrate how interventions at an aggregate level affect the dynamics of linguistic variables that are captured by fuzzy agents and how interactions among fuzzy agents, at the same time, affect the formation of different clusters(groups) that are targeted by specific interventions. Results:)
Word sense disambiguation (WSD) is the process of identifying an appropriate sense for an ambiguous word. With the complexity of human languages in which a single word could yield different meanings, WSD has been utilized by several domains of interests such as search engines and machine translations. The literature shows a vast number of techniques used for the process of WSD. Recently, researchers have focused on the use of meta-heuristic approaches to identify the best solutions that reflect the best sense. However, the application of meta-heuristic approaches remains limited and thus requires the efficient exploration and exploitation of the problem space. Hence, the current study aims to propose a hybrid meta-heuristic method that consists of particle swarm optimization (PSO) and simulated annealing to find the global best meaning of a given text. Different semantic measures have been utilized in this model as objective functions for the proposed hybrid PSO. These measures consis)
Comprehending natural language quantifiers (like many, all, or some) involves linguistic and numerical abilities. However, the extent to which both factors play a role is controversial. In order to determine the specific contributions of linguistic and number skills in quantifier comprehension, we examined two groups of participants that differ in their language abilities while their number skills appear to be similar: Participants with Down syndrome (DS) and participants with Williams syndrome (WS). Compared to rather poor linguistic skills of individuals with DS, individuals with WS display relatively advanced language abilities. Participants with WS also outperformed participants with DS in a quantifier comprehension task while number knowledge did not differ between the two groups. When compared to typically developing (TD) children of the same mental age, participants with WS displayed similar levels regarding quantifier abilities, but participants with DS performed worse than th)
Most language users agree that some words sound harsh (e.g. grotesque) whereas others sound soft and pleasing (e.g. lagoon). While this prominent feature of human language has always been creatively deployed in art and poetry, it is still largely unknown whether the sound of a word in itself makes any contribution to the word’s meaning as perceived and interpreted by the listener. In a large-scale lexicon analysis, we focused on the affective substrates of words’ meaning (i.e. affective meaning) and words’ sound (i.e. affective sound); both being measured on a two-dimensional space of valence (ranging from pleasant to unpleasant) and arousal (ranging from calm to excited). We tested the hypothesis that the sound of a word possesses affective iconic characteristics that can implicitly influence listeners when evaluating the affective meaning of that word. The results show that a significant portion of the variance in affective meaning ratings of printed words depends on a number of spe)
Application of a phonological rule is often conditioned by prosodic structure, which may create a potential perceptual ambiguity, calling for phonological inferencing. Three eye-tracking experiments were conducted to examine how spoken word recognition may be modulated by the interaction between the prosodically-conditioned rule application and phonological inferencing. The rule examined was post-obstruent tensing (POT) in Korean, which changes a lax consonant into a tense after an obstruent only within a prosodic domain of Accentual Phrase (AP). Results of Experiments 1 and 2 revealed that, upon hearing a derived tense form, listeners indeed recovered its underlying (lax) form. The phonological inferencing effect, however, was observed only in the absence of its tense competitor which was acoustically matched with the auditory input. In Experiment 3, a prosodic cue to an AP boundary (which blocks POT) was created before the target using an F0 cue alone (i.e., without any temporal cue)
Existing approaches to describe social interactions consider emotional states or use ad-hoc descriptors for microanalysis of interactions. Such descriptors are different in each context thereby limiting comparisons, and can also mix facets of meaning such as emotional states, short term tactics and long-term goals. To develop a systematic set of concepts for second-by-second social interactions, we suggest a complementary approach based on practices employed in theater. Theater uses the concept of dramatic action, the effort that one makes to change the psychological state of another. Unlike states (e.g. emotions), dramatic actions aim to change states; unlike long-term goals or motivations, dramatic actions can last seconds. We defined a set of 22 basic dramatic action verbs using a lexical approach, such as ‘to threaten’–the effort to incite fear, and ‘to encourage’–the effort to inspire hope or confidence. We developed a set of visual cartoon stimuli for these basic dramatic action)
Sequence-to-sequence constituency parsing casts the tree structured prediction problem as a general sequential problem by top-down tree linearization,and thus it is very easy to train in parallel with distributed facilities. Despite its success, it relies on a probabilistic attention mechanism for a general purpose, which can not guarantee the selected context to be informative in the specific parsing scenario. Previous work introduced a deterministic attention to select the informative context for sequence-to-sequence parsing, but it is based on the bottom-up linearization even if it was observed that top-down linearization is better than bottom-up linearization for standard sequence-to-sequence constituency parsing. In this paper, we thereby extend the deterministic attention to directly conduct on the top-down tree linearization. Intensive experiments show that our parser delivers substantial improvements over the bottom-up linearization in accuracy, and it achieves 92.3 Fscore on the Penn English Treebank section 23 and 85.4 Fscore on the Penn Chinese Treebank test dataset, without reranking or semi-supervised training.
Sarcasm is a sophisticated form of sentiment expression where speaker express their opinions opposite of what they mean. Sarcasm detection and Emotion detection from social net-working sites has been a great field of study. With the growth of e-services such as e-commerce, e-tourism and e-business, the companies are very keen on exploiting emotion and sarcasm analysis for their marketing strategies in order to evaluate the public attitudes towards their brand. Thus efficient emotion and sarcasm modeling system can be a good solution to the above problem. This work aims at developing a system that groups posts based on emotions, sentiment and find sarcastic posts, if present. The proposed system is to develop a prototype that help to come to an inference about the emotions of the posts namely anger, surprise, happy, fear, sorrow, trust, anticipation and disgust with three sentic levels in each. This helps in better understanding of the posts when compared to the approaches which senses the polarity of the posts and gives just their sentiments i.e., positive, negative or neutral. The posts handling these emotions might be sarcastic too. The Sentiment & emotion identification module identifies the sentiment or emotion of the post by evaluating score of each word in the comment which is used by different sarcasm detection methods to detect sarcasm. The emotion identification module uses the lexical databases WordNet, SentiWordNet to find the right sentiment scores for the words with respect to each emotion. It also uses Sarcasm detection algorithms like Emoticon sarcasm detection, Hybrid sarcasm detection, Hashtag Processing, Interjection Word Start (IWT).
Despite increased use of behavioral analogues to identify casual mechanisms of self-injurious behavior (e.g., suicide attempts; non-suicidal self-injury), little is known about the impact on participants. The current study examined the impact of a specific behavior analogue, Self-Aggressive Paradigm (SAP), on participant affect. Community participants (n = 507) reported several affective ratings before and after completing SAP task procedures. Following the SAP, participants reported reductions in nervousness and fear and increases in calmness and anger (d =.21). Participants with a current anxiety disorder reported greater increases in happiness; those with a suicide attempt history reported greater increases in sadness. Findings demonstrate the SAP has no adverse mood effects, supporting its use in experimental research.
For nearly 50 years, psychologists have studied prospective memory, or the ability to execute delayed intentions. Yet, there remains a gap in understanding as to whether initial encoding of the intention must be elaborative and strategic, or whether some components of successful encoding can occur in a perfunctory, transient manner. In eight studies (N = 680), we instructed participants to remember to press the Q key if they saw words representing fruits (cue) during an ongoing lexical decision task. They then typed what they were thinking and responded whether they encoded fruits as a general category, as specific exemplars, or hardly thought about it at all. Consistent with the perfunctory view, participants often reported mind wandering (42.9%) and hardly thinking about the prospective memory task (22.5%). Even though participants were given a general category cue, many participants generated specific category exemplars (34.5%). Bayesian analyses of encoding durations indicated tha)
At the interface between scene perception and speech production, we investigated how rapidly action scenes can activate semantic and lexical information. Experiment 1 examined how complex action-scene primes, presented for 150 ms, 100 ms, or 50 ms and subsequently masked, influenced the speed with which immediately following action-picture targets are named. Prime and target actions were either identical, showed the same action with different actors and environments, or were unrelated. Relative to unrelated primes, identical and same-action primes facilitated naming the target action, even when presented for 50 ms. In Experiment 2, neutral primes assessed the direction of effects. Identical and same-action scenes induced facilitation but unrelated actions induced interference. In Experiment 3, written verbs were used as targets for naming, preceded by action primes. When target verbs denoted the prime action, clear facilitation was obtained. In contrast, interference was observed when)
This paper analyzes distributional properties that facilitate the categorization of words into lexical categories. First, word-context co-occurrence counts were collected using corpora of transcribed English child-directed speech. Then, an unsupervised k-nearest neighbor algorithm was used to categorize words into lexical categories. The categorization outcome was regressed over three main distributional predictors computed for each word, including frequency, contextual diversity, and average conditional probability given all the co-occurring contexts. Results show that both contextual diversity and frequency have a positive effect while the average conditional probability has a negative effect. This indicates that words are easier to categorize in the face of uncertainty: categorization works best for words which are frequent, diverse, and hard to predict given the co-occurring contexts. This shows how, in order for the learner to see an opportunity to form a category, there needs to)
Web genre detection is a task that can enhance information retrieval systems by providing rich descriptions of documents and enabling more specialized queries. Most of previous studies in this field adopt the closed-set scenario where a given palette comprises all available genre labels. However this is not a realistic setup since web genres are constantly enriched with new labels and existing web genres are evolving in time. Open-set classification, where some pages used in the evaluation phase do not belong to any of the known genres, is a more realistic setup for this task. In this case, all pages not belonging to known genres can be seen as noise. This paper focuses on systematic evaluation of open-set web genre identification when the noise is either structured or unstructured. Two open-set methods combined with alternative text representation schemes and similarity measures are tested based on two benchmark corpora. Moreover, we adopt the openness test for web genre identification that enables the observation of effectiveness for a varying number of known/unknown labels.
Lexical resources are fundamental to tackle many tasks that are central to present and prospective research in Text Mining, Information Retrieval, and connected to Natural Language Processing. In this article we introduce COVER, a novel lexical resource, along with COVERAGE, the algorithm devised to build it. In order to describe concepts, COVER proposes a compact vectorial representation that combines the lexicographic precision characterizing BabelNet and the rich common-sense knowledge featuring ConceptNet. We propose COVER as a reliable and mature resource, that has been employed in as diverse tasks as conceptual categorization, keywords extraction, and conceptual similarity. The experimental assessment is performed on the last task: we report and discuss the obtained results, pointing out future improvements. We conclude that COVER can be directly exploited to build applications, and coupled with existing resources, as well.
Larysa Kolibaba PhD in Philology, Senior Research Scientist of the Department of Grammar and Scientific Terminology, Institute of the Ukrainian Language of National Academy of Sciences of Ukraine 4 Hrushevskyi St., Kyiv 01001, Ukraine Е-mail: kolibaba.lm@meta.ua Heading: Researches Language: Ukrainian Abstract: In this article the problem of fixing of morphological forms of nouns in the Ukrainian dictionaries of different time and its […]
Different approaches to the interpretation of the concept “error in speech in a non-native language” are analyzed: psychological, psycholinguistic, linguistic, methodical. The possibility to use the results of the analysis of errors in speech in a non-native language to study the processes of learning the language is proved. It is emphasized that the conclusions should be taken into account when developing a complex of productive and receptive lexical exercises. The conclusion about the impropriety of dividing errors in foreign speech to “communicatively significant and insignificant” is made. The most frequent lexical errors in the Russian speech of Greek students and deviations from the lexical norms of the Russian language are revealed. The features of the influence of errors on the nature of communication are described. The article deals with the types of lexical errors that Greek students make when mastering the system of the Russian language. The classification of lexical errors in the Russian speech of Greek students is made on the basis of two criteria: the reasons for the deviation from the lexical norm and the consequences of the deviation. Recommendations for the prevention of lexical level errors when teaching the Greeks the Russian language are given.
In the modern world it can be easy to encounter advertising slogans that are supposed to grab the attention of a potential recipient as well as affect their imagination and sensitivity. This communication tends to convince the client that buying the advertised product would make them feel exceptional. This article focuses on linguistic gimmicks that are used by marketing experts in order to convince the client to a certain product. The analysis in based on the cosmetics sector. The lexical database consists of websites, television commercials and Internet advertisements.
Implicit discourse relation recognition is a challenging task as the relation prediction without explicit connectives in discourse parsing needs understanding of text spans and cannot be easily derived from surface features from the input sentence pairs. Thus, properly representing the text is very crucial to this task. In this paper, we propose a model augmented with different grained text representations, including character, subword, word, sentence, and sentence pair levels. The proposed deeper model is evaluated on the benchmark treebank and achieves state-of-the-art accuracy with greater than 48% in 11-way and F1 score greater than 50% in 4-way classifications for the first time according to our best knowledge.
Abstract This paper illustrates how enriched diachronic treebank data can shed new light on an old and vexed topic, even when that topic is primarily morphological and semantic in nature rather than syntactic. The topic is the rise of the Russian po delimitatives, a change seen as crucial in most accounts of the history of Russian aspect, since it represents a major step in generalising the derivational aspect system. Earlier accounts concur that the po delimitatives spread fairly recently, too recently for the development to be connected to the loss of the aorist tense, which also had delimitative readings with atelic verbs. Using treebank data from the Tromsø Old Russian and Old Church Slavonic Treebank, enriched with tags for derivational morphology and semantics, I show that the po delimitatives were not marginal even in the earliest Slavic sources, either in terms of frequency or semantics, and that they first complemented and then competed with the delimitative aorists. It can thus be claimed that the exotic po delimitatives grew organically out of the old Indo-European inflectional aspect system.