Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Existing research shows that “pleasant” or “unpleasant” moods can be primed by presenting participants with “pleasant” or “unpleasant” images (Avero & Calvo, 2006), and that stronger priming effects are induced by images as opposed to text (Powell et al., 2015). However, no previous research shows whether or not mood induction effects may differ based on image presentation format. Therefore, the present work aimed to test this hypothesis, by presenting participants (N = 145) with either standalone or grouped images, displaying either positive or negative facial expressions. We found that both facial expression and image presentation had a significant effect on participants’ average ratings of the emotional valence of the images, including a significant interaction effect. However, only facial expression had a significant effect on mood change. We found a slight correlation (r =.298) between image rating and mood change, suggesting that image presentation may have a slight effect on mood change that was unable to be observed in this small-scale study.
Background: Obesity is a chronic disease characterized by a reduction in life expectancy. Bariatric surgery has been an alternative to conventional treatments, but lifestyle changes such as increased physical activity are crucial for achieving body image and health outcomes. Objective: Test the hypothesis that physical activity is associated with satisfaction with body image in obese individuals undergoing bariatric surgery. Methods: Cross-sectional study conducted in Salvador, BA, Brazil. All participants at 3 and 6 months postoperatively answered the Stunkard Image Rating Scales as the main outcome, in addition to the International Physical Activity Questionnaire to assess the physical activity profile. Results: Physically active individuals were 80% more likely to report body satisfaction when compared with physically inactive individuals. According to multivariate analysis, the adjusted physical activity for females and advanced age increased by 83% the chance of reporting satisfaction with body image and, when adjusted for marital status, body mass index, and surgery time, the strength of association increased to 89% the chance to refer body satisfaction when compared with physically inactive people. Conclusion: Physical activity was associated with better body image. These results indicate opportunities to improve outcomes in patients after bariatric surgery through counseling and treatment intervention.
Abstract Kankana-ey is a widely used dialect in the northern region of the Philippines. Unfortunately, there are documented studies on the syntactic rules of this dialect. This study explored the development of a corpus for the Kankana-ey dialect. Further, the corpus was then used to establish the syntactic rules of Kankana-ey. A Kankana-ey version of the bible, dictionaries, news articles, songs and various online resources were used to collect words for the corpus of the Kankana-ey dialect. These identified words were also tagged using the parts of speech tags of the Penn TreeBank. Using the corpus and TensorFlow, 320 Kankana-ey sentences were analysed to determine the syntactic rules. In addition, 80 sentences were used to test the accuracy of the identified rules. At the end of the study, the created corpus has 3,412 tagged Kankana-ey words, while the analysis of the syntactic rules resulted to 1,722 rules. Testing also showed a 60% accuracy of the syntactic rules. In conclusion, the high number of identified rules from the 320 sentences was due to multiple Kankana-ey words having different possible tags. This also resulted to the low accuracy of the syntactic rules.
This article discusses one of the forms of machine translation, the Instagram translation feature called “see translation”. The research is focused on the translation techniques applied by the machine in translating Banyumas batik motifs from Indonesian to English found in @batikantodjamil and @batk_rd. This topic is worth discussing since machine translation is now getting more developed and is projected to replace human translator. However, in some cases, for example in dealing with culturally-bound terms, machine translation cannot perform contextual knowledge as well as the human translator. this mini research was conducted by applying qualitative research with purposive sampling technique in which the researchers obtain the data by selecting two batik center Instagram accounts containing batik motif names in the captions. The result shows that there are three translation techniques applied by the Instagram translation features, namely literal, borrowing, and particularization. The most dominant technique to use is borrowing technique, and it shows a tendency that such cultural terms in the source language do not have one-to-one correspondence in the target language. In other words, the touch of human translator is very important in the post-editing process of translation by machine to make the translation more acceptable. However, if it is impossible to involve human translator, the Instagram administrator should enrich the machine with more contextual linguistic database to provide the users with better translation results.
How do each of us come to view the world uniquely? An emerging theory of microvalence proposes that subtle feelings of reward and punishment derived from individualized experiences with basic everyday objects help determine how we later attend and behave towards them. These objects that are part of our more mundane experiences are thought to be given attentional priority similar to objects that evoke stronger emotional responses. However, this relationship between preferences guided by daily experience and attention has not been tested. I introduced a novel paradigm to induce microvalences by simulating real life experience paired with an interocular suppression technique (bCFS) to explore its role in attention. Consistent with the theory of microvalence, affective ratings indicated that our novel shapes possessed pre-existing affective properties by which they are evaluated, giving rise to preferences. Unexpectedly, we observed a unifying effect of experience, blurring perceived differences between novel shapes, thus collapsing initial preferences (feelings of like or dislike). Results showed, however, that microvalences were not prioritized in attention. Our findings place emphasis on the role of experience in shifting automatic preferences to create unbiased representations of the world.
Literary works are classified as works of imagination in the form of fictional or imaginary experiences. The messages to be conveyed through literary works must be creative so that they appear attractive to read and listen to, so it is necessary to have a stile from the authorship itself to make his work beautiful and attractive. There are many ways to enjoy, understand and appreciate the work of the author Tenas Effendy, one of which is by studying the Stile of Tenas Effendy's authorship in Tunjuk Ajar Melayu. This study aims to analyze and interpret Tenas Effendy's stile authorship in Tunjuk Ajar Melayu. This needs to be examined because the existence of a literary work can be seen from how the author packs his work so that he can create his own stile from the author's side. The method used in this research is content analysis method. The source of data in this research is Tunjuk Ajar Melayu Karya Tenas Effendy in 2013 which has been recorded. The data collection technique is done by applying the hermeneutic technique. After the author of the analysis, Tunjuk Ajar Melayu by Tenas Effendy has a unique and distinctive authorship stile seen from the stylistic aspect, namely the stile as a pack of thoughts, the stile as a deviation from linguistic norms and the stile as a collection of personal characteristics.
The article highlights monitoring and dynamics of language development in post-soviet Kyrgyzstan. The crucial politico-social changes after the collapse of the USSR affected greatly the language situation of the country. Hence, this article’s aim is to give some sociolinguistic analysis to the following issues: • To monitor country’s up-to date language situation • To speak on the new language policy of Kyrgyzstan in modern stage • To encounter new socio-economic and political challenges in the process of language lawmaking • To define the interrelationships between dominant languages and vernaculars of minority communities This article also explains the reasons of granting Russian language the status of Lingua Franca, immediately after Kyrgyzstan’s becoming an independent state. One of the acute problems today is corpus planning, which means codification of newly coined or borrowed words. The flow of new terminology from the other languages, especially “Americanisms” and “Englishisisms” which replaced so-called “Sovietisms”, needs to be standardized according to the linguistic norms of state (Kyrgyz) language. The article also reveals “hierarchical” disposition of main languages due to their functional load in 20 domains of Kyrgyzstan. For the last 28 years’ tremendous changes have happened in language space of Kyrgyzstan, directly touching upon the positions of Kyrgyz, Russian, Uzbek, English, Turkish and other languages. Therefore, language system of Kyrgyzstan nowadays presents complex, intertwined interrelations between all nationalities residing in Kyrgyzstan and language cooperation among them. It happened because of some socio-political reasons: 1) The flow of Russian speaking population from industrial areas to the Russian Federation. 2) Changing demographic situation on the country. 3) Inner immigration process when many unemployed people from the distant regions came to Bishkek and Chui valley to find job possibilities.
We analyse and explain the increased generalisation performance of iterate averaging using a Gaussian process perturbation model between the true and batch risk surface on the high dimensional quadratic. We derive three phenomena \latestEdits{from our theoretical results:} (1) The importance of combining iterate averaging (IA) with large learning rates and regularisation for improved regularisation. (2) Justification for less frequent averaging. (3) That we expect adaptive gradient methods to work equally well, or better, with iterate averaging than their non-adaptive counterparts. Inspired by these results\latestEdits{, together with} empirical investigations of the importance of appropriate regularisation for the solution diversity of the iterates, we propose two adaptive algorithms with iterate averaging. These give significantly better results compared to stochastic gradient descent (SGD), require less tuning and do not require early stopping or validation set monitoring. We showcase the efficacy of our approach on the CIFAR-10/100, ImageNet and Penn Treebank datasets on a variety of modern and classical network architectures.
'Synonym' is an imperative instrument of commonsense knowledge that we apply to make a good sense and sound judgement of our reading. To investigate the ability of machine comprehension models in handling the synonym commonsense knowledge, we developed an innovative approach to automatically generate a dataset based on the Stanford Question Answering Dataset (SQuAD 2.0). The brand-new dataset consists of additional distracting sentences or questions spawned using synonym commonsense knowledge. We formulated new questions by replacing noun entities of the original ones in SQuAD 2.0 with their synonyms. This approach followed the two fundamental principles of SQuAD 2.0 dataset: relevancy and plausibility (incorrect answers are more challenging if they are relevant and plausible). It improves the robustness/abstraction of the question set. To improve the synonym selection strategy in Word Sense Disambiguation (WSD) problem, we designed a new algorithm Multiple Source Adapted Lesk Algorithm (MSALA). Rather than only using WordNet as the source of gloss for adapted Lesk algorithm, we used both lexical database WordNet and commonsense database ConceptNet. This fusion provides a rich hierarchy of semantic relations for the MSALA algorithm. Using this method, we devised 11,000 questions and evaluated the performance of the state-of-the-art question answering system-BERT. Our result shows that the accuracy of the contemporary BERT-Base model dropped from 74.98% to 63.24%. This 10+% accuracy drop revealed the limitations of BERT in handling synonym commonsense knowledge.
Abstract In this paper, A shorter version of the paper appeared in German in the final report of the Digital Plato project which was funded by the Volkswagen Foundation from 2016 to 2019. [35], [28]. we present a method for paraphrase extraction in Ancient Greek that can be applied to huge text corpora in interactive humanities applications. Since lexical databases and POS tagging are either unavailable or do not achieve sufficient accuracy for ancient languages, our approach is based on pure word embeddings and the word mover’s distance (WMD) [20]. We show how to adapt the WMD approach to paraphrase searching such that the expensive WMD computation has to be computed for a small fraction of the text segments contained in the corpus, only. Formally, the time complexity will be reduced from <m:math xmlns:m="http://www.w3.org/1998/Math/MathML"><m:mi>O</m:mi><m:mo>(</m:mo><m:mi>N</m:mi><m:mo>·</m:mo><m:msup><m:mrow><m:mi>K</m:mi></m:mrow><m:mrow><m:mn>3</m:mn></m:mrow></m:msup><m:mo>·</m:mo><m:mo>log</m:mo><m:mi>K</m:mi><m:mo>)</m:mo></m:math> \mathcal{O}(N\cdot {K^{3}}\cdot \log K) to <m:math xmlns:m="http://www.w3.org/1998/Math/MathML"><m:mi>O</m:mi><m:mo>(</m:mo><m:mi>N</m:mi><m:mo>+</m:mo><m:msup><m:mrow><m:mi>K</m:mi></m:mrow><m:mrow><m:mn>3</m:mn></m:mrow></m:msup><m:mo>·</m:mo><m:mo>log</m:mo><m:mi>K</m:mi><m:mo>)</m:mo></m:math> \mathcal{O}(N+{K^{3}}\cdot \log K), compared to the brute-force approach which computes the WMD between each text segment of the corpus and the search query. N is the length of the corpus and K the size of its vocabulary. The method, which searches not only for paraphrases of the same length as the search query but also for paraphrases of varying lengths, was evaluated on the Thesaurus Linguae Graecae ® (TLG ® ) [25]. The TLG consists of about <m:math xmlns:m="http://www.w3.org/1998/Math/MathML"><m:mn>75</m:mn><m:mo>·</m:mo><m:msup><m:mrow><m:mn>10</m:mn></m:mrow><m:mrow><m:mn>6</m:mn></m:mrow></m:msup></m:math> 75\cdot {10^{6}} Greek words. We searched the whole TLG for paraphrases for given passages of Plato. The experimental results show that our method and the brute-force approach, with only very few exceptions, propose the same text passages in the TLG as possible paraphrases. The computation times of our method are in a range that allows its application in interactive systems and let the humanities scholars work productively and smoothly.
With the popularity of the Internet, the public can get news from the recent hottest news, and express their opinion in time in the network of social media, such as microblog, twitter. With the help of sentiment analysis of the comments, the government or the media can inform of the public opinion, and make corresponding decisions, so that they are able to get the positive feedback. Also, the sentiment analysis can help people block the detrimental comments. Since the sentiment analysis is useful in the daily life, the author made an experiment about the sentiment analysis of three models, namely, Naive Byes, Maximum Entropy, and SVM, to compare the results' accuracy of them. And the dataset used in this experiment is Stanford Twitter Sentiment (STS). Besides, the reference is made to Stanford Sentiment Treebank and IMDB. In addition, the determination of the emotional tone is based on emotional dictionaries like GI (General Inquirer) and How Net. By comparing the accuracy and training time of different models, SVM is selected to be the optimal model with the F-scores of 82%.
This paper presents the first treebank for the Laz language, which is also the first Universal Dependencies Treebank for a South Caucasian language. This treebank aims to create a syntactically and morphologically annotated resource for further research. We also aim to document an endangered language in a systematic fashion within an inherently cross-linguistic framework: the Universal Dependencies Project (UD). As of now, our treebank consists of 576 sentences and 2,306 tokens annotated in light with the UD guidelines. We evaluated the treebank on the dependency parsing task using a pretrained multilingual parsing model, and the results are comparable with other low-resourced treebanks with no training set. We aim to expand our treebank in the near future to include 1,500 sentences. The bigger goal for our project is to create a set of treebanks for minority languages in Anatolia.
Electroencephalography (EEG)-based emotion recognition has advanced the field in affective computing and has enabled applications in human-computer interactions. Despite significant progress has been made in decoding emotion using supervised machine-learning methods, few studies applied data-driven, unsupervised approaches to explore the underlying EEG dynamics during an emotion experiment and examine how such dynamics correlate with subjective reports of emotion. This study employs the adaptive mixture independent component analysis (AMICA), an unsupervised approach, to EEG data from the DEAP dataset where 32 subjects watched emotional videos. Empirical results showed that AMICA could learn distinct models that separated EEG date collected in the emotion experiment. The identified changes in EEG patterns were weakly-correlated with the four reported emotion scales, indicating the underlying EEG dynamics partially reflected the emotional activities as well as the emotion-irrelevant brain dynamics. Further, the correlations between EEG dynamics and individuals' subjective emotional ratings were significantly higher than those between the EEG and the average ratings from online raters. Finally, building an emotion-decoding model based on the EEG dynamics revealed a significantly better classification performance for valence ratings compared to arousal. This study demonstrated the use of AMICA in characterizing the EEG dynamics in emotion experiments and provided insight into the relationship between EEG and the reported emotional experiences. The unsupervised learning approach can be applied to studying emotion and other confounding factors such as emotion irrelevant EEG artifacts, thereby improving the performance of emotion decoding for EEG-based affective computing.
Abstract This article presents a novel method to determine particular syntactical attributes of Ancient Greek oratory to quantitatively compare the orators of the Classical to those of the Imperial era based on their style of writing. The study first provides a philological overview of Classical Atticism and its Imperial counterpart and argues that the latter is the product of creative mimēsis and not a mere reproduction of archetypes. Then the article briefly explains a node-based metric method that was developed to quantify the morphology of a syntactically annotated Treebank that led to a more thorough weighting scheme using Haar Wavelets. Wavelets were then used to capture both the linear topology of a sentence and the tree network topology of the corresponding syntactical tree. The results were subsequently processed using principal component analysis to analyze and visualize the data. The method is demonstrated using a database of syntactically annotated sentences from six Attic orators.
The performance of a long short-term memory (LSTM) recurrent neural network (RNN)-based language model has been improved on language model benchmarks. Although a recurrent layer has been widely used, previous studies showed that an LSTM RNN-based language model (LM) cannot overcome the limitation of the context length. To train LMs on longer sequences, attention mechanism-based models have recently been used. In this paper, we propose a LM using a neural Turing machine (NTM) architecture based on localized content-based addressing (LCA). The NTM architecture is one of the attention-based model. However, the NTM encounters a problem with content-based addressing because all memory addresses need to be accessed for calculating cosine similarities. To address this problem, we propose an LCA method. The LCA method searches for the maximum of all cosine similarities generated from all memory addresses. Next, a specific memory area including the selected memory address is normalized with the softmax function. The LCA method is applied to pre-trained NTM-based LM during the test stage. The proposed architecture is evaluated on Penn Treebank and enwik8 LM tasks. The experimental results indicate that the proposed approach outperforms the previous NTM architecture.
Text segmentation is a fundamental task in natural language processing. Depending on the levels of granularity, the task can be defined as segmenting a document into topical segments, or segmenting a sentence into elementary discourse units (EDUs). Traditional solutions to the two tasks heavily rely on carefully designed features. The recently proposed neural models do not need manual feature engineering, but they either suffer from sparse boundary tags or cannot efficiently handle the issue of variable size output vocabulary. In light of such limitations, we propose a generic end-to-end segmentation model, namely <inline-formula><tex-math notation="LaTeX">${\mathrm{S}\scriptstyle{\mathrm{EG}}}{\mathrm{B}\scriptstyle{\mathrm{OT}}}$</tex-math></inline-formula>, which first uses a bidirectional recurrent neural network to encode an input text sequence. <inline-formula><tex-math notation="LaTeX">${\mathrm{S}\scriptstyle{\mathrm{EG}}}{\mathrm{B}\scriptstyle{\mathrm{OT}}}$</tex-math></inline-formula> then uses another recurrent neural networks, together with a pointer network, to select text boundaries in the input sequence. In this way, <inline-formula><tex-math notation="LaTeX">${\mathrm{S}\scriptstyle{\mathrm{EG}}}{\mathrm{B}\scriptstyle{\mathrm{OT}}}$</tex-math></inline-formula> does not require any hand-crafted features. More importantly, <inline-formula><tex-math notation="LaTeX">${\mathrm{S}\scriptstyle{\mathrm{EG}}}{\mathrm{B}\scriptstyle{\mathrm{OT}}}$</tex-math></inline-formula> inherently handles the issue of variable size output vocabulary and the issue of sparse boundary tags. In our experiments, <inline-formula><tex-math notation="LaTeX">${\mathrm{S}\scriptstyle{\mathrm{EG}}}{\mathrm{B}\scriptstyle{\mathrm{OT}}}$</tex-math></inline-formula> outperforms state-of-the-art models on two tasks: document-level topic segmentation and sentence-level EDU segmentation. As a downstream application, we further propose a hierarchical attention model for sentence-level sentiment analysis based on the outcomes of <inline-formula><tex-math notation="LaTeX">${\mathrm{S}\scriptstyle{\mathrm{EG}}}{\mathrm{B}\scriptstyle{\mathrm{OT}}}$</tex-math></inline-formula>. The hierarchical model can make full use of both word-level and EDU-level information simultaneously for sentence-level sentiment analysis. In particular, it can effectively exploit EDU-level information, such as the inner properties of EDUs, which cannot be fully encoded in word-level features. Experimental results show that our hierarchical model achieves new state-of-the-art results on the Movie Review and Stanford Sentiment Treebank benchmarks.
This paper investigates whether typical stress patterns in English nouns and verbs are available as a prosodic cue for categorisation and accelerated word learning during first language acquisition. The stress typicality hypothesis states that left-stressed nouns and right-stressed verbs should be acquired earlier than the reverse configurations if stress effectively signals lexical class membership. In this view, class-typical stress patterns are expected to facilitate learning of novel items. A series of generalized additive models (GAMs) based on a comprehensive set of lexical data (CELEX) as well as a large set of age-of-acquisition (AoA) and concreteness ratings reveals that stress typicality plays a minor role in early acquisition, as it is generally superseded by a preference for left-hand (or 'trochaic') patterns in both nouns and verbs. This may be explained by general cognitive constraints (perceptual salience and recency) or exposure to the dominant pattern in the ambient language.
An interesting and frequent type of multiword expression (MWE) is the headless MWE, for which there are no true internal syntactic dominance relations; examples include many named entities ("Wells Fargo") and dates ("July 5, 2020") as well as certain productive constructions ("blow for blow", "day after day").Despite their special status and prevalence, current dependency-annotation schemes require treating such flat structures as if they had internal syntactic heads, and most current parsers handle them in the same fashion as headed constructions.Meanwhile, outside the context of parsing, taggers are typically used for identifying MWEs, but taggers might benefit from structural information.We empirically compare these two common strategies-parsing and tagging-for predicting flat MWEs.Additionally, we propose an efficient joint decoding algorithm that combines scores from both strategies.Experimental results on the MWE-Aware English Dependency Corpus and on six non-English dependency treebanks with frequent flat structures show that: (1) tagging is more accurate than parsing for identifying flat-structure MWEs, (2) our joint decoder reconciles the two different views and, for non-BERT features, leads to higher accuracies, and (3) most of the gains result from feature sharing between the parsers and taggers.
Sentiment analysis, especially for long documents, plausibly requires methods\ncapturing complex linguistics structures. To accommodate this, we propose a\nnovel framework to exploit task-related discourse for the task of sentiment\nanalysis. More specifically, we are combining the large-scale,\nsentiment-dependent MEGA-DT treebank with a novel neural architecture for\nsentiment prediction, based on a hybrid TreeLSTM hierarchical attention model.\nExperiments show that our framework using sentiment-related discourse\naugmentations for sentiment prediction enhances the overall performance for\nlong documents, even beyond previous approaches using well-established\ndiscourse parsers trained on human annotated data. We show that a simple\nensemble approach can further enhance performance by selectively using\ndiscourse, depending on the document length.\n
In neural machine translation (NMT), sequence distillation (SD) through creation of distilled corpora leads to efficient (compact and fast) models.However, its effectiveness in extremely low-resource (ELR) settings has not been well-studied.On the other hand, transfer learning (TL) by leveraging larger helping corpora greatly improves translation quality in general.This paper investigates a combination of SD and TL for training efficient NMT models for ELR settings, where we utilize TL with helping corpora twice: once for distilling the ELR corpora and then during compact model training.We experimented with two ELR settings: Vietnamese-English and Hindi-English from the Asian Language Treebank dataset with 18k training sentence pairs.Using the compact models with 40% smaller parameters trained on the distilled ELR corpora, greedy search achieved 3.6 BLEU points improvement in average while reducing 40% of decoding time.We also confirmed that using both the distilled ELR and helping corpora in the second round of TL further improves translation quality.Our work highlights the importance of stage-wise application of SD and TL for efficient NMT modeling for ELR settings.
The position of the word final –s, after a weakening in archaic Latin, seems to be fixed in the spoken language in the classical period. Then, it partially disappeared in the Romance languages: in modern languages, it is conserved only north and west of the Massa–Senigallia line, while we cannot find it neither in the eastern regions nor in South Italy. Based on this fact, linguists generally claim that the weakening of the final –s started only after the intensive dialectal diversification of Latin, simultaneously with the evolution of the Romance languages. However, the data of the Computerized Historical Linguistic Database of Latin Inscriptions of the Imperial Age (LLDB) do not verify this generally accepted opinion. We can find almost as many examples of the lack of word final –s as that of –m also from the earlier centuries of the Imperial age. The aim of this paper is to explore the reasons behind the inconsistencies between the scholarly consensus and the epigraphical data.
We present a method for conducting morphological disambiguation for South Sámi, which is an endangered language. Our method uses an FST-based morphological analyzer to produce an ambiguous set of morphological readings for each word in a sentence. These readings are disambiguated with a Bi-RNN model trained on the related North Sámi UD Treebank and some synthetically generated South Sámi data. The disambiguation is done on the level of morphological tags ignoring word forms and lemmas; this makes it possible to use North Sámi training data for South Sámi without the need for a bilingual dictionary or aligned word embeddings. Our approach requires only minimal resources for South Sámi, which makes it usable and applicable in the contexts of any other endangered language as well.
Background: Virtual reality (VR) allows people to embody avatars that are different from themselves in appearance and ability. These experiences provide opportunities to challenge bodily perceptions. We devised a novel VR Body Image Training (VR-BIT) approach to target self-perceptions and pain in people with persistent pain. Methods: A 45-year old male with a 5-year history of disabling chronic low back pain participated in a four week VR-BIT intervention. Pain began following a fall from a first-floor deck. Pain was central and on the right side of his lower back, radiating to his right buttock and thigh. Pain was constant and varying at a 5/10 average intensity. The 4-week intervention consistent of three face-to-face sessions one week apart, followed by 1-week of in-home VR-BIT. During the first face-to-face session, the participant embodied three athletic avatars: a superhero (Incredible Hulk), a boxer, and a rock climber. Since the participant strongly identified with the boxer, only boxing experiences were subsequently used. Primary outcomes relating to body image (self-perceived strength, vulnerability, agility and confidence with activity) and pain intensity were assessed using numerical rating scales (0 to 10 NRS). Disabilility, kinesiophobia, overall change, and self-efficacy were assessed as secondary outcomes. Outcomes were assessed during each face-to-face session, and at 1-week and 3-month follow-up. Results: The participant reported a high degree of engagement. Positive changes were noted during and after VR for all body image and pain assessments. Improvements were retained at 3-months for body image ratings (mean change: 4.5/10 NRS) and average pain intensity (change: 2/10 NRS)). Improvements in disability (45% improvement); self-efficacy (pre: 2/12; post: 10/12); and overall change (‘Very much improved’) were noted at 3-month follow-up. No change in kinesiophobia was detected. No adverse advents were recorded. Conclusion: The participant engaged strongly with the intervention and showed clinically meaningful changes in body image, pain, disability and self-efficacy. Despite his long history of pain and rapid improvements, reported changes may be due to non-treatment effects. Nonetheless, VR-BIT clearly warrants further investigation as a potential addition to usual care.
The author reflects upon the place of the Russian language in modern Russia, its distribution in the world, the importance of basic Russian studies for the development of science and Russian society, and the activity of the Russian Academy of Sciences in maintaining the stability of linguistic norms and the culture of Russian speech. On the one hand, scientific research in the field of the Russian language is oriented at obtaining basic theoretical knowledge, which favors comprehensive study of man and society. On the other hand, there is a social order, which is formulated by society proceeding from the need to document language resources and adapt them to topical communicative requirements. The Russian Academy of Sciences carries out expert assessment of speech innovations and codification of the norms of the literary language in normative dictionaries, grammars, and reference books on the culture of speech. The current state of research on the Russian language is analyzed, special attention being paid to problems of the codification of the norms of Russian speech and related tasks.
The concept of Ontologies has been used in a wide range of application domains, due to the fact that ontologies provide a useful mean for establishing a formal, shared and collective understanding of the concepts and their underlying relations at a certain domain of interest, which allows for interoperability and information exchange in a formal an understandable way for both humans and machines. In Cultural Heritage (CH) domain, ontologies serve as a fundamental building block for the traceability of the cultural heritage objects, especially with the increasing demand of providing digital formats for cultural objects and make them available for public. In this paper we implement OntoM; an Ontology model that incorporates the relevant concepts of the Cultural Heritage (CH) domain in Qatar. Then, we will use such an ontology to perform inferences about cultural object classifications via two approaches: string matching, that allows for direct matching between the object and the ontology concepts, and semantic matching, in which we use WordNet lexical database to find all possible synonyms for properties of a given anonymous object.
This paper describes the automatic construction of FinnMWE: a lexicon of Finnish Multi-Word Expressions (MWEs). In focus here are syntactic frames: verbal constructions with arguments in a particular morphological form. The verbal frames are automatically extracted from FinnWordNet and English Wiktionary. The resulting lexicon interoperates with dependency tree searching software so that instances can be quickly found within dependency treebanks. The extraction and enrichment process is explained in detail. The resulting resource is evaluated in terms of its coverage of different types of MWEs. It is also compared with and evaluated against Finnish PropBank.
Many government schemes were unsuccessful because lack of proper feedback on the ongoing schemes, where billion dollars investment is going to be in vain. Sentiment analysis is one of best approach to analyse opinions of the peoples on various government schemes. Sentiment analysis and machine learning techniques emerged to analyse huge social media corpora to track people's views on government policies, products and services. Sentiment analysis process consists of various phases which include data discovery, data collection, data pre-processing, and data analysis. Stemming is a process to generate the morphemes in natural language sentences for various applications such as sentiment analysis, information retrieval, and domain analysis. The stemming process involved two major errors, which are over-stemming and under-stemming errors. Most of sentiment analysis natural languages processing applications used Lancaster and Porter stemming algorithms where more than one word inflected into same morpheme, which causes the etymology behaviour of the stemming word and prone to classify the tweets false positives and false negative. The proposed un-prejudice light stemming algorithm prevent etymology behaviour of morpheme and sustain its meaning during stemming process by selecting a word which has maximum number of synonyms in lexical database.
The present paper seeks to explore the phenomenon of linguistic creativity. Over the years there have been numerous theoretical and experimental studies on the topic of language and creativity. However, among the research papers few discuss linguistic creativity in cognitive and communicative aspects, which are of particular relevance to the current study. The findings are discussed in the light of cognitive-discursive approach. In today’s world violations of norms are manifested at all levels of a language and in almost all types of discourse. The existing standards determine the use of language tools in accordance with the rules of a language, its laws of register, genre, code, function, rules regarding the appropriateness of language units, their collocability, derivation, etc., as well as, in a broader sense, with the objectives of communication. Taking into account the latter statement, the question arises if a linguistic personality should prioritize the choice of preserving the linguistic norm or violate it in their lingua-creative activity to achieve a particular goal of communication. A lingua-creative personality while searching for a name to some innovative mental formations, those that have not yet been verbalized by linguistic means, either produces novel linguistic units and categories by further exploiting the productive potential of a language; or rethinks the existing models, bending the rules and norms of a language; or violates those rules and norms.
OBJECTIVE: To compare "virtual" unenhanced (VUE) computed tomography (CT) images, reconstructed from rapid kVp-switching dual-energy computed tomography (DECT), to "true" unenhanced CT images (TUE), in clinical abdominal imaging. The ability to replace TUE with VUE images would have many clinical and operational advantages. METHODS: VUE and TUE images of 60 DECT datasets acquired for standard-of-care CT of pancreatic cancer were retrospectively reviewed and compared, both quantitatively and qualitatively. Comparisons included quantitative evaluation of CT numbers (Hounsfield Units, HU) measured in 8 different tissues, and 6 qualitative image characteristics relevant to abdominal imaging, rated by 3 experienced radiologists. The observed quantitative and qualitative VUE and TUE differences were compared against boundaries of clinically relevant equivalent thresholds to assess their equivalency, using modified paired t-tests and Bayesian hierarchical modeling. RESULTS: Quantitatively, in tissues containing high concentrations of calcium or iodine, CT numbers measured in VUE images were significantly different from those in TUE images. CT numbers in VUE images were significantly lower than TUE images when calcium was present (e.g. in the spine, 73.1 HU lower, p < 0.0001); and significantly higher when iodine was present (e.g. in renal cortex, 12.9 HU higher, p < 0.0001). Qualitatively, VUE image ratings showed significantly inferior depiction of liver parenchyma compared to TUE images, and significantly more cortico-medullary differentiation in the kidney. CONCLUSIONS: Significant differences in VUE images compared to TUE images may limit their application and ability to replace TUE images in diagnostic abdominal CT imaging.
With the tremendous success of deep learning models on computer vision tasks, there are various emerging works on the Natural Language Processing (NLP) task of Text Classification using parametric models. However, it constrains the expressability limit of the function and demands enormous empirical efforts to come up with a robust model architecture. Also, the huge parameters involved in the model causes over-fitting when dealing with small datasets. Deep Gaussian Processes (DGP) offer a Bayesian non-parametric modelling framework with strong function compositionality, and helps in overcoming these limitations. In this paper, we propose DGP models for the task of Text Classification and an empirical comparison of the performance of shallow and Deep Gaussian Process models is made. Extensive experimentation is performed on the benchmark Text Classification datasets such as TREC (Text REtrieval Conference), SST (Stanford Sentiment Treebank), MR (Movie Reviews), R8 (Reuters-8), which demonstrate the effectiveness of DGP models. © European Language Resources Association (ELRA), licensed under CC-BY-NC
This paper attempts to measure the similarity of frequently occurring modal auxiliaries in both L1 and L2 writings. The modal auxiliaries are known to be difficult areas of study for both L1 and L2, due to the overlapping in their use. A subset of TOEFL11 corpus and a subset of PTB(Penn Treebank) corpus were used to capture the similarities among modal auxiliaries within L1 and L2, respectively, and also across L1 and L2. Based on the hypothesis of distributional representation, similarities of modal auxiliaries were computed by applying cosine similarity and mutual information to every pair of words in the corpus, and then finally reducing dimensions with Singular Value Decomposition. The results show that the modals are found to be similar to other modals, and the distribution of modals in the reduced dimensions show that the modals in L1 and those in L2 exhibits different patters of similarities, while the distance among the modals in L1 is further away than the distance among those L2 modals.
BACKGROUND: How dental education influences students' dental and dentofacial esthetic perception has been studied for some time, given the importance of esthetics in dentistry. However, no study before has studied this question in a large sample of students from all grades of dental school. This study sought to fill that gap. The aim was to assess if students' dentofacial esthetic autoperception and heteroperception are associated with their actual stage of studies (grade) and if autoperception has any effect on heteroperception. METHODS: Between October 2018 and August 2019, a questionnaire was distributed to 919 dental students of all 5 grades of dental school at all four dental schools in Hungary. The questionnaire consisted of the following parts (see also the supplementary material): 1. Demographic data (3 items), Self-Esthetics I (11 multiple- choice items regarding the respondents' perception of their own dentofacial esthetics), Self-Esthetics II (6 Likert-type items regarding the respondents' perception of their own dentofacial esthetics), and Image rating (10 items, 5 images each, of which the respondents have to choose the one they find the most attractive). Both the self-esthetics and the photo rating items were aimed at the assessment of mini- and microesthetic features. RESULTS: The response rate was 93.7% (861 students). The self-perception of the respondents was highly favorable, regardless of grade or gender. Grade and heteroperception were significantly associated regarding maxillary midline shift (p < 0.01) and the relative visibility of the arches behind the lips (p < 0.01). Detailed analysis showed a characteristic pattern of preference changes across grades for both esthetic aspects. The third year of studies appeared to be a dividing line in both cases, after which a real preference order was established. Association between autoperception and heteroperception could not be verified for statistical reasons. CONCLUSION: Our findings corroborate the results of most previous studies regarding the effect of dental education on the dentofacial esthetic perception of students. We have shown that the effect can be demonstrated on the grade level, which we attribute to the specific curricular contents. We found no gender effect, which, in the light of the literature, suggests that the gender effect in dentofacial esthetic perception is highly culture dependent. The results allow no conclusion regarding the relation between autoperception and heteroperception.
The article analyzes differences in the description of discourse relations in corpus research, in particular with the reference to the use of discourse markers – expressions that tie together subsequent fragments of the text and provide information about the nature of these relations. The text presents three concepts of the description of explicitness and implicitness of the content: Rhetorical Structure Theory, Penn Discourse Treebank and the author’s original proposal and indicates the consequences of each solution. The analysis of relations with particles as metatexual expressions defined in accordance with The Nest Dictionary of Polish reveals the possibility of expressing explicitness as a representation of elements of informational structure shaped by the use of a given particle, and implicitness as a lack of representation of certain elements of this type.
Neural machine translation (NMT) models are typically trained using a softmax cross-entropy loss where the softmax distribution is compared against smoothed gold labels. In low-resource scenarios, NMT models tend to over-fit because the softmax distribution quickly approaches the gold label distribution. To address this issue, we propose to divide the logits by a temperature coefficient, prior to applying softmax, during training. In our experiments on 11 language pairs in the Asian Language Treebank dataset and the WMT 2019 English-to-German translation task, we observed significant improvements in translation quality by up to 3.9 BLEU points. Furthermore, softmax tempering makes the greedy search to be as good as beam search decoding in terms of translation quality, enabling 1.5 to 3.5 times speed-up. We also study the impact of softmax tempering on multilingual NMT and recurrently stacked NMT, both of which aim to reduce the NMT model size by parameter sharing thereby verifying the utility of temperature in developing compact NMT models. Finally, an analysis of softmax entropies and gradients reveal the impact of our method on the internal behavior of NMT models.
This journal article follows the research line opened on the search for semantic primes’ exponents in Old English within the frame of the Natural Semantic Metalanguage theory (Goddard 1997, 2012; Goddard and Wierzbicka 2002). The aim of this study is to complete the line of research on prime identification opened on the category Actions, events, movement, contact by establishing the Old English exponent of the prime DO. With this purpose, this paper discusses the adequacy of different OE verbs as possible prime exponent on the basis of textual frequency, morphology, semantics and syntactic complementation. Relevant data of analysis have been retrieved mainly from the lexical database of Old English Nerthus, the Dictionary of Old English (Healey et al. 2018) and the Dictionary of Old English Corpus (Healey et al. 2009).
The aim of this paper is to retrieve the most relevant expansion words for expanding the initial query of the user in order to enhance the outcomes of web search results. Query expansion plays a major role in reformulating a user’s initial query to a one more pertinent to the user’s intended meaning. The reformulated query is then used to obtain more appropriate outcomes from a large amount of information on the web. The proposed semantic query expansion technique uses Wikipedia and WordNet as data sources. Wikipedia is taken as a base for all query expansions because it is one of the most diversified and relevant databases available on the web. To further improve the proposed query expansion technique,WordNet—a lexical database—is used as the as another data source because the synonyms (synsets) of the query term provided by it can be quite useful for query expansion. The proposed expansion technique successfully combines the two data sources to retrieve the most relevant expansion terms from the data sources in response to the user’s original query. The proposed work has been divided into four phases: (1) extraction of relevant words from Wikipedia (2) extraction of relevant words from WordNet (3) merging of the expansion terms obtained from Wikipedia and WordNet, and (4) query formulation by combining the expansion terms using Boolean operators. This reformulated query is then fired on the web to find the desired result. The Experimental result shows a significant improvement in information retrieval using query expansion.
Using incorrect worked examples during mathematics instruction can improve student learning. However, teachers worry that students may confuse correct and incorrect examples over time, and memory research supports this fear. To examine if this forgetting occurs, we had undergraduates rate the correctness of correct and incorrect worked examples immediately and one week later (Experiment 1). Previously studied incorrect examples were rated as slightly more correct after the delay, but this did not affect ratings of unstudied examples or problem-solving accuracy. In Experiment 2, we more closely mimicked how incorrect worked examples are used in classroom settings. Again, we found only small changes in students’ memory for studied worked examples after the delay, and no changes for unstudied examples or problem-solving accuracy. Our findings suggest the costs of teaching with incorrect worked examples are limited to the specific studied problems, and do not affect learning of the underlying mathematical rule.
BACKGROUND: Instrumental activities of daily living (IADL) impairment can begin in mild cognitive impairment (MCI), and is the core criteria for diagnosing dementia in both Alzheimer's (AD) and Parkinson's (PD) diseases. The Functional Activities Questionnaire (FAQ) has high discriminative power for dementia and MCI in older age populations, but is influenced by demographic factors. It is currently unclear whether the FAQ is suitable for assessing cognitive-associated IADL in non-demented PD patients, as motor disorders may affect ratings. OBJECTIVE: To compare IADL profiles in MCI patients with PD (PD-MCI) and AD (AD-MCI) and to verify the discriminative ability of the FAQ for MCI in patients with (PD-MCI) and without (AD-MCI) additional motor impairment. METHODS: Data of 42 patients each of PD-MCI, AD-MCI, PD cognitively normal (PD-CN), and healthy controls (HC), matched according to age, gender, education, and global cognitive impairment were analyzed. ANCOVA and binary regressions were used to examine the relationship between the FAQ scores and groups. FAQ cut-offs for PD-MCI (versus PD-NC) and AD-MCI (versus HC) were separately identified using receiver operating characteristic analyses. RESULTS: FAQ total score did not differentiate between MCI groups. PD-MCI subjects had greater difficulties with tax records and traveling while AD-MCI individuals were more impaired in managing finances and remembering appointments. Classification accuracy of the FAQ was good for diagnosing AD-MCI (69%, cut-off ≥1) compared to HC, and sufficient for differentiating PD-MCI (38.1%, cut-off ≥3) from PD-CN. CONCLUSION: The FAQ task profiles and classification accuracy differed between MCI related to PD and AD.
This paper analyses data to address a specific linguistic problem, i.e. the acquisition of the modification potential of the three more or less synonymous Dutch degree modifiers heel, erg and zeer, all meaning 'very', which show syntactic differences in modification potential. It continues the research reported on in The analysis makes crucial use of linguistic applications developed in the CLARIN infrastructure, in particular the treebank search applications PaQu (Parse and Query) and GrETEL Version 4.00. The analysis benefits from the use of parsed corpora (treebanks) in combination with the search and analysis options offered by PaQu and GrETEL. Earlier work showed that despite little data for zeer modifying adpositional phrases adult speakers end up with a generalised modification potential for this word. In this paper, I extend the dataset considered, and find more (but still little) data for this phenomenon. However, I also find a similar amount of data that form counterexamples to the non-generalisation of the modification potential of heel. I argue that the examples with heel concern constructions with idiosyncratic semantics and therefore are not counted as evidence for the general rule of modification. I suggest a simple statistical analysis to account for the fact that children 'learn' that heel cannot modify verbs or adpositions though there is no explicit evidence for this and they are not explicitly taught so.
Scene graph is a graph representation that explicitly represents high-level semantic knowledge of an image such as objects, attributes of objects and relationships between objects. Various tasks have been proposed for the scene graph, but the problem is that they have a limited vocabulary and biased information due to their own hypothesis. Therefore, results of each task are not generalizable and difficult to be applied to other down-stream tasks. In this paper, we propose Entity Synset Alignment(ESA), which is a method to create a general scene graph by aligning various semantic knowledge efficiently to solve this bias problem. The ESA uses a large-scale lexical database, WordNet and Intersection of Union (IoU) to align the object labels in multiple scene graphs/semantic knowledge. In experiment, the integrated scene graph is applied to the image-caption retrieval task as a downstream task. We confirm that integrating multiple scene graphs helps to get better representations of images.
This paper represents the development of the Myanmar Named Entity Recognition (NER) system using Conditional Random Fields (CRFs). In order to develop the system, a manually annotated Named Entities (NEs) corpus - collected from Myanmar news websites and Asia Language Treebank(ALT)-Parallel-Corpus has been used. We compare the performance of the system getting syllable-based input to the one getting character-based input. We observed that training data has more impact on the performance of the system. The experimental results show that the syllable-based system performs better than the character-based system. It achieves that Precision, Recall and F1-score values of 93.62%, 91.64% and 92.62% respectively.
OBJECTIVE: To design and evaluate the effectiveness of a stimulus material in eliciting the N400 event related potential (ERP). DESIGN: A set of 700 semantically congruent and incongruent sentences was developed in accordance with current linguistic norms, and validated with an electroencephalography (EEG) study, in which the influence of age and gender on the N400 ERP magnitude was analysed. STUDY SAMPLE: Forty-five normal-hearing subjects (19-57 years, 21 females) participated in the EEG study. RESULTS: The stimulus material used in the EEG study elicited a robust N400 ERP, with a morphology consistent with the literature. Results also showed no statistically significant effect of age or gender on the N400 magnitude. CONCLUSIONS: The material presented in this paper constitutes the largest complete stimulus set suitable for both auditory and text-based N400 experiments. This material may help facilitate the efficient implementation of future N400 ERP studies, as well as promote standardisation and consistency across studies.
The research work is dealt with the culture of speech. Culture of speech is identified by language of speech, social surrounding, language and psycology, language and pragmatics. As a result of culture of speech, personal culture, human quality, linguistic knowledge of a person is realized. Several types of antropolinguistic analyses are mentioned. Public speech (in auditorium, in crowd) shows the social aspect of communicative, linguistic norms and humans morality in speaking expresses wisdom and psycholinguistic aspect of speaker. Literary norm, functional grammar, cognitive pragmatics, literary language are thoroughly explained.
Increasing popularity of electronic dictionaries, ontologies, thesauri and lexical databases makes them an effective tool for language learning purposes (Dash, 2013; Fellbaum, 2010; Miller & Fellbaum, 1992; Shimodaira et al, 2006; Sun et al, 2011). The aim of this research is to study the educational potential of electronic lexical database for the English language WordNet (Miller, 1995; Fellbaum, 1998) and electronic thesauri for the Russian language RuWordNet (Loukachevitch, 2011; Loukachevitch, Lashevich, 2016) in teaching English as a foreign language. In this research the authors focus on teaching colours, in particular a colour term white, as colours represent complex linguistic and culture-specific phenomena, reflected in the “cultural memory” of people, the very concept of ‘colour’ being a function of language and culture (Wierzbicka, 2006). Thus, understanding colours helps students both study a foreign language and learn its history and culture.This is a mixed method study based on comparison of synonyms for colour term white in WordNet and RuWordNet. The main relation among words in these thesauri is synonymy. Based on their meanings, words are grouped into unordered sets of synonyms expressing one underlying concept (synsets). This allows to consider specific senses of words and semantic relations. Firstly, Russian students studying English as a foreign language were asked to analyse the meanings and synonyms for colour term white in RuWordNet. Secondly, the students compared the representation of colour term white in WordNet. Additionally, they examined set phrases with the adjective white in English dictionaries. Then a qualitative method (interviewing) was used to reveal the students’ perceptions of studying English by means of electronic dictionaries, thesauri and lexical databases.The study allowed to claim that electronic thesauri and lexical databases are of educational value, increasing students’ linguistic awareness and language proficiency.
BACKGROUND: Analyze intrarater and interrater reliability for evaluating endoscopic images of velopharyngeal (VP) physiology. METHOD: Speakers produced 9 speech stimuli representing 4 stimulus types: sustained phonemes, repetitions of "puh," single words, and short phrases. The 37-speaker participants included 16 patients with VP dysfunction and 21 control participants. Five raters independently rated the video images for degree of VP opening, location of opening, and pattern of closure. Outcome measures included intrarater and interrater measures of reliability and the effects of raters and stimulus type on ratings. RESULTS: Intrarater reliability was acceptable, and ratings were logically consistent. Fixed effects regression coefficients for the patient and the control groups showed that raters were a significant source of variability for degree of opening and pattern of closing. Stimulus type was not a significant source of variation for any metric for the controls, but stimulus type was a significant determinant for degree of opening for patients. The degree of opening was larger for sustained phonemes than for the other speech stimuli. Ratings for degree of opening were most similar for repeated "puh." CONCLUSIONS: Interrater reliability needs to be improved so that the assessment procedure produces more consistent findings among clinicians, thus strengthening our evidence base for this procedure. Interrater additional research is needed to understand how the stimulus affects ratings of VP physiology, to identify stimuli that yield the most useful clinical information, and to understand how training affects the ratings of VP physiology.
Our work on the automatic detection of English discourse connectives in the Penn Discourse Treebank (PDTB) shows that syntactic information from the Universal Dependencies (UD) framework is a viable alternative to that from the Penn Treebank (PTB) framework. In fact, we found minor increases when comparing between the use of gold standard PTB part-of-speech (POS) tag information and automatically parsed UD information. The former has traditionally been used for the task but there are now much more UD corpora and in many more languages than that available in the PTB framework. As such, this finding is promising for areas in discourse parsing such as in multilingual as well as under production settings, where gold standard PTB information may be scarce.
In order to effectively respond to the increased linguistic and cultural diversity in the U.S. schools and close the consistently documented achievement gap between culturally and linguistically diverse (CLD) students and mainstream students, teachers need to take an asset-based approach and be able to draw on CLD students’ entire funds of linguistic knowledge. However, few studies have examined CLD students’ linguistic choices in multiple discursive spaces with different linguistic norms, values and practices. This article addresses this research gap through a case study of Elif, a Turkish American student and her linguistic boundary crossing experiences within and across three discursive spaces: her home, her Turkish heritage language school, and her mainstream school. Through in-depth analysis of interviews, observations, and field notes, the study revealed that Elif experienced different linguistic environments and boundary types. She negotiated experiences that ranged from smooth to managed to insurmountable boundaries. Finally, translanguaging practices acted as a key boundary object that mediated sociocultural discontinuities in the Turkish heritage language school, and facilitated Elif’s experiences between Turkish dominant and English dominant discursive spaces.
Corporate credit ratings (CRs) are closely related to companies’ cost of debt financing. Recent research has drawn wide attention to how nonfinancial as well as financial factors may affect ratings. By manually collecting information about the profiles of chief financial officers (CFOs) of US companies, we examine the effect of CFOs’ accounting expertise on corporate CRs. The results show that firms with accounting expert CFOs are more likely to receive higher CRs and that the effect of CFOs’ accounting expertise on the ratings is more pronounced for firms with higher default risk, suggesting that the accounting expertise of CFOs may be an important factor that affects CRs. Moreover, we find a dynamic relation between accounting expert CFOs and CRs such that a downgrade in a firm’s CR in a prior year affects the subsequent selection of an accounting expert CFO.