Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
The main objective of this research was to study the comprehension level of heritage speakers of Turkish with regard to Turkish proverbs. The familiarity factor in relation to its role in comprehending proverbs is also examined. Familiarity judgments of proverbs were made by heritage speakers of Turkish to determine whether the familiarity ratings of the heritage speakers of Turkish could be associated with their understanding of proverbs. The results of the proverb comprehension tests indicate that the difference between the bilingual heritage speakers of Turkish and baseline monolingual speakers in their comprehension of proverbs was significant. Yet, when the arithmetic mean rank is taken into consideration, both heritage speakers of Turkish and monolinguals display high level performance in their comprehension of proverbs. However, there appears to be a noteworthy difference between the frequency level of heritage speakers’ and monolinguals’ in their encounter with proverbs. Monolinguals outperformed bilinguals in comprehension of proverbs and familiarity rating. Results also showed that familiarity had a nonsignificant correlation with participants’ performance on proverb comprehension.
Most fetal brain MRI reconstruction algorithms rely only on brain tissue-relevant voxels of low-resolution (LR) images to enhance the quality of inter-slice motion correction and image reconstruction. Consequently the fetal brain needs to be localized and extracted as a first step, which is usually a laborious and time consuming manual or semi-automatic task. We have proposed in this work to use age-matched template images as prior knowledge to automatize brain localization and extraction. This has been achieved through a novel automatic brain localization and extraction method based on robust template-to-slice block matching and deformable slice-to-template registration. Our template-based approach has also enabled the reconstruction of fetal brain images in standard radiological anatomical planes in a common coordinate space. We have integrated this approach into our new reconstruction pipeline that involves intensity normalization, inter-slice motion correction, and super-resolution (SR) reconstruction. To this end we have adopted a novel approach based on projection of every slice of the LR brain masks into the template space using a fusion strategy. This has enabled the refinement of brain masks in the LR images at each motion correction iteration. The overall brain localization and extraction algorithm has shown to produce brain masks that are very close to manually drawn brain masks, showing an average Dice overlap measure of 94.5%. We have also demonstrated that adopting a slice-to-template registration and propagation of the brain mask slice-by-slice leads to a significant improvement in brain extraction performance compared to global rigid brain extraction and consequently in the quality of the final reconstructed images. Ratings performed by two expert observers show that the proposed pipeline can achieve similar reconstruction quality to reference reconstruction based on manual slice-by-slice brain extraction. The proposed brain mask refinement and reconstruction method has shown to provide promising results in automatic fetal brain MRI segmentation and volumetry in 26 fetuses with gestational age range of 23 to 38 weeks.
The SCATE (Smart Computer-Assisted Translation Environment) research project is a multidisciplinary co-operation with the objective of improving translators’ efficiency and consistency through a better integration of existing translation technologies and exploitation of resources. One aspect of this project was to investigate the process of terminology extraction by humans and to automate the process of terminology extraction from comparable corpora. The methods consist of three main tasks: (1) study of translators’ methods to acquire domain knowledge and terminology, (2) determining comparable corpora, and (3) automatic terminology extraction from comparable text. In this presentation we will focus on the results of task (1) which has been completed, and present the preliminary results of task (3). In task (1) we collected the data through an international survey and observations of 16 translators and terminologists at their workplaces across a period of 8 months in 2015. The study revealed information about translators’ web search behaviour and usage of online linguistic resources to solve terminological problems. Besides these, we have identified needs and shortcomings of different CAT tools regarding the terminology management component, integration with online databases and exchange of terminological data. In task (3) we identified two subtasks: monolingual term extraction and term linking (i.e., linking terms to their corresponding translation). For the term extraction, we applied a hybrid approach combining linguistic and statistical information (Macken et al., 2013). Subsequently, two different techniques were investigated to link the translation equivalents: probabilistic topic models as well as different neural network architectures. The best results were obtained with a neural network model. In order to evaluate the performance of the different modules, we created a gold standard for three different domains (heart failure, wind energy, corruption) in three different languages (English, French, Dutch). References Poly-GrETEL. Available online at https://clarin.eu/showcase/poly-gretel-search-engine-querying-syntactic-constructions-parallel-treebanks Macken, L., Lefever, E. & Hoste, V. (2013). TExSIS: Bilingual Terminology Extraction from Parallel Corpora Using Chunk-based Alignment. Terminology, 19 (1), 1-30. John Benjamins Publishing Company, Amsterdam, Netherlands. van den Bergh, Jan et al. (2015). Recommendations for Translation Environments to Improve Translators' workflows. Translating and the Computer 37. Asling. London. van der Lek-Ciudin, Iulianna, Tom Vanallemeersch and Ken de Wachter (2015). Contextual Inquiries at translators’ workplaces. In Proceedings of the 1st TAO-CAT, Angers 2015 TermWise: Resources for Specialised Language Use. Information available online at http://liir.cs.kuleuven.be/projects.php?project=177
The PARSEME shared task aims at identifying verbal MWEs in running texts. Verbal MWEs include idioms (let the cat out of the bag), light verb constructions (make a decision), verb-particle constructions (give up), and inherently reflexive verbs (se suicider 'to suicide' in French). VMWEs were annotated according to the universal guidelines in 18 languages. The corpora are provided in the parsemetsv format, inspired by the CONLL-U format. For most languages, paired files in the CONLL-U format - not necessarily using UD tagsets - containing parts of speech, lemmas, morphological features and/or syntactic dependencies are also provided. Depending on the language, the information comes from treebanks (e.g., Universal Dependencies) or from automatic parsers trained on treebanks (e.g., UDPipe). This item contains training and test data, tools and the universal guidelines file.
This paper is concerned with whether deep syntactic information can help surface parsing, with a particular focus on empty categories. We design new algorithms to produce dependency trees in which empty elements are allowed, and evaluate the impact of information about empty category on parsing overt elements. Such information is helpful to reduce the approximation error in a structured parsing model, but increases the search space for inference and accordingly the estimation error. To deal with structure-based overfitting, we propose to integrate disambiguation models with and without empty elements, and perform structure regularization via joint decoding. Experiments on English and Chinese TreeBanks with different parsing models indicate that incorporating empty elements consistently improves surface parsing.
Scientific research within the humanities is different from what it was a few decades ago. For instance, new sources of information, such as digital grammars, lexical databases and large corpora of real-language data offer new opportunities for linguistics. The Taalportaal grammatical database, with its links to other linguistic resources via the CLARIN infrastructure, is a prime example of a new type of tool for linguistic research.
Following We trained our transition-based projective parser in UD version 2.0 datasets without any additional data. The parser is fast, lightweight and effective on big treebanks.
Identifying the sense of a word within a context is a challenging problem and has many applications in natural language processing. This assignment problem is called word sense disambiguation (WSD). Many papers in the literature focus on English language and data. Our dataset consists of 1400 sentences translated to Turkish from the Penn Treebank Corpus. This paper seeks to address and discuss 6 different feature extraction methods and its classification performances using C4.5, Random Forests, Rocchio, Naive Bayes, KNN, Linear and multilayer Perceptron. This paper calls into question how the described features perform on a morphologically rich language (Turkish) with several classifiers.
We propose a shared task on multilingual Surface Realization, i.e., on mapping unordered and uninflected universal dependency trees to correctly ordered and inflected sentences in a number of languages. A second deeper input will be available in which, in addition, functional words, fine-grained PoS and morphological information will be removed from the input trees. The first shared task on Surface Realization was carried out in 2011 with a similar setup, with a focus on English. We think that it is time for relaunching such a shared task effort in view of the arrival of Universal Dependencies annotated treebanks for a large number of languages on the one hand, and the increasing dominance of Deep Learning, which proved to be a game changer for NLP, on the other hand.
A word may have multiple meanings or senses, it could be modeled by considering that words in a sentence have a fuzzy set that contains words with similar meaning, which make detecting plagiarism a hard task especially when dealing with semantic meaning, and even harder for cross language plagiarism detection. Arabic is known by its richness, word’s constructions and meanings diversity, hence changing texts from/to Arabic is a complex task, and therefore adopting a fuzzy semantic-based approach seems to be the best solution. In this paper, we propose a detailed fuzzy semantic-based similarity model for analyzing and comparing texts in CLP cases, in accordance with the WordNet lexical database, to detect plagiarism in documents translated from/to Arabic, a preprocessing phase is essential to form operable data for the fuzzy process. The proposed method was applied to two texts (Arabic/English), taking into consideration the specificities of the Arabic language. The result shows that the proposed method can detect 85% of the plagiarism cases.
In this work, we propose a novel, implicitly-defined neural network architecture and describe a method to compute its components. The proposed architecture forgoes the causality assumption used to formulate recurrent neural networks and instead couples the hidden states of the network, allowing improvement on problems with complex, long-distance dependencies. Initial experiments demonstrate the new architecture outperforms both the Stanford Parser and baseline bidirectional networks on the Penn Treebank Part-of-Speech tagging task and a baseline bidirectional network on an additional artificial random biased walk task.
A recently proposed encoding for noncrossing digraphs can be used to implement generic inference over families of these digraphs and to carry out first-order factored dependency parsing. It is now shown that the recent proposal can be substantially streamlined without information loss. The improved encoding is less dependent on hierarchical processing and it gives rise to a high-coverage boundeddepth approximation of the space of noncrossing digraphs. This subset is presented elegantly by a finite-state machine that recognizes an infinite set of encoded graphs. The set includes more than 99.99% of the 0.6 million noncrossing graphs obtained from the UDv2 treebanks through planarisation. Rather than taking the low probability of the residual as a flat rate, it can be modelled with a joint probability distribution that is factorised into two underlying stochastic processes -the sentence length distribution and the related conditional distribution for deep nesting. This model points out that deep nesting in the streamlined code requires extreme sentence lengths. High depth is categorically out in common sentence lengths but emerges slowly at infrequent lengths that prompt further inquiry.
This paper extends literature showing that gender status beliefs differentially shape the evaluations of men and women to show that these gender status beliefs transfer to the evaluations of products made by men and women. Two online experiments were conducted to simulate male-typed and female-typed product markets (craft beer and cupcakes, respectively). In the male-typed product market, a craft beer described as produced by a woman is evaluated at a discount vis-à-vis the same product described as produced by a man. The gender discount is only observed among evaluators with little knowledge of craft beer, and the evaluation of the beer made by a woman improved if the beer was conferred external status via an award. In the female-typed product market of cupcakes, producer’s gender does not affect ratings, nor are the ratings significantly affected by winning an award. Together, the two studies provide evidence of an asymmetric negative bias: products made by women are disadvantaged in male-typed markets, but products made by men are not disadvantaged in female-typed markets. We draw out the implications of these findings and suggest ways that gender biases in product markets can be reduced.
We propose a shared task on multilingual SurfaceRealization, i.e., on mapping unorderedand uninflected universal dependency trees tocorrectly ordered and inflected sentences in anumber of languages. A second deeper inputwill be available in which, in addition,functional words, fine-grained PoS and morphologicalinformation will be removed fromthe input trees. The first shared task on SurfaceRealization was carried out in 2011 witha similar setup, with a focus on English. Wethink that it is time for relaunching such ashared task effort in view of the arrival of UniversalDependencies annotated treebanks fora large number of languages on the one hand,and the increasing dominance of Deep Learning,which proved to be a game changer forNLP, on the other hand.
The paper examines the multifunctionality of the word-formation suffixes used in the formation of Old and Middle Czech deverbal nouns and the possibility of using semantic maps to analyze this multifunctionality. This methodology is first briefly introduced, then a semantic map of six onomasiological categories used in the word-formation patterns is presented (based on extensive data from Old and Middle Czech dictionaries and lexical databases), showing which combinations of functions are possible in a single suffix and which are not. The map is compared with maps found in the literature and several methodological issues are discussed.
Abstract This paper addresses the feasibility of cross-lingual parsing with Universal Dependencies (UD) between Romance languages, analyzing its performance when compared to the use of manually annotated resources of the target languages. Several experiments take into account factors such as the lexical distance between the source and target varieties, the impact of delexicalization, the combination of different source treebanks or the adaptation of resources to the target language, among others. The results of these evaluations show that the direct application of a parser from one Romance language to another reaches similar labeled attachment score (LAS) values to those obtained with a manual annotation of about 3,000 tokens in the target language, and unlabeled attachment score (UAS) results equivalent to the use of around 7,000 tokens, depending on the case. These numbers can noticeably increase by performing a focused selection of the source treebanks. Furthermore, the removal of the words in the training corpus (delexicalization) is not useful in most cases of cross-lingual parsing of Romance languages. The lessons learned with the performed experiments were used to build a new UD treebank for Galician, with 1,000 sentences manually corrected after an automatic cross-lingual annotation. Several evaluations in this new resource show that a cross-lingual parser built with the best combination and adaptation of the source treebanks performs better (77 percent LAS and 82 percent UAS) than using more than 16,000 (for LAS results) and more than 20,000 (UAS) manually labeled tokens of Galician.
Because of the importance of the information conveyed by the clinical documents and owing to the large quantity of raw texts produced in the healthcare system, it became a determinant challenge, in the NLP research field, to arrange the extraction and the management of meaningful data, starting from real text occurrences. In this paper we approach a corpus of 5000 medical diagnoses with sophisticated linguistic and computational devices, which are able to access the semantic dimension of words and sentences contained in it. Our morphosemantic method is grounded on a list of neoclassical formative elements pertaining to the medical domain which has been used for the automatic creation and population of medical lexical resources. The outcomes of this work are automatically built electronic dictionaries and thesauri and an annotated corpus for the NLP in the medical domain.
Word segmentation is a basic problem in natural language processing. With the languages having the complex writing system like the Khmer language in Southern of Vietnam, this problem really very intractable, posing the significant challenges. Although there are some experts in Vietnam as well as international having deeply researched this problem, there are still no reasonable results meeting the demand, in particular, no treated thoroughly the ambiguous phenomenon, in the process of Khmer language processing so far. This paper present a solution based on the syllable division into component clusters using two syllable models proposed, thereby building a Khmer syllable database, is still not actually available. This method using a lexical database updated from the online Khmer dictionaries and some supported dictionaries serving role of training data and complementary linguistic characteristics. Each component cluster is labelled and located by the first and last letter to identify entirety a syllable. This approach is workable and the test results achieve high accuracy, eliminate the ambiguity, contribute to solving the problem of word segmentation and applying efficiency in Khmer language processing.
Because of the importance of the information conveyed by the clinical documents and owing to the large quantity of raw texts produced in the healthcare system, it became a determinant challenge, in the NLP research field, to arrange the extraction and the management of meaningful data, starting from real text occurrences. In this paper we approach a corpus of 5000 medical diagnoses with sophisticated linguistic and computational devices, which are able to access the semantic dimension of words and sentences contained in it. Our morphosemantic method is grounded on a list of neoclassical formative elements pertaining to the medical domain which has been used for the automatic creation and population of medical lexical resources. The outcomes of this work are automatically built electronic dictionaries and thesauri and an annotated corpus for the NLP in the medical domain.
Recurrent neural networks (RNNs), such as long short-term memory networks (LSTMs), serve as a fundamental building block for many sequence learning tasks, including machine translation, language modeling, and question answering. In this paper, we consider the specific problem of word-level language modeling and investigate strategies for regularizing and optimizing LSTM-based models. We propose the weight-dropped LSTM which uses DropConnect on hidden-to-hidden weights as a form of recurrent regularization. Further, we introduce NT-ASGD, a variant of the averaged stochastic gradient method, wherein the averaging trigger is determined using a non-monotonic condition as opposed to being tuned by the user. Using these and other regularization strategies, we achieve state-of-the-art word level perplexities on two data sets: 57.3 on Penn Treebank and 65.8 on WikiText-2. In exploring the effectiveness of a neural cache in conjunction with our proposed model, we achieve an even lower state-of-the-art perplexity of 52.8 on Penn Treebank and 52.0 on WikiText-2.
Motivated form-meaning mappings are pervasive in signlanguages, and iconicity has recently been shown to facilitatesign learning from early on. This study investigated the role oficonicity for language acquisition in Turkish Sign Language(TID). Participants were 43 signing children (aged 10 to 45months) of deaf parents. Sign production ability was recordedusing the adapted version of MacArthur Bates CommunicativeDevelopmental Inventory (CDI) consisting of 500 items forTID. Iconicity and familiarity ratings for a subset of 104 signswere available. Our results revealed that the iconicity of a signwas positively correlated with the percentage of childrenproducing a sign and that iconicity significantly predicted thepercentage of children producing a sign, independent offamiliarity or phonological complexity. Our results areconsistent with previous findings on sign language acquisitionand provide further support for the facilitating effect of iconicform-meaning mappings in sign learning.
It is widely believed that different parts of a classical Chinese poem vary in syntactic properties. The middle part is usually parallel, i.e. the two lines in a couplet have similar sentence structure and part of speech; in contrast, the beginning and final parts tend to be non-parallel. Imagistic language, dominated by noun phrases evoking images, is concentrated in the middle; propositional language, with more complex grammatical structures, is more often found at the end. We present the first quantitative analysis on these linguistic phenomena—syntactic parallelism, imagistic language, and propositional language—on a treebank of selected poems from the Complete Tang Poems. Written during the Tang Dynasty between the 7th and 9th centuries CE, these poems are often considered the pinnacle of classical Chinese poetry. Our analysis affirms the traditional observation that the final couplet is rarely parallel; the middle couplets are more frequently parallel, especially at the phrase rather than the word level. Further, the final couplet more often takes a non-declarative mood, uses function words, and adopts propositional language. In contrast, the beginning and middle couplets employ more content words and tend toward imagistic language.
The purpose of this study was to describe the improvement of student learning outcomes IPA with Model make a match at SDN 12 Api-Api. This type of research is the Classroom Action Research (PTK) is conducted in two cycles. The data source is the fourth grade students of SDN 12 Api-Api numbered 17 people. The instrument used in this study is the assessment sheet affective student, teacher activity sheet and test the students' understanding. Based on analysis of the affective ratings of students In the first cycle the change in student behavior responsibilities increased 51.0% to 83.3% in the second cycle and the change in behavior of the cooperation of students in the first cycle of 44.1% increased to 74.5% in the second cycle. The results of cognitive learning that an understanding had also increased. In the first cycle student comprehension 58.8% increase to 82.4 in the second cycle. From the data obtained it can be concluded that there is a learning outcome IPA fourth grade students of SDN 12 Api-Api after using the model make a match.
With the advent of word embeddings, lexicons are no longer fully utilized for sentiment analysis although they still provide important features in the traditional setting. This paper introduces a novel approach to sentiment analysis that integrates lexicon embeddings and an attention mechanism into Convolutional Neural Networks. Our approach performs separate convolutions for word and lexicon embeddings and provides a global view of the document using attention. Our models are experimented on both the SemEval'16 Task 4 dataset and the Stanford Sentiment Treebank and show comparative or better results against the existing state-of-the-art systems. Our analysis shows that lexicon embeddings allow building high-performing models with much smaller word embeddings, and the attention mechanism effectively dims out noisy words for sentiment analysis.
This paper presents a new method with which to assist individuals with no background in linguistics to create monolingual dictionaries such as those used by the morphological analysers of many natural language processing applications. The involvement of non-expert users is especially critical for under-resourced languages which either lack or cannot afford the recruitment of a skilled workforce. Adding a word to a morphological dictionary usually requires identifying its stem along with the inflection paradigm that can be used in order to generate all the word forms of the new entry. Our method works under the assumption that the average speakers of a language can successfully answer the polar question “is x a valid form of the word w to be inserted?”, where x represents tentative alternative (inflected) forms of the new word w. The experiments show that with a small number of polar questions the correct stem and paradigm can be obtained from non-experts with high success rates. We study the impact of different heuristic and probabilistic approaches on the actual number of questions.
When teaching a foreign language to law students the teacher has to place work emphasis on some peculiarities of the translation of legal terminology.Legal documents have a clearly defined form, which must be preservedduring the translation.Therefore, one of the important matters in the process of students trainingfor professional activity is their appropriate mastering of legal terminology and ability to translatecorrectly.In English there are enough terms having a large number of synonyms.Sometimes there are situations when arises the problem of the translation of nonequivalent lexicon.In the UK there are norms and concepts that are the specifics of this country and the terms, respectively, have no analogues in other languages.Being translated the English terms undergo such types of transformation as: differentiation of the meanings, a specification of the meanings, contents development, the antonymous translation, complete transformation, compensation of losses in the translation process etc. [1, p. 41].The following principles of termscreationshould be mentioned: the principle of the translated terminology, the use of the specific opportunities of the target language, terms formed by a terminologization of common lexicon, the principle of association.When translating the nonequivalent legal terms it is possible to use also a transcoding method (for example, solicitor -,; auditor -; motive ).The descriptive translation is also possible (for example, misdirection; depositions -, ) [2, p. 57].When training law students a foreign language it must be kept in mind that the modern specialist needs to have the level which would allow them to communicate if necessary with the specialists from other countries.For this purposethey have to know the fundamentals of grammar, but, the main thing, they have to know is legal lexicon.That is whyan important role in language training of students is provided to the mastering of professional vocabulary.Mastering of professional lexical units
Large-scale distributed training requires significant communication bandwidth for gradient exchange that limits the scalability of multi-node training, and requires expensive high-bandwidth network infrastructure. The situation gets even worse with distributed training on mobile devices (federated learning), which suffers from higher latency, lower throughput, and intermittent poor connections. In this paper, we find 99.9% of the gradient exchange in distributed SGD is redundant, and propose Deep Gradient Compression (DGC) to greatly reduce the communication bandwidth. To preserve accuracy during compression, DGC employs four methods: momentum correction, local gradient clipping, momentum factor masking, and warm-up training. We have applied Deep Gradient Compression to image classification, speech recognition, and language modeling with multiple datasets including Cifar10, ImageNet, Penn Treebank, and Librispeech Corpus. On these scenarios, Deep Gradient Compression achieves a gradient compression ratio from 270x to 600x without losing accuracy, cutting the gradient size of ResNet-50 from 97MB to 0.35MB, and for DeepSpeech from 488MB to 0.74MB. Deep gradient compression enables large-scale distributed training on inexpensive commodity 1Gbps Ethernet and facilitates distributed training on mobile. Code is available at: https://github.com/synxlin/deep-gradient-compression.
In this study, we presented pictorial representations of happy, neutral, and fearful expressions projected in the eye regions to determine whether the eye region alone is sufficient to produce a context effect. Participants were asked to judge the valence of surprised faces that had been preceded by a picture of an eye region. Behavioral results showed that affective ratings of surprised faces were context dependent. Prime-related ERPs with presentation of happy eyes elicited a larger P1 than those for neutral and fearful eyes, likely due to the recognition advantage provided by a happy expression. Target-related ERPs showed that surprised faces in the context of fearful and happy eyes elicited dramatically larger C1 than those in the neutral context, which reflected the modulation by predictions during the earliest stages of face processing. There were larger N170 with neutral and fearful eye contexts compared to the happy context, suggesting faces were being integrated with contextual threat information. The P3 component exhibited enhanced brain activity in response to faces preceded by happy and fearful eyes compared with neutral eyes, indicating motivated attention processing may be involved at this stage. Altogether, these results indicate for the first time that the influence of isolated eye regions on the perception of surprised faces involves preferential processing at the early stages and elaborate processing at the late stages. Moreover, higher cognitive processes such as predictions and attention can modulate face processing from the earliest stages in a top-down manner.
We systematically explore regularizing neural networks by penalizing low\nentropy output distributions. We show that penalizing low entropy output\ndistributions, which has been shown to improve exploration in reinforcement\nlearning, acts as a strong regularizer in supervised learning. Furthermore, we\nconnect a maximum entropy based confidence penalty to label smoothing through\nthe direction of the KL divergence. We exhaustively evaluate the proposed\nconfidence penalty and label smoothing on 6 common benchmarks: image\nclassification (MNIST and Cifar-10), language modeling (Penn Treebank), machine\ntranslation (WMT'14 English-to-German), and speech recognition (TIMIT and WSJ).\nWe find that both label smoothing and the confidence penalty improve\nstate-of-the-art models across benchmarks without modifying existing\nhyperparameters, suggesting the wide applicability of these regularizers.\n
Virtual reality (VR) has been proposed as a methodological tool to study the basic science of psychology and other fields. One key advantage of VR is that sharing of virtual content can lead to more robust replication and representative sampling. A database of standardized content will help fulfill this vision. There are two objectives to this study. First, we seek to establish and allow public access to a database of immersive VR video clips that can act as a potential resource for studies on emotion induction using virtual reality. Second, given the large sample size of participants needed to get reliable valence and arousal ratings for our video, we were able to explore the possible links between the head movements of the observer and the emotions he or she feels while viewing immersive VR. To accomplish our goals, we sourced for and tested 73 immersive VR clips which participants rated on valence and arousal dimensions using self-assessment manikins. We also tracked participants' rotational head movements as they watched the clips, allowing us to correlate head movements and affect. Based on past research, we predicted relationships between the standard deviation of head yaw and valence and arousal ratings. Results showed that the stimuli varied reasonably well along the dimensions of valence and arousal, with a slight underrepresentation of clips that are of negative valence and highly arousing. The standard deviation of yaw positively correlated with valence, while a significant positive relationship was found between head pitch and arousal. The immersive VR clips tested are available online as supplemental material.
The Web has accumulated a rich source of information, such as text, image, rating, etc, which represent different aspects of user preferences. However, the heterogeneous nature of this information makes it difficult for recommender systems to leverage in a unified framework to boost the performance. Recently, the rapid development of representation learning techniques provides an approach to this problem. By translating the various information sources into a unified representation space, it becomes possible to integrate heterogeneous information for informed recommendation.
Recent papers have shown that neural networks obtain state-of-the-art performance on several different sequence tagging tasks. One appealing property of such systems is their generality, as excellent performance can be achieved with a unified architecture and without task-specific feature engineering. However, it is unclear if such systems can be used for tasks without large amounts of training data. In this paper we explore the problem of transfer learning for neural sequence taggers, where a source task with plentiful annotations (e.g., POS tagging on Penn Treebank) is used to improve performance on a target task with fewer available annotations (e.g., POS tagging for microblogs). We examine the effects of transfer learning for deep hierarchical recurrent networks across domains, applications, and languages, and show that significant improvement can often be obtained. These improvements lead to improvements over the current state-of-the-art on several well-studied tasks.
Discourse parsing is an integral part of understanding information flow and argumentative structure in documents. Most previous research has focused on inducing and evaluating models from the English RST Discourse Treebank. However, discourse treebanks for other languages exist, including Spanish, German, Basque, Dutch and Brazilian Portuguese. The treebanks share the same underlying linguistic theory, but differ slightly in the way documents are annotated. In this paper, we present (a) a new discourse parser which is simpler, yet competitive (significantly better on 2/3 metrics) to state of the art for English, (b) a harmonization of discourse treebanks across languages, enabling us to present (c) what to the best of our knowledge are the first experiments on crosslingual discourse parsing.
Word embeddings that can capture semantic and syntactic information from contexts have been extensively used for various natural language processing tasks. However, existing methods for learning contextbased word embeddings typically fail to capture sufficient sentiment information. This may result in words with similar vector representations having an opposite sentiment polarity (e.g., good and bad), thus degrading sentiment analysis performance. Therefore, this study proposes a word vector refinement model that can be applied to any pre-trained word vectors (e.g., Word2vec and GloVe). The refinement model is based on adjusting the vector representations of words such that they can be closer to both semantically and sentimentally similar words and further away from sentimentally dissimilar words. Experimental results show that the proposed method can improve conventional word embeddings and outperform previously proposed sentiment embeddings for both binary and fine-grained classification on Stanford Sentiment Treebank (SST).
STYX 1.0 is a corpus of Czech sentences selected from the Prague Dependency treebank. The criterion for including sentences into STYX was their suitability for practicing Czech morphology and syntax in elementary schools. The sentences contain both the PDT annotations and the school sentence analyses. The school sentence analyses were created by transforming the PDT annotations using handcrafted rules. Altogether the STYX 1.0 corpus contains 11 655 sentences. Originally, the STYX 1.0 corpus was an inseparable part of the Styx system (http://hdl.handle.net/11858/00-097C-0000-0001-48FB-F)
This paper presents a novel neural machine translation model which jointly learns translation and source-side latent graph representations of sentences. Unlike existing pipelined approaches using syntactic parsers, our end-to-end model learns a latent graph parser as part of the encoder of an attention-based neural machine translation model, and thus the parser is optimized according to the translation objective. In experiments, we first show that our model compares favorably with state-of-the-art sequential and pipelined syntax-based NMT models. We also show that the performance of our model can be further improved by pretraining it with a small amount of treebank annotations. Our final ensemble model significantly outperforms the previous best models on the standard Englishto-Japanese translation dataset.
Participants explored a representative set of 47 solid, fluid and granular materials and rated them according to a list of 32 perceptual and 20 affective attributes. In a principal component analysis (PCA) of the perceptual ratings, we extracted six dimensions: Fluidity, Roughness, Deformability, Fibrousness, Heaviness, and Granularity explained 88% variance. A PCA on affective ratings revealed the dimensions: Valence, Arousal, and Dominance, explaining 92% variance. Greater Valence was significantly associated with reduced Roughness, greater Arousal with more Fluidity and greater Dominance with decreasing Deformability and decreasing Heaviness. Overall, the present study demonstrates that the range of affective responses to touched material is broader than previously assumed, and that these affective responses are systematically associated with certain perceptual dimensions.
This study examined to what extent children and adults differ in how they process negative emotions during reading, and how they rate their own and protagonists’ emotional states. Results show that both children’s and adults’ processing of target sentences was facilitated when they described negative emotions. Processing of spill-over sentences was facilitated for adults but inhibited for children, suggesting children needed additional time to process protagonists’ emotional states and integrate them into coherent mental representations. Children and adults were similar in their valence and arousal ratings as they rated protagonists’ emotional states as more negative and more intense than their own emotional states. However, they differed in that children rated their own emotional states as relatively neutral, whereas adults’ ratings of their own emotional states more closely matched the negative emotional states of the protagonists. This suggests a possible difference between children and adults in the mechanism underlying emotional inferencing.
OBJECTIVE: New MRI sequences based on rapid radial acquisition have reduced gradient noise. The purpose of this study was to compare Silent T1-weighted and unenhanced MR angiography (MRA) against conventional sequences in a clinical population. MATERIALS AND METHODS: The study cohort consisted of 40 patients with suspected brain metastases (median age, 60 years; range, 23-91 years) who underwent T1-weighted contrast-enhanced MRI and 51 patients with suspected vascular lesions or cerebral ischemia (median age, 60 years; range, 16-94 years) who underwent unenhanced intracranial MRA. Three neuroradiologists reviewed the images blindly and rated several measures of image quality on a 5-point Likert scale. Reviewers recorded the number of enhancing lesions and whether Silent images were better than, worse than, or equivalent to conventional images. RESULTS: For T1-weighted MR images, ratings were slightly lower for Silent versus conventional images, except for diagnostic confidence. Although more lesions were detected on conventional images, this difference was not statistically significant; agreement was seen in 88% of cases. In 48% of cases, T1-weighted scans were deemed equivalent, but when a preference existed, it was usually for conventional images (38% vs 14%). Conventional MRA images were rated higher on all image quality metrics and were strongly preferred (reviewers preferred conventional images in 69% of cases, rated the images as equivalent in 27% of cases, and preferred Silent images in 4% of cases). In some cases, artifacts on Silent images caused reduced vessel caliber, vessel irregularities, and even absent vessels. CONCLUSION: Although conventional T1-weighted images were preferred overall, most Silent T1-weighted images were rated as equivalent to or better than conventional images and represent a potential alternative for imaging of noise-averse patients. Silent MRA scored significantly worse and could not be recommended at this time, suggesting that it requires additional refinement before routine clinical use.
BACKGROUND: Formulaic expressions, including idioms and other fixed expressions, comprise a significant proportion of discourse. Although much has been written about this topic, controversy remains about their psychological status. An important claim about formulaic expressions, that they are known to native speakers, has seldom been directly demonstrated. This study tested the hypothesis that formulaic expressions are known and stored as whole unit mental representations by performing three perceptual experiments. METHOD: Listeners transcribed two kinds of spectrally-degraded spoken sentences, half formulaic, and half novel, newly created expressions, matched for grammar and length. Two familiarity ratings, usage and exposure, were obtained from listeners for each expression. Text frequency data for the stimuli and their constituent words were obtained using a spoken corpus. RESULTS: Participants transcribed formulaic more successfully than literal utterances. Usage and familiarity ratings correlated with accuracy, but formulaic utterances with low ratings were also transcribed correctly. Phrase types differed significantly in text frequency, but word frequency counts did not differentiate the two kinds of expressions. DISCUSSION: These studies provide new converging evidence that formulaic expressions are encoded and processed as whole units, supporting a dual-process model of language processing, which assumes that grammatical and formulaic expressions are differentially processed.
Generative neural models have recently achieved state-of-the-art results for constituency parsing. However, without a feasible search procedure, their use has so far been limited to reranking the output of external parsers in which decoding is more tractable. We describe an alternative to the conventional action-level beam search used for discriminative neural models that enables us to decode directly in these generative models. We then show that by improving our basic candidate selection strategy and using a coarse pruning function, we can improve accuracy while exploring significantly less of the search space. Applied to the model of Choe and Charniak (2016), our inference procedure obtains 92.56 F1 on section 23 of the Penn Treebank, surpassing prior state-of-the-art results for single-model systems.
Abstract CzeDLex is a new electronic lexicon of Czech discourse connectives, planned for publication by the end of this year. Its data format and structure are based on a study of similar existing resources, and adjusted to comply with the Czech syntactic tradition and specifics and with the Prague approach to the annotation of semantic discourse relations in text. In the article, we first put the lexicon in context of related resources and discuss theoretical aspects of building the lexicon – we present arguments for our choice of the data structure and for selecting features of the lexicon entries, while special attention is paid to a consistent and (as far as possible) uniform encoding of both primary (such as in English because, therefore ) and secondary connectives (e.g. for this reason, this is the reason why ). The main principle adopted for nesting entries in the lexicon is – apart from the lexical form of the connective – a discoursesemantic type (sense) expressed by the given connective, which enables us to deal with a broad formal variability of connectives and is convenient for interlinking CzeDLex with lexicons in other languages. Second, we introduce the chosen technical solution based on the Prague Markup Language, which allows for an efficient incorporation of the lexicon into the family of Prague treebanks – it can be directly opened and edited in the tree editor TrEd, processed from the command line in btred, interlinked with its source corpus and queried in the PML Tree Query engine. Third, we describe the process of getting data for the lexicon by exploiting a large corpus manually annotated with discourse relations – the Prague Discourse Treebank 2.0: we elaborate on the automatic extraction part, post-extraction checks and manual addition of supplementary linguistic information.