Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
This paper presents the first dependency treebank for Bhojpuri, an Indo-Aryan language. Bhojpuri is one of the resource-poor Indian languages. The objective of the Bhojpuri Treebank (BHTB) project is to provide a substantial, syntactically annotated treebank for Bhojpuri which helps in building language technological tools. This project will also help in cross-lingual learning and typological research. Currently, the treebank consists of 4,881 tokens using the annotation scheme of Universal Dependencies (UD). We develop a Bhojpuri tagger and parser using the machine learning approach. The accuracy of the model is 57.49% UAS, 45.50% LAS, 79.69% UPOS accuracy and 77.64% XPOS accuracy. Finally, we discuss linguistic analysis and annotation process of the Bhojpuri UD treebank.
This paper explores the possibility of improving the performance of specialized parsers for pre-modern Slavic by training them on data from different related varieties. Because of their linguistic heterogeneity, pre-modern Slavic varieties are treated as low-resource historical languages, whereby cross-dialectal treebank data may be exploited to overcome data scarcity and attempt the training of a variety-agnostic parser. Previous experiments on early Slavic dependency parsing are discussed, particularly with regard to their ability to tackle different orthographic, regional and stylistic features. A generic pre-modern Slavic parser and two specialized parsers -- one for East Slavic and one for South Slavic -- are trained using jPTDP (Nguyen & Verspoor 2018), a neural network model for joint part-of-speech (POS) tagging and dependency parsing which had shown promising results on a number of Universal Dependency (UD) treebanks, including Old Church Slavonic (OCS). With these experiments, a new state of the art is obtained for both OCS (83.79\% unlabelled attachment score (UAS) and 78.43\% labelled attachement score (LAS)) and Old East Slavic (OES) (85.7\% UAS and 80.16\% LAS).
The current study looked at the impact of British regional accents on evaluations of eyewitness testimony in criminal trials. Ninety participants were randomly presented with one of three video recordings of eyewitness testimony manipulated to be representative of Received Pronunciation (RP), Multicultural London English (MLE) or Birmingham accents. The impact of the accent was measured through eyewitness (a) accuracy, (b) credibility, (c) deception, (d) prestige, and (e) trial outcome (defendant guilt and sentence). RP was rated more favourably than MLE on accuracy, credibility and prestige. Accuracy and prestige were significant with RP rated more highly than a Birmingham accent. RP appears to be viewed more favourably than the MLE and Birmingham accents although the witness's accents did not affect ratings of defendant guilt. Taken together, these findings show a preference for eyewitnesses to have RP speech over some regional accents.
The Neural Machine-Parsed IcePaHC is a machine-parsed treebank which consists of Icelandic texts from the 13th to 20th century, mostly Icelandic sagas. The texts were parsed using the IceNeuralParsingPipeline, a parsing pipeline which includes an Icelandic model of the Berkeley Neural Parser along with pre- and postprocessing steps. The parser was trained on IcePaHC and the parsing scheme of the treebank is therefore the same, although the treebank does not include empty phrases or lemmas. The treebank includes 52 texts. The total word count is 1,716,429 and the total number of clauses is 167,815.
Among the unusually high number of variants in the three surviving texts of the Old English Life of Saint Mary of Egypt are many instances in which a scribe has changed an inherited reading by substituting one word for another. Many of the substitutions are the result of error or unconscious scribal preference but this article demonstrates that all three texts of the Old English Life, which is of likely Anglian origin, also reveal a pattern of deliberate rewording. This rewording arises from a desire to regularise and “improve” the language of the Life, bringing it more into line with the norms of Late West Saxon, the literary language generally in use in the period when our scribes were at work. No such pattern of substitution is evident in other hagiographical texts in the same manuscripts. The Life of Saint Mary of Egypt was clearly viewed by compilers of late Anglo-Saxon hagiographical manuscripts as a work worthy of inclusion but, unlike other lives, as one in need of some linguistic revision to make it fit in with accepted literary standards.
We present a bracketing-based encoding that can be used to represent any 2-planar dependency tree over a sentence of length n as a sequence of n labels, hence providing almost total coverage of crossing arcs in sequence labeling parsing. First, we show that existing bracketing encodings for parsing as labeling can only handle a very mild extension of projective trees. Second, we overcome this limitation by taking into account the well-known property of 2-planarity, which is present in the vast majority of dependency syntactic structures in treebanks, i.e., the arcs of a dependency tree can be split into two planes such that arcs in a given plane do not cross. We take advantage of this property to design a method that balances the brackets and that encodes the arcs belonging to each of those planes, allowing for almost unrestricted non-projectivity ( 99.9% coverage) in sequence labeling parsing. The experiments show that our linearizations improve over the accuracy of the original bracketing encoding in highly non-projective treebanks (on average by 0.4 LAS), while achieving a similar speed. Also, they are especially suitable when PoS tags are not used as input parameters to the models.
This paper is part of the project Between Lexicon and Grammar (2016–2018), supported by the Grant Agency of the Czech Republic, reg. no. 16-07473S. This project is a follow-up of the project entitled The Grammar-Based Treebank of Czech (2013–2015,cf. Skoumalova et al. 2014; Petkevic et al. 2015a, 2015b) and devoted to automatic parsing driven by a formal HPSG-like grammar of Czech.
Cet article propose d’analyser les apports d’un modele de langue pre-entraine de type BERT (bidirectional encoder representations from transformers) a l’analyse syntaxique en constituants discontinus en anglais (PTB, Penn Treebank). Pour cela, nous realisons une comparaison des erreurs d’un analyseur syntaxique dans deux configurations (i) avec un acces a BERT affine lors de l’apprentissage (ii) sans acces a BERT (modele n’utilisant que les donnees d’entrainement). Cette comparaison s’appuie sur la construction d’une suite de tests que nous rendons publique. Nous annotons les phrases de la section de validation du Penn Treebank avec des informations sur les phenomenes syntaxiques a l’origine des discontinuites. Ces annotations nous permettent de realiser une evaluation fine des capacites syntaxiques de l’analyseur pour chaque phenomene cible. Nous montrons que malgre l’apport de BERT a la qualite des analyses (jusqu’a 95 en F1 ), certains phenomenes complexes ne sont toujours pas analyses de maniere satisfaisante.
This chapter provides an account of the Epicurean theory of language, focusing in particular on Epicurus’ account of the origins of language, as detailed at <italic>Ep. Hdt</italic>. 75–6. It identifies two forms of linguistic naturalism (‘functional’ and ‘referential’) in Epicurus’ account of the first stage of linguistic phylogeny. It goes on to describe the implications of the advent of the second, conventionalist stage for Epicurus’ linguistic naturalism. The chapter suggests that Lucretius (like Epicurus before him) may be considered a latter-day συνειδών, enlarging and improving the language via the introduction and development of new expressions for new philosophical concepts. Finally, it considers how, if at all, Epicurean linguistic norms may have been grounded in Epicurean linguistic naturalism.
Some fragments of Epicharmus are examined in the context of contemporary scholarly discourse, especially with regard to textual and literary criticism, grammar and stylistics as they developed in the Sicily of his time. The interaction of the comedy of Epicharmus with scholarship is complex and involves a reflection and revision of contemporary ideas, but also a process of literary differentiation. This new genre was created through a critical interaction with other genres as part of a process of reinterpretation and literary exegesis. Furthermore, the fragments discussed offer the possibility to speculate on the handling of linguistic norms and also literary standards in pre- and early classical Sicily. They illustrate the interaction between different genres and the way they were integrated into text, metatext and performance to create tension and comic effects.
The article attempts to conduct a primary analysis of the consequences of digital transformation for heritage languages which make up the cultural and historical legacy of individual ethnic communities. In a multilingual society, such a study requires an integrated approach, which takes into account the sociolinguistic parameters of various target audiences, communication channels aimed to disseminate and transfer information, discourse analyses of lin-guistic means, as well as extralinguistic factors impacting the development of different environments. It is equally important to study the specificity of the socio-cultural interaction between communicants in the professional sphere, which primarily indicates the institutional status of participants in communication, as well as their observance / nonobservance of linguistic norms. The latter seems extremely important with regard to heritage languages and their linguistic status in institutional discourse. In many respects, observance / nonobservance of linguistic norms makes it possible, on the one hand, to define the linguistic portrait of the communicant and, on the other hand, to assess the survival of national identity. Both aspects are central across various types of institutional discourse, including political, marketing, ad-vertising discourse etc. The analysis of the institutional aspects of cross-cultural and cross-lingual communication is carried out using an etiological approach that allows to determine the degree of importance of sociolinguistic parameters to achieve adequacy of socio-cultural interaction of representatives of different linguocultures. It is performed indirectly using vari-ous language pairs, in the context of heritage bilingualism, as well as interpersonal interaction. The article also expounds consequences of the global turn towards digital transformation affecting the overall knowledge in liberal arts and human sciences in general and cross-lingual and cross-cultural communication in particular. The study discusses areas of application of heritage language resources such as locus branding, image making, reports of scientific and technical achievements, etc. The article concludes by inferring the need to preserve linguistic diversity and its teleological use in various types of institutional discourse.
In this paper, we aim at improving the study of Latin in three ways: 1) by providing better visualizations of syntagma and structure for both research and the classroom, 2) by supporting a high-level search interface for corpus exploration, and 3) by improving the accuracy of taggers and parsers. To achieve this, we introduce a new linguistic description called Intelligenti Pauca, an alternative to Universal Dependencies for under-resourced languages. We show the key differences between the two linguistic descriptions, how the structure of Intelligenti Pauca favours our goals, and the effect it has on parsing accuracy for the Index Tomisticus Treebank.
The UD framework defines guidelines for a crosslingual syntactic analysis in the framework of dependency grammar, with the aim of providing a consistent treatment across languages that not only supports multilingual NLP applications but also facilitates typological studies. Until now, the UD framework has mostly focussed on bilexical grammatical relations. In the paper, we propose to add a constructional perspective and discuss several examples of spoken-language constructions that occur in multiple languages and challenge the current use of basic and enhanced UD relations. The examples include cases where the surface relations are deceptive, and syntactic amalgams that either involve unconnected subtrees or structures with multiply-headed dependents. We argue that a unified treatment of constructions across languages will increase the consistency of the UD annotations and thus the quality of the treebanks for linguistic analysis.
Objectives. The article deals with semantic and motivational features, an attempt is made to classify pharmacy names in Chernivtsi. The aim of research is to study the names of pharmacies in Chernivtsi as an important component of ergonomics in the onomastic view of the city. Research methods are predetermined by its goals and objectives. The following methods are used in the work: descriptive, analytical, semantic-motivational, component analysis, classification, elements of statistical and method of quantitative calculations. The topicality of the article is determined by the need to study the names of pharmacies in Chernivtsi as an important component of ergonomics throughout the city's onomastic. The scientific novelty of the work is that for the first time the names of pharmacies in Chernivtsi were analyzed as part of the ergonomic picture of the city in terms of semantics, motivation, structure, adherence to linguistic norms in the modern chronological section. Conclusions. Eight main semantic-motivational groups of pharmacists have been identified, the names of pharmacies have been analyzed in terms of the specifics of their formation and functioning, semantic analysis, systematization, motivation, nomination, compliance with language norms. Further perspectives of the study can be seen in the complex study of pharmaconyms and finding out their place in the entire ergonomics of the city.
The study investigates the rule of spelling the root -ravn-/-rovn- and is considered to be a fragment of the academic description of Russian spelling, which is currently being under investigation at the Russian Language Institute of the Russian Academy of Sciences. The authors clarify the meanings that determine the spelling of the unstressed root, supplement the lists of exceptions, denote words with meanings not corresponding to the given values-criteria, and, for the first time in linguistics, investigate the words that can be correlated with different values-criteria, that is, they have double motivation. The rule codifies the spelling of words that have double motivation and fluctuate in usus, dictionaries, study guides and reference books. Spelling recommendations for these words correspond to the current linguistic norm and were approved by the Spelling Commission of the Russian Academy of Sciences in 2019. The linguistic commentary to the rule contains the most significant etymological facts concerning the root -ravn-/-rovn- and summarises the scientific and methodological attempts to figure out the distribution of vocabulary with root -ravn-/-rovn- based on the meanings selected in the spelling rules. In the paper it is shown that the instability in spelling of various verbs with the root -ravn-/-rovn- in modern writing and dictionaries is determined by the double motivation of words, as well as contradictory recommendations and gaps in the rules.
Text discourse parsing plays an important role in understanding information flow and argumentative structure in natural language. Previous research under the Rhetorical Structure Theory (RST) has mostly focused on inducing and evaluating models from the English treebank. However, the parsing tasks for other languages such as German, Dutch, and Portuguese are still challenging due to the shortage of annotated data. In this work, we investigate two approaches to establish a neural, cross-lingual discourse parser via: (1) utilizing multilingual vector representations; and
The article deals with semantic and motivational features, an attempt is made to classify pharmacy names in Chernivtsi. The aim of re- search is to study the names of pharmacies in Chernivtsi as an important component of ergonomics in the onomastic view of the city. Research methods are predetermined by its goals and objecti- ves. The following methods are used in the work: descriptive, anal- ytical, semantic-motivational, component analysis, classification, elements of statistical and method of quantitative calculations. The topicality of the article is determined by the need to study the names of pharmacies in Chernivtsi as an important component of ergonomics throughout the city's onomastic. The scientific novelty of the work is that for the first time the names of pharmacies in Chernivtsi were analyzed as part of the ergonomic picture of the city in terms of semantics, motivation, structure, adherence to linguistic norms in the modern chronologi- cal section. Conclusions. Eight main semantic-motivational groups of pharmacists have been identified, the names of pharmacies have been analyzed in terms of the specifics of their formation and func- tioning, semantic analysis, systematization, motivation, nomina- tion, compliance with language norms. Further perspectives of the study can be seen in the complex study of pharmaconyms and find- ing out their place in the entire ergonomics of the city.
Writing systems play a very important role in human languages, but the mathematical nature of writing systems remainsunderstudied. Here, we conduct a case study of an open-class writing system Chinese characters, which consists of aset of expandable basic units, in contrast to most other writing systems whose basic units form closed sets, or closed-class systems. We demonstrate that probabilistic context-free grammars underlie the representation of Chinese writing, byformalizing Chinese characters as a grammar with character shapes, as nonterminal rules, and components. as terminalnodes. Rule probabilities are estimated from a character treebank of the most frequent 3500 characters. Exploratoryanalysis reveals Zipfian distributions of both shapes and components. Our experiments also demonstrate that Chinesewriting system shows generative powers similar to PCFG, with 78% of the noncharacters generated from our grammarjudged acceptable, which suggests fundamental differences between open-class and closed-class writing systems.
This chapter seeks to evaluate how student users of English are viewed beyond the English-as-school-subject curriculum, both in and out of classrooms. In particular, it exposes some of the tangible effects of ontologies of English in the education context, with important implications for education policy. Despite extensive scholarly work in Applied Linguistics offering positive reconceptualisations of language use in a variety of approaches, such as World Englishes, English as a Lingua Franca, and translanguaging (e.g. Creese and Blackledge, 2010; Hornberger and Link, 2012; García and Wei, 2014; García and Kleyn, 2016), the notion of a 'target' for the learning and teaching of 'good English' for most monolingual mainstream teachers in the United Kingdom remains based on the norms of Standard English, or N-English (to adopt the categorisation terminology proposed by Hall, this volume). For more discussion on the nature of linguistic norms, see Harder (this volume). In this chapter, I show how the presentation of Standard English as the ideal, on the assumption that it is "the language we have in common" (DES, 1988, p. 14), alienates not just multilingual learners of English but also many school children who would regard themselves as first-language English speakers.
Abstract This paper presents a treebank-based study of the effect the text form (prose vs. verse) has on the course of two grammatical changes in Medieval French: the loss of null subjects and the loss of OV word order. By means of statistical analysis, we demonstrate that naive estimates of the spread of overt subjects and VO orders give the impression that there is a significant difference between the rates of development in prose vs. verse. By contrast, estimates based on an abstract grammar competition model which distinguishes between grammar-ambiguous surface forms (overt personal subjects, null subjects in coordination contexts) and grammar-unambiguous surface forms (overt expletive subjects, null subjects in non-coordination contexts) show prose-verse parallelism, prose having an earlier change onset, in line with traditional intuitions. At a more general level, these results suggest that the product of the interaction of a particular grammar with universal pragmatic laws is constant, which can be observed if the factors responsible for variation in grammatical choices are controlled for.
In recent years 360$^{\circ}$ videos have been becoming more popular. For traditional media presentations, e.g., on a computer screen, a wide range of assessment methods are available. Different constructs, such as perceived quality or the induced emotional state of viewers, can be reliably assessed by subjective scales. Many of the subjective methods have only been validated using stimuli presented on a computer screen. This paper is using 360$^{\circ}$ videos to induce varying emotional states. Videos were presented 1) via a head-mounted display (HMD) and 2) via a traditional computer screen. Furthermore, participants were asked to rate their emotional state 1) in retrospect on the self-assessment manikin scale and 2) continuously on a 2-dimensional arousal-valence plane. In a repeated measures design, all participants (N = 18) used both presentation systems and both rating systems. Results indicate that there is a statistically significant difference in induced presence due to the presentation system. Furthermore, there was no statistically significant difference in ratings gathered with the two presentation systems. Finally, it was found that for arousal measures, a statistically significant difference could be found for the different rating methods, potentially indicating an underestimation of arousal ratings gathered in retrospect for screen presentation. In the future, rating methods such as a 2-dimensional arousal-valence plane could offer the advantage of enabling a reliable measurement of emotional states while being more embedded in the experience itself, enabling a more precise capturing of the emotional states.
The research paper aims to give an accurate account of how Kirpal Singh/Kip in The English Patient by Michael Ondaatje copies the socio-cultural and linguistic norms of the Europeans (colonizers) unlike Kipling’s Kim who emulates the Eastern people (colonized) and their culture. They are examples of going through a long drawn process of growing up, looking into the mirror of mimicry. Kip joins the English army as a grown up, learns the need to show affinity to the new culture by way of imitation, adopting their ways to weave a comfort zone. Being different could be an assaulting fact for both sides, Kip is quick to realize that. But his childish view of looking down upon his native culture is the irony of mimicry. It wipes out the original being to rewrite a new identity. Kip leaves the small community sprouted accidentally in the Italian monastery, showing traces of a stricken conscience. Kim, by the virtue of living in close company of Indians, adopts their habits and manners without any qualm, in a most unconscious manner. He never worries to look or sound his original self which he has not experienced for long. Thus, a kind of reverse mimicry is his fate and character when we look at him as an outsider living as an Indian native. The ambivalence of their characters, presented by both, is an interesting aspect of mimicry. In the paper, we have used the views of postcolonial and cultural literary theorists on mimicry, deliberating upon how with the effect of both the processes, Kip and Kim, consciously or unconsciously, get their national identity peeled off, affixing new hybrid identity.
The present study is an approach to Aymara ethnogeometry that aims to identify mathematical terminology on Aymara geometry, under the epistemological support of Ethnomathematics and interculturality, hoping that it can have a positive impact on the learning and identity of students from rural schools in Puno, where there are problems due to linguistic interference. Within the framework of the ethnographic method, the information was obtained through visits and interviews with the Aymara speakers of the communities of the Moho and El Collao province of the Puno region - Peru, contrasted and supplemented with a documentary source Vocabulary of the Aymara language from 1612 and current specialized literature (books and dictionaries). The geometric terms were identified by equivalence and conceptual approximation, and this process showed that the Aymara language has a rich and flexible mathematical lexical background of its own to adapt to current scientific and pedagogical requirements, whether by creating neologisms or borrowing from raw or foreign languages, in case of gaps. The terminology presented in tables was written respecting the linguistic norms of Aymara, and it is expected that it will be standardized and socialized by the authorities of the Ministry of Education, and serve as the basis for future discussions and research.
Following highly negative events, people are deemed resilient if they maintain psychological stability and experience fewer mental health problems. The current research investigated how trait resilience (Block & Kremen, 1996, ER89) influences recovery from anticipated threats. Participants viewed cues (‘aversive’, ‘threat’, ‘safety’) that signified the likelihood of an upcoming picture (100% aversive, 50/50 aversive/neutral, or 100% neutral; respectively), and provided continuous affective ratings during the cue, picture, and after picture offset (recovery period). Participants high in trait resilience (HighR) exhibited more complete affective recovery (compared to LowR) after viewing a neutral picture that could have been aversive. Although other personality traits previously associated with resilience (i.e., optimism, extraversion, neuroticism) predicted affective responses during various portions of the task, none mediated the influence of trait resilience on affective recovery.
Fault diagnosis can be supported by conversational case-based reasoning, but case descriptions are often incomplete. However, users might be able to infer missing information when the system provides context information. We investigated how such information affects ratings, solution times, and learning. Using a simulated case-based reasoning system, participants could ask about the presence of symptoms to find out how well each case matched the current situation. The system answered with yes or no, simply stated that it did not know the answer, or provided three or 15 pieces of context information (suggesting similarity, dissimilarity, or including no relevant information). Only context suggesting similarity increased accuracy, while all types of context increased subjective certainty. The amount of context had little impact: a high amount decreased speed and learning, but much less than expected. Taken together, context information can be helpful, but performance outcomes depend on the type of context provided.
This paper describes the development of Turkic Morpheme web portal, a toolkit that takes into account core features of Turkic languages and meets the requirements for research activities in computational linguistics and typology. This portal was created on the basis of the structural-parametric functional model of the Turkic morpheme and contains special linguistic databases that describe the categories of Turkic languages at different levels: morphological, syntactic, and semantic. The portal can also be used in educational process as a reference system for Turkic languages.
Practical significance: a comparative analysis of the lexical speech development of preschoolers with mental disorders and preschoolers with the norm of speech development is carried out.The following features were found in children with mental disorders: very low cognitive activity, the predominance of substitution and suffixing in the formation of possessive adjectives, difficulties in selecting the appropriate word, ignoring less pronounced features when describing the subject; poor passive and active vocabulary, inaccuracy of the subject and verb vocabulary, difficulties in updating the dictionary, frequent use of neologisms.
The <strong>PTNews Corpus</strong> is a collection of over 19 million tokens extracted from 10 years of political news articles (in Portuguese) from the <strong>Portuguese</strong> newspaper PÚBLICO. The corpus is available under the Creative Commons Attribution-NonCommercial-ShareAlike Licence. The material contained on the PTNews Corpus is © 2010-2020 PÚBLICO Comunicação Social SA. The corpus sizes between the preprocessed version of Penn Treebank (PTB) and WikiText-103. Similarly to WikiText, PTNews has a larger vocabulary than PTB and retains the original case, punctuation and numbers. This corpus contains over 31000 publicly available full articles which makes it well suited for models that can take advantage of long-term dependencies.<br> <br> The corpus is available as a <strong>word-level</strong> collection of articles in two version: the first (ptnews_origin) contains a single file with all the articles in the form: <strong>title</strong>, <strong>URL</strong>, <strong>date</strong>, <strong>body</strong>; the second, contains only the <strong>title</strong> and <strong>body</strong> of the news articles and it is split into <strong>train</strong>, <strong>test</strong>, <strong>validation</strong> sets. In this processed version, the words with less than 3 occurrences are mapped to the <em><unk></em> token. Each sentence in an article body occupies a single line of the dataset and the end of paragraph is marked with the <em><eop></em> tag at the end of a sentence. Portuguese words resulting from contractions like <em>"desta"</em>, ou <em>"nesta"</em> are separated into <em>"d"</em>, <em>"esta"</em>, <em>"n"</em>, <em>"esta",</em> respectively.<br> <strong>Sample article</strong>:<br> <pre><code class="language-bash">Carlos César: Cavaco " cansado e sem entusiasmo " quis afastar responsabilidades sobre a crise https://publico.pt/2010/06/10/politica/noticia/carlos-cesar-cavaco-cansado-e-sem-entusiasmo-quis-afastar-responsabilidades-sobre-a-crise-1441369 2010-06-10 15:38:00 O presidente do Governo Regional dos Açores, Carlos César, considerou hoje que Cavaco Silva esteve " cansado e sem entusiasmo " no discurso do Dia de Portugal, onde afastou responsabilidades sobre a actual crise. <eop> " O país ouviu um Presidente cansado e sem entusiasmo, que andou às voltas com os papéis para dizer que não tinha nada a ver com as razões da crise ", afirmou Carlos César, num comentário à Lusa sobre o discurso do Presidente da República na cerimónia oficial do 10 de Junho, realizada em Faro. <eop> Carlos César considerou, no entanto, " positivo " que Cavaco Silva tenha feito " um discurso alinhado com um tema recorrente na apreciação do momento que vivemos, o da coesão e da corresponsabilização ". <eop> No mesmo sentido, manifestou concordância com o apelo que Cavaco Silva fez " à responsabilidade dos empregadores e empregados ", mas deixou um alerta relativamente à referência do Presidente da República à necessidade de " limpar Portugal ". <eop> Para Carlos César, se essa referência " for despida de conteúdo institucional útil, tratou-se de mais um discurso que se perderá na babugem política d aquilo que Cavaco Silva entendeu recordar como o ' rectângulo ' ". <eop></code></pre> <strong>Reporting Results</strong><br> If you wish to report results or other resources obtained on the PTNews contact Davide Nunes with the following information: <strong>Task</strong>: e.g. Language Modelling, Semantic Similarity, etc; <strong>Publication URL</strong>: url to published article or preprint; <strong>Type of Model</strong>: LSTM Neural Network, n-grams, GloVe vectors, etc; <strong>Evaluation Metrics</strong>: e.g. validation and testing perplexities in the case of language modelling. They will be displayed here <strong>Preprocessed Corpus Statistics</strong> articles: 31.919 articles by split: train: 25.537 test: 3.191 val: 3.191 unique tokens: 68.318 unique OoV Tokens: 76.157 total tokens: 19.021.661 total OoV tokens: 95.043 OoV rate: 0.5% tokens by split: train: 15.242.995 test: 1.895.184 val: 1.883.482 <strong>Contact Information</strong> If you have questions about the corpus or want to report benchmark results, contact Davide Nunes. <br>
Conceptual concreteness and categorical specificity are two continuous variables that allow distinguishing, for example, justice (low concreteness) from banana (high concreteness) and furniture (low specificity) from rocking chair (high specificity). The relation between these two variables is unclear, with some scholars suggesting that they might be highly correlated. In this study, we operationalize both variables and conduct a series of analyses on a sample of > 13,000 nouns, to investigate the relationship between them. Concreteness is operationalized by means of concreteness ratings, and specificity is operationalized as the relative position of the words in the WordNet taxonomy, which proxies this variable in the hypernym semantic relation. Findings from our studies show only a moderate correlation between concreteness and specificity. Moreover, the intersection of the two variables generates four groups of words that seem to denote qualitatively different types of concepts, which are, respectively, highly specific and highly concrete (typical concrete concepts denoting individual nouns), highly specific and highly abstract (among them many words denoting human-born creation and concepts within the social reality domains), highly generic and highly concrete (among which many mass nouns, or uncountable nouns), and highly generic and highly abstract (typical abstract concepts which are likely to be loaded with affective information, as suggested by previous literature). These results suggest that future studies should consider concreteness and specificity as two distinct dimensions of the general phenomenon called abstraction.
In the history of databases, eXtensible Markup Language (XML) has been thought of as the standard format to store and exchange semi-structured data. With the advent of IoT, XML technologies can play an important role in addressing the issue of processing a massive amount of data generated from heterogeneous devices. As the number and complexity of such datasets increases there is a need for algorithms which are able to index and retrieve XML data efficiently even for complex queries. In this context twig pattern matching, finding all occurrences of a twig pattern query (TPQ), is a core operation in XML query processing. Until now holistic joins have been considered the state-of-the-art TPQ processing algorithms, but they fail to guarantee an optimal evaluation except at the expense of excessive storage costs which limit their scope in large datasets. In this article, we introduce a new approach which significantly outperforms earlier methods in terms of both the size of the intermediate storage and query running time. The approach presented here uses Child Prime Labels (Alsubai & North, 2018) to improve the filtering phase of bottom-up twig matching algorithms and a novel algorithm which avoids the use of stacks, thus improving TPQs processing efficiency. Several experiments were conducted on common benchmarks such as DBLP, XMark and TreeBank datasets to study the performance of the new approach. Multiple analyses on a range of twig pattern queries are presented to demonstrate the statistical significance of the improvements.
The article deals with the study of the relevance of the concept “Home” as an integral part of the daily communication of English people during the Late Middle Ages. The purpose of the article is the etymological analysis of the concept and its nominee — “home” and “house”, as well as highlighting the key conceptual lexemes, which formed fixed combinations with the concept in the 14th and 15th centuries. In the study, the authors used descriptive and comparative-historical methods in a diachronic approach. According to A. Jucker’s theory of “Collective Characteristics of the Concept”, each historical period includes a set of concepts capable of revealing both the social structure and the linguistic norms of the culture studied. This confirms the practical significance of the study: by revealing one of the most significant concepts of medieval society, it is possible to immerse deeper into its customs and informal laws. The authors have examined the etymology of the concept on the basis of ancient English, ancient German and Gothic languages. Thanks to this, it is possible to establish the cultural and linguistic field of the concept. The concept “home’’ includes lexemes of figurative, conceptual and value aspects, but it is conceptual that allows, even with a severe lack of material, to observe the verbalization of the concept. By producing an analysis of “home” and “house” based on Oxford and Cambridge dictionaries, the modern meaning of words has been identified. Next, the lexemes applied to the concept in the 14th and 15th centuries, namely, nouns, verbs, and adjectives have been examined and the significance of the diachronic approach to the study of language concepts was confirmed. The study was supported by examples from A. Jucker’s medieval legal and artistic works — “The History of English language and English Historical Lin-guistics” and J. Firbas’ — “On Defining a Theme in Functional Analysis of Proposals”. The perspective of the study is to make a comparison between all concepts in the English culture of the Late Middle Ages and to fully disclose the social dynamics of this historical period.
Przedmiotem badań jest jednojęzyczna leksykografia elektroniczna. Celem artykułujest ukazanie wpływu technik komputerowych na organizację, rozmiar, przeznaczeniei zawartość słowników. W swych badaniach autorka koncentruje się na elektronicznychbazach danych. Definiuje, czym są, oraz objaśnia, jak ich budowa i sposób organizacjizgromadzonych w nich danych wpływają na postać słowników elektronicznych. W artykulezostały poddane analizie trzy współczesne słowniki języka polskiego: Uniwersalny słownikjęzyka polskiego PWN, Wielki słownik języka polskiego PAN oraz Słownik gramatycznyjęzyka polskiego. Autorka dowodzi, że sposób organizacji i prezentacji wiedzy w omówionychdziełach umożliwia użytkownikom korzystanie z nich w sposób zaawansowany,co oznacza sprawne dotarcie do szczegółowych informacji o jednostkach leksykalnych,grupowanie ich, jak również doraźne kompilowanie „podsłowników”, spełniających określoneoczekiwania odbiorców.
As in any field of inquiry that depends on experiments, the verifiability of experimental studies is important in computational linguistics. Despite increased attention to verification of empirical results, the practices in the field are unclear. Furthermore, we argue, certain traditions and practices that are seemingly useful for verification may in fact be counterproductive. We demonstrate this through a set of multi-lingual experiments on parsing Universal Dependencies treebanks. In particular, we show that emphasis on exact replication leads to practices (some of which are now well established) that hide the variation in experimental results, effectively hindering verifiability with a false sense of certainty. The purpose of the present paper is to highlight the magnitude of the issues resulting from these common practices with the hope of instigating further discussion. Once we, as a community, are convinced about the importance of the problems, the solutions are rather obvious, although not necessarily easy to implement.
Focus of the CONcreTEXT task is conceptual concreteness: systems were solicited to compute a value expressing to what extent target concepts are concrete (i.e., more or less perceptually salient) within a given context of occurrence. To these ends, we have developed a new dataset which was annotated with concreteness ratings and used as gold standard in the evaluation of systems. Four teams participated in this first edition of the task, with a total of 15 runs submitted.Interestingly, these works extend information on conceptual concreteness available in existing (non contextual) norms derived from human judgments with new knowledge from recently developed neural architectures, in much the same multidisciplinary spirit whereby the CONcreTEXT task was organized.
This article is an investigation of teachers’ corrections and comments on graduation essays from the 1870s. Our aim is to learn more about the linguistic norms in Swedish secondary schools at the time, such as they appear in this specific context. The study is based on 13 graduation essays (in total 7,184 words) from one secondary school (Högre Allmänna Läroverket i Falun). There are 532 corrections and comments in the texts, corresponding to an average of 7.4 corrections/comments per 100 words. Corrections and comments are thus very common in the texts that we have investigated. Many of the corrections focus on textual details, such as the choice of words and sentence structure, but by no means are all of the corrections due to incorrect language use. Rather, it seems that many of the corrections target the style of essays. Corrections or comments concerning the structure of the texts are very rare. Most of the comments made by teachers are in the category called focusing, i.e. co-appearing with a correction in the text. None of the essays contain positive feedback from the teacher, which we believe could be significant for the correction norm at the time.
Abstract In this paper, A shorter version of the paper appeared in German in the final report of the Digital Plato project which was funded by the Volkswagen Foundation from 2016 to 2019. [35], [28]. we present a method for paraphrase extraction in Ancient Greek that can be applied to huge text corpora in interactive humanities applications. Since lexical databases and POS tagging are either unavailable or do not achieve sufficient accuracy for ancient languages, our approach is based on pure word embeddings and the word mover’s distance (WMD) [20]. We show how to adapt the WMD approach to paraphrase searching such that the expensive WMD computation has to be computed for a small fraction of the text segments contained in the corpus, only. Formally, the time complexity will be reduced from <m:math xmlns:m="http://www.w3.org/1998/Math/MathML"><m:mi>O</m:mi><m:mo>(</m:mo><m:mi>N</m:mi><m:mo>·</m:mo><m:msup><m:mrow><m:mi>K</m:mi></m:mrow><m:mrow><m:mn>3</m:mn></m:mrow></m:msup><m:mo>·</m:mo><m:mo>log</m:mo><m:mi>K</m:mi><m:mo>)</m:mo></m:math> \mathcal{O}(N\cdot {K^{3}}\cdot \log K) to <m:math xmlns:m="http://www.w3.org/1998/Math/MathML"><m:mi>O</m:mi><m:mo>(</m:mo><m:mi>N</m:mi><m:mo>+</m:mo><m:msup><m:mrow><m:mi>K</m:mi></m:mrow><m:mrow><m:mn>3</m:mn></m:mrow></m:msup><m:mo>·</m:mo><m:mo>log</m:mo><m:mi>K</m:mi><m:mo>)</m:mo></m:math> \mathcal{O}(N+{K^{3}}\cdot \log K), compared to the brute-force approach which computes the WMD between each text segment of the corpus and the search query. N is the length of the corpus and K the size of its vocabulary. The method, which searches not only for paraphrases of the same length as the search query but also for paraphrases of varying lengths, was evaluated on the Thesaurus Linguae Graecae ® (TLG ® ) [25]. The TLG consists of about <m:math xmlns:m="http://www.w3.org/1998/Math/MathML"><m:mn>75</m:mn><m:mo>·</m:mo><m:msup><m:mrow><m:mn>10</m:mn></m:mrow><m:mrow><m:mn>6</m:mn></m:mrow></m:msup></m:math> 75\cdot {10^{6}} Greek words. We searched the whole TLG for paraphrases for given passages of Plato. The experimental results show that our method and the brute-force approach, with only very few exceptions, propose the same text passages in the TLG as possible paraphrases. The computation times of our method are in a range that allows its application in interactive systems and let the humanities scholars work productively and smoothly.
Attention mechanisms have improved the performance of NLP tasks while allowing models to remain explainable. Self-attention is currently widely used, however interpretability is difficult due to the numerous attention distributions. Recent work has shown that model representations can benefit from label-specific information, while facilitating interpretation of predictions. We introduce the Label Attention Layer: a new form of self-attention where attention heads represent labels. We test our novel layer by running constituency and dependency parsing experiments and show our new model obtains new state-of-the-art results for both tasks on both the Penn Treebank (PTB) and Chinese Treebank. Additionally, our model requires fewer self-attention layers compared to existing work. Finally, we find that the Label Attention heads learn relations between syntactic categories and show pathways to analyze errors.
Morphologically Rich Languages (MRLs) such as Arabic, Hebrew and Turkish often require Morphological Disambiguation (MD), i.e., the prediction of the correct morphological decomposition of tokens into morphemes, early in the pipeline. Neural MD may be addressed as a simple pipeline, where segmentation is followed by sequence tagging, or as an end-to-end model, predicting morphemes from raw tokens. Both approaches are suboptimal; the former is heavily prone to error propagation, and the latter does not enjoy explicit access to the basic processing units called morphemes. This paper offers an MD architecture that combines the symbolic knowledge of morphemes with the learning capacity of neural end-to-end modeling. We propose a new, general and easy-to-implement Pointer Network model where the input is a morphological lattice and the output is a sequence of indices pointing at a single disambiguated path of morphemes. We demonstrate the efficacy of the model on segmentation and tagging, for Hebrew and Turkish texts, based on their respective Universal Dependencies (UD) treebanks. Our experiments show that with complete lattices, our model outperforms all shared-task results on segmenting and tagging these languages. On the SPMRL treebank, our model outperforms all previously reported results for Hebrew MD in realistic scenarios.
The lack of large and diverse discourse treebanks hinders the application of\ndata-driven approaches, such as deep-learning, to RST-style discourse parsing.\nIn this work, we present a novel scalable methodology to automatically generate\ndiscourse treebanks using distant supervision from sentiment-annotated\ndatasets, creating and publishing MEGA-DT, a new large-scale\ndiscourse-annotated corpus. Our approach generates discourse trees\nincorporating structure and nuclearity for documents of arbitrary length by\nrelying on an efficient heuristic beam-search strategy, extended with a\nstochastic component. Experiments on multiple datasets indicate that a\ndiscourse parser trained on our MEGA-DT treebank delivers promising\ninter-domain performance gains when compared to parsers trained on\nhuman-annotated discourse corpora.\n
This paper attempts to measure the similarity of frequently occurring modal auxiliaries in both L1 and L2 writings. The modal auxiliaries are known to be difficult areas of study for both L1 and L2, due to the overlapping in their use. A subset of TOEFL11 corpus and a subset of PTB(Penn Treebank) corpus were used to capture the similarities among modal auxiliaries within L1 and L2, respectively, and also across L1 and L2. Based on the hypothesis of distributional representation, similarities of modal auxiliaries were computed by applying cosine similarity and mutual information to every pair of words in the corpus, and then finally reducing dimensions with Singular Value Decomposition. The results show that the modals are found to be similar to other modals, and the distribution of modals in the reduced dimensions show that the modals in L1 and those in L2 exhibits different patters of similarities, while the distance among the modals in L1 is further away than the distance among those L2 modals.
Випуск 13.Том 2 and pragmatics is only a small sample of the coinages based on the analysed non-combined and combined word-formation models, used in Modern English media discourse.Media is an amazing source of data for any linguistic research area.It reflects the dynamic changes occurring in the language.An increasing number of lexical innovations, customary found in Modern English media to represent the totality of facts in various spheres of life, are used to organize the message in the media discourse.These findings provide the following insights for further research: to establish the cognitive mechanisms in the process of forming lexical innovations involving a particular word-formation model; to determine the role of lexical innovations in organization of other types of discourse. ФОНЕТИЧНА СТРУКТУРА ФРАНЦУЗЬКИХ ЗАПОЗИЧЕНЬ В АСПЕКТІ БРИТАНСЬКОЇ ВИМОВНОЇ НОРМИ
OBJECTIVE: Superselective pseudocontinuous arterial spin labeling (ss-pCASL) is an MRI technique in which individual vessels are labeled to trace their perfusion territories. In this study, the authors assessed its merit in defining feeding vessels and gauging preoperative embolization feasibility for patients with meningioma, using digital subtraction angiography (DSA) as the reference method. METHODS: Thirty-one consecutive patients with meningiomas were prospectively recruited, each undergoing DSA (and embolization, if feasible) before resection. All ss-pCASL imaging studies were performed 1 day prior to DSA. Two neuroradiologists independently reviewed ss-pCASL images, rating the contribution of each labeled vessel to tumor blood supply as none, minor, or major. Two neuroradiologists also gauged the feasibility of embolization in each patient, based on ss-pCASL images. Interobserver and intermodality agreement were determined using Cohen's kappa statistic. The diagnostic performance of ss-pCASL was assessed in terms of discerning tumor blood supply and the potential for embolization. RESULTS: Interobserver agreement in the rating of blood supply by ss-pCASL was very good (κ = 0.817, 95% CI 0.771-0.863), and intermodality agreement (consensus ss-pCASL readings vs DSA findings) was good (κ = 0.688, 95% CI 0.632-0.744). In delineating tumor blood supply, ss-pCASL showed high sensitivity (87.1%) and specificity (87.2%). The positive and negative predictive values for embolization feasibility were 85.2% and 100%, respectively. CONCLUSIONS: In patients with meningiomas, feeding vessels are reliably predicted by ss-pCASL. This noninvasive approach, involving no iodinated contrast or radiation exposure, is particularly beneficial if there are no prospects of embolization.
The fight against HIV is one of the targets in our century. Thus, among the HIV-infected patients, one of the most dangerous and outstanding with its complications is those with lung pathologies. According clinical staging of the disease, such patients may present Tuberculosis, Pneumocystis jirovecii, Cytomegaloviruses, Candidiasis, Toxoplasmosis etc. The research by scientific research institute of lung disease was carried out among the inpatient individuals in amount of 48.37 (77%) of them were presented with tuberculosis and 11 (23%) with Interstitial Lung Disease (ILD). Studies were presented on HIVpositive patients who were divided by the randomization techniques. Among 37 patients with tuberculosis, 29 (78%) had AFB (acid fast bacillius) with Gexpert, HAIN methods, 6 (22%) were diagnosed by imaging methods (HRCT, chest X-ray) and serum ADA level. According to previous studies, there were no correlations between serum ADA level elevations at HIV-positive patients (p value 0.05). Among 11 patients presented with ILD Pneumocystis jirovecii were detected at 5 (45%), 3 (27.5%) were presented with daily mortality, 3 took a Co-Trimaxozole therapy diagnosed by imaging methods. Clinical effectiveness was approved by the presence of pneumocystis origin. At the second stage of the study was found a correlation between different Cd4 cell count and imaging rating. Thus, among total number of 119 HIVpositive patients, 38 (32%) had infiltration zones, 53 (44%) had a destruction, 20 (17%) dissemination, 8 (7%) mediastinal lymphadenopathy. Statistic results p value 0.000424, thus there is direct correlation. The range of HIV-related lung diseases is wide, and many HIV-related infectious and non-infectious complications are identified. Since HIV infection Pulmonary Complications more than 15 years ago, the life expectancy of antiretroviral therapy ( ART) increased, the existence and incidence of lung complications were not fully characterized. In the current age of ART our knowledge of the global epidemiology of these illnesses is minimal and mechanisms are not completely understood for rises in noninfectious conditions. This article addresses our existing awareness about the effects of HIV on pulmonary health and pulmonary growth. Latest global figures show that about 33 million people live with HIV (1). In low- and middle-income countries, a disproportionate number of people live with the highest prevalence of HIV in subsaharan Africa recorded (1). Within these resource-limited areas, many people with HIV suffer poorly defined severe or fatal complications of their lung. HIV-related lungs are also common causes of disease and death in developed countries, where HIV infection has become a chronic disease because of the increased availability and affordability to selective antiretroviral therapy ( ART). Among ART people, more and more people survive and develop co-morbid diseases that affect mortality considerably, with serious 'non-AIDS' diseases accounting for the majority of deaths in recent studies, in proportion (2-6). This paper discusses the existing awareness of the impact of HIV on lung health, highlights ongoing work in this area and presents the following papers in this issue of the American Thoracik Society proceedings, each of which focusses on common pulmonary and essential complications of HIV infection. HIV infected individuals undergo a gradual but persistent loss of host immunity from infection, leading to immune deregulation, dysfunction and deficiency syndrome (Figure 1). HIV leads to the massive loss of CD4 + effector memory cells from mucosal tissue (7) after initial infection. After this, the effector memory is associated with HIV. The generalized immune activation occurs during the chronic process of HIV infection, and gradually a steady decline in the naive and memory of the T-cell pool contributes to systemic CD4 + lymphocyte depletion (7, 8). T-cells are unstable and they also mount abnormal host reactions to T-cell antigenes. Furthermore, B cell dysfunction results in polyclones, hypergammaglobulinemia and the lack of specific This work is partly presented at 6th International Conference on Chronic Obstructive Pulmonary Disease on May 17-18, 2018 held at Tokyo, Japan Short communication Vol.5, Iss.1 2020 Insights in Chest Diseases antimicrobial responses. Combined, these factors lead to immune dysfunctions, deregulation and degradation of CD4 + lymphocytes with a significant increase in chance infections and other complications during HIV infection. While in those who do not use ART the most abnormalities are immunological, recent data show that while ART restores immune function, inflammation and immunodeficiency that persist, particularly in the case of patients initiating ART in lower lymphocyte numbers of CD4 +. A host defense in the lung, and respiratory tract, that contributes to a greater risk for lung complications is caused by HIV infection in several lines. These changes include defensins in respiratory secretions and defects in the mucociliar structure and soluble protection molecules. In the lung parenchyma, the pathogens can be affected by innate and adaptive immune responses. For example, HIV-infected alveolar macrophages have shown deficiency in pathogen detection. HIV also contributes to chronic stimulation and inflammatory cell activation in the alveolar space. HIV-associated lung complications have a wide scope, considering both infectious and non-infectious complications. Such risks include AIDS-defining diseases and related HIV-related conditions ( e.g. pneumonia[PCP], tuberculosis[TB] or bacterial pneumonia), non-AIDS defining illnesses, but more severe in patients with HIV infection (e.g., lung disease, pulmonary arterial hypertension and chronic obstructive pulmonary diseases) and conditions associated with HIV-related disorder (HIV). There are myriad HIV-related pulmonary disorders, from acute infections to chronic non-communicable diseases. In the age of widespread antiretroviral therapy the epidemiology of these diseases has changed significantly. Evaluating the patient diagnosed with HIV involves determining the severity of the disease and carrying out a detailed but effective definitive diagnosis involving a variety of etiologies at the same time. Substance usage, sexual behaviour, domiciliary and prison status provide the demographic data, including travel and geographic location, important clues to a diagnosis. Many laboratory adjunctive studies may indicate or exclude different diagnoses. Testing pulmonary function (PFT) can help classify many chronic HIVaccelerated non-infected diseases. Chest x-ray and CT scans allow for classifying diseases by pattern of pathognomonic illustration, but there are many infectious disorders, with lower CD4 numbers, in particular. Ultimately, it is also important to make certain diagnoses for sputum, bronchoscopy or lung tissue.
The study was a comparison of general students of promise affect and mathematical students of promise affect after doing a mathematical modeling activity. Participants’ gender (n=160), in grades 7-8, were nearly equal in number (81 girls & 79 boys). After completing a Model-eliciting Activity (MEA) in groups of three, participants completed the 31-item Chamberlin Affective Instrument for Mathematical Problem Solving, hereafter referred to as CAIMPS (Chamberlin, Moore, & Parks, 2017). Using four subconstructs, it was determined that the only statistically significant difference in student affect among the groups was self-esteem and self-efficacy (SS) with the general students of promise group having a mean of 3.43 and the mathematical students of promise group having a mean of 3.76. Implications are that the difference in SS may have surfaced because of the mathematical demands of the problems that ultimately influenced participants’ ratings. Three subconstructs (Attitude Value Interest [AVI], Anxiety [ANX], and Aspiration [ASP]) may not have realized a statistically significant difference because they were not as contingent upon mathematical content knowledge as was SS. The final implication is that similar affective ratings may be an indication that MEAs are similarly suitable for use with groups containing individuals with varying talents.
'Synonym' is an imperative instrument of commonsense knowledge that we apply to make a good sense and sound judgement of our reading. To investigate the ability of machine comprehension models in handling the synonym commonsense knowledge, we developed an innovative approach to automatically generate a dataset based on the Stanford Question Answering Dataset (SQuAD 2.0). The brand-new dataset consists of additional distracting sentences or questions spawned using synonym commonsense knowledge. We formulated new questions by replacing noun entities of the original ones in SQuAD 2.0 with their synonyms. This approach followed the two fundamental principles of SQuAD 2.0 dataset: relevancy and plausibility (incorrect answers are more challenging if they are relevant and plausible). It improves the robustness/abstraction of the question set. To improve the synonym selection strategy in Word Sense Disambiguation (WSD) problem, we designed a new algorithm Multiple Source Adapted Lesk Algorithm (MSALA). Rather than only using WordNet as the source of gloss for adapted Lesk algorithm, we used both lexical database WordNet and commonsense database ConceptNet. This fusion provides a rich hierarchy of semantic relations for the MSALA algorithm. Using this method, we devised 11,000 questions and evaluated the performance of the state-of-the-art question answering system-BERT. Our result shows that the accuracy of the contemporary BERT-Base model dropped from 74.98% to 63.24%. This 10+% accuracy drop revealed the limitations of BERT in handling synonym commonsense knowledge.
International audience
We analyse and explain the increased generalisation performance of iterate averaging using a Gaussian process perturbation model between the true and batch risk surface on the high dimensional quadratic. We derive three phenomena \latestEdits{from our theoretical results:} (1) The importance of combining iterate averaging (IA) with large learning rates and regularisation for improved regularisation. (2) Justification for less frequent averaging. (3) That we expect adaptive gradient methods to work equally well, or better, with iterate averaging than their non-adaptive counterparts. Inspired by these results\latestEdits{, together with} empirical investigations of the importance of appropriate regularisation for the solution diversity of the iterates, we propose two adaptive algorithms with iterate averaging. These give significantly better results compared to stochastic gradient descent (SGD), require less tuning and do not require early stopping or validation set monitoring. We showcase the efficacy of our approach on the CIFAR-10/100, ImageNet and Penn Treebank datasets on a variety of modern and classical network architectures.
Existing research shows that “pleasant” or “unpleasant” moods can be primed by presenting participants with “pleasant” or “unpleasant” images (Avero &amp; Calvo, 2006), and that stronger priming effects are induced by images as opposed to text (Powell et al., 2015). However, no previous research shows whether or not mood induction effects may differ based on image presentation format. Therefore, the present work aimed to test this hypothesis, by presenting participants (N = 145) with either standalone or grouped images, displaying either positive or negative facial expressions. We found that both facial expression and image presentation had a significant effect on participants’ average ratings of the emotional valence of the images, including a significant interaction effect. However, only facial expression had a significant effect on mood change. We found a slight correlation (r =.298) between image rating and mood change, suggesting that image presentation may have a slight effect on mood change that was unable to be observed in this small-scale study.