Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Dependency parsing aims to identify syntactic dependencies between words in a sentence.Dependency parsing can provide syntactic features and improve model performance for tasks such as information extraction,automatic question answering and machine translation.The training data size has an significant impact on the performance of the dependency parsing model.The lack of training data will cause serious unknown word problems and model over-fitting problems.This paper proposes various data augment strategies for the problem of low-resource dependency parsing.The proposed method effectively expands the training data by synonym substitution and alleviates the unknown words problem.The data augment strategies of multiple Mixups effectively alleviate the model overfitting problem and improve the generalization ability of the model.Experimental results on the universal dependencies treebanks(UD treebanks) dataset show that the proposed methods effectively improve the performance of Thai,Vietnamese and English dependency parsing under small-scale training corpus conditions.
The advancement of machine learning and artificial intelligence has created potential for the automation of the time-consuming process of analyzing, evaluating, and interpreting large quantities of imagery. In order for the large-scale deployment of artificial intelligence in image analysis to be made possible, a new image rating system akin to the more traditional National Interpretability Rating Scale (NIIRS) must be created, so as to aid in the defining of system requirements in the development of new sensors. Through prototyping models using the open-source object detection library Detectron2 alongside proxy imagery from the Overhead Imagery Research Data Set (OIRDS), current limitations faced by machine learning algorithms in the context of target identification in aerial imagery are brought to light, including how they may affect a new image-rating scale, and insights into how such limitations could be overcome.
The workshop will explore the potential of Knowledge Organization Systems (KOS), such as classification systems, taxonomies, thesauri, ontologies, and lexical databases, in the context of current developments and possibilities. These tools help model the underlying semantic structure of a domain for purposes of information retrieval, knowledge discovery, language engineering, etc. The workshop provides an opportunity to discuss projects, research and development activities, evaluation approaches, lessons learned, and research findings. The main theme of the workshop is Designing for Cultural Hospitality and Indigenous Knowledge in KOS.
A previous experiment tested subjects' new/old judgments of previously-studied faces, distractors, and morphs between pairs of studied parents. We examine the extent to which models based on principal component analysis (eigenfaces) can predict human recognition of studied faces and false alarms to the distractors and morphs. We also compare eigenface models to the predictions of previous models based on the positions of faces in a multidimensional "face space" derived from a multidimensional scaling (MDS) of human similarity ratings. We find that the error in reconstructing a test face from its position in an "eigenface space" provides a good overall prediction of human familiarity ratings. However, the model has difficulty accounting for the fact that humans false alarm to morphs with similar parents more frequently than they false alarm to morphs with dissimilar parents. We ascribe this to the limitations of the simple reconstruction error-based model. We then outline preliminary wo...
The Universal Knowledge Core (UKC) is a large multilingual lexical database with a focus on language diversity and covering over a thousand languages. The aim of the database, as well as its tools and data catalogue, is to make the somewhat abstract notion of diversity visually understandable for humans and formally exploitable by machines. The UKC website lets users explore millions of individual words and their meanings, but also phenomena of cross-lingual convergence and divergence, such as shared interlingual meanings, lexicon similarities, cognate clusters, or lexical gaps. The UKC LiveLanguage Catalogue, in turn, provides access to the underlying lexical data in a computer-processable form, ready to be reused in cross-lingual applications.
The processability account anticipates that learners will make more underpassivization errors than overpassivization errors since passivization entails more processing. Although one study on psych-verbs and a few on unaccusatives examined Turkish L2 learners’ acquisition, no research compared a single set of learners’ acquisitions of these verbs together from a processing point of view. In this regard, the current study aims to investigate whether the processing complexity of passivization influences acquisition of psych and unaccusative verbs. It also questions whether general accuracy levels in Grammaticality Judgement Task (GJT) and degree of familiarity with target verbs are related to their level of accuracy with individual psych and unaccusative verbs. 33 undergraduate-level university students performed on the GJT and a Word Familiarity Rating Task (WFRT). The GJT included 38 items with 12 sentences for psych-verbs, 12 sentences for unaccusative-verbs, 12 sentences for distracters and 2 sentences for examples. The WFRT was a survey questioning familiarity with 6 psych and 6 unaccusative verbs. To analyse the data, a set of nonparametric tests and descriptive statistics were used. The results revealed that learners performed mor The processability account anticipates that learners will make more underpassivization errors than overpassivization errors since passivization entails more processing. Although one study on psych-verbs and a few on unaccusatives examined Turkish L2 learners’ acquisition, no research compared a single set of learners’ acquisitions of these verbs together from a processing point of view. In this regard, the current study aims to investigate whether the processing complexity of passivization influences acquisition of psych and unaccusative verbs. It also questions whether general accuracy levels in Grammaticality Judgement Task (GJT) and degree of familiarity with target verbs are related to their level of accuracy with individual psych and unaccusative verbs. 33 undergraduate-level university students performed on the GJT and a Word Familiarity Rating Task (WFRT). The GJT included 38 items with 12 sentences for psych-verbs, 12 sentences for unaccusative-verbs, 12 sentences for distracters and 2 sentences for examples. The WFRT was a survey questioning familiarity with 6 psych and 6 unaccusative verbs. To analyse the data, a set of nonparametric tests and descriptive statistics were used. The results revealed that learners performed more accurately on unaccusatives than on psych-verbs. They did more underpassivization errors by accepting ungrammatical active constructions of psych verbs. Their performances on psych and unaccusative verbs went parallel with their general accuracy levels in GJT while their degree of familiarity with and accuracy level for two verbs do not correlate with each other.The results suggest that such factors as processability and L1 transfer seem to impact the acquisition. Keywords:Second language acquisition; psych verbs; unaccusative verbs; underpassivization; overpassivization. e accurately on unaccusatives than on psych-verbs. They did more underpassivization errors by accepting ungrammatical active constructions of psych verbs. Their performances on psych and unaccusative verbs went parallel with their general accuracy levels in GJT while their degree of familiarity with and accuracy level for two verbs do not correlate with each other.The results suggest that such factors as processability and L1 transfer seem to impact the acquisition. Keywords:Second language acquisition; psych verbs; unaccusative verbs; underpassivization; overpassivization.
TIME is a highly abstract concept and prevalent in languages worldwide. Cross-cultural and cross-linguistic research suggests that TIME is embodied dissimilarly in different languages. Still the literature has not received sufficient attention in examining the differences. This study aimed to identify and compare how TIME is metaphorically represented and embodied worldwide. We investigated 14 languages; Arabic, Assamese, Chinese, English, Finnish, French, German, Japanese, Kikuyu, Persian, Polish, Russian, Spanish, and Swedish, which represent nine language families. The metaphors were categorized conceptually as TIME IS AN ORGANISM, TIME IS MOTION, TIME IS SPACE, and TIME IS VALUABLE to see how universally these concepts were embodied. We employed a two-part paper-based task. The first part consisted of generation of metaphor items and the second part consisted of a valence rating task. 'The key variables considered were 'metaphor category' and 'language family' while controlling for demographic variables such as gender, age and handedness. Data from 513 participants were collected. Results showed a significant association between language categories and the valences of time metaphors. The data of this study suggest that within the languages of a certain category, there might be some similarity between the valences of words that are used to realize a given conceptual metaphor.
The aim of this paper is to present a comprehensive description of the ideophones of Teko, a Tupi language spoken in French Guiana. This word class, previously only briefly described, is defined in this paper through a systematic comparison to nouns and verbs, at various levels: phonology, word structure, prosody, semantic, morphology, syntax and discourse use. When relevant, it offers quantitative analyses with statistical tests to support the comparison. The analysis is based on a lexical database of 177 ideophones, 420 occurrences in texts and a subset of 101 tokens with audio-recording. It particularly investigates in detail various aspects of prosody, including syllabic structure, pitch, intensity and duration, and pauses. Contrary to the common view on ideophones that postulates a rather marginal status of the latter, this paper shows that ideophones are in fact rather well integrated in the linguistic system of Teko. They yet show regularities that require to consider them a distinct word category. This paper aims at contributing to the growing literature on the cross-linguistic definition and description of ideophones.
A large number of tone languages are distributed throughout Asia. However, despite an abundant amount of descriptive studies on tone inventories of individual languages in the region, our understanding of the characteristics of tones in Asian languages still remains limited due to the paucity of comparative research. Adopting phonological typology as a tool to measure the complexity of tone systems, this chapter presents quantitative analyses of tones collected from the cross-linguistic database and published data on Asian tone languages. It finds that while the distribution of tone types is inversely related with markedness, Asian tone languages generally have more contour tones with larger tone inventories, and that a higher pitch register tends to be utilized more frequently in various tone types. This chapter further addresses the tonal diversity of less known Asian languages that awaits further research.
The language errors could potentially hinder the transmission of information and disturb theconcretization of the act of communication. In the contemporary context, many factors exercise influence on thelanguage especially the social media triggering disregard of the linguistic norms, devaluing the correct use of thelanguage and disrupting the normal flow of communication. These irregularities in the use of the Spanish languageare known as language vices. The subject of this paper is a broad view of the irregularities (that could be of differentgrammatical nature) that could result from the improper use of the Modern Spanish language and encompasses thedefinition, review of the most common types of language vices (pleonasm, solecism, amphibology, barbarism,archaism, and neologism) supported by a series of examples excerpted from relevant sources among whichOrtografía de la lengua Española (RAE:2010). Diccionario panhispánico de dudas (RAE:2005), Gran diccionariode sinónimos, voces afines e incorrecciones (Fernando Corripio, Bruguera, 1979), etc. The purpose of the paper is,in absence of Spanish grammar(s) in Macedonian language and published comparative papers in particular in thefield of lexicology, to fill to some extant the vacuum and shed light on the language vices in the Modern Spanishlanguage, thus arousing the interest of the Macedonian linguists, especially those with knowledge of Spanish forfuture more profound research and their systematization.
Abstract Face context effect refers to the effects of emotional information from the surrounding context on the face perception. Numerous studies investigated the face context effects by exploring the effect of suprathreshold or subthreshold emotional context on the perception of neutral face, but no consistent conclusions have been drawn. Hence, we explored cognitive mechanisms underlying face context effects by comparing the effects of suprathreshold and subthreshold emotional contexts on neutral face perception. In Experiment 1, we investigated the mechanisms underlying the valence-based face context effect by comparing the effect between suprathreshold (1a) and subthreshold (1b) emotional contexts with different valences on neutral faces. In Experiment 2, we investigated the mechanisms underlying the type-base face context effect by comparing the effect between suprathreshold (2a) and subthreshold (2b) emotional contexts with different emotional types on neutral faces. The results of experiment 1 revealed significant differences in valence ratings of neutral faces under suprathreshold and subthreshold emotional contexts with different valences. The results of experiment 2 showed that the emotional-dimension ratings of neutral faces was significantly different under suprathreshold emotion-specific contexts but not subthreshold emotion-specific contexts. We concluded that the mechanism of the valence-based face context effect is different from that of the type-based face context effect. The former is more automatic, and the latter is more non-automatic.
The relationship between emotion and cognition is a central topic of interest in psychology. In the present study we focus specifically on the relationship of the affect dimensions of valence and activation with construal level. Although several studies have looked into this relationship, findings have been inconsistent. We propose that these inconsistencies may result from the fact that valence and activation, which have been typically linked with construal level separately, are not independent of each other. Specifically, we propose that emotion valence and activation interact in their effects on construal level, such that activation moderates the relationship between emotion valence and construal level. The relationship between the two increases the higher the emotion activation. We test this interaction effect through a text analysis of real-life descriptions of events. The content of 330,853 TripAdvisor online hotel reviews were analyzed, using the: (a) Affective Norms for English Words, (b) Warriner et al.’s Emotion Norms, and (c) Brysbaert et al.’s Concreteness Ratings. Results confirm the predicted U-shaped relationship between valence and activation, and the interaction effect between valence and activation on construal level.
Philosophers of science have long argued that when evaluating explanations, we do not consider ideas in isolation. Instead, we possess an integrated web of information that comprises the context we consider when weighing evidence about any component of this web. In this paper, we provide empirical evidence that theories are considered in context by demonstrating that non-scientists change the strength of their belief in both of two alternative theories, even when only given information about one of these hypotheses. In addition, we seek to identify and describe some of the types of information people use in evaluating theories. Information about mechanism, inferences that discriminate between two explanations, and information about closely related situations in which the target factor operates as a mechanism can all significantly affect ratings of two rival explanations.
Corpus linguistics and computational approaches to language constitute an important trend in today’s linguistics, and Slavic historical linguistics is no exception. This chapter serves as an empirical touchstone for the entire volume. Using parallel Greek and Old Church Slavonic data from the PROIEL/ TOROT treebanks, the first attested state of the phenomena covered in the volume is explored, including their relationship to the Greek sources. The chapter covers accusatives with infinitives (Gavrančić this volume, Tomelleri this volume), absolute constructions (Mihaljević 2017), deverbal nouns (Tomelleri this volume), prepositional phrase connectors (Kisiel & Sobotka this volume), numeral syntax (Słoboda this volume), the ordering of pronominal clitics (Kosek, Čech & Navratilova this volume), tense use in performative declaratives (Dekker this volume) and relative clauses (Sonnenhauser & Eberle this volume; Podtergera 2020). The chapter presents corpus statistics on each of the phenomena, and a brief discussion of the possibility of influence from Greek. The chapters that provide their own studies of Old Church Slavonic data (Fuchsbauer this volume on “mock” articles, Pichkhadze this volume on syntactic blocking and Šimić this volume on negative concord), are not replicated, but brought into the discussion when relevant.
In the educational process, the effects of the communication function on the change of meaning of the ideas expressed in teaching are important in today's teaching and learning. Considering that sometimes misunderstandings arise between the parties involved in the learning process (teacher and student, teacher and learner). The article is about the phenomenon of the “language game”, which is one of the topical paradigms of recent times. The famous philosopher L.Wittgenstein was the author of this term, which is included in the science of linguistics. Of course, similar terms can be found in the works of other famous philosophers, linguists, psychologists. But with the appearance of the term “language game”, L.Wittgenstein seems to have restored the bridge between language and philosophy, which existed for a long time, but which remained due to some problems since centuries. Since pragmatics, which is one of the aspects of the sign system, examines the points related to the activity of the language, the “language game” is also considered one of the main points in the center of its research. Although “language games” by their general appearance and even their origin resemble the structural form of ordinary speech acts, they are completely distinguished from them in terms of a number of features. The main characteristic of “language games” is the disruption they cause in linguistic norms. Deliberately, purposefully violated rules have a special effect on the semantic load of each expression in the encoding-decoding process. As a result, the illocutionary force of the locative act formed in the form of a “language game” differs from the illocutionary force of ordinary speech acts. An element of illocutionary force that is normally present acquires a dual character. The illocutionary act, which has a double power, makes the communication process very interesting on the one hand, and complicated on the other hand, compared to ordinary speech acts. In this article, we will try to extensively investigate the effects of the communication function in the educational process on the change of meaning of the ideas expressed in the teaching, based on the educational materials.
Previous research suggested the possibility of establishing systematic links between some intrinsic features of a presupposition and textual functions that it can carry out with greater probability. This study starts from an organic analysis of the semantics of presupposition triggers, thanks to a lexical database comprising 19,500 entries. The database was used to investigate a corpus of chat conversations that included about 200,000 tokens, with the general objective of exploring possible pragmatic values inside bidirectional interactions, thereby verifying the effects on the audience and the negotiation of the presupposed content. In the corpus, triggers mainly occur as noninformative, maintaining information already known by all those participating in the communication; but, depending on their different features, some of them are systematically associated to a function of anaphora and textual cohesion, while others strengthen social conventions and stereotypes. The informative uses, although in a minority proportion, are the most interesting category, where the presupposed content corresponds to the 'New' in the Given/New dichotomy, thus generating a misalignment between what is actually implicit and what is conveyed as such. This misalignment appears to be an error in just a few cases and it corresponds to a low number of reader reactions: particular conditions in pragmatics and in the social context must be called into question to interpret the data.
Part-of-speech (PoS) tagging constitutes a common task in Natural Language Processing (NLP) given its widespread applicability. However, with the advance of new information technologies and language variation, the contents and methods for PoS-tagging have changed. The majority of Italian existing data for this task originate from standard texts, where language use is far from multifaceted informal real-life situations. Automatic PoS-tagging models trained with such data do not perform reliably on non-standard language, like social media content or language learners’ texts. Our aim is to provide additional training and evaluation data from language learners tagged in Universal Dependencies (UD), as well as testing current automatic PoS-tagging systems and evaluating their performance on such data. We use Italian texts from a multilingual corpus of young language learners, LEONIDE, to create a tagged gold standard for evaluating UD PoS-tagging performance on non-standard language. With the 3.7 version of Stanza, a Python NLP package, we apply available automatic PoS-taggers, namely ISDT, ParTUT, POSTWITA, TWITTIRÒ and VIT, trained with diversified data, on our dataset. Our results show that the above taggers, trained on non-standard data or multilingual treebanks, can achieve up to 95% of accuracy on young multilingual learner data, if combined.
The quality of features is one of the main factors that affect classification performance. Feature selection aims to remove irrelevant and redundant features from data in order to increase classification accuracy. However, identifying these features is not a trivial task due to a large search space. Evolutionary algorithms have been proven to be effective in many optimization problems, including feature selection. These algorithms require an initial population to start their search mechanism, and a poor initial population may cause getting stuck in local optima. Diversifying the initial population is known as an effective approach to overcome this issue; yet, it may not suffice as the search space grows exponentially with increasing feature sizes. In this study, we propose an enhanced initial population strategy to boost the performance of the feature selection task. In our proposed method, we ensure the diversity of the initial population by partitioning the candidate solutions according to their selected number of features. In addition, we adjust the chances of features being selected into a candidate solution regarding their information gain values, which enables wise selection of features among a vast search space. We conduct extensive experiments on many benchmark datasets retrieved from UCI Machine Learning Repository. Moreover, we apply our algorithm on a real-world, large-scale dataset, i.e., Stanford Sentiment Treebank. We observe significant improvements after the comparisons with three off-the-shelf initialization strategies.
One of the main goals of the national curriculum is to develop oral speech culture in students, to develop basic speech skills - writing, reading, listening, speaking. In the standard of Georgian language and literature, one of the most important directions of teaching is the development of oral speech, because this direction combines two closely related speech behaviors: listening and speaking.The consistent development of skills related to these speech behaviors is aimed at forming a person who is ready for free, modern communication. This involves the development of oral communication, expression skills, listening skills, social communication and interactive skills. At school, the student learns the results of social experience through textbooks, where the necessary information is given in linguistic form. Language is needed so that a person can express his feelings, emotions, mysterious thoughts with its help.A modern school should provide the upbringing of such a person who possesses not only knowledge, but is able to use this knowledge in life. The goal is not for the student to know as much as possible, but to be able to act adequately in any situation and solve problems.We have to teach the child the spoken language, to master the linguistic norms that have been established in the language throughout history, a child's speech and views are formed by the adoption of these norms, a person needs speech at every step, in every business. After all, speech is a necessary means of expressing thoughts, communicating between people, and perfecting the inner world of a person; It is a type of human action that realizes thinking based on the use of linguistic means; Performs the function of communication, emotional self-expression and influence on other people. Well-developed speech is an important means of active human activity in modern society, and for a student it is a necessary condition for successful learning at school.
Trees are among the most studied data structures and several techniques have consequently been developed for comparing two trees belonging to the same category. Until the end of year 2020, there was a serious lack of suitable metrics for comparing two weighted trees or two trees from different categories. The problem of comparing two tree sets was not also specifically addressed. These limitations have been overcome in a paper published in 2021 where a customizable metric based on hidden Markov models has been proposed for comparing two tree sets, each containing a mixture of trees belonging to various categories. Unfortunately, that metric does not allow the use of non metric-dependent classifiers which take descriptor vectors as inputs. This paper addresses this drawback by deriving a descriptor vector for each tree set using meta-information related to its corresponding models. The comparison between two tree sets is then realized by comparing their associated descriptor vectors. Classification experiments carried out on the databases FirstLast-L (FL), FirstLast-LW (FLW) and Stanford Sentiment Treebank (SSTB) respectively showed best accuracies of 99.75%, 99.75% and 87.22%. These performances are respectively 40.75% and 20.52% better than the tree Edit distance respectively for FLW and SSTB. Additional clustering experiments exhibited 54.25%, 98.75% and 75.53% of correctly clustered instances for FL, FLW and SSTB. No clustering was performed in existing work.
Introduction The American Academy of Ophthalmology (AAO) Task Force on Academic Global Ophthalmology was convened to catalyze development of AAO guidance and perspectives related to clinical service, education, and research initiatives within the field of Global Ophthalmology. The term “Global Ophthalmology” represents an expansive, cross-cutting, yet growing field with initiatives that may range from teaching to clinical service, to research facilitated through in-country partnerships. This article highlights the why, what, and how related to research and investigation in international settings. While much of this article focuses on low-income and middle-income countries (LMICs), multiple principles may be appropriate to research in both resource-limited and resource-replete settings (eg, Ministry of Health and Sanitation partnerships; engagement with local stakeholders; ethical and sociocultural assessment before the initiation of a project). Prior seminal studies in ophthalmology have highlighted the impact of vision science research in LMICs on both vision and systemic health, particularly where leading discoveries have directly led to actions. Examples include seminal work in vitamin A deficiency and mortality,1 ivermectin for onchocerciasis,2,3 azithromycin for the management of trachoma,4 and deployment of manual small incision cataract surgery at scale.5,6 While these bodies of work represent significant advances in our understanding of the burden and treatment of eye disease in LMICs, many vital research questions remain. The Lancet Global Health Commission on Global Eye Health defined eye health as “maximized vision, ocular health and functional ability, thereby contributing to overall health and wellbeing, social inclusion, and quality of life.”7 The contribution of vision health to individual well-being includes improved educational outcomes, work productivity, and reduced disparities. Moreover, numerous studies have shown an association between vision impairment and an increased mortality risk.8–10 To improve eye health globally, survey data and robust indicator data are needed, as well as both discovery and implementation science research.7 In this article, we synthesize recent literature related to the topic of global vision science research drawing from the experience of other surgical disciplines and public health experts to describe principles of research in LMICs and global settings, learnings from experiences in the field, and unmet needs. Case-based examples highlight some of the principles that may help to facilitate engagement for trainees, ophthalmologists, and health care providers who aim to answer important questions related to global ophthalmic health. Sociocultural Context and Ethics Performing culturally competent global vision research demands a clear understanding and appreciation of sociocultural context, local governmental and ethics regulations, and a commitment to local partnerships and engagement. There is no established set of guidelines to direct well-intentioned vision researchers in these pursuits. Rather, culturally competent and ethical global vision research relies on strong local partnerships built on mutual trust. Sociocultural context includes the set of cultural and linguistic norms that shape how ethical research is to be conducted, how research is perceived locally, and the willingness of individuals to participate in research, among myriad of other factors. The context in which research is conducted may also shape bias, including social desirability, acquiescence, and observation biases. Thus, an inadequate understanding of context may not only threaten the ethical and cultural competence of global vision research, but also its validity. Global vision research must include local partners in all stages, from initial needs assessments and project planning to publication. An emphasis on partnership highlights not only the involvement of local researchers, but also the need for equity between colleagues who may hail from the Global South and North. The long history of colonialism and neocolonialism that impacts these partnerships cannot be ignored. There is a growing body of literature that points to inequities in scientific authorship and recognition. Accordingly, issues like authorship should be addressed at the outset of a collaboration and guided by each individual’s role in the project, including the contribution of invaluable local knowledge. Local partners are best positioned to devise and conduct research that is contextually competent. Too often, well-intentioned investigators may plan to address research questions with minimal local relevance. In such cases, the burden on participants and any risks may not be justifiable. It is obligatory to obtain local regulatory approval before commencing global vision research. Foreign ethics review boards are not likely to be familiar with local norms, acceptable practices, or local research priorities. In most cases, foreign researchers should also obtain regulatory approval from their own institution after local approvals are in place. These procedures exist to ensure that the ethics of the research and the collaboration meet institutional and community standards. Implementation Implementation of research requires multidisciplinary teams and partnerships between local researchers, key community stakeholders, governing bodies (eg, Ministry of Health), and community members to create effective research groups with a mutual understanding and shared goal to improving the health outcomes of key beneficiaries.11 Laying the groundwork with strong equitable partnerships is key to successful research initiatives. When developing and designing the study it is important to include local stakeholders and to maintain these relationships along with “shared decision making.” Shared decision making upholds the value that everyone on the team has an equal voice despite power disparities, as well as the need for joint consensus12 when decisions are made. Another key ideal is creating equity among partnerships. This encompasses not only shared decision making, but also a framework that builds capacity, ensures development of local stakeholders, and takes into account inclusive decision making with viewpoints from all.13 Measuring equity can occur through practices such as capacity building, assessing tasks within a team, data sharing, practices of dissemination of research, authorship, and funding transparency.13 Effective team building and collaboration are paramount for implementing a study, but there are other practical considerations to conducting research. Funding needs to be secured and must be adequate to conduct all phases of the study including capacity building. Flexibility with budgets and approval from all partners is imperative. A research study needs to be approved by local governing institutional review board as noted above. Drafting of protocols should include local stakeholders and key researchers to ensure the protocol and consent fit within the norms of the community. To promote respect, it is important to understand the economic, social, and political climate where research activities take place. This includes challenging locations such as outbreak or conflict zones, as well as an understanding of environmental patterns and optimal timings for meetings and travel for study staff. Other local matters that might impact research include political elections, religious celebrations, and holidays. Logistic considerations include laboratory capacity, equipment, electrical needs, and transportation. Research workflows and training should promote skill transfer, with mutual benefits for all stakeholders and the long-term impact of building capacity and sustainability.14 Culture, Language, and Communication There is enormous diversity of languages around the world, and this has important implications for conducting global health research. Communications within the research team may be made challenging when some members are more facile in either the dominant scientific language (often English or another European language) or the local language. Thus, linguistic challenges may arise for both local and foreign team members. In fact, such issues may exist even among study team members from the same country, since in many places there are regional or local languages and dialects that may be used to communicate with research participants or even between study team members, but that are not universally spoken nationally. The implications of these challenges should be addressed early on to mitigate any impact on scientific rigor, contributions to the project, authorship, comfort, and collegiality. Language is also an issue that arises when study instruments or protocols are being implemented in a new context and require linguistic translation. To ensure appropriate cultural and linguistic adaptation of surveys, teams should follow best-practices and procedures,15 including forward translation of instruments, followed by back-translation to the original language to ensure that the intended meanings are retained. In some cases (eg, survey research), cognitive interviewing should then be used to evaluate whether study participants perceive the same intended meaning. While linguistic translation is a complex undertaking, cultural translation may be even more complex. Culture is highly variable even within a single country and it may be challenging to adapt measures from other contexts. Team members with deep local knowledge are best equipped to provide the relevant insights and guide the process of cultural adaptation. Cognitive interviewing can also be a useful tool to gauge cultural appropriateness. Finally, some quantitative methods may be useful for evaluating validity (eg, construct and content validity of survey measures) in a specific study population.16 Capacity Building Global vision research may pose a distinct set of challenges compared with domestic research, and may encompass both addressing a research question, as well as developing local capacity during the process of operationalizing the research methodologies. Infrastructure growth, equipment procurement, and maintenance may also be required, depending on the scope of the project. The overarching objective of global vision research is to improve the vision and eye health of individuals and populations worldwide. This is ideally accomplished by generating generalizable knowledge on vision and eye health; addressing scientific questions with local relevance; and building local research capacity. The importance of growing research capacity in locations where there is inadequate infrastructure, knowledge, or resources to conduct vision research cannot be overstated. Fortifying local capacity will not decrease the relevance of global collaborations, however it does aim to enable greater South-South collaboration, equitable relationships between colleagues, and to open doors for researchers, particularly in resource-limited settings. In fact, across all global settings, there is a need to strengthen vision research capacity among groups and institutions that are historically underrepresented in vision research.17 Research capacity building can take many forms. Dedicated mentorship is often a necessary component of such efforts. In addition, formal predoctoral and postdoctoral research fellowships provide an opportunity for dedicated aspiring researchers to gain skills in key areas like scientific writing, grantsmanship, study design, and biostatistics. Shorter-term workshops and opportunities to participate in mentored research may provide less intensive and time-consuming opportunities to build some of these skills. Capacity building also involves working with stakeholders to ensure that infrastructure exists locally to conduct research, thus decreasing reliance on foreign entities. For example, laboratory capacity to carry out genotyping and complex assays, as well as data science capacity to construct large databases and carry out complex analyses are key resources for many vision research projects. With appropriate resources and capacity locally, researchers need not rely on the resources and priorities of external collaborators to undertake the research that they deem important in their own contexts. Some of these key principles are illustrated in the following case study describing the development of a retinopathy of prematurity (ROP) program and the related infrastructural growth. Case Study: Development of a ROP program in Mongolia The “third epidemic” of ROP has taken hold in LMICs because of increase in neonatal survival. Studies have demonstrated that screening guidelines established in high-income countries do not adequately encompass at-risk infants in LMICs, where infants that develop ROP have been demonstrated to have greater birth weights and gestational ages.18 In 2011, in collaboration with ORBIS International, an international group of investigators worked with local partners at the National Center for Maternal and Child Health in Ulaanbaatar, Mongolia to assess the ROP needs in the country. During this initial screening program, several children with high risk for developing severe ROP were identified, in addition to many children who had stage 4 and stage 5 ROP. At that time, screening protocols and treatment of ROP were lacking. Infrastructure development was needed for screening and the clinical care of children at-risk of ROP. This was coupled with the development of data management systems, imaging of the fundus in at-risk infants, and tele-education for local physicians. Following the development of clinical infrastructure, a study in Mongolia was conducted to evaluate screening guidelines utilizing a web-based data management system, which has since been expanded to screening programs in Kathmandu, Nepal, and Coimbatore, India. Investigation of ROP guidelines in Mongolia was facilitated by a data management system that allowed for data management and remote expert reading. An advantage of this system was that international experts could remotely access the clinical data and images and corroborate ROP diagnoses in challenging cases. As a result of the lessons learned from this screening program, iTeleGEN, which is a web-based platform that integrates data management, tele-education modules, and telemedicine, was developed. Pilot projects in Kathmandu and Nepal in ROP screening were conducted, followed by expansion for its used at Aravind Eye Hospital (Coimbatore, India) for both telemedicine and tele-screening of ROP, and has promising utility in the adoption of artificial intelligence-assisted screening programs.19–22 The current screening guidelines utilized in Mongolia are gestational age <34 weeks and birth weight <2000 g, which are evidence-based guidelines developed from this screening program. In the Mongolian cohort, 18 infants (9.3%), including 8 with type 1, were outside of US screening guidelines, demonstrating that guidelines must be specific to the region in which the screening takes place.23 Local providers, NGOs, and international partners were instrumental in clinical program development and research programs that were scalable to different country settings. How to get Involved Participating in international research, teaching, or capacity building can be a vital part of one’s career whether in academic medicine or private practice. While many avenues exist for involvement in global initiatives, finding mentorship is one of the most important aspects for successful individual and program development. A mentor may be valuable in providing introductions and collaborations with pre-existing partnerships. The mentor will know how to navigate a new setting that you may be less familiar with and serve as a guide for your participation. For ophthalmology trainees there are multiple existing learning opportunities. Many residency programs have a global experience built into the residency training program with an increasing number of residency programs with a global track, If this is an aspect of training that one values, it may be ideal to find a program that provides a global research or learning experience. There are also a few ophthalmology residencies that have a global track or curriculum in place with didactics, training, and specific experiences to provide the knowledge to navigate local and global projects, which are described in more detail in this issue of International Ophthamology Clinics. There are currently nine year-long academic global ophthalmology fellowship programs. Programs with active fellowship programs include the Emory Eye Center (Emory University), Kellogg Eye Center (University of Michigan), Dean McGee Eye Institute (University of Oklahoma), Stanford University, Wills Eye Center (Wills Eye Hospital), Illinois Eye and Ear Infirmary (University of Illinois), the Moran Eye Center (University of Utah), Seva Foundation, and the Truhlsen Eye Institute (University of Nebraska). These programs offer intensive 1-year training experiences in global ophthalmology that are unique to each institution. Other places to network are the young ophthalmologist at the AAO as well as the Global Ophthalmology Symposium at the AAO annual meeting. A newly formed meeting, the Global Ophthalmology Summit, brings together key global stakeholders in advocacy, education, and research and will be an opportunity for all to network and engage. Working with foundations, nongovernmental organizations, or academic institutions whose mission aligns with the work you would like to be a part of or the service you want to provide are other ways to participate. Funding Funding for global ophthalmology is often challenging and varies depending on the stage of the project, stakeholders involved, and the scope of work involved. For clinical service delivery, self-funded projects may include short-term visits and service delivery (eg, partnering with local care providers for eye screening or surgical services). Work with nongovernment organizations may involve a combination of philanthropy and self-funding. For research projects that evolve into programs, answering specific research questions in partnership with in-country partners may evolve from the pilot phase of funding (eg, pilot grants ranging from $10,000 to $50,000) and require additional funding through federal grants (eg, National Institutes of Health Fogarty International Center and other National Institutes of Health entities), United States Agency for International Development (USAID), and various foundations (eg, Bill and Melinda Gates Foundation). When considering the program’s funding needs, a range of considerations may need to be accounted for, but broad categories include salary support for investigators and staff, travel costs including air and ground transportation, visa fees, housing, equipment and supplies, facility fees, and administrative costs of local regulatory agencies (e.g., Institutional Review Board fees). Depending on the environment where field research is conducted, other country-specific requirements may also exist including fuel costs, patient transportation fees, interpretation support, and local security. During coronavirus disease-2019, additional budget requirements include costs of laboratory testing before inbound and outbound flights, which varied depending on country-specific requirements (implicit in this budget would be additional housing fees should an individual test positive for coronavirus disease-2019 during travel). Conclusion Through equitable partnerships and collaboration, global vision research has the potential to greatly improve the vision and eye health of people worldwide. In this article we have sought to illustrate some of the key considerations for researchers beginning to undertake collaborations with colleagues from distinct settings. We have also highlighted opportunities to optimize equity in global collaborations, and to ensure that global vision research adheres to the highest standards of ethics and cultural competence.
Implicit discourse relation classification refers to a task of automatically determining relationships between arguments. It has been widely proven that, in a neural classification architecture, decoding discourse relations heavily relies on the reliable semantic representations of arguments. In addition, our previous survey shows that, for a target argument, the external semantic information hidden in the accompanying argument benefits the encoding of the target, either wholly or partially. Moreover, dependency structure appears as the crucial feature for synthesizing word senses of the entire words in arguments. Accordingly, we propose a novel method to enhance the current representation learning of pairwise arguments, which takes into consideration both external semantic information and internal dependency structure. In particular, we inject external semantic information into the Long-Short Term Memory (LSTM) unit of Recurrent Neural Network (RNN) through the input and forget gates. Different from the existing one-off interactive learning models, our method allows the neuronal memory of internal argument semantics to be affected by external information at each encoding step. On the basis, we apply the parser-based Graph Convolutional Networks (GCN) over the semantic presentations of words, so as to accumulate the closely-related semantic information in terms of dependency structures. We conduct experiments on Penn Discourse TreeBank Corpus of version 2.0 (PDTB 2.0). The test results illustrate that the proposed method enhances the baseline significantly, and it obtains comparable performance compared to the state of the art.
In our research we have found that the Pronouns of address in Costa Rican Spanish are in a constant struggle between standardization and linguistic change, as well as between the norm and variation. For example, in chapter 1 of the thesis, regarding the diachronic approach, we can conclude that the pronoun usted has always been linked to an explanatory and descriptive analysis in Costa Rican Spanish. The studies point out that there is objectivity and neutrality when analyzing the functional and structural categorization of the pronoun. Meanwhile, vos and tú have been the object of a struggle between the prescriptive and the descriptive. In relation to chapter 2, regarding the chapter on attitudes, perceptions and linguistic evaluations concerning the Pronouns of address in Costa Rican Spanish, informants assign different positive evaluations to usted. Respondents focus mostly on clarifying how they use usted to mark positioning and acts of identity. On the other hand, for the same informants, different stereotypes continue to circulate around vos and tú. In the following chapters (3 and 4), in terms of the written language, newspapers as mass media are opting more for tú as part of the use of a standardized modality. On the other hand, in the spoken language, in terms of oral advertising both on television and radio, Costa Rica is following its own linguistic norm. The vos above all and the usted in a certain way are the instruments of expression of the mass media. Likewise, outside the advertising space, tú is the normative pronoun in electronic writing (chapter 5). In summary, there are different statuses of pronouns because the rules are different. For example, at the diacritical and perceptual level usted is the point of reference in usage. While in oral advertising it is vos and in written advertising and electronic writing it is tú. From this perspective, we can confirm, once again, the complexity concerning the Pronouns of address in Costa Rican Spanish, since it moves between different regulations depending on the register.
The article presents the first attempt to study the Hungarian translations of Fyodor Dostoevsky’s The Gambler. The research aims at determining the completeness of the translation of sensemaking elements in the two most popular Hunagarian translations by Endre Szabo (1900) and Erzsebet Guthi Devecserine (1957) to assess the impact of the translation shifts on the preservation of the novel’s idea. The method of studying the original and the translations is based on the concept analysis. Studying the original, the authors have revealed that the concept of passion (composed of such concepts as passion in love, passion for gambling, greed, and pride) is one of the sensemaking elements in the novel. The article focuses on passion in love in three episodes of the introduction which verbalise this concept in the image of Alexey Ivanovich, thus establishing his psychological portrait and describing his attitude to Polina. The worldview in Dostoevsky’s novels is built upon the orthodox values, which define the dominants of the novel axiology. The Hungarian culture is catholic. The contradictions between the Orthodox and Catholic interpretation of passion in general allow hypothesizing that the reproduction of some features of passion in love in translation may be challenging. The analysis of the translations has revealed that the translators rendered some features of the concept practically without loss (appetence, hatred, murder, jealousy, desire, appetite, agony / excruciation, suicide, disease). Theidentified losses (pleasure, extinction of appetence, loss of control) do not distort the sense of the episodes studied as well as the portrait of the character. The authors believe that it was possible to preserve the concept due to a number of factors. Firstly, the translators focused on the similarities rather that differences in the Orthodox and Catholic interpretations of passion. Secondly, the approach of the Hungarian translators is distinguished by an extremely careful attitude to the original: Szabo adheres to the literal reproduction of the original style and Devecserine aspires after the balance between the original and the Hungarian linguistic norm. Contribution of the authors: the authors contributed equally to this article. The authors declare no conflicts of interests.
The convergence of artificial intelligence (AI) technology and natural language processing (NLP) has rapidly increased the demands for an analysis on the natural language that involves plenty of ambiguities not present in formal language. For this reason, the language model (LM), a statistical approach, has been used as a key role in this area. Recently, the emerging field of deep learning, which applies complex deep neural networks for machine learning tasks, has been applied to language modeling and achieved more remarkable results than traditional language models. One of the important techniques that have led neural network-based LM success is the attention mechanism. Attention mechanism makes neural networks pay attention to specific words in the input sentence when generating the output words. However, although the attention mechanism has improved the performance of many neural network models, it requires tons of parameters to achieve the state-of-art level performance. This is because attention mechanism encodes the context of a word by simply accumulating the outputs from the network for all the input words, which may cause information loss. To compensate for this limitation, we propose an extension of attention mechanism by adopting a convolutional neural network to replace the accumulation. With only far fewer parameters, our model achieved comparable performance to the recent state-of-the-art models on the very popular benchmark datasets, yielding perplexity scores of 58.4 on the Penn Treebank dataset and 50.1 on the Wikitext-2 dataset, respectively.
In this article, we used methods of analysis, synthesis, comparison, and systemization. The generalization of theoretical positions of scientific work has been developed and the feasibility of the study of the selected problem has been demonstrated; based on a modern methodology, theoretical and methodological principles and specificities of linguistic norms and rules in the production of practical skills and skills from the Ukrainian language in the fifth-grade students are substantiated.An effective method of their formation is proposed, the components of which are defined: innovative approaches, general-practice, and lingua didactic principles, the use of didactic material, the study of spelling of the Ukrainian language as a dynamic system, the assimilation of spelling norms based on knowledge about the structure of the word, systematic assimilation of spelling material, a skillful combination of theory and practice in learning, the selection of effective exercises, methods, and techniques that will increase the literacy of five-graders. It was established that when learning language material and developing language skills, students need to name new terms, update the content of the grammatical concepts that they already have, so that, by merging them, forming a new rule. Specific examples of this approach are given when studying the suffix (“a significant part of the word that stands at the root and serves to create new words”). It is emphasized that this rule should be formulated completely, precisely, without rotation, using accepted motivation. The role of additional questions has been highlighted in studying the substance of itsformulation. It is determined that among the most effective ways of conscious mastery of grammatical concepts and rules is the reception of comparison, which develops logical thinking, teaches to highlight their characteristic features in comparing phenomena.The expediency of using such a technique as comparing certain grammatical phenomena with the corresponding rules, such as to find in the text presented in visual or auditory form, spellings studied in order to avoid possible spelling mistakes is proved. As a result of the conducted research, it was found out that the collective understanding of the studied linguistic material in the lesson promotes a better understanding of the mastered terms. Keywords: spelling; grammar; orphogram and rules; spelling skills and abilities; cognitive activities; word structure; prefix; suffix; derived and non-derived words; comparison.
Human visual attention is highly structured around gathering relevant information on the underlying goals or sub-goals one wishes to accomplish. Typically, this has been modelled using either qualitative top-down saliency models, or through the use of highly reductionist psychophysical experiments. Modelling an information-gathering process in these ways however often ignores the rich and complex repertoire of behaviour that make up ecologically-valid gaze. We propose a new way of analysing natural data, suitable for analysing temporal structure of visual attention in complex, freely moving environments. To achieve this, we capture visual information from subjects performing an unconstrained task in the real world - in this case, cooking in a kitchen. We use eye-tracking glasses with a built-in scene camera (SMI ETG 2W @ 120Hz) to record n=15 subjects, setting up, cooking breakfast and eating in a real-world kitchen. We process the visual data using a deep-learning based pipeline (Auepanwiriyakul et al., 2018, ETRA), to obtain the stream of objects in the field of view. Eye-tracking gives us a sequence of objects people are focussing their overt attention on throughout the task. We resolve ambiguities using pixel-level object segmentation and classification techniques. We analyse these sequences using HMM and context-free grammar induction models (IGGI), revealing a potential hierarchical structure which is invariant across subjects. We compare this grammatical structure against a “ground-truth”, the WordNet lexical database of semantic relationships between objects (Miller et al, 1995), and find some surprising similarities and counterintuitive differences between attention-derived and textual base structure of objects, suggesting that the differences between how we look and how we verbally reason about tasks is an open question of cognition and attention.
Improving methods of stylometrics and classification so that they give good results with small texts is the focus of much research in the digital humanities and in the NLP community more generally. Recent work has suggested that an approach using combinations of shallow and deep morpho-syntactic information can be quite successful. But because the data in that study were taken from hand annotated dependency treebanks, the wider applicability of such an approach remains in question. The present paper seeks to answer this question by using machine-generated morphological and syntactic annotations as the basis for a closed-set classification experiment. Texts were parsed according to the Universal Dependency schema using the udpipe package for R. Experiments were carried out on data from several languages covering a range of morphological complexity. To limit confounders, consideration of vocabulary was excluded. Results were quite promising, and, not surprisingly, a more complex morphology correlates with better accuracy (e.g., 100-token texts in Polish: 88% correct; 100-token texts in English: 74%).The method presented here has particular advantages for stylometrics as practiced in literary analysis and other fields in the humanities. The Universal Dependency annotation categories are generally similar to those used in traditional grammars. Thus, the variables which serve to distinguish the style of a given author are relatively easier to interpret and understand than, for example, are character n-grams or function words. This fact, combined with the availability of easy-to-use dependency parsers, opens up the study of a syntax-centered stylometrics to persons with a wide range of expertise. Even students at the early stages of their studies can identify and investigate the morpho-syntactic signature of a particular author. Therefore, the characterization of texts based on computational annotation of this type deserves a place in classification studies because of its combination of good results and good interpretability.
The book entitled New insights into the mental lexicon provides significant insights into meaning creation by unfolding a wide-ranging array of current preoccupations in cognitive linguistics, psycholinguistics, corpus research and media studies. Notably, the present volume pursues the subtle issues of word or concept meanings, of the argument over the likely interface between meaning fixedness and fuzziness or fluidity, of the (im)possibility to deal with these ‘slippery customers’ (Labov 1973), and the ‘vague boundaries and fuzzy edges’ (Lakoff 1972) of natural language concepts. It is organised into three main parts, the first dedicated to the functioning and processing of meaning in both mother tongue (Romanian) and foreign language (English), the second concentrates on meaning creation in various discourses (cultural, political, business journalese, social media), whereas the third focuses on the subject of meaning construction in literary productions and film adaptations. As for the research methodology, the book is based on a large mixture of frameworks: Pragglezaj method (2007), MIPVU technique (2010), Charteris- Black’s (2004) critical metaphor analysis framework, Alice Deignan’s corpus-based metaphor analysis (2005), Halliday and Hasan’s Systemic Functional Linguistics (SFL) framework (1989), Forceville’s multimodal metaphor analysis theory (2009), parallel text analysis, Isabela and Norman Fairclough’s political discourse analysis (2012). Qualitative as well as quantitative analysis of data was carried out, coupled with semi-automatic treatment of text, using LancsBox and ConcApp concordancing software. The mental lexicon is mainly tackled from a cognitive-linguistic perspective rather than from a psycholinguistic one, with a conspicuous focus on abstract features of the lexicon, applying linguistic grammars and dictionaries, lexical databases (WordNet) and word categorisation in corpora. It is also worth mentioning that experimental data was formed from elicited behaviour of Romanian learners of English. Last but not least, limitations are imposed by the scope of research, carried out within a limited time frame and involving the construction of a small size of corpora.
both University of Amsterdam) and E. Hoekstra (Fryske Akademy).It contains eight chapters (an introduction, theoretical framework, five case studies and final discussion), supplemented by appendices with survey questionnaires, additional analyses of data from MAND, and summaries in English, Dutch and Frisian.In the following, I first summarise the main empirical questions and findings before then discussing the methods and theoretical framework. 1 The overall aim of the study is to empirically describe, theoretically model and explain the ongoing changes in the conjugation system of West Frisian in order to gain insight into the mechanisms of morphological change, in particular, relating to why certain changes occur and others do not (Ch. 1 Introduction, p. 1).Merkuur considers Frisian, due to its low degree of standardisation and close contact to Dutch, as a fruitful testing ground for theoretical notions of morphological variation and change.In contrast to Dutch and most other West Germanic varieties, West Frisian has retained two weak conjugation classes.This study focuses on phenomena occurring within the largest of them, weak class II, which incorporates 81% of types listed in the lexical databases of the Fryske Akademy (p.57).The class features discussed include the lack of a dental suffix in the preterite and past participle and syncretism in the category tense in the 2nd person singular.Chapter 2 then introduces the notion of morphological change adopted in the study, which essentially focuses on generalisations based on ambiguous input in acquisition, and the theoretical framework, which combines two formal approaches, Distributed Morphology and the Tolerance Principle.The subsequent chapters provide in-depth case studies.Chapters 3 and 4 address the relationship between weak class I (with infinitives in -ə, and preterites and past participles with a dental suffix) and weak class II (with infinitives in -jə, and preterites and past participles without a dental suffix) within the overall conjugation system.Merkuur investigates which classes and subclasses can be considered rule-based from a Distributed Morphology perspective (Ch.3).On this basis, she tests which of these rules are then identified as productive generalisations when the Tolerance Principle, a measure of inflectional productivity first proposed by Yang, is applied to them (Ch.4).Contributing to a long-standing discussion in the field, Merkuur then focuses on the question of whether one of the 1.I am grateful to Julie
Aesthetic principle not only concerns the perception of visual beauty, but most crucially, the cognitive-affective responses derived from experiencing such stimuli. Despite this knowledge, user experience (UX) designers often prioritise beauty over intended end-user affect. This reversed order of prioritisation contradicts the sequential UX design process, leading to unexpected end-user perceptions and responses. The challenge in shaping UX before user interface (UI) design is that there first must be prior knowledge of aesthetic affect. A study with 1,782 worldwide participants was conducted evaluating affective user-responses to 43 atomic aesthetics using 153,252 data points and presented as affect ratings (ARs). Results demonstrated high affective resonance amongst aesthetics evaluated, suggesting aesthetic AR may be a viable method of improving the UX design process by influencing user perceptions, responses and actions at a non-conscious level. This is the first of a series of studies in the direction of Aesthetic Semantics®.
Dependency parsing has become the norm for its advantages of representing syntactic information for numerous tasks of natural language processing (NLP). In Vietnamese, a challenging problem which arises in this domain is the insufficiency of the training resource. Our work presents a new method to automatically convert a Vietnamese constituency treebank into dependency trees. We designed new dependency labels for Vietnamese treebank. Furthermore, in this research, we proposed new head-percolation rules and dependency relations. The experimental results on two state-of-the-art parsers, MaltParser and MSTParser, indicated that our treebank were roughly 13% UAS and 21% LAS higher than previous works.
In this study, we aim to offer linguistically motivated solutions to resolve the issues of the lack of representation of null morphemes, highly productive derivational processes, and syncretic morphemes of Turkish in the BOUN Treebank without diverging from the Universal Dependencies framework. In order to tackle these issues, new annotation conventions were introduced by splitting certain lemmas and employing the MISC (miscellaneous) tab in the UD framework to denote derivation. Representational capabilities of the re-annotated treebank were tested on a LSTM-based dependency parser and an updated version of the BoAT Tool is introduced.
Modern Irish is a minority language lacking sufficient computational resources for the task of accurate automatic syntactic parsing of usergenerated content such as tweets. Although language technology for the Irish language has been developing in recent years, these tools tend to perform poorly on user-generated content. As with other languages, the linguistic style observed in Irish tweets differs, in terms of orthography, lexicon, and syntax, from that of standard texts more commonly used for the development of language models and parsers. We release the first Universal Dependencies treebank of Irish tweets, facilitating natural language processing of user-generated content in Irish. In this paper, we explore the differences between Irish tweets and standard Irish text, and the challenges associated with dependency parsing of Irish tweets. We describe our bootstrapping method of treebank development and report on preliminary parsing experiments.
Treebanks are one of the most needed and used linguistic resources in the fields of Natural language processing (NLP) and Natural language understanding (NLU). Arabic has only two constituency-based treebanks and a number of dependency treebanks. The current research presents the guidelines for building a parsed Arabic treebank for Modern Standard Arabic (MSA). The guidelines show, firstly the choice of the grammar formalism, then the genre and size of the treebank, and finally the annotation layers of the treebank. The study also shows that using the traditional Arabic grammar syntactic theory to describe the Arabic syntax has proven to be more suitable than using any of the modern syntax theories. Working with the traditional Arabic grammar also helps avoid the errors that the available treebank fell in as a result of using guidelines that don't suit the Arabic grammar. The study adopts three layers of annotations: the morphological layer, the syntactic layer, and the grammatical function layer. The resultant tree is a very detailed and rich syntactic tree, which is preferable by the researcher over having a huge amount of data poorly and shallowly annotated.
Abstract Prior work suggests that more frequent or higher exposure to stressors relates to less positive affect and more negative affect in daily life. Limited knowledge exists about whether subjective appraisals of such stressors (i.e., perceived negative impacts on daily routine, personal health and safety, and finances) also have negative links to daily well-being. This study examines this link using data from an 8-day daily dairy study (n=675 days) in an online sample of older adults (n = 110 people, ages 60-90). We also explored potential psychological moderators particularly relevant to the experience of aging (i.e., self-views of aging, S-VOA). Results from multilevel models indicate that people reported more negative affect and less positive affect on days with more negative appraisals, especially on those days when they also had more negative self-views of aging. These findings highlight S-VOA as psychological resources that help people cope with stressful events in everyday life.
Discourse parsing is an essential upstream task in Natural Language Processing with strong implications for many real-world applications. Despite its widely recognized role, most recent discourse parsers (and consequently downstream tasks) still rely on small-scale human-annotated discourse treebanks, trying to infer general-purpose discourse structures from very limited data in a few narrow domains. To overcome this dire situation and allow discourse parsers to be trained on larger, more diverse and domain-independent datasets, we propose a framework to generate "silver-standard" discourse trees from distant supervision on the auxiliary task of sentiment analysis.
Visual processing of emotional words modulates early event-related potentials (ERPs) such as the early posterior negativity (EPN). Questions remain as to whether this modulation reflects modality-specific processing, preferentially elicited by emotional words of the native language (L1). This study investigates the modulation of early ERPs during rapid serial visual presentation (RSVP) of adjectives or nouns referring to emotional feeling states, neutral traits, to an overweight or lean body or to concrete body parts or neutral objects, presented in the L1 and the second language (L2). Word ratings in the L2 were assessed in a pilot study. The N100 and the EPN were modulated by the emotional valence of the stimuli irrespective of the word class or the task (silent reading vs. word counting). The results suggest that early affective appraisal is obligatory, not restricted to privileged categories of linguistic information (emotions) or solely found for the embodied language (L1).
= 192, 179) were assigned to the certain (CG) or uncertain group (UG) and presented with 100% (CG) or 50% (UG) S1-S2 congruency between visual stimuli. During the test phase, participants were presented with a new 75% S1-S2 paradigm and visual (Experiment 1) or auditory (Experiment 2) S2s. Participants were asked to rate the expected valence of upcoming S2s (expectancy ratings) or valence and arousal to S2s. In both experiments, the CG reported more extreme expectancy ratings than the UG, suggesting that experiencing previous reliable S1-S2 associations led CG participants to subsequently predict similar associations. No group differences emerged on valence and arousal ratings, which were more prominently influenced by the new 75% contingencies of the test phase rather than by previous learned contingencies. Last, comparing the two experiments, no significant group by experiment interaction was found, supporting the hypothesis of cross-modality generalization at the subjective level. Overall, our results advance knowledge about the mechanisms by which previous learned contingencies shape subjective affective experience. (PsycInfo Database Record (c) 2023 APA, all rights reserved).
Neuroscientists have formulated the model of emotional intelligence (EI) based on brain imaging findings of individual differences in EI. The main objective of our study was to operationalize the advantage of high EI individuals in emotional information processing and regulation both at behavioral and neural levels of investigation. We used a self-report measure and a cognitive reappraisal task to demonstrate the role of EI in emotional perception and regulation. Participants saw pictures with negative or neutral captions and shifted (reappraised) from negative context to neutral while we registered brain activation. Behavioral results showed that higher EI participants reported more unpleasant emotions. The Utilization of emotions scores negatively correlated with the valence ratings and the subjective difficulty of reappraisal. In the negative condition, we found activation in hippocampus (HC), parahippocampal gyrus, cingulate cortex, insula and superior temporal lobe. In the neutral context, we found elevated activation in vision-related areas and HC. During reappraisal (negative-neutral) condition, we found activation in the medial frontal gyrus, temporal areas, vision-related regions and in cingulate gyrus. We conclude that higher EI is associated with intensive affective experiences even if emotions are unpleasant. Strong skills in utilizing emotions enable one not to repress negative feelings but to use them as source of information. High EI individuals use effective cognitive processes such as directing attention to relevant details; have advantages in allocation of cognitive resources, in conceptualization of emotional scenes and in building emotional memories; they use visual cues, imagination and executive functions to regulate negative emotions effectively.
When inferring emotions, humans rely on a number of cues, including not only facial expressions, body posture, but also expressor-external, contextual information. The goal of the present study was to compare the impact of such contextual information on emotion processing in humans and two deep neural network (DNN) models. We used results from a human experiment in which two types of pictures were rated for valence and arousal: the first type depicted people expressing an emotion in a social context including other people; the second was a context-reduced version in which all information except for the target expressor was blurred out. The resulting human ratings of valence and arousal were systematically decreased in the context-reduced version, highlighting the importance of context. We then compared human ratings with those of two DNN models (one trained on face images only, and the other trained also on contextual information). Analyses of both categorical and the valence/arousal ratings showed that although there were some superficial similarities, both models failed to capture human rating patterns both in context-rich and context-reduced conditions. Our study emphasizes the importance of a more holistic, multi-modal training regime with richer human data to build better emotion-understanding systems in the area of affective computing.
How words are associated within the linguistic environment conveys semantic content; however, different contexts induce different linguistic patterns. For instance, it is well known that adults speak differently to children than to other adults. We present results from a new word association study in which adult participants were instructed to produce either unconstrained or child-oriented responses to each cue, where cues included 672 nouns, verbs, adjectives, and other word forms from the McArthur-Bates Communicative Development Inventory (CDI; Fenson et al., 2006). Child-oriented responses consisted of higher frequency words with fewer letters, earlier ages of acquisition, and higher contextual diversity. Furthermore, the correlations among the responses generated for each pair of cues differed between unconstrained (adult-oriented) and child-oriented responses, suggesting that child-oriented associations imply different semantic structure. A comparison of growth models guided by a semantic network structure revealed that child-oriented associations are more predictive of early lexical growth. Additionally, relative to a growth model based on a corpus of naturalistic child-directed speech, the child-oriented associations explain added unique variance to lexical growth. Thus, these new child-oriented word association norms provide novel insight into the semantic context of young children and early lexical development.
Abstract Previous work on norm orientations in the Caribbean Englishes has focussed largely on phonological norms, such as accents, and, to a lesser extent, grammatical norm orientation. Outside of the publication of dictionaries, however, lexical norms and their spread have received little attention. This paper examines lexical norm orientations in Trinidadian English, presenting the results of a corpus‐based study and survey study of the lexical preferences of speakers of Trinidadian English. The findings suggest that while there is evidence of American influence on Trinidadian lexicon, British variants persist. Indeed, British and American variants often coexist, albeit with different connotations in Trinidadian English. From a methodological standpoint, this paper demonstrates the benefits of using a mixed‐methods approach in looking at norms, particularly with regard to lexicon.
Olfactory perception, and especially affective responses of odors, is highly flexible, but some mechanisms involved in this flexibility remain to be elucidated. This study investigated the odor perceptions of several essential oils used in aromatherapy with emotion regulation functions among college students. The influences of people's characteristics including gender, hometown region, and fragrance usage habit on odor perception were further discussed. Odor perception of nine essential oils, which can be divided into the ester-alcohol type (e.g., lavender oil) and terpene type (e.g., lemon oil) were evaluated under three odor concentrations. The results indicated that chemical type, but not concentration, significantly influenced the odor perception and there was no interaction between the two factors in this study. The arousal and emotional perception scores of odors with terpene-type oil were significantly higher than odors with ester-alcohol type. In terms of people's characteristics, participants from the southern Yangtze river gave a higher familiarity rating to almost all of these odors. The habits of fragrance usage also significantly influenced some of the odors' subjective intensity and emotional perception ratings. However, there were no significant gender differences in most of the odor perceptions. In addition, familiarity and pleasantness were positively correlated, and emotional perception and subjective intensity also showed a weak correlation. These results suggested that users' cultural characteristics could be considered to be important factors that affect the essential oil's odor perception in aromatherapy.
The article is devoted to the study of the process of borrowing and adaptation of English loan words in the Ukrainian language. It is established that in the lexical system of the Ukrainian language foreign words make up about 10%, 70 – 80% of which are English loan words. The presence of a significant share of English loan words in the Ukrainian language is due to a number of extra- and intralinguistic factors: the development of economic, cultural and political ties; quantitative and qualitative complication of various spheres of language communication; diversity of norms of speech behavior; expansion of regulatory limits; achievements of English-speaking countries in certain fields of activity; striving for linguistic economy; the need to replenish the composition of expressive language means; the need to clarify and detail the concepts available in the language; «Americanization»; imitation of fashion. The study systematizes English loan words in the Ukrainian language and distributes them by spheres of use (sociopolitical, financial and economic, culture and art, technical, mass communication, sports, science, and education). Three stages in the process of adaptation of English loan words into the Ukrainian language are distinguished: 1) the initial stage, which is characterized by a change in the morpheme structure of English loan words; 2) in-depth, related to the selection of the same components in groups of English loan words based on the similarity of final elements and the development of new suffixes of English origin in the Ukrainian language; 3) the stage of full adaptation, which is characterized by participation of English loan words in the process of word formation through the mediation of Ukrainian language suffixes, the formation of new complex words based on English loan words and Ukrainian or previously borrowed words, as well as the consolidation of the spelling form of complex words. It is established that the inclusion of English loan words in the lexical structure of the Ukrainian language and their active use in oral and written speech leads to the formation of synonymous pairs containing proper Ukrainian counterparts.
L’objet de cet article est, en premier lieu, de faire apparaître des marqueurs de la modalité déontique dans des textes appartenant au discours juridique en russe. Autrement dit, il s’agit des formes grammaticales et lexicales qui sont employées afin d’exprimer l’obligation. Notre corpus de textes juridiques est organisé selon la hiérarchie des normes proposée par Hans Kelsen, théoricien du droit et fondateur de l’école normativiste. Cette hiérarchie se présente sous la forme d’une pyramide dont le sommet se réfère au niveau juridique le plus élevé, c’est-à-dire le bloc constitutionnel, alors que sa base implique le niveau le plus bas, le bloc contractuel. Selon notre analyse, dans le bloc constitutionnel, le marqueur de l’obligation le plus fréquent est le présent imperfectif, tandis que dans le bloc contractuel on peut constater une variété de marqueurs modaux exprimant l’obligation, tels que должен, обязан, подлежит, etc. Ainsi, dans une approche jurilinguistique, nous révèlerons, en deuxième lieu, des liens entre la nature juridique de chaque type de textes appartenant à différents blocs de la pyramide et le choix des marqueurs de la modalité déontique qui y expriment l’obligation. Dans une perspective sémasiologique, nous posons qu’une telle répartition des marqueurs de l’obligation permet de reconstruire les traces d’une forme d’échange communicationnel entre l’énonciateur et les destinataires dans le discours qui est traditionnellement considéré comme discours non communicationnel avec effacement énonciatif de toute voix parlante.Notre étude n’est pas inscrite dans un cadre théorique précis, mais elle est inspirée, néanmoins, par la grammaire cognitive, l’analyse du discours ainsi que par la linguistique de corpus.
The article analyzes the epistolary discourse of Ján Kollár using the categories and concepts of modern linguistic pragmatics. The subject of the study was 16 letters addressed to V. Ganka, V. Kopitar, K. Ya. Erben, S. Grobon, V. A. Maciejowski in the period from 1824 to 1851. The study of the material is based on the method of subjectobject interpretation of the addressee’s communication. The analysis of Ján Kollár’s texts showed that his epistolary discourse is a speechlanguage work that was created with regard to the chronologically determined national epistolary tradition. Contacts and the content of Ján Kollár’s communication concerned primarily the socio-political and cultural sphere related to his professional activities. The socio-cultural dimension, within which the communication took place, determined the addressee’s socio-pragmatic role, and their civil position, influenced the topics of the correspondence, its personal and subject components, and assumed dialogue. The components of the studied epistolary discourse are characterized by the following features: formality — informality, distance — non-distance of communication, the hierarchy of relations, and observance of etiquette norms with their mainly emotional verbal design. Markers of politeness in Ján Kollár’s letters are etiquette lexical means, as well as tropes metaphors, and phraseology, which are methods of expressing the author’s empathic moods. Individual features include numerous author innovations. Thus, the social and socio-psychological characteristics of the author affected the language, structure, and content of his epistolary discourse. Kollár’s epistolary legacy is also important in the extralingual aspect. It serves as an additional source for studying the role of the author in social, literary, and scientific circles of his time.