Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
An increasing amount of research has recently focused on dimensional sentiment analysis that represents affective states as continuous numerical values on multiple dimensions, such as valence-arousal (VA) space. Compared to the categorical approach that represents affective states as distinct classes (e.g., positive and negative), the dimensional approach can provide more fine-grained (real-valued) sentiment analysis. However, dimensional sentiment resources with valence-arousal ratings are very rare, especially for the Chinese language. Therefore, this study aims to: (1) Build a Chinese valence-arousal resource called Chinese EmoBank, the first Chinese dimensional sentiment resource featuring various levels of text granularity including 5,512 single words, 2,998 multi-word phrases, 2,582 single sentences, and 2,969 multi-sentence texts. The valence-arousal ratings are annotated by crowdsourcing based on the Self-Assessment Manikin (SAM) rating scale. A corpus cleanup procedure is then performed to improve annotation quality by removing outlier ratings and improper texts. (2) Evaluate the proposed resource using different categories of classifiers such as lexicon-based, regression-based, and neural-network-based methods, and comparing their performance to a similar evaluation of an English dimensional sentiment resource.
Pavlovian fear conditioning is widely used to study mechanisms of fear learning, but high-throughput studies are hampered by the labor-intensive nature of examining participants in the lab. To circumvent this bottle-neck, fear conditioning tasks have been developed for remote delivery. Previous studies have examined remotely delivered fear conditioning protocols using expectancy and affective ratings. Here we replicate and extend these findings using an internet-delivered version of the Screaming Lady paradigm, evaluating the effects on negative affective ratings and response time to an auditory probe during stimulus presentation. In a sample of 80 adults, we observed clear evidence of both fear acquisition and extinction using affective ratings. Response times were faster when probed early, but not later, during presentation of stimuli paired with an aversive scream. The response time findings are at odds with previous lab-based studies showing slower as opposed to faster responses to threat-predicting cues. The findings underscore the feasibility of employing remotely delivered fear conditioning paradigms with affective ratings as outcome. Findings further highlight the need for research examining optimal parameters for concurrent response time measures or alternate non-verbal indicators of conditioned responses in Pavlovian conditioning protocols.
Abstract This study examines the distinction between knowing the meaning of a word and experiencing the feelings associated with it. We collected affective ratings for a set of emotional and neutral English words from a group of English native speakers and a group of European Portuguese–English bilinguals. Half of the emotional words named emotions (emotion words) and the other half did not name emotions but could provoke them (emotion-laden words). Some participants were asked to focus on the meaning of words while others were asked to focus on the feeling produced by the words. Native speakers of English produced more intense affective ratings that Portuguese–English bilinguals. Such difference was larger when participants focused on their feelings than when they focused on the words’ meaning. Accordingly, such distinction should be considered in the study of bilingual affective language processing. Finally, the type of emotional word (emotion vs. emotion-laden) had only modest effects.
OBJECTIVES: Gender has been suggested to play a critical role in how facial expressions of pain are perceived by others. With the present study we aim to further investigate how gender might impact the decoding of facial expressions of pain, (i) by varying both the gender of the observer as well as the gender of the expressor and (ii) by considering two different aspects of the decoding process, namely intensity decoding and pain recognition. METHODS: In two online-studies, videos of facial expressions of pain as well as of anger and disgust displayed by male and female avatars were presented to male and female participants. In the first study, valence and arousal ratings were assessed (intensity decoding) and in the second study, participants provided intensity ratings for different affective states, that allowed for assessing intensity decoding as well as pain recognition. RESULTS: The gender of the avatar significantly affected the intensity decoding of facial expressions of pain, with higher ratings (arousal, valence, pain intensity) for female compared to male avatars. In contrast, the gender of the observer had no significant impact on intensity decoding. With regard to pain recognition (differentiating pain from anger and disgust), neither the gender of the avatar, nor the gender of the observer had any affect. CONCLUSIONS: Only the gender of the expressor seems to have a substantial impact on the decoding of facial expressions of pain, whereas the gender of the observer seems of less relevance. Reasons for the tendency to see more pain in female faces might be due to psychosocial factors (e.g., gender stereotypes) and require further research.
The value of quality treebanks is steadily increasing due to the crucial role they play in the development of natural language processing tools. The creation of such treebanks is enormously labor-intensive and time-consuming. Especially when the size of treebanks is considered, tools that support the annotation process are essential. Various annotation tools have been proposed, however, they are often not suitable for agglutinative languages such as Turkish. BoAT v1 was developed for annotating dependency relations and was subsequently used to create the manually annotated BOUN Treebank (UD_Turkish-BOUN). In this work, we report on the design and implementation of a dependency annotation tool BoAT v2 based on the experiences gained from the use of BoAT v1, which revealed several opportunities for improvement. BoAT v2 is a multi-user and web-based dependency annotation tool that is designed with a focus on the annotator user experience to yield valid annotations. The main objectives of the tool are to: (1) support creating valid and consistent annotations with increased speed, (2) significantly improve the user experience of the annotator, (3) support collaboration among annotators, and (4) provide an open-source and easily deployable web-based annotation tool with a flexible application programming interface (API) to benefit the scientific community. This paper discusses the requirements elicitation, design, and implementation of BoAT v2 along with examples.
Fully data-driven, deep learning-based models are usually designed as language-independent and have been shown to be successful for many natural language processing tasks. However, when the studied language is not high-resource and the amount of training data is insufficient, these models can benefit from the integration of natural language grammar-based information.We propose two approaches to dependency parsing especially for languages with restricted amount of training data. Our first approach combines a state-of-the-art deep learning-based parser with a rule-based approach and the second one incorporates morphological information into the parser. In the rule-based approach, the parsing decisions made by the rules are encoded and concatenated with the vector representations of the input words as additional information to the deep network. The morphology-based approach proposes different methods to include the morphological structure of words into the parser network. Experiments are conducted on three different Turkish treebanks and the results suggest that integration of explicit knowledge about the target language to a neural parser through a rule-based parsing system and morphological analysis leads to more accurate annotations and hence, increases the parsing performance in terms of attachment scores. The proposed methods are developed for Turkish, but can be adapted to other languages as well.
In recent years, cross-lingual transfer learning has been gaining positive trends across NLP tasks. This research aims to develop a dependency parser for Indonesian using cross-lingual transfer learning. The dependency parser uses a Transformer as the encoder layer and a deep biaffine attention decoder as the decoder layer. The model is trained using a transfer learning approach from a source language to our target language with fine-tuning. We choose four languages as the source domain for comparison: French, Italian, Slovenian, and English. Our proposed approach is able to improve the performance of the dependency parser model for Indonesian as the target domain on both same-domain and cross-domain testing. Compared to the baseline model, our best model increases UAS up to 4.31% and LAS up to 4.46%. Among the chosen source languages of dependency treebanks, French and Italian that are selected based on LangRank output perform better than other languages selected based on other criteria. French, which has the highest rank from LangRank, performs the best on cross-lingual transfer learning for the dependency parser model.
OBJECTIVE: Negative affect and food insecurity have been proposed to impede adherence to weight loss interventions. Therefore, this study examined the role of these variables on dietary adherence using Ecological Momentary Assessment. METHODS: ) received a weight-maintaining energy needs (WMEN) diet, and participants with obesity (BMI ≥ 30) were randomized to receive either a WMEN diet (n = 14) or a 35% calorie-reduced diet (n = 14). Food insecurity was measured, and, twice daily, Ecological Momentary Assessment captured real-time affect ratings and adherence. Between-person (trait-level) and lagged within-person (state-level) scores were calculated. RESULTS: Greater food insecurity and trait-level negative affect were associated with reduced adherence (p = 0.0015, p = 0.0002, respectively), whereas higher trait-level positive affect was associated with greater adherence (p < 0.0001). Significant interactions between affect and food insecurity revealed an association between higher trait positive affect and increased adherence at lower levels of food insecurity. Higher trait negative affect was more strongly associated with decreased adherence in participants with greater levels of food insecurity (-1 SD: B = -0.21, p = 0.22; mean: B = -0.46, SE = 0.13, p = 0.0004; +1 SD: B = -0.71, SE = 0.17, p < 0.0001). CONCLUSIONS: Trait-level affect may be crucial in predicting dietary adherence, especially in those with greater food insecurity.
Children have an intuitive predilection to play with language and respond to language play. Multilingual children may demonstrate additional talents and characteristics in using language playfully as a result of being able to access multiple cultural and linguistic resources. This chapter provides an overview of the effects of multilingualism on children’s language play by first addressing how language play is used by children who are learning a new language in the classroom environment. It then presents a case study of how two simultaneous trilingual siblings displayed their dexterity in the use of ludic language in the everyday context. The evidence suggests that multilingual children tend to use language play to transcend the linguistic norms of their ambient languages to negotiate meaning, leverage their communicative intents, and develop their unique multilingual identity. The chapter suggests that multilingual children tend to use language play to synthesize hybrid elements from their languages and cultures and create a wide variety of new meanings that no single linguistic system can offer. This syncretic nature of multilingual children’s language play enables them to develop a nuanced and creative manner of communication.
Numerous traditional medical imaging methods, including computed tomography with X‐rays, positron emission tomography (PET), and magnetic resonance imaging (MRI), are utilized frequently in medical settings to screen for illnesses, diagnose patients, and track the effectiveness of treatments. When examining bone protrusions, CT is preferred over MRI for scanning connective tissue. Although the picture quality of PET is inferior to that of CT and MR, it is outstanding for detecting the molecular markers and metabolic functions of illnesses. To give high‐resolution structural pictures and improved ailment sensitivity and specificity within another image, multimodal data and substantial therapeutic influence on advanced diagnostics and therapeutics have been used. The goal was to evaluate the clinical significance of multimodal photoacoustic/ultrasound (PA/US) articular imaging scoring, a cutting‐edge image technique that may show the microvessels and oxygen levels of rheumatoid arthritis‐related inflamed joints (RA). The PA/US imaging technology analyzed seven tiny joints. The PA and power Doppler (PD) impulses were semiquantified using a 0–3 grading scale, and the averages of the PA and PD scores for the seven joints are computed. Three PA+SO 2 types were found determined by the relative oxygen levels (SO 2 ) measurements of the affected joints. Researchers evaluated the relationships between the disease activity ratings and the PA/US imaging ratings. The PA scores and medical ratings that reflect the extent of the pain have strong relationships with each other, as do the PA+SO 2 combinations. PA may be clinically useful in assessing RA. Thus, the research evaluated the clinical symptoms of inflammatory arthritis using a multimodal photoacoustic image process.
Abstract Older adults are more dependent on their surrounding environment. Extensive research has demonstrated beneficial effects of both nature and built environment on mental health of older people. However, most previous research used cross-sectional designs failed to test the intraindividual variability between environment, behavior, and mental wellness in daily life. We used ecological momentary assessment (EMA), activity sensors, and GPS tracking to examine the association between real-time environment, mobility and activity, and momentary affect among older adults in Hong Kong. Data collection and data processing was conducted from December 2021 to May 2022. 168 older adults aged 65 to 84 received seven EMA prompts per day during a fifteen-day period, and completed a total of 17,345 momentary assessments of affective states, mobility, and activities. A set of GPS-derived indicators were used to measure the real-time environment. To disaggregate the between- and within-person effects, we used multilevel models to estimate three dimensions of affect, i.e., valence, calmness, and energetic arousal, in EMA observations, nested within individual participants. Preliminary results indicate significant concurrent associations between environmental attributes and momentary affect at the within-person level, while the between-person differences appear to be either null or modest. Being out of home is associated with higher valence ratings (b=0.04, p=0.0427), while exposure to green is associated with a lower level of energetic arousal (b=-0.03, p=0.0163). Greater walkability is consistently associated with higher momentary affect ratings in three dimensions, but these associations are not statistically significant. Implications of these findings for promoting healthy aging will be discussed.
Public and private decision-making on health problems relies on scientific evidence. However, scientific knowledge includes uncertainty, as does knowledge about COVID-19. In an experimental study, we tested how the trustworthiness (on the three dimensions expertise, integrity, and benevolence) of a source of information (either a scientist or a politician), was affected when messages were either two-sided (including arguments pro and contra the effectiveness of mask-wearing) or one-sided (only pro arguments). Results showed that scientists were ascribed more expertise and integrity compared to politicians, and both sources were ascribed more expertise when they gave two-sided (instead of one-sided) information. Moreover, trustworthiness ratings on all three dimensions were affected by participants' prior topic attitudes and epistemic certainty beliefs. These findings underline that when a source provides two-sided information, this may increase people's willingness to trust that source. To use this strategy most effectively in health communication, more research should be done on how many and what types of counterarguments to include.
Physical proximity is important in social interactions. Here, we assessed whether simulated physical proximity modulates the perceived intensity of facial emotional expressions and their associated physiological signatures during observation or imitation of these expressions. Forty-four healthy volunteers rated intensities of dynamic angry or happy facial expressions, presented at two simulated locations, proximal (0.5 m) and distant (3 m) from the participants. We tested whether simulated physical proximity affected the spontaneous (in the observation task) and voluntary (in the imitation task) physiological responses (activity of the corrugator supercilii face muscle and pupil diameter) as well as subsequent ratings of emotional intensity. Angry expressions provoked relative activation of the corrugator supercilii muscle and pupil dilation, whereas happy expressions induced a decrease in corrugator supercilii muscle activity. In proximal condition, these responses were enhanced during both observation and imitation of the facial expressions, and were accompanied by an increase in subsequent affective ratings. In addition, individual variations in condition related EMG activation during imitation of angry expressions predicted increase in subsequent emotional ratings. In sum, our results reveal novel insights about the impact of physical proximity in the perception of emotional expressions, with early proximity-induced enhancements of physiological responses followed by an increased intensity rating of facial emotional expressions.
"Words and pictures stimuli are often used in study of perception, language, and memory. More and more studies are being done on how emotional words or pictures influence different cognitive processing. However, the emotional rating process of these stimuli has rarely been studied in young children. Especially, no study has investigated emotional rating process on pre-schoolers. This research examines how young children process emotional words and pictures stimuli. More precisely, we measured age (4, 5, and 6-years-old) and sex differences (girls and boys) in emotional valence rating of pictures and words. A corpus of 90 words and 90 pictures was selected from among the emotional databases compiled by Alario & Ferrand (1999), Bonin et al. (2003), Cannard et al. (2006) Syssau & Monnier (2009). This corpus was rated by 92 French children (28 four-years-old children, 16 girls and 12 boys; 34 five-years-old children, 14 girls and 20 boys; and 30 six-years-old children, 13 girls and 17 boys). These ratings were made using a three points emotional valence rating scale (negative, neutral, and positive) based on AEJE scale (Largy, 2018). To keep the rating task simple for the children, the scale labels were using drawings of faces. The 90 Words and 90 pictures were divided in sets of 15 stimuli. Each child rated all sets of stimuli in separate sessions. These sessions were in a random order between words and pictures stimuli sets. Good response reliability was observed in the three age groups. We assessed age differences in the valence ratings: Four-year-old children shown lower mean scores in valence rating (positive, neutral, and negative) than did five-year-old ones who shown lower mean scores in valence rating than did six-year-old ones. Despite a lack of consensus in the literature, we found sex differences in the valence ratings. Girls in each age groups shown higher mean scores in valence rating than did boys. Moreover, results shown a significant difference between pictures and words ratings. Children better rated words than pictures in each age group and sex. Besides, analyses revealed significant differences in emotional valence rating between negative, neutral, and positive words and pictures stimuli. Positive words and pictures stimuli were better rated by children than negative ones which were better rated than neutral ones. Future research will compile this corpus in a database, and it could become a worthwhile tool to control emotional verbal and visual stimuli in experimental design for children."
While face masks provide necessary protection against disease spread, they occlude the lower face parts (chin, mouth, nose) and consequently impair the ability to accurately perceive facial emotions. Here we examined how wearing face masks impacted making inferences about emotional states of others (i.e., affective theory of mind; Experiment 1) and sharing of emotions with others (i.e., affective empathy; Experiment 2). We also investigated whether wearing transparent masks ameliorated the occlusion impact of opaque masks. Participants viewed emotional faces presented within matching positive (happy), negative (sad), or neutral contexts. The faces wore opaque masks, transparent masks, or no masks. In Experiment 1, participants rated the protagonists' emotional valence and intensity. In Experiment 2, they indicated their empathy for the protagonist and the valence of their emotion. Wearing opaque masks impacted both affective theory of mind and affective empathy ratings. Compared to no masks, wearing opaque masks resulted in assumptions that the protagonist was feeling less intense and more neutral emotions. Wearing opaque masks also reduced positive empathy for the protagonist and resulted in more neutral shared valence ratings. Wearing transparent masks restored the affective theory of mind ratings but did not restore empathy ratings. Thus, wearing face masks impairs nonverbal social communication, with transparent masks able to restore some of the negative effects brought about by opaque masks. Implications for the theoretical understanding of socioemotional processing as well as for educational and professional settings are discussed.
Blocking facial mimicry can disrupt recognition of emotion stimuli. Many previous studies have focused on facial expressions, and it remains unclear whether this generalises to other types of emotional expressions. Furthermore, by emphasizing categorical recognition judgments, previous studies neglected the role of mimicry in other processing stages, including dimensional (valence and arousal) evaluations. In the study presented herein, we addressed both issues by asking participants to listen to brief non-verbal vocalizations of four emotion categories (anger, disgust, fear, happiness) and neutral sounds under two conditions. One of the conditions included blocking facial mimicry by creating constant tension on the lower face muscles, in the other condition facial muscles remained relaxed. After each stimulus presentation, participants evaluated sounds' category, valence, and arousal. Although the blocking manipulation did not influence emotion recognition, it led to higher valence ratings in a non-category-specific manner, including neutral sounds. Our findings suggest that somatosensory and motor feedback play a role in the evaluation of affect vocalizations, perhaps introducing a directional bias. This distinction between stimulus recognition, stimulus categorization, and stimulus evaluation is important for understanding what cognitive and emotional processing stages involve somatosensory and motor processes.
We describe Turkish Discourse Bank 1.2, the latest version of a discourse corpus annotated for explicitly or implicitly conveyed discourse relations, their constitutive units, and senses in the Penn Discourse Treebank style. We present an evaluation of the recently added tokens and examine three commonly occurring dependency patterns that hold among the constitutive units of a pair of adjacent discourse relations, namely, shared arguments, full embedding and partial containment of a discourse relation. We present three major findings: (a) implicitly conveyed relations occur more often than explicitly conveyed relations in the data; (b) it is much more common for two adjacent implicit discourse relations to share an argument than for two adjacent explicit relations to do so; (c) both full embedding and partial containment of discourse relations are pervasive in the corpus, which can be partly due to subordinator connectives whose preposed subordinate clause tends to be selected together with the matrix clause rather than being selected alone. Finally, we briefly discuss the implications of our findings for Turkish discourse parsing.
Online lexicographical resources for the morphologically rich Indigenous languages in Canada use a wide range of strategies for conveying their language’s morphological system, i.e. how words are inflected and derived, which this paper illustrates in a survey of seventeen bilingual online resources. The strategies these resources employ boil down to two basic approaches to the underlying structure of the resource: 1) a lexical database, or 2) a computational model. Most resources we surveyed are constructed around lexical databases. These assume the word(form) as the basic unit, an assumption that makes it difficult to incorporate the language’s sub-word, morphological structure in full detail. However, one resource uses a computational morphological model to bring the language’s morphology into the core of the lexicon – this proved to be a “low-hanging fruit” in the application of language technology that had been accomplished within a reasonable time-frame, as has been advocated by Trond Trosterud. We discuss the value created and questions raised by this approach and argue that it successfully overcomes the traditional Boasian three-way partition of dictionary, grammar, and text, creating integrated language resources that meet the modern needs of low-resource endangered languages and their communities.
The in-out effect refers to the tendency that novel words whose consonants follow an inward-wandering pattern (e.g., P-T-K) are rated more positively than stimuli whose consonants follow an outward-wandering pattern (e.g., K-T-P). While this effect appears to be reliable, it is not yet clear to what extent it generalizes to existing words in a language. In two large-scale studies, we sought to extend the in-out effect from pseudowords to real words and from perception to production. In Study 1, we investigated whether previously collected affective ratings for English and Dutch words were more positive for inward-wandering words and more negative for outward-wandering words. No systematic relationship between wandering direction and affective valence was found. In Study 2, we investigated whether inward-wandering words are more likely to occur in positive online consumer restaurant reviews written in English and Dutch, compared to negative reviews, and whether this association was stronger for food ratings than for decor ratings. Again, no systematic relationship between wandering direction and review rating emerged. We suggest that the affective states triggered by different consonantal wandering directions might be used as a cue for forming judgments in the absence of other information, but that wandering direction is too low in salience to drive the shape of words in the lexicon.
This paper aims to describe the building-up of SciE-Lex (http://www.ub.edu/grelic/eng/scielex2/scielex.html), a collocational database of non-specialized terms in biomedical English, which was primarily conceived as a response to the lack of reference tools accounting for the lexicogrammatical patterning associated with non-technical terms frequently used in the health science discourse. SciE-Lex thus serves the purpose of assisting L2 English writers from the health science discourse community in their production of biomedical texts in English. This collocational database is the result of a lexicographic project carried out by the GreLiC research group at the University of Barcelona, and as such it has undergone various developmental stages since its inception. In order to evaluate its adequacy as a writing tool addressed to the Spanish biomedical community and confirm the appropriateness of the combinatorial patterning and phraseological information included in each entry, a group of language experts were asked to assess the dictionary by stressing both its weaknesses and strengths. Their feedback has stressed the suitability of SciE-Lex as a lexicographic resource and yielded significant improvement of the tool. Last but not least, SciE-Lex has also been successfully tested with targeted users in a series of “Writing for publication” workshops held at the University of Barcelona, and taught by the author, the results of which have corroborated the usefulness of this lexical database to enhance Spanish users’ biomedical English published writing.
Abstract We present an event structure classification empirically derived from inferential properties annotated on sentence- and document-level Universal Decompositional Semantics (UDS) graphs. We induce this classification jointly with semantic role, entity, and event-event relation classifications using a document-level generative model structured by these graphs. To support this induction, we augment existing annotations found in the UDS1.0 dataset, which covers the entirety of the English Web Treebank, with an array of inferential properties capturing fine-grained aspects of the temporal and aspectual structure of events. The resulting dataset (available at decomp.io) is the largest annotation of event structure and (partial) event coreference to date.
This dataset contains all image, complexity and distinctiveness data that was used for: Han, S. J, Kelly, P., Winters, J., & Kemp, C. (2022). Simplification is not dominant in the evolution of Chinese characters. <em>Open Mind</em>. The code for this project can be found here. The file uploaded here is intended to replace the sample data folder that is available in the code repository. Our dataset includes data scraped from hanziyuan.net, as well as data from the following sources: Sun, C. C., Hendrix, P., Ma, J., & Baayen, R. H. (2018). Chinese lexical database (CLD): A large-scale lexical database for simplified Mandarin Chinese. Behavior Research Methods, 50(6), 2606–2629. Wikimedia Commons. (2021). Chinese characters decomposition. https://commons.wikimedia.org/wiki/Commons:Chinese_characters_decomposition Liu, C.-L., Yin, F., Wang, D.-H., & Wang, Q.-F. (2011). CASIA online and offline Chinese handwriting databases. In 2011 international conference on document analysis and recognition (pp. 37–41). https://doi.org/10.1109/ICDAR.2011.17 Chen, P.-C. (2020). Traditional Chinese handwriting dataset. GitHub. https://github.com/AI-FREE-Team/Traditional-Chinese-Handwriting-Dataset
TuLaR (Tupian Language Resources) is a project for collecting, documenting, analyzing, and developing computational and pedagogical material for low-resource Brazilian indigenous languages. It provides valuable data for language research regarding typological, syntactic, morphological, and phonological aspects. Here we present TuLaR's databases, with special consideration to TuDeT (Tupian Dependency Treebanks), an annotated corpus under development for nine languages of the Tupian family, built upon the Universal Dependencies framework. The annotation within such a framework serves a twofold goal: enriching the linguistic documentation of the Tupian languages due to the rapid and consistent annotation, and providing computational resources for those languages, thanks to the suitability of our framework for developing NLP tools. We likewise present a related lexical database, some tools developed by the project, and examine future goals for our initiative.
AbstractThe most important objective of our study was to build and construct a complete and comprehensive morphological analyzing scheme, tagging, and parsing system which can be used for annotating Arabic corpora. In our dissertation, we did an analytical study, implementation, and evaluation of Arabic morphological analysis, tagging, and syntactic analysis starting from raw text. The three different systems were implemented in various methods; for morphological analysis, we use finite-state automaton as discussed in chapter four, after doing the process of tokenization and segmentation of the raw text as explained in chapter three. The tagging system was implemented under a new very rich tag set, which was designed and developed by us. It consists of 30 tags in addition to some other features and linguistic information as described in chapter five. An appropriate set of tags has a direct influence on the accuracy and the usefulness of tagging system. So, the smaller the tag set, the higher the accuracy, and the larger the tag set, the lower the accuracy. Regarding syntactic annotation, we started with developing a syntactic tag for the Arabic language as explained in chapter six. Based on the syntactic tag we started developing chunker by analyzing sentence structure categories that are identified using some grammatical theory. And the last step is parsing which is very common for sentence structure annotation. We tried to bring in as many details as possible for sentence analysis as described in chapter six as well. This thesis provides an overview and annotation scheme for representing the NLP analysis, at the most important stages, such as lexical internal structure level, part-of-speech level, chunk level, and parsing at the sentence level. The main reason for our dissertation was to annotate Arabic corpora morpho-syntactically. The annotation was done semi-automatically which is more efficient and significant to tackle and present solutions for most of the Arabic language problems. The importance of annotated corpora in the present day of natural language processing is widely known. The main idea behind corpus annotation is to add value to a corpus and that will help to be a source of linguistic information for future research and development. Annotated corpora serve as an important tool for investigators of natural language processing, speech recognition, information retrieval, and other related areas. It proves to be a basic building block for developing and constructing different models, tools, and programs for automatic and semi-automatic processing of natural languages.It is very hard and difficult to encode manually all the information needed to encode the natural language to develop and build a tool or program that will annotate text with some necessary information (Brill, E, 1994). Annotation of corpora can be done at various levels such as part of speech, phrase/clause level, dependency level, etc. Part of speech tagging forms the basic step towards building an annotated corpus. Chunking can form the next level of tagging and then parsing the full sentence. Parsing is usually performed after basic morphosyntactic categories have been identified in a text, it brings these categories into higher-level syntactic relationships with one another. Corpora that have been parsed are sometimes known as treebanks as discussed in chapter six.There are various advantages of studying corpus, particularly for the areas of NLP; Machine Translation, Computational Lexicography, Stylistics, etc. It is extremely useful not only in building and formulating rules but also in testing them. Studies on a corpus with a focus on a particular linguistic structure can lead to a better understanding of the problems involved. Corpus can be used as an immediate source to extract and retrieve useful information in building natural language applications.
Cite the source of the dataset as: Fabrício Ferraz Gerardi, Stanislav Reichert, Carolina Aragon, Johann-Mattis List, & Tim Wientzek. (2021). TuLeD: Tupían lexical database. Max Planck Institute for Evolutionary Anthropology: Leipzig
2), designed by Dennis Shasha (NYU, Computer Science).
Collection of Ancient Greek annotated trees of Artemidorus' Oneirocritica Book 5. Part of the Open Projects in Digital Classics at the College of Letters and Sciences of the State University of São Paulo in Araraquara, São Paulo, Brazil. The trees were annotated manually on Perseids Platform using the Arethusa tool. The treebank tagset and guidelines used were those from The Ancient Greek Dependency Treebank with the morphological and syntactic layer, which was based on Bamman's and Crane's 2008 Guidelines for the Syntactic Annotation of the Ancient Greek Dependency Treebank (1.1). We translated it into Portuguese with a few additions from some specifications provided by a forum maintained in 2013 by Alpheios.net, which is no longer online. The trees are visible in Perseids Collection as UNESP-trees at https://perseids-publications.github.io/unesp-trees.
Analogies between 4 sentences, “a is to b as c is to d”, are usually defined between two pairs of sentences (a, b) and (c, d) by constraining a relation R holding between the sentences of the first pair, to hold for the second pair. From a theoretical perspective, three postulates define an analogy - one of which is the “central permutation” postulate which allows the permutation of central elements b and c. This postulate is no longer appropriate in sentence analogies since the existence of R offers no guarantee in general for the existence of some relation S such that S also holds for the pairs (a, c) and (b, d). In this paper, the “central permutation” postulate is replaced by a weaker “internal reversal” postulate to provide an appropriate definition of sentence analogies. To empirically validate the aforementioned postulate, we build a LSTM as well as baseline Random Forest models capable of learning analogies based on quadruplets. We use the Penn Discourse Treebank (PDTB), the Stanford Natural Language Inference (SNLI) and the Microsoft Research Paraphrase (MSRP) corpora. Our experiments show that our models trained on samples of analogies between (a, b) and (c, d), recognize analogies between (b, a) and (d, c) when the underlying relation is symmetrical, validating thus the formal model of sentence analogies using “internal reversal” postulate. © 2022 Copyright for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0).
Reviewed by: Postclassical Greek: Contemporary Approaches to Philology and Linguistics ed. by Dariya Rafiyenko and Ilja A. Seržant William A. Ross dariya rafiyenko and ilja a. seržant (eds.), Postclassical Greek: Contemporary Approaches to Philology and Linguistics (Trends in Linguistics: Studies and Monographs 335; Berlin: de Gruyter, 2020). Pp. viii + 339. $114.99. This volume brings together a slate of scholars to examine linguistic aspects of postclassical Greek, including the biblical corpus, in an interdisciplinary context. It begins with an essay by the editors Dariya Rafiyenko and Ilja A. Seržant that offers an overview of postclassical Greek itself, which they define as "the entire set of spoken and written varieties of the period from 323 BC up to 1453 AD" (p. 1). Their essay briefly surveys issues of periodization and grammar, with an emphasis on the importance of variation in linguistic description, for which reason Rafiyenko and Seržant join others—rightly so in my opinion—in questioning the separation of linguistics from philology that has prevailed throughout much of the last century. They advocate instead what they call the "rephilologization of historical linguistics" (p. 12). The volume is divided into two sections. The first is "Grammatical Categories," which begins with "Purpose and Result Clauses: ἵνα-hína and ὥστε-hṓste in the Greek Documentary Papyri of the Roman Period," by Giuseppina di Bartolo, who surveys the disappearance of certain semantic distinctions in different moods and a reduction of conjunctions to introduce purpose and result clauses. In "Syntactic Factors in the Greek Genitive-Dative Syncretism: The Contribution of New Testament Greek," Chiara Gianollo considers postnominal genitives in the NT, which allowed for the propagation of an external possession construction. Then, in "Future Periphrases in John Malalas," Daniel Kölligan looks at the earliest Byzantine world chronicler, in whose work certain novel, but not idiosyncratic, expressions of future reference appear alongside older patterns. The next essay is "Combining Linguistics, Paleography and Papyrology: The Use of eis, pros and epí in Greek Papyri," by Joanne Stolk, who contrasts these prepositions with the bare dative in expressions of animate goal of motion and transfer verbs, providing new interpretations for uses otherwise considered exceptional. In "Future Forms in Postclassical Greek: Some Remarks on the Septuagint and the New Testament," Liana Tronci examines diachronic morphology to mark future tense, taking external factors such as register variation within the biblical corpus into account. The next essay, by Brian D. Joseph, is entitled "Greek Infinitive-Retreat versus Grammaticalization: An Assessment"; he argues that the morphosyntactic changes in infinitive use represent degrammaticalization and thus that grammatical change is not always unidirectional. The last essay in the first section of this volume is "Postclassical Greek and Treebanks for a Diachronic Analysis," by Nikolaos Lavidas and Dag Trygve Truslew Haug. They employ diachronic analysis to examine syntactic phenomena such as backward control and thus build a linguistic profile of earlier texts. The second section in this volume includes topics related to sociolinguistic aspects and variation in postclassical Greek. The first essay, by Marina Benedetti, is "The Perfect [End Page 538] Paradigm in Theodosius' Κανόνες: Diathetically Indifferent and Diathetically Non-Indifferent Forms," in which she examines the discussion of the perfect in a fourth-century c.e. grammatical treatise. Next, in "Forms of the Directive Speech Act: Evidence from Early Ptolemaic Papyri," Carla Bruno considers strategies available for inducing others to action, including performative utterances and indirect implicatures, among other possibilities. The next essay is one of the longest in the book: "What's in a (personal) Name? Morphology and Identity in Jewish Greek Literature in the Hellenistic and Roman Periods," by Robert Crellin. He challenges the notion that non-nativized morphology of Hebrew personal names implies low-level Greek, arguing instead that sociolinguistic and literary considerations are at work in the decision to adapt personal names or not. The next essay is "Confusion of Mood or Phoneme? The Impact of L1 Phonology on Verb Semantics," by Sonja Dahlgren and Martti Leiwo, who question whether the extensive nonstandard vowel use in Egyptian Greek texts is evidence of poor command of Greek. They suggest instead that such spellings represent underdifferentiated Greek phonemes and transfer elements from non-Greek prosodic systems. The last...
The present paper represents an attempt to shed some light upon the phenomenon of language death in Iraqi dialect in relation to the social factors of age, gender and region. For this purpose, 400 informants (200 males and 200 females) have been approached so as to elicit relevant data. They are subdivided according to age group, gender as well as region. The sample used in this investigation is native speakers of Iraqi Arabic. The model is used according to Trudgill (1974). It could be said that the age of the speaker is a social factor with which linguistic variables have been found to correlate very closely. The normal pattern of style and age differentiation occurs when the youngest and the oldest represent the highest scores. When they become older, they will be more influenced by the values of the society. So, they will be affected by their social networks. Thus, they are linguistically more influenced by the standard language. On the other hand, older and retired people are away from the standard forms due to the effect of their peer group. Again, their social networks are narrower than that of the middle aged groups.There is another social pressure on women to acquire correct linguistic norms. Men tend to be favorably disposed to low status speech forms than women. This is may be because of the connotations of 'toughness' and 'masculinity' associated with working class language, so it could be said that woman don't use the dying word.
In parsing phrase structures, supertagging achieves a symbiosis between the interpretability of formal grammars and the accuracy and speed of more recent neural models.The approach was only recently transferred to parsing discontinuous constituency structures with linear context-free rewriting systems (LCFRS).We reformulate and parameterize the previously fixed extraction process for LCFRS supertags with the aim to improve the overall parsing quality.These parameters are set in the context of several steps in the extraction process and are used to control the granularity of extracted grammar rules as well as the association of lexical symbols with each supertag.We evaluate the influence of the parameters on the sets of extracted supertags and the parsing quality using three treebanks in the English and German language, and we compare the best-performing configurations to recent state-of-the-art parsers in the area.Our results show that some of our configurations and the slightly modified parsing process improve the quality and speed of parsing with our supertags over the previous approach.Moreover, we achieve parsing scores that either surpass or are among the state-of-the-art in discontinuous constituent parsing.
This paper presents and discusses a discourse relation annotation scheme for the MUCH corpus of academic writing, based on Rhetorical Structure Theory (RST). The set of proposed relational tags takes into regard both distinctiveness, pedagogical needs and implementability with automatic rules. We show how a pilot grammar with 180 rules can map discourse relations between existing syntactic nodes, exploiting lower-level grammatical/treebank markup and surface clues such as connectives (e.g. conjunctions and prepositions). In an evaluation of a live run on student essays from teacher training courses, the average false positive rate across the most frequent 21 categories was 26.7 % for tags and 17.1 % for relation links. Performance was best for categories with a high percentage of rules using surface connectives and, for in-sentence relations, their corresponding dependency links.
The first time such a newly emerged linguistic variety has been documented at NUML, Islamabad. This language variation means morphological, semantic, connotative deviations from standard linguistic norms. Contact linguistics, code meshing, diversified linguistic backgrounds and their colloquialism play their vital roles in shaping Numlianlect of BS English. The research problem faced by learners of super standard language is that they do not comprehend the newly code meshed discourse in the informal settings of NUML. So, this research raises questions about the exploration of linguistic variations, their formulation process and their denotative and connotative semantic shades. To answer these questions, qualitative and survey-based research design has been employed. The collected data have been analyzed with Labov and Weinreich’s Variation Theory (1960). Major findings reveal that students mix English, Punjabi, Sindhi, Pashto and Urdu words to produce variant code-meshed linguistic patterns. Connotation, semantic shades and contexts of words have been changed by students.
Abstract The affective variability of Bipolar Disorder (BD) is thought to qualitatively differ from that of Borderline Personality Disorder (BPD), with changes in affect persisting for longer in BD. However, quantitative studies have not been able to confirm this distinction. It has therefore not been possible to accurately quantify how treatments like lithium influence affective variability in BD. We assessed the affective variability associated with BD and BPD as well as the effect of lithium using a novel computational model that defines two subtypes of variability: affective changes that persist (volatility) and changes that do not (noise). We hypothesized that affective volatility would be raised in the BD group, noise would be raised in the BPD group and that lithium would impact affective volatility. Daily affect ratings were prospectively collected for up to 3 years from patients with BD, BPD and non-clinical controls. In a separate experimental-medicine study, patients with BD were randomized to receive lithium or placebo, with affect ratings collected from week -2 to +4. We found a diagnostically specific pattern of affective variability. Affective volatility was raised in patients with BD whereas affective noise was raised in patients with BPD. Rather than suppressing affective variability, lithium increased the volatility of positive affect in both studies. These results provide a quantitative measure of the affective variability associated with BD and BPD. They suggest a novel mechanism of action for lithium, whereby periods of persistently low or high affect are avoided by increasing the volatility of affective responses.
Familiarity could be a key attribute of attention for visual perception of objects in our environment.Subjective ratings of familiarity with visual objects could be influenced by individuals’ attention and judgment ability. The current study investigated the relationship between subjective familiarity perception of visual objects with the objective measures of individuals’ attention and judgment ability in forty-seven healthy participants. Familiarity ratings of the visual objects belonging to four categories (Human, Inanimate, Animal, Plant) were collected from the participants. Attention and judgment ability of the participants were assessed objectively using attentional blink task and relative area estimation task respectively. Individuals’ who were objectively less attentive and prone to judgment errors manifested a high propensity to subjectively rate the visual objects as more familiar. The current study findings indicate the importance of objective cognitive assessment for fitness testing of jobs that involve one’s subjective visual perception such as pilots, air traffic controllers and navy personnel.
Abstract How we perceive the world is not solely determined by our experiences at a given moment in time, but also by what we have experienced in our immediate past. Here, we investigated whether such sequential effects influence the affective appraisal of food images. Participants from 16 different countries ( N = 1278) watched a randomly presented sequence of 60 different food images and reported their affective appraisal of each image in terms of valence and arousal. For both measures, we conducted an inter-trial analysis, based on whether the rating on the preceding trial(s) was low or high. The analyses showed that valence and arousal ratings for a given food image are both assimilated towards the ratings on the previous trial (i.e., a positive serial dependence). For a given trial, the arousal rating depends on the arousal ratings up to three trials back. For valence, we observed a positive dependence for the immediately preceding trial only, while a negative (repulsive) dependence was present up to four trials back. These inter-trial effects were larger for males than for females, but independent of the participants’ BMI, age, and cultural background. The results of this exploratory study may be relevant for the design of websites of food delivery services and restaurant menus.
This is a small hand-annotated partial treebank of Tibetan, primarily in CoNLL-U format. It builds upon the following corpus: Hill, Nathan W., & Garrett, Edward. (2017). A part-of-speech (POS) tagged corpus of Classical Tibetan [Data set]. Zenodo. http://doi.org/10.5281/zenodo.574878 This corpus differs from the above in three ways: The tagset has been converted from the SOAS tag system to the Universal Dependency part-of-speech tagset. We have added dependency relations between verbs and their argument. For some of the texts, English translations were available in digital form. These translations were manually aligned to the Tibetan texts and included in the CoNLL-U files. It was created as part of the AHRC-funded project <em>Lexicography in Motion </em>(PI Ulrich Pagel, 2017-2021).
The categorical approach to cross-cultural emotion perception research has mainly relied on constrained experimental tasks, which have arguably biased previous findings and attenuated cross-cultural differences. On the other hand, in the constructionist approach, conclusions on the universal nature of valence and arousal have mainly been indirectly drawn based on participants' word-matching or free-sorting behaviors, but studies based on participants' continuous valence and arousal ratings are very scarce. When it comes to self-reports of specific emotion perception, constructionists tend to rely on free labeling, which has its own limitations. In an attempt to move beyond the limitations of previous methods, a new instrument called the Two-Dimensional Affect and Feeling Space (2DAFS) has been developed. The 2DAFS is a useful, innovative, and user-friendly instrument that can easily be integrated in online surveys and allows for the collection of both continuous valence and arousal ratings and categorical emotion perception data in a quick and flexible way. In order to illustrate the usefulness of this tool, a cross-cultural emotion perception study based on the 2DAFS is reported. The results indicate the cross-cultural variation in valence and arousal perception, suggesting that the minimal universality hypothesis might need to be more nuanced.
The gleam-glum effect is a novel sound symbolic finding that words with the /i:/-phoneme (like gleam) are perceived more positive emotionally than matched words with the /Λ/-phoneme (like glum). We provide data that not only confirm the effect but also are consistent with an explanation that /i:/ and /Λ/ articulation tend to co-occur with activation of positive versus negative emotional facial musculature respectively. Three studies eliminate selection bias by including all applicable English words from the English Lexicon Project (Balota et al., 2007) and the Warriner et al. (2013) database and every possible Mandarin Pinyin combination that differ only in the middle phoneme (/i:/ vs /Λ/). In Study 1, 61 U.S. undergraduates rated monosyllabic English /i:/ words as robustly more positive than matched /Λ/ words. Study 2 analyzed the Warriner et al. (2013) valence ratings, extending the gleam-glum effect to all applicable words in the database. In Study 3, 38 U.S. participants (using English) and 37 participants in China (using Mandarin Pinyin) rated word pairs under three conditions that moderate musculature activity: Read aloud (Enhance), read silently (Control), and read silently while chewing gum (Interfere). Indeed, the effect was both replicated and was significantly larger when facial musculature was enhanced than when interfered with, and the two language populations did not significantly differ. These findings confirm a robust gleam-glum effect, despite semantic noise, in English and Mandarin Pinyin. Furthermore, these data are consistent with the hypothesis that this type of sound symbolism arises from the overlap in muscles used both in articulation and emotion expression. (PsycInfo Database Record (c) 2021 APA, all rights reserved).