Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Individuals with High functioning autism (HFA) are distinguished by relative preservation of linguistic and cognitive skills. However, problems with pragmatic language skills have been consistently reported across the autistic spectrum, even when structural language is intact. Our main goal was to investigate how highly verbal individuals with autism process figurative language and whether manipulation of the stimuli presentation modality had an impact on the processing. We were interested in the extent to which visual context, e.g., an image corresponding either to the literal meaning or the figurative meaning of the expression may facilitate responses to such expressions. Participants with HFA and their typically developing peers (matched on intelligence and language level) completed a cross-modal sentence-picture matching task for figurative expressions and their target figurative meaning represented in images. We expected that the individuals with autism would have difficulties in)
We adapted the adult French version of the Basic Empathy Scale to French children aged 6–11 years, in order to probe the factorial structure underlying empathy. A total of 410 children (189 girls and 221 boys) were instructed to fill out the resulting Basic Empathy Scale in Children (BES-C). Results showed that, as in adulthood, the three-factor model of empathy (i.e., emotional contagion, cognitive empathy, and emotional disconnection) was more relevant than the one- and two-factor ones. This means that as early as 6 years of age, children’s responses should reflect the same organization of the three components of empathy as those of adults. In line with the literature, cognitive empathy increased and emotional disconnection decreased in middle childhood, while emotional contagion remained stable. Moreover, girls exhibited greater emotional contagion than boys, with the reverse pattern being observed for emotional disconnection. No sex difference was found regarding cognitive empathy.
Previous research (e.g., cultural consensus theory (Romney, Weller, & Batchelder, American Anthropologist, 88, 313–338, 1986); cultural mixture modeling (Mueller & Veinott, 2008)) has used overt response patterns (i.e., responses to questionnaires and surveys) to identify whether a group shares a single coherent attitude or belief set. Yet many domains in social science have focused on implicit attitudes that are not apparent in overt responses but still may be detected via response time patterns. We propose a method for modeling response times as a mixture of Gaussians, adapting the strong-consensus model of cultural mixture modeling to model this implicit measure of knowledge strength. We report the results of two behavioral experiments and one simulation experiment that establish the usefulness of the approach, as well as some of the boundary conditions under which distinct groups of shared agreement might be recovered, even when the group identity is not known. The results reveal that the ability to recover and identify shared-belief groups depends on (1) the level of noise in the measurement, (2) the differential signals for strong versus weak attitudes, and (3) the similarity between group attitudes. Consequently, the method shows promise for identifying latent groups among a population whose overt attitudes do not differ, but whose implicit or covert attitudes or knowledge may differ.
The proposed paper details a contrastive interlanguage analysis (i.e. Granger, 1996) of metalinguistic features of certainty and doubt including 'hedges' and 'boosters' (following Hyland, 2000) and 'epistemic stance nouns' (Jiang, 2015) in a 350,000 word corpus of L2 written essays and reports collected at three data points (pre-training, post-training and final assessment) during a 6-credit mandatory freshmen English for academic purposes (EAP) course. The paper explores to what extent freshman undergraduate students are more or less certain in their treatment of theirs' or others' claims via the linguistic devices used prior to their EAP training, and what happens to their use of these linguistic devices as a result of their EAP training. Data was collected from 87 participants spread across five classes with the same participants submitting data at each data point. The results suggest significant impacts of time and task-type on the normalised frequencies and individual wordings of hedging and boosting devices, with pre-training data suggesting significantly more overt hedging and boosting devices used than in the final assessment data and with more epistemic nouns used post-training, and with differences in frequencies and wordings of individual devices across essay and report task types. The longitudinal trend in particular is characterised by a reduction in the use of modals for hedging (‘May’, ‘Would’ etc.) and an increase in lexical means, and a drop in categorical/assumption based statements (‘Undeniably’, ‘Obviously’) to a more academic tone. These findings suggest a positive effect of EAP training on L2 writer’s presentation of their stance on their own or others’ claims, towards the linguistic norms of an academic register.
We compared the quality of prediction of word variables based on a Dutch word association and text corpus. Wederived estimates for: valence, arousal, dominance, concreteness and age of acquisition (AoA) for 2831 words. Based on thesimilarity between words we: (1) used projections on a dimension identified as the variable in question in a multidimensionalrepresentation, (2) used the k-nearest neighbors values, weighted according to their proximity. Estimates prevailed when basedon word associations. Differences between the predictions of the two methods were small. Based on the word association corpusit yielded correlations of.92,.85, and.85, for valence, arousal, and dominance, respectively. Its corresponding correlationsbased on the text corpus were.80,.74, and.67. For concreteness and AoA, both the association and the text corpus yieldedcorrelations of.88 and.73, respectively. This suggests word associations are better at capturing human ratings of affective wordvariables.
Word ratings on affective dimensions are an important tool in psycholinguistic research. Traditionally, they are obtained by asking participants to rate words on each dimension, a time-consuming procedure. As such, there has been some interest in computationally generating norms, by extrapolating words’ affective ratings using their semantic similarity to words for which these values are already known. So far, most attempts have derived similarity from word co-occurrence in text corpora. In the current paper, we obtain similarity from word association data. We use these similarity ratings to predict the valence, arousal, and dominance of 14,000 Dutch words with the help of two extrapolation methods: Orientation towards Paradigm Words and k-Nearest Neighbors. The resulting estimates show very high correlations with human ratings when using Orientation towards Paradigm Words, and even higher correlations when using k-Nearest Neighbors. We discuss possible theoretical accounts of our results and compare our findings with previous attempts at computationally generating affective norms.
Tiedemann K. Linguistic norms in mathematics lessons. In: Krainer K, Vondrová N, eds. <em>Proceedings of the Ninth Congress of the European Society for Research in Mathematics Education</em>. 2016: 1503-1509.
The analysis of vocal expression is a critical endeavor for psychological and clinical sciences and is an increasingly popular application for computer–human interfaces. Despite this, and despite advances in the efficiency, affordability, and sophistication of vocal analytic technologies, there is considerable variability across studies regarding what aspects of vocal expression are studied. Vocal signals can be quantified in a myriad of ways, and their underlying structure, at least with respect to “macroscopic” measures from extended speech, is presently unclear. To address this issue, we evaluated the psychometric properties—notably, the structural and construct validity—of a systematically defined set of global vocal features. Our analytic strategy focused on (a) identifying redundant variables among this set, (b) employing principal components analysis (PCA) to identify nonoverlapping domains of vocal expression, (c) examining the degrees to which the vocal variables are modulated as a function of changes in speech task, and (d) evaluating the relationship between the vocal variables and cognitive (i.e., verbal fluency) and clinical (i.e., depression, anxiety, and hostility) variables. Spontaneous speech samples from 11 independent studies of young adults (>60 s in length), employing one of three different speaking tasks, were examined (N = 1,350). Confounding variables (i.e., sex, ethnicity) were statistically controlled for. The PCA identified six distinct domains of vocal expression. Collectively, vocal expression (defined in terms of these domains) was modulated as a function of speech task and was related to the cognitive and clinical variables. These findings provide empirically grounded implications for the study of vocal expression in psychological and clinical sciences.
This paper describes the SimpleNLG-IT realiser, i.e. the main features of the porting of the SimpleNLG API system The paper gives some details about the grammar and the lexicon employed by the system and reports some results about a first evaluation based on a dependency treebank for Italian. A comparison is developed with the previous projects developed for this task for English and French, which is based on the morpho-syntactical differences and similarities between Italian and these languages.
Editorial I am delighted to announce the successful publication of Volume 26, 2020 of our esteemed journal, Lagos Notes and Records. This current edition is made up of thirteen well-researched articles across the various disciplines of the Humanities and Social Sciences namely History, Philosophy, Creative Arts, Language Studies, Literature, Communication Studies, and Linguistics. Lynn Schler in the first article, ‘The Local and the Global in African Studies: An Essay in Honour of Prof. Ayodeji Olukoju @ 60’, argues that in every geographic context, African studies evolved as an intersection between local and global flows of ideas, politics and capital. She concludes that the future of African studies requires scholars to view Africa as both a singular idea and a conglomeration of vastly diverse cultural contexts. Scholars must be aware of what is distinctive in local contexts and also take cognizance of global solutions. In the second article, ‘Identity and Ideological Positioning in Popular Nigerian Ethnic Jokes’, ’Rotimi Taiwo and David Dontele examine the discursive constructions of selected jokes to determine their expression of attitudinal and ideological dispositions of the ethnic groups within the multilingual/multicultural context of Nigeria. They argue that ethnic jokes in Nigeria construct stereotypes about linguo-cultural signs, and that the jokes have been stripped of their stigmatizing effects owing to the ability of Nigerians to laugh collectively at their perceived prejudices and stereotypes. In a related article, ‘Impression Management and Face Sensitivities in Delta State Courtroom Interactions’, Olasimbo Takpor and Felix Ogoanah investigate impression management and courtroom interactions in High Court proceedings in Delta State of Nigeria within the theoretical framework of Rapport Management Model (Spencer Oatey). They conclude that to manage face sensitivities, courtroom interactions create diverse impressions of themselves or others by deploying impression management strategies such as self-promotion, intimidation, apologies, ingratiation and conformity as determined by the peculiarities of legal procedures and cultural norms, which mediate judicial proceedings, interpretations and decisions. Felix Ajiola’s ‘Colonial Capitalism and the Structure of the Nigerian Cocoa Marketing Board, 1947-1960’ examines the origin, structure and impact of the Nigerian Cocoa Marketing Board (NCMB) from its inauguration in 1947 up to 1960. The author argues that the NCMB served various interests and purposes, which hardly benefitted cocoa producers, but rather exploited them through intolerable taxes, harmful price regulations and unfavourable grading policies. In another article, ‘The Language Factor and Internet Penetration in Nigeria: A Practical Assessment’, Olushola Are examines all the unstated assumptions behind quests for more language options on the internet with specific reference to Nigeria. The author concludes that the provision of Nigerian language options online would not significantly enhance internet penetration in the country without broader adjustments to the roles and status of indigenous languages as well as greater socio-economic and political reforms to fight general social exclusion for which linguistic exclusion of any form may be merely symptomatic. In the sixth article, ‘Theatrical Intervention towards “Birth Preparedness and Complication Readiness”’, Oluwatoyin Olokodana-James examines Birth Preparedness and Complication Readiness (BPCR) strategies. She argues that BPCR reduces the risks of complications in that it helps health practitioners to detect danger signs from both mother and the newborn early enough. Using qualitative research approach, the author employs theatre and dance as interventionist tools to educate women within Ifako-Ijaye LGA in Lagos State on the usefulness of BPCR. In a different article on ‘Stress Patterning in Polysyllabic Words among Educated Yoruba Speakers of English in Lagos’, Emmanuel Osifeso investigates one hundred (100) undergraduate and post-graduate students across Lagos State to underscore the role of stress patterning of polysyllabic words among educated Yoruba speakers of English in Lagos (EYSEL). He concludes that EYSEL have a propensity for shifting the main stress in English polysyllabic words rightward. Victor Ariole’s article, ‘Peul (Fulani) Worldview as seen in Ba’s Work: A Critique’, identifies the cultural integration constraints in Africa using Ba’s discussion of the Peul/Fulani as a case study. He concludes that Ba’s thought patterns are quite relevant in understanding the Peul’s worldview which sees probity and constituents’ responsibilities as inalienable with peaceful living or existence. Babatunji Adepoju in the ninth article, ‘Cohesion in English Biblical Narratives: A Study of “The Prodigal Son”’, examines the different methods that writers/speakers employ in making English narratives coherent. He discusses the reasons why many texts are considered disjointed/disorganised thereby making such texts lose the desired radiance. He concludes that the unity of a text is enhanced by adherence to the appropriate usage of grammatical and lexical ties in English narratives. Ayodele Shotunde in ‘A Discourse on the Nature of Crime and Punishment in the Administration of Social Justice in an African Culture’ evaluates the nature of crime and punishment among the Yoruba of Nigeria. Adopting the critical and prescriptive methodology, he concludes that it is important to take an insightful look at the traditional Yoruba conception of crime and punishment given its embedded spirit of forgiveness because such has the potential of fostering better social ethics in contemporary Nigeria. In the next article, ‘China-Hong Kong Dual System: Twenty-Three Years of Uncertainty and Broken Promises’, Henry Ogunjewo argues that the relationship between China and Hong Kong in the last twenty-three years have been characterised by broken promises, failed covenants, unnecessary political meddling, judicial undercutting, press gagging and restrictions on freedom of expressions, leading to protests and political tension in Hong Kong. He concludes that the United Kingdom, former colonial administrator of Hong Kong, needed to bring international pressure on China to protect the interests of Hong Kong. Bisoye Eleshin’s ‘High-Toned Vowel Prefix in Yoruba’ examines prefixation as it relates to gerund derivation in Yoruba. He uses the morpho-syntactic approach to establish the claim that there actually exists a high-toned vowel prefix i- in Yoruba and that the class of noun it derives is gerund. The last paper by Mosunmola Ogunmolaji and Oyinade Adekunle ‘‘Madam Due Process’: The Public Life of Obiageli Ezekwesili’ is a biography of Obiageli Ezekwesili. The authors analyse the public life of Obiageli Ezekwesili providing insights into her lifestyle, especially the major forces that spurred her interest in politics and public administration. They conclude that Ezekwesili is an intellectual who has broken gender barriers in Nigeria. She possesses pragmatic understanding of the yearnings of Nigerians through deliberate identification of their problems, acquisition of necessary problem-solving tools, and swift responses to the problems whether or not she stepped on toes in the process. I hereby warmly recommend these articles to the academic community with the hope that scholars will find them interesting and useful. I congratulate the Editorial Team for a job well done despite the constraints of the COVID era! Professor Olufunke Adeboye Dean, Faculty of Arts Editor-in-Chief
Abstract We are investigating methods by which data from dependency syntax treebanks of ancient Greek can be applied to questions of authorship in ancient Greek historiography. From the Ancient Greek Dependency Treebank were constructed syntax words (sWords) by tracing the shortest path from each leaf node to the root for each sentence tree. This paper presents the results of a preliminary test of the usefulness of the sWord as a stylometric discriminator. The sWord data was subjected to clustering analysis. The resultant groupings were in accord with traditional classifications. The use of sWords also allows a more fine-grained heuristic exploration of difficult questions of text reuse. A comparison of relative frequencies of sWords in the directly transmitted Polybius book 1 and the excerpted books 9–10 indicate that the measurements of the two texts are generally very close, but when frequencies do vary, the differences are surprisingly large. These differences reveal that a certain syntactic simplification is a salient characteristic of Polybius’ excerptor, who leaves conspicuous syntactic indicators of his modifications.
This paper reports on the development of a French FrameNet, within the ASFALDA project. While the first phase of the project focused on the development of a French set of frames and corresponding lexicon (Candito et al., 2014), this paper concentrates on the subsequent corpus annotation phase, which focused on four notional domains (commercial transactions, cognitive stances, causality and verbal communication). Given full coverage is not reachable for a relatively " new " FrameNet project, we advocate that focusing on specific notional domains allowed us to obtain full lexical coverage for the frames of these domains, while partially reflecting word sense ambiguities. Furthermore, as frames and roles were annotated on two French Treebanks (the French Treebank (Abeillé and Barrier, 2004) and the Sequoia Treebank (Candito and Seddah, 2012), we were able to extract a syntactico-semantic lexicon from the annotated frames. In the resource's current status, there are 98 frames, 662 frame-evoking words, 872 senses, and about 13000 annotated frames, with their semantic roles assigned to portions of text. The French FrameNet is freely available at alpage.inria.fr/asfalda.
The article discusses the methodology and the preliminary results of the research project entitled “Latvian language in Monolingual and Bilingual Acquisition: tools, theories and applications” (LAMBA). The project involves 25 researchers – linguists, educators, psychologists – from five institutions in Latvia and Norway, and focuses on phonological, lexical and morphosyntactic acquisition of Latvian as a native language in monolingual and bilingual settings. One of the main goals of the project is to develop a set of norm-referenced language assessment tools that would allow for accurate and time-efficient evaluation of language development in pre-school children.The article will focus specifically on the Latvian adaptation of MacArthur-Bates Communicative Development Inventories – a parental report tool that assesses the development of receptive and productive vocabulary, and certain aspects of grammar. Two CDI forms were adapted in the project: CDI Words and Gestures designed for use with children between 8 and 16 months of age, and CDI Words and Sentences designed for 16- to 36-month old children. Each CDI form contains extensive and language-specific checklists of lexical items, communicative gestures and grammatical constructions.
Is there a single trajectory to third-language (L3) communicative proficiency in proficient, adult bilingual speakers? Parsimony favours such a possibility but in this theoretical paper we argue that multiple trajectories will be the norm. We focus on the processes of language control. These processes mediate the initial transfer of syntactic forms, entrain processes that change the language network and govern the selection of L3 syntactic structures and lexical items. Theoretical models of initial transfer differ in terms of their demands on top-down and bottom-up processes of language control. L3 learning, though, requires both types of process, yielding potential variability in the syntactic structures that populate the landscape of transfer. A language network can capture that landscape by tagging any existing structure (whether from the first language or from the second language) for use in the L3 by linking it to a L3 language node. Representational change incurs further processing costs because speakers must select L3 syntactic forms and lexical items in the face of competition. In line with earlier research, we propose that top-down control processes external to the language network help select outputs for speech production but these processes themselves must adapt to the demands of selecting amongst three rather than two languages. In a final section we review the nature of variability in language control processes and the processes they entrain. Such variability strongly predicts multiple trajectories to L3 proficiency. Exploring the nature of such variety, using converging methods in longitudinal designs, provides an opportunity for theoretical and practical advance.
We present a study on two key characteristics of human syntactic annotations: anchoring and agreement. Anchoring is a well known cognitive bias in human decision making, where judgments are drawn towards pre-existing values. We study the influence of anchoring on a standard approach to creation of syntactic resources where syntactic annotations are obtained via human editing of tagger and parser output. Our experiments demonstrate a clear anchoring effect and reveal unwanted consequences, including overestimation of parsing performance and lower quality of annotations in comparison with human-based annotations. Using sentences from the Penn Treebank WSJ, we also report the first systematically obtained inter-annotator agreement estimates for English syntactic parsing. Our agreement results control for anchoring bias, and are consequential in that they are \emph{on par} with state of the art parsing performance for English. We discuss the impact of our findings on strategies for future annotation efforts and parser evaluations.
The conventional randomized response design is unidimensional in the sense that it measures a single dimension of a sensitive attribute, like its prevalence, frequency, magnitude, or duration. This paper introduces a multidimensional design characterized by categorical questions that each measure a different aspect of the same sensitive attribute. The benefits of the multidimensional design are (i) a substantial gain in power and efficiency, and the potential to (i i) evaluate the goodness-of-fit of the model, and (i i i) test hypotheses about evasive response biases in case of a misfit. The method is illustrated for a two-dimensional design measuring both the prevalence and the magnitude of social security fraud.
The pol-nkjp1m-pargram-dev structure bank was created using POLFIE: an LFG grammar of Polish. This structure bank contains sentences from the NKJP1M subcorpus of NKJP which were not included in Skladnica treebank. The pol-nkjp1m-pargram-dev structure bank can be accessed via INESS treebanking system in two ways: • use the direct link: http://clarino.uib.no/iness/lfg-sentences?&treebank=pol-nkjp1m-pargram-dev • go to http://iness.uib.no --> choose Treebank Selection in the menu on the left-hand side --> choose POLFIE in Treebank Collections --> choose pol-nkjp1m-pargram-dev
Abstract: This technical note describes the US Army Research Laboratory (ARL) Arabic Dependency Treebank (AADT) for the purpose of documenting its release. The AADT was derived from existing Arabic treebanks distributed by the Linguistic Data Consortium using constituent-to-dependency conversion software written at ARL. Earlier versions of the AADT, as well as parsers trained from it, have been used in several published ARL research efforts, and, by releasing the data, we hope to facilitate additional Arabic language processing research by the greater community.
Background Conventionally, it is believed that high-frequency auditory information is important for speech understanding. This is only partly true, as recent studies have demonstrated the importance of low-frequency information. This research was taken up to develop, standardize, and validate auditory low-frequency word lists in Hindi, an Indian language. Material and Methods The first phase of the study involved collection of bisyllabic words followed by verification by a native linguist. Words were then short-listed based on familiarity ratings given by 10 adult native speakers; those words were recorded and the best recorded words selected through subjective and objective analysis. Then, using Fast Fourier Transform and k-means clustering, words with more energy below 1.5 kHz were isolated. Finally, equally difficult 10 word lists were generated by obtaining psychometric function curves. Finally, lists were administered on 40 adult normal hearing particip Results Results showed a similar trend of increase in speech identification scores with increase in SL across all lists except list 4. During the final phase, developed lists were validated on 10 simulated low-frequency cochlear hearing loss participants. Hearing loss was simulated using Matlab and National Institute for Occupational Safety and Health (NIOSH) software. Results of validation revealed that auditory low-frequency word lists were sensitive enough to tap the speech understanding difficulty in the simulated condition. Conclusions The developed word lists can be used clinically to assess communication ability in individuals with rising hearing loss. The word lists also have the potential to assess the performance after amplification provided to individuals with rising hearing loss.
Research finds we make spontaneous trait inferences from facial appearance, even after brief exposures to a face (i.e., less than or equal to 100 ms). We examined spontaneous impressions of criminality from facial appearance, testing whether these impressions persist after repeated presentation (i.e., one to three exposures) and increased exposure duration (100, 500, or 1,000 ms) to the face. Judgement confidence and response times were recorded. Other participants viewed the faces for an unlimited period of time, rating trustworthiness, dominance and criminal appearance. We found evidence that participants spontaneously make criminal appearance attributions. These inferences persisted with repeated presentation and increased exposure duration, were related to trustworthiness and dominance ratings, and were made with high confidence. Implications are discussed.
International audience
In this paper the systems submitted by the joint team of Dublin City University and National Taiwan University to the IALP 2016 Shared Task: Dimensional Sentiment Analysis for Chinese Words are presented. The systems learn the vector representation using Word2Vec algorithm for each Chinese word for sentiment analysis. The corpus used for the calculation of vector representation is 5 years (2006 to 2010) of the LDC Chinese Gigaword Fifth Edition corpus. The systems calculated similarities between a test Chinese word and each word in training corpus of the shared task with human annotation and took the valence-arousal ratings of the most similar words as the ratings of the test word. The performance of the submitted systems are around the same level of the shared task's baseline system. We will be looking at the performance gap with top-ranked systems in several aspects including corpus used for training and methodology.
In this paper, we describe the compilation of the Slovene Lexical Database; main focus being on developing the methodology to improve the tools used for lexicographic analysis and to introduce automatic data extraction in the lexicographic process. The semi-automated approach, which was devised in the last stages of database compilation, involved extracting corpus data, i.e. grammatical relations, collocations, examples, and grammatical labels, and conducting lexicographic analysis in the dictionary-writing system rather than in the corpus tool. An evaluation that compared the manual approach with the semi-automatic approach showed that the semi-automatic approach is much quicker and presents the lexicographers with almost all the information they identified as relevant during the manual analysis, as well as additional potentially relevant information for the dictionary entry. The final section of the paper proposes a few avenues for improvement of the semi-automated approach, including the implementation of crowdsourcing and additional post-processing of automatically extracted data.
Various treebanks have been released for dependency parsing. Despite that treebanks may belong to different languages or have different annotation schemes, they contain common syntactic knowledge that is potential to benefit each other. This paper presents a universal framework for transfer parsing across multi-typed treebanks with deep multi-task learning. We consider two kinds of treebanks as source: the multilingual universal treebanks and the monolingual heterogeneous treebanks. Knowledge across the source and target treebanks are effectively transferred through multi-level parameter sharing. Experiments on several benchmark datasets in various languages demonstrate that our approach can make effective use of arbitrary source treebanks to improve target parsing models.
We introduce an approach to train lexicalized parsers using bilingual corpora obtained by merging harmonized treebanks of different languages, producing parsers that can analyze sentences in either of the learned languages, or even sentences that mix both. We test the approach on the Universal Dependency Treebanks, training with MaltParser and MaltOptimizer. The results show that these bilingual parsers are more than competitive, as most combinations not only preserve accuracy, but some even achieve significant improvements over the corresponding monolingual parsers. Preliminary experiments also show the approach to be promising on texts with code-switching and when more languages are added.
In this paper, we propose a neural network model for graph-based dependency parsing which utilizes Bidirectional LSTM (BLSTM) to capture richer contextual information instead of using high-order factorization, and enable our model to use much fewer features than previous work. In addition, we propose an effective way to learn sentence segment embedding on sentence-level based on an extra forward LSTM network. Although our model uses only first-order factorization, experiments on English Peen Treebank and Chinese Penn Treebank show that our model could be competitive with previous higher-order graph-based dependency parsing models and state-of-the-art models.
Affective computing is a very important issue. An increasing amount of research focused on representing affective states as continuous numerical values on multiple dimensions. Such as the emotional space, which is about the valence and arousal. Due to the affective dimension representation can be useful to sentiment analysis, building dimensional sentiment resources with valence-arousal ratings are very important. Therefore, this study proposes a method to automatically obtain the valence-arousal ratings of affective words. Experiment results using the evaluation metrics to get the error rates about the mean absolute error and pearson correlation coefficient.
We describe the Corpus of Spoken Icelandic (ÍS-TAL) which is made up of 15 hours of spontaneous naturally occurring conversa-tions, 31 conversations in all. The corpus comprises 184,080 tokens, 14,297 types and 9,221 lemmas. It has been transcribed using standard orthography. We present a list of the 30 most common lemmas in the corpus and compare it to a list of the most frequent lemmas in the written language, concluding that the differences between the two lists are smaller than expected. We have tagged the corpus morphologically with a statistical tagger that had been trained on written texts. The results are much better than we expected, and the tagging accuracy is as least as high as for the written texts. The final part of the paper is a report on a work in progress. We have been experimenting with converting the morphological tagging into a shallow syntactic markup by applying a few simple hand-written rules. Even though the analysis we get by using this procedure is bound to be incomplete and contain several errors, we conclude that the results are promising and we can use this method to build a simple yet useful treebank with minimal effort. 1.
This work elaborates the semi-semantic part of speech annotation guidelines for the URDU.KON-TB treebank: an annotated corpus. A hierarchical annotation scheme was designed to label the part of speech and then applied on the corpus. This raw corpus was collected from the Urdu Wikipedia and the Jang newspaper and then annotated with the proposed semi-semantic part of speech labels. The corpus contains text of local & international news, social stories, sports, culture, finance, religion, traveling, etc. This exercise finally contributed a part of speech annotation to the URDU.KON-TB treebank. Twenty-two main part of speech categories are divided into subcategories, which conclude the morphological, and semantical information encoded in it. This article reports the annotation guidelines in major; however, it also briefs the development of the URDU.KON-TB treebank, which includes the raw corpus collection, designing & employment of annotation scheme and finally, its statistical evaluation and results. The guidelines presented as follows, will be useful for linguistic community to annotate the sentences not only for the national language Urdu but for the other indigenous languages like Punjab, Sindhi, Pashto, etc., as well.
Treebanks are curial for natural language processing (NLP). In this paper, we present our work for annotating a Chinese treebank in scientific domain (SCTB), to address the problem of the lack of Chinese treebanks in this domain. Chinese analysis and machine translation experiments conducted using this treebank indicate that the annotated treebank can significantly improve the performance on both tasks. This treebank is released to promote Chinese NLP research in scientific domain.
This article proposes an ontology design pattern for leading knowledge providers to represent knowledge in more normalized, precise and interrelated ways, hence in ways that help the matching and exploitation of knowledge from different sources. This pattern is a knowledge sharing best practice that is domain and language independent. It can be used as a criteria for measuring the quality of an ontology. This pattern is: using binary relation types directly derived from concept types, especially role types or types of process. The article explains and illustrates this pattern, and relates it to other patterns and general ontology quality criteria. It also provides an ontology for automatically deriving relation types from concept types (e.g., those from lexical ontologies such as those derived from the WordNet lexical database). This derivation helps normalizing knowledge, reduces having to introduce new relation types and helps keeping all the types organized.
Two of the major problems in social media message classification are the data sparseness issue and the high degree of lexical variation. Paraphrases, or synonyms, are alternative ways of expressing the same meaning using different lexical variations. In this study, we try to use paraphrases to improve tweet topic classification performance. We explored two approaches to generating paraphrases, WordNet, which is a lexical database grouping English words into sets of synonyms, and word embeddings, which are learned from millions of tweets and billions of words. Our experiment shows that using paraphrases can improve the topic classification task, and the word embedding approach outperforms the WordNet method. To our knowledge, this is the first study exploiting paraphrases for tweet classification.
This paper presents the conversion of Syn-TagRus dependency structures into Penn Treebank style phrase structures, whose resulting data will be used to train a statistical constituency parser for Russian and create a large-scale constituency-parsed corpus. The implemented conversion includes various innovative features in order to create phrase structure trees that are closest to Penn Treebank style while optimally preserving information of the original dependency structure annotations. We believe the newly converted phrase structure treebank will be not only an adequate training dataset for our ongoing project but also a valuable resource for traditional and computational linguistic research.
In the recent years, sentiment analysis has emerged as a major research problem in the field of Natural Language Processing. Here, the problem is to identify the sentiment/emotion in given sentence/paragraph. Usually it is positive, negative and neutral. Here, we consider only binary classification task (positive and negative). We have considered the best performing sentiment analysis model which is a ensemble of NB-SVM, Paragraph2Vec and RNN. We added CNN into this stacking model and showed that our ensemble model perform better than the existing one. We achieved the state of the art performance on IMDB Movie review dataset, Stanford sentiment treebank dataset (SST) and Elec reviews dataset.
Neural network training has been shown to be advantageous in many natural language processing \napplications, such as language modelling or machine translation. In this paper, we describe in \ndetail a novel domain adaptation mechanism in neural network training. Instead of learning \nand adapting the neural network on millions of training sentences – which can be very timeconsuming or even infeasible in some cases – we design a domain adaptation gating mechanism \nwhich can be used in recurrent neural networks and quickly learn the out-of-domain knowledge \ndirectly from the word vector representations with little speed overhead. In our experiments, \nwe use the recurrent neural network language model (LM) as a case study. We show that the \nneural LM perplexity can be reduced by 7.395 and 12.011 using the proposed domain adaptation \nmechanism on the Penn Treebank and News data, respectively. Furthermore, we show that using \nthe domain-adapted neural LM to re-rank the statistical machine translation n-best list on the \nFrench-to-English language pair can significantly improve translation quality
This journal article carries out a structural-functional analysis of the formation of Old English nouns by means of affixation. The data comprise a total of 4,370 nouns which result from either prefixation or suffixation, retrieved from the lexical database of Old English Nerthus. Twenty-five derivational functions, inspired by functional grammars and Pounder’s (2000) paradigmatic morphology are proposed to explain the relationship holding between affixes and their bases of derivation. These functions have been divided into split and unified, the former being realized by both prefixes and suffixes and the latter by either prefixation or suffixation. The conclusion is reached that the main target of prefixation is the modification of meaning, in such a way that the meaning of the derivative is less predictable from the input category whereas the main target of suffixation is the change of lexical category, given that the meaning of the derivative is more predictable from the the input category.
Resumen El presente trabajo tiene como objetivo analizar con qué elementos textuales y mediante qué procedimientos y estrategias se explicita la norma especialmente en el diccionario bilingüe de Lucio Ambruzzi: Nuovo dizionario spagnolo-italiano e italiano-spagnolo (1948-49). Adquiere, para ello, un gran relieve el análisis de las dos secciones del diccionario, en las que se observa la presencia del legado de la tradición nacional en la que se colocan: el diccionario académico usual de 1925 y el manual e ilustrado de 1927 en el ámbito español; los diccionarios de Panzini, Migliorini, Monelli e Jàcono en el ámbito italiano. En el diccionario bilingüe de Lucio Ambruzzi (DBA) se analizarán las marcas de usos y los comentarios normativos en la microestructura de los neologismos, especialmente de los extranjerismos. El trabajo, en su desarrollo, tratará de indicar las características de los comentarios normativos explícitos, situándolos en el contexto histórico y cultural de la época a la que pertenecen, para lo cual se analizarán los diccionarios españoles e italianos consultados por Ambruzzi, que reflejan la mentalidad general y la ideología vigente respecto a los extranjerismos. Palabras clave: Ambruzzi, diccionario bilingüe, norma, neologismos, extranjerismos. Abstract This paper aims to analyse the textual elements and the procedures and strategies that Lucio Ambruzzi’s bilingual dictionary [Nuovo dizionario spagnolo-italiano e italiano-spagnolo (1948-49)] uses to specify the linguistic norm. For this reason, it is very important to study the two dictionary sections, where we can notice the national tradition legacy in which are placed (Spanish academic diccionary: “usual” 1925 and “manual e ilustrado” 1927. Italian dictionaries: Panzini, Migliorini, Monelli and Jàcono). In Lucio Ambruzzi’s bilingual dictionary (DBA) we will analyse the usage labels and the linguistic norm comments in the microstructure of neologisms, specially foreign terms. The paper points to indicate and catalogue the labels and explicit linguistic norm comments characteristics, establishing the historical and cultural context of the period they belong to. For this purpose, we will study the Spanish and Italian dictionaries Ambruzzi consulted, which reflect the general mind-set and the author’s ideology regarding foreign words. Keywords: Ambruzzi, bilingual dictionary, linguistic norm, neologism, foreignisms.
Recently, there has been an explosion in the availability of large, good-quality cross-linguistic databases such as WALS (Dryer & Haspelmath, 2013), Glottolog (Hammarstrom et al., 2015) and Phoible (Moran & McCloy, 2014). Databases such as Phoible contain the actual segments used by various languages as they are given in the primary language descriptions. However, this segment-level representation cannot be used directly for analyses that require generalizations over classes of segments that share theoretically interesting features. Here we present a method and the associated R (R Core Team, 2014) code that allows the exible denition of such meaningful classes and that can identify the sets of segments falling into such a class for any language inventory. The method and its results are important for those interested in exploring cross-linguistic patterns of phonetic and phonological diversity and their relationship to extra-linguistic factors and processes such as climate, economics, history or human genetics.
Recurrent neural networks have been very successful at predicting sequences of words in tasks such as language modeling. However, all such models are based on the conventional classification framework, where the model is trained against one-hot targets, and each word is represented both as an input and as an output in isolation. This causes inefficiencies in learning both in terms of utilizing all of the information and in terms of the number of parameters needed to train. We introduce a novel theoretical framework that facilitates better learning in language modeling, and show that our framework leads to tying together the input embedding and the output projection matrices, greatly reducing the number of trainable variables. Our framework leads to state of the art performance on the Penn Treebank with a variety of network models.
OBJECTIVES: Theoretical models of adult development suggest changes in emotion systems with age. This study determined how younger and older adults judged and classified 70 emotion terms that varied in valence and arousal, and that have been used in previous studies of adult aging and emotion. The terms were from the Positive and Negative Affect Schedule - Expanded (PANAS-X) and the (KS) affect scales. METHOD: Older (n = 32) and younger adults (n = 111) engaged in a card sort task which determined how the 70 emotion terms were classified (i.e. grouped) in relation to one another. Activation and valence ratings of emotion terms were collected. RESULTS: There were 17 age group differences in item ratings for activation and 19 for valence. Older adults tended to rate emotion terms and scales as more positive and activating than younger persons. Card sort data indicated similarity in conceptualizations of emotion terms across groups with exceptions for serene, sad, and lonely. CONCLUSIONS: Research that utilizes self-report emotion data from older and younger persons should consider how perceptions of emotion terms may vary systematically with age. The constructs of sadness, loneliness, and serene may be age-variant and necessitate age-based adjustments in assessment and intervention. Further, older adults may perceive some emotion terms to be more activating and positive than younger persons.
Objectives: This study investigated the role of response style biases in the assessment of positive and negative affect in aging research; it addressed whether response styles (a) are associated with age-related changes in cognitive abilities, (b) lead to distorted conclusions about age differences in affect, and (c) reduce the convergent and predictive validity of affect measures in relation to health outcomes. Method: A multidimensional item response theory model was used to extract response styles from affect ratings provided by respondents to the psychosocial questionnaire (n = 6,295; aged 50-100 years) in the Health and Retirement Study (HRS). Results: The likelihood of extreme response styles (disproportionate use of "not at all" and "very much" response categories) increased significantly with age, and this effect was mediated by age-related decreases in HRS cognitive test scores. Removing response styles from affect measures did not alter age patterns in positive and negative affect; however, it consistently enhanced the convergent validity (relationships with concurrent depression and mental health problems) and predictive validity (prospective relationships with hospital visits, physical illness onset) of the affect measures. Discussion: The results support the importance of detecting and controlling response styles when studying self-reported affect in aging research.
Treebanks have recently been released for a number of languages with the harmonized annotation created by the Universal Dependencies project. The representation of certain constructions in UD are known to be suboptimal for parsing and may be worth transforming for the purpose of parsing. In this paper, we focus on the representation of verb groups. Several studies have shown that parsing works better when auxiliaries are the head of auxiliary dependency relations which is not the case in UD. We therefore transformed verb groups in UD treebanks, parsed the test set and transformed it back, and contrary to expectations, observed significant decreases in accuracy. We provide suggestive evidence that improvements in previous studies were obtained because the transformation helps disambiguating POS tags of main verbs and auxiliaries. The question of why parsing accuracy decreases with this approach in the case of UD is left open.
This study investigates whether individual differences in attachment status can be detected by electrophysiological responses to loss-themed pictures. The Adult Attachment Interview (AAI) was used to identify discourse/reasoning lapses during the discussion of loss experiences via death that place speakers in the Unresolved/disorganized AAI category. In parents, Unresolved AAI status has been associated with Disorganized infant Strange Situation response, a known risk factor for psychopathology (e.g., internalizing/externalizing/dissociation). This association has been related to anomalous frightening (FR) parental behavior in the infant's presence, behavior presumed to be instigated by vulnerability to trauma-related fright. Here, psychophysiological methods were utilized to examine whether Unresolved AAI status could be detected in brain responses to subtle/symbolic reminders of loss. One year after AAI administration, 31 undergraduate women who had experienced loss (16 Unresolved) underwent continuous electroencephalogram (EEG) recording during a picture-viewing, valence-rating task. Picture onset-locked event-related potentials (ERPs) revealed millisecond responses to 4 picture categories: pleasant people, pleasant nature, cemetery (symbolic death), and gruesome death (dead or dying people). Participants' valence ratings did not differ between groups across picture categories. However, the N2 ERP, implicated in detecting stimulus salience, was selectively greater in Unresolved participants viewing cemetery scenes; it was in fact as high as the N2 for gruesome death images observed throughout the sample. Additionally, Unresolved participants exhibited a right-hemispheric P3 asymmetry across picture categories, suggestive of continuously heightened vigilance/arousal. Together, these results suggest that Unresolved AAI status is associated with greater neurophysiological sensitivity to subtle reminders of loss that may disrupt ongoing mental function. (PsycINFO Database Record
The present study evaluated the efficacy of adding a virtual reality (VR) component to the treatment of compulsive hoarding (CH), following inference-based therapy (IBT). Participants were randomly assigned to either an experimental or a control condition. Seven participants received the experimental and seven received the control condition. Five sessions of 1 h were administered weekly. A significant difference indicated that the level of clutter in the bedroom tended to diminish more in the experimental group as compared to the control group F(2,24) = 2.28, p = 0.10. In addition, the results demonstrated that both groups were immersed and present in the environment. The results on posttreatment measures of CH (Saving Inventory revised, Saving Cognition Inventory and Clutter Image Rating scale) demonstrate the efficacy of IBT in terms of symptom reduction. Overall, these results suggest that the creation of a virtual environment may be effective in the treatment of CH by helping the compulsive hoarders take action over their clutter.